Sovereign AI Inference. Faster. Cheaper. Ready now.
ASIC-accelerated inference with OpenAI-compatible APIs, predictable STU pricing, and on-shore governance for enterprise and government workloads.
Why this is different
Why teams choose SCX.ai
Performance that scales
Up to 10x more tokens per watt vs GPU-only clouds, with sub-100ms latency for real-world apps.
Lower $/inference by design
Efficient silicon + orchestration means you do more with every dollar (and every kWh).
Sovereign by default
Data, models, and logs remain in-region with auditability and strict egress control.
Model choice, one API
Access open and commercial models, plus private fine-tunes and RAG in a single workflow.
How STUs work
STUs are standard compute units that make usage predictable. Each workload “spends” STUs at a known rate. Your plan includes a monthly STU allowance; if you run out, we throttle instead of charging surprise overage. Add a top-up anytime to restore full speed.
| Workload | Volume | Cost |
|---|---|---|
| Band-L LLM | 1M tokens | 1 STU |
| Band-S LLM (70B-class) | 1M tokens | 2 STU |
| Band-P (premium / MoE) | 1M tokens | 8 STU |
| STT (batch) | 1 hour | 0.40 STU |
| TTS (standard) | 1M chars | 0.25 STU |
| Embeddings (small) | 1M tokens | 0.033 STU |
Plans & pricing
No automatic overage — safe throttling if you run out; instant top-up available.
Starter
400 STU/mo
Best for pilots and early launches
- OpenAI-compatible API
- Band-L & Band-S models
- Standard support
Growth
5,000 STU/mo
Scale to meaningful production; Band-P unlock available
- Everything in Starter
- Band-P unlock available
- RAG Vault add-on ready
- Priority support
- Fine-tune access
Enterprise
15,000 STU/mo
Large workloads, highest caps, Band-P native
- Everything in Growth
- Band-P native access
- Secure tool runner
- Audit export & compliance
- Dedicated support
- Custom SLAs
30-second calculator
Estimate your STU usage and recommended plan
Security & compliance
Sovereign & sustainable
Proven Tier-3+ colo
Proven infrastructure with specialised silicon optimised for efficiency.
Lower gCO2e/inference
Efficient ASIC silicon reduces energy and water intensity per inference call.
Renewable-forward
Strategy to progressively reduce carbon intensity across operations.
Developer-friendly
Drop-in replacement for OpenAI. Use your existing code, SDKs, and workflows — just point to SCX.ai.
curl https://api.scx.ai/v1/chat/completions \
-H "Authorization: Bearer $SCX_API_KEY" \
-d '{
"model": "scx-band-l",
"messages": [{"role": "user", "content": "Hello"}]
}'