Per-token pricing with residency enforcement included
Every tier includes jurisdiction enforcement, refusal semantics and residency records.
Developer
Per token
billed from the first request
Full catalogue access with jurisdiction enforcement on. Billing is per token from the first request. There is no free tier.
- All catalogue models, one key
- Jurisdiction checks on every request
- Prefix caching
- Community support
Residency posture
Single-jurisdiction requests
Contact us to create an accountScale
Most usedPer token
usage-based
Production traffic with residency enforcement on each request, spend ceilings and the reports your reviewers ask for.
- Everything in Developer
- Residency policy per key and team
- Spend limits with overrides
- Request-level residency logs
- Priority serving paths
Residency posture
Multi-jurisdiction, policy-enforced
Contact us to create an accountSovereign
Contract
annual
For organisations that must show every routing decision to a regulator or an internal review board.
- Everything in Scale
- Bring your own provider keys
- Residency attestations and export
- Fail-closed checks in your CI
- Named engineer
Residency posture
Attested, exportable and fail-closed
Book a walkthroughRepresentative models, with the regions that serve them
These six catalogue entries show the range, rather than the cheapest case.
Representative per-million-token rates| Model | Provider | Input / 1M | Output / 1M | Serving regions |
|---|
| DeepSeek V4 Pro | DeepSeek | $0.30 | $1.20 | CN · cn-beijing · ap-southeast-1 |
| Kimi K3 | Moonshot AI (Kimi) | $3.00 | $15.00 | cn-beijing · cn-hongkong · ap-southeast-1 |
| GLM-5.2 | Z.ai (GLM) | $1.40 | $4.40 | ap-east-1 |
| Qwen3.8-Max | Alibaba Qwen | $2.00 | $6.00 | ap-southeast-1 · ap-east-1 · ap-southeast-5 |
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 | us-east-1 · ap-east-1 |
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | ap-east-1 |
- These prices change when provider rates change.
- Region lists show only where enforcement can serve a model. An unenforced request may land elsewhere.
- Bring-your-own-key traffic uses the same policy layer and carries no token markup.
03Embeddings and rerankers
Embedding and reranking rates
Embedding and reranking rates| Band | Rate |
|---|
| Small embeddings (up to 350M parameters) | $0.01 / 1M input tokens |
| Mid-size embeddings | $0.02 / 1M input tokens |
| Rerankers | $0.05 / 1M input tokens |
04Costs of residency enforcement
Cache costs, token prices and refusals
Residency constraints affect cache hit rates, token prices and request availability.
- 01
Cache hit rate drops
Prefixes are cached by jurisdiction instead of globally, so hit rate is lower than it could be. The cost of partitioning is measurable.
- 02
You may pay more per token
Within a small eligible set, the cheapest global endpoint may be ineligible. The rate you pay is the best eligible one, and it can exceed the global best.
- 03
You may get a refusal
Where no eligible endpoint serves a model, the request fails. For a product team, that is a visible lost request. Coverage planning is the mitigation for that case, rather than a fallback.
- 04
Bring your own key removes the markup
Traffic on your provider keys goes through the same policy layer and carries no token markup. You pay for routing, while the provider contract remains yours.
Credits
Free credits and a promotional tier are not offered. Billing starts with the first token for every account. Enforcement is part of every tier, with no add-on charge. This site does not take a card.
Contact us for a quotation
Start
Per-token billing
Every tier includes enforcement, refusal semantics and residency records.