The serving layer beneath the router
Menu Items does not operate GPUs. It reaches models through provider endpoints and presents them on one OpenAI-compatible surface. Every call carries the jurisdiction decision.
- 1
- Endpoint for every provider
- Per-request
- Jurisdiction enforcement
- Partitioned
- Cache scope
- 0
- Silent cross-region fallbacks
OpenAI-compatible
Never shared across a border
Surface
Endpoint, providers, policy
An existing client points here with a base-URL change.
| Property | What you get | What it means in practice |
|---|---|---|
| Endpoint shape | OpenAI-compatible | Chat completions and embeddings use the same surface. An existing client changes only its base URL for integration |
| Streaming | Server-sent events | Streaming never permits a request to leave the declared boundary |
| Retries | Confined to eligible providers | Retries stay within eligible providers. Load can shrink the candidate set, but it cannot widen past a border |
| Caching | Partitioned by jurisdiction | A warm prefix in one jurisdiction is not reused in another |
| Rate limits | Per key and per workspace | Limits and jurisdiction are set at the same layer. Cost controls cannot break the residency guarantee |
| Concurrency | Governed by provider capacity | Headroom is planned in-region. Spillover is not offered as a capacity planning lever |
Caching
Partitioned caching
Prefixes are cached separately per jurisdiction.
Example global cache pool
- Highest hit rate across all traffic
- A warm prefix can pull work into another region
- Hit rate measured per jurisdiction
- Warm prefixes never cross a boundary
- The cost uplift of partition is measured
Limits
Serving limits
Coverage is bounded by the providers we reach
A model with no endpoint inside a jurisdiction cannot be served there. The catalogue states coverage per model and region.
We depend on provider availability
Availability follows provider health. When one degrades, the eligible set shrinks.
Provider-disclosed regions only
Only provider-disclosed regions qualify. Where a provider publishes no serving region, the endpoint is ineligible for jurisdiction-bound traffic. We do not infer a location from latency, support pages or marketing claims.
Latency
Routing inside a small eligible set can add latency compared with routing to the nearest global endpoint.
Start