Jurisdiction-enforced inferenceAccounts are reviewed per jurisdiction and billed from the first token. No free credits, no promotional tier.Contact us for a quotation

Product · Inference

The serving layer beneath the router

Menu Items does not operate GPUs. It reaches models through provider endpoints and presents them on one OpenAI-compatible surface. Every call carries the jurisdiction decision.


1
Endpoint for every provider

OpenAI-compatible

Per-request
Jurisdiction enforcement
Partitioned
Cache scope

Never shared across a border

0
Silent cross-region fallbacks
01

Surface

Endpoint, providers, policy

An existing client points here with a base-URL change.

Serving properties
PropertyWhat you getWhat it means in practice
Endpoint shapeOpenAI-compatibleChat completions and embeddings use the same surface. An existing client changes only its base URL for integration
StreamingServer-sent eventsStreaming never permits a request to leave the declared boundary
RetriesConfined to eligible providersRetries stay within eligible providers. Load can shrink the candidate set, but it cannot widen past a border
CachingPartitioned by jurisdictionA warm prefix in one jurisdiction is not reused in another
Rate limitsPer key and per workspaceLimits and jurisdiction are set at the same layer. Cost controls cannot break the residency guarantee
ConcurrencyGoverned by provider capacityHeadroom is planned in-region. Spillover is not offered as a capacity planning lever
02

Caching

Partitioned caching

Prefixes are cached separately per jurisdiction.

Example global cache pool

  • Highest hit rate across all traffic
  • A warm prefix can pull work into another region

Menu Items partitioned cache

  • Hit rate measured per jurisdiction
  • Warm prefixes never cross a boundary
  • The cost uplift of partition is measured
03

Limits

Serving limits

  1. 01

    Coverage is bounded by the providers we reach

    A model with no endpoint inside a jurisdiction cannot be served there. The catalogue states coverage per model and region.

  2. 02

    We depend on provider availability

    Availability follows provider health. When one degrades, the eligible set shrinks.

  3. 03

    Provider-disclosed regions only

    Only provider-disclosed regions qualify. Where a provider publishes no serving region, the endpoint is ineligible for jurisdiction-bound traffic. We do not infer a location from latency, support pages or marketing claims.

  4. 04

    Latency

    Routing inside a small eligible set can add latency compared with routing to the nearest global endpoint.

In-region servingProvider capacityNo cross-border spillover


Start

Base URL integration with residency enforcement