Jurisdiction-enforced inferenceAccounts are reviewed per jurisdiction and billed from the first token. No free credits, no promotional tier.Contact us for a quotation
4 Sep 2026Platform

A cache across borders is a transfer you did not plan

Prefix reuse is the largest single cost lever in inference. It also moves data, so it belongs on the same side of the boundary as the rest of the request.

Prefix caching is one of the strongest cost controls in modern inference. Long shared system prompts and reused context are billed once, then reused, so a stable workload can save substantially.

A prefix is stored near the accelerator that will reuse it, and a matching request is served from that copy. Global cache pools maximise hit rate by ignoring location, which creates the problem.

With a global cache, a warm prefix can pull work toward it. A request that policy would confine to one jurisdiction may then be served from another. The dashboard reports the saving but often omits the transfer.

Partition cache scope by jurisdiction. Shared prefixes are cached once per boundary rather than once globally. The hit rate falls without relaxing the residency guarantee. The resulting cost increase is measured.

When teams compare routers on price, they should ask whether caching exists and what its scope is.



01

More notes

Next in the set