The economics of routing under a boundary
Inference cost is decided before the model-provider invoice. A routing layer chooses, request by request, where a workload may run.
Contents
Seven sections
Written for the person who has to defend an inference bill to a reviewer and a finance team in the same meeting.
Cost per request
Cost per token is quoted, but cost per request is decided earlier. It depends on eligible endpoints and warm caches. It also depends on the regions your policy has already excluded.
Eligible endpoint pricing
The gap between the cheapest global endpoint and the best eligible endpoint is a policy outcome.
Cache partition costs
Jurisdiction-scoped caching lowers hit rate.
The cost of a refusal
A fail-closed router produces refusals, and refusals carry a price. Work goes unserved or gets retried. Engineers also spend time on coverage.
In-region open-weight hosting
When provider coverage is thin, hosting open-weight models inside your boundary is the only option independent of a provider's region list.
East Asian cost structure
The East Asian markets in this catalogue produce a different cost structure from the US and EU. The paper explains the difference.
Planned operating metrics
The metrics we intend to publish include refusal rate by jurisdiction and the measured cost of cache partition.
- Format
- Typeset PDF, 7 sections
- Author
- Menu Items engineering
- Revision
- Third edition, August 2026
- Scope
- Cost mechanics of residency-constrained inference. Not legal advice; jurisdiction rules are documented separately in the catalogue.
We store only your email address and the organisation, jurisdiction and monthly token volume band you submit here. This form does not take a card.
Download
Download access
Costs covered
The paper covers reduced cache hit rates, higher eligible token prices and request refusals.
Revisions and measurements
It carries a date and a revision number. We intend to publish measured operating numbers against it.
Catalogue first
Start