The whole design, in writing
Learn AI system design by building an LLM gateway and model router step by step. An interactive guide to a multi-provider control plane — a unified API, provider adapters, policy-based routing, fallback and retries, per-tenant rate limits and budgets, response caching, and cost/latency observability — so one endpoint serves many models reliably and affordably.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
Why put a gateway in front of the models?
Your product calls an LLM. Easy — until you have a dozen services each hard-coding one vendor’s SDK, key, and quirks. Then that vendor has an outage, or triples its price, or a cheaper model ships, or one feature floods your shared rate limit at 2am. Every one of those is now a code change in a dozen places. How do you make "which model, from whom, under what limits" a runtime decision instead of a compile-time one?
Put an LLM Gateway in front of every provider: one stable, provider-agnostic API your apps code against, with a model router, fallback, rate limits, caching, and cost accounting behind it. Swap models, add vendors, and enforce budgets centrally — without touching a single app.
Step 1 · The skeleton
One API for every model
Apps need completions from many models over time. What should they actually call, so that changing the underlying model never changes their code?
What’s the right contract between your apps and the models?
That scatters vendor keys, retry logic, and model choices across every service — the exact coupling a gateway removes. A model swap becomes a dozen deploys.
A single unified API (often OpenAI-compatible) lets apps say "complete this" without naming a vendor. Routing, keys, limits and fallback all live behind it.
A library still runs in every app’s process — you can’t change routing, revoke a key, or enforce a global budget without redeploying them all. Control has to be a network hop, not a dependency.
Apps call one LLM Gateway endpoint with a provider-agnostic request (messages, params, maybe a task hint). The gateway authenticates the caller and owns everything downstream. Because it’s a network boundary, you change routing, keys, limits and providers centrally — the apps never notice.
Why this piece earns its place
The unified API is a product you now own and have to version. Being OpenAI-compatible buys instant adoption — every SDK already speaks it — and quietly commits you to a schema you do not control, so the first time a vendor ships something the canonical shape has no field for, every app that wants it is blocked on a gateway release. The usual escape hatch is a passthrough bag of vendor-specific options the adapter forwards untouched. Take it deliberately: anything in that bag is a request the router can no longer freely move elsewhere, so the field that unblocks a team today is the field that breaks failover tomorrow. Mark those requests as pinned rather than pretending they are portable. The other thing that happens here and nowhere else is identity. The tenant has to be derived from the credential presented at the edge, never read out of a field the caller sets, because that derived tenant is the subject every later step keys on — buckets, budgets, cache scope, cost attribution. Get it from a header an app can lie about and all four inherit the lie at once.
What the new pieces do
- Applicationclient
- Any service or product surface that wants a completion. It codes against one stable, provider-agnostic API and never talks to a model vendor directly.
- LLM Gatewaybackend
- The single control-plane endpoint. Authenticates the caller, enforces limits, checks the cache, asks the router which model to use, calls the provider (with fallback), and records cost + latency — behind one OpenAI-compatible contract.
Step 2 · Speak every dialect
Provider adapters normalize the vendors
Provider A, Provider B and your self-hosted model each have different request shapes, auth, streaming formats, error codes, and token accounting. The gateway promises one contract. How does it talk to all of them?
How does one API front three incompatible provider APIs?
You don’t control vendors’ APIs — they’ll never conform. The translation has to live on your side, per provider.
Then apps must special-case each vendor again — you’ve moved the coupling, not removed it. The gateway’s value is a single normalized shape.
A thin per-provider adapter translates the gateway’s canonical request into the vendor’s format (and its response, streaming chunks, errors and token counts back). Add a vendor = add an adapter.
Each provider gets an adapter that maps the gateway’s canonical request ↔ that vendor’s native API — including auth, streaming format, error codes, and token accounting. The core stays vendor-neutral; onboarding a new model is writing one adapter, not rewriting the gateway.
Why this piece earns its place
Step 4's retry policy cannot branch on vendor error codes, so what the adapter has to hand inward is a class, not a message: retryable on this provider, retryable only somewhere else, never retryable, the caller's own fault. Only the adapter can decide which is which — one vendor signals an over-length context with a 400 while another buries it in a 200 carrying an error object in the body, and one status code can mean "slow down for a second" at one vendor and "you have no credit left" at the next, which are a wait and a dead end respectively. Token accounting is the other thing that stays invisible until finance asks. Usage counts arrive differently everywhere, and some vendors send them only in a final streaming frame — which never arrives if the client disconnects or the call times out, so the ledger loses precisely the requests that cost the most. The adapter owes an estimate when the true number is missing, flagged as an estimate. And because adapters are the one layer a vendor can break while you deploy nothing, they need contract tests run against the live provider on a schedule; recorded mocks keep passing happily for months after the thing they recorded changed.
- 1 adapterper provider
- canonical shapeone request/response
- add vendor= add adapter, no core change
What the new pieces do
- Provider Amodel
- A hosted frontier model (e.g. a top-tier API). Strongest, priciest — the router sends hard tasks here.
- Provider Bmodel
- A second vendor’s smaller, faster, cheaper model. Good for easy tasks and a fallback when Provider A is down or rate-limited.
- Self-hostedmodel
- An open-weights model you run yourself. No per-token vendor cost and full data control — another lane for routing and failover.
Step 3 · Pick the right model
The router chooses per request
Now several models are reachable through one API. But which one should a given request use? "Always the biggest" is slow and expensive; "always the cheapest" is the trap you’ll feel later. How does the gateway decide, per request?
What should the router optimize when choosing a model?
Sending a one-line classification to a frontier model burns money and latency for no gain. Overkill is as wrong as underkill; the point is to match task to model.
The router reads a policy: route by required capability (reasoning, tools, context length, JSON strictness), then let cost and latency choose among the models that can actually do the job.
Cost-only routing sends hard tasks to weak models and silently tanks quality — no error, just wrong answers. (That’s literally the chaos button on this page.)
A Model Router chooses per request from a policy in the Model Registry: match the required capability first (reasoning depth, tool use, context length, structured-output reliability), then let cost and latency break ties within a capability tier. Pins and per-tenant overrides sit on top. The registry updates without a redeploy.
Why this piece earns its place
Hot-reloadable policy means the router's input is mutable, and that has a consequence you meet during an incident rather than at design time: with no version stamped on the decision, "why did this request go to that model" is unanswerable a week later. Record the registry version and the rule that matched beside the model id, so the ledger carries a decision you can replay instead of a name you can only guess at. Pin that snapshot for the life of a request, too — policy that reloads mid-flight can route a second attempt by different rules than the first, and a request explained by two policies is a request nobody can explain. The flip side of "no redeploy" is that a routing change also skips everything a deploy brings with it: review, a staging run, a rollback somebody has practised. Edit a tier late in the afternoon and it is live for every tenant in seconds with no artifact to revert to, so the registry wants versioned writes and a one-command rollback long before it wants more rules. And keep the decision cheap. The router sits in every request's critical path, so evaluate policy in memory against a cached snapshot with a background refresh; fetching config per request adds a network hop to the one component whose job is to add none, and turns the registry into a second thing that can take the gateway down.
- capabilitythe primary key
- cost / latencytie-breakers within a tier
- registrypolicy, hot-reloaded
What the new pieces do
- Model Routerservice
- Chooses a target model per request from a policy: task/capability first, then cost and latency as tie-breakers, honoring pins and overrides. The brain of the gateway.
- Model Registrystore
- The config the router reads: available models, their capabilities, prices, context limits, health, and the routing rules — updatable without a redeploy.
Back of the envelope
- classify the task
- a hint or a tiny classifier picks the tier
- pin when needed
- some flows must use a specific model
- A/B & canary
- route a % to a new model, compare quality
- policy as data
- change routing in the registry, not in code
Step 4 · When a provider fails
Fallback, retries, and timeouts
Providers rate-limit you, time out, 500, or go fully down — regularly. Your gateway is now a single point every app depends on. How do you keep serving when the chosen model can’t?
The chosen provider returns 429/503. What should the gateway do?
That pushes retry and failover logic back into every app — the coupling the gateway exists to remove. Resilience belongs in the control plane.
Hammering a rate-limited or dying provider makes it worse (a retry storm) and stalls the caller. Retries need backoff and a limit — and a Plan B.
Bounded retries with jittered backoff handle blips; on repeated failure the router fails over to a capability-equivalent model elsewhere, and a circuit breaker stops sending to a provider that’s down.
On failure the gateway does bounded, jittered retries, then fails over to a capability-equivalent model on another provider. A circuit breaker stops routing to a provider that’s persistently failing (and probes for recovery). Hard timeouts and hedging cap tail latency. Multi-provider isn’t just for cost — it’s your redundancy.
Why this piece earns its place
The safety here comes from a budget shared across the whole request rather than per hop: one deadline and one total attempt count, set at the edge, decremented before each attempt, with any attempt that cannot finish inside what is left simply not started. Count attempts per hop instead and they compose by multiplication — the caller's own retries wrapping the gateway's, wrapping the adapter's — and the surge lands exactly when providers are least able to absorb it. Two operational consequences follow. A circuit breaker's counters are per-process unless you deliberately make them otherwise, so every replica has to learn independently that a provider is down, and the fleet keeps feeding it long after any single replica would have tripped; either share that state or accept that the real trip threshold is the configured one times your replica count. And a fallback lane that has never carried traffic is not a fallback: the spare provider's key usually still holds the quota you asked for while you were evaluating it, so the first genuine failover moves the entire load onto the one lane provisioned for none of it. Keep a slice of production flowing there continuously. Last, tell the caller which model actually served the request — a path that parses strict JSON has a real interest in knowing it was answered by the backup, and downgrading silently is how a provider incident becomes a data incident.
- backoff + jitterabsorb transient errors
- failoverto an equivalent model
- circuit breakerisolate a down provider
Back of the envelope
- timeouts everywhere
- never wait forever on a provider
- idempotency keys
- so a retry doesn’t double-bill or double-act
- hedge tail latency
- a second request cancels the slow one
- degrade, don’t 500
- a smaller model beats no answer
Step 5 · Protect the shared resource
Rate limits, quotas, and budgets
Every app shares your provider quotas and your bill. One buggy loop or one heavy tenant can exhaust the rate limit for everyone — or spend a month’s budget by lunch. How do you isolate them?
How do you stop one caller from starving the rest (or blowing the bill)?
Each key gets its own rate (requests + tokens) and a spend budget. Over the limit → throttle, queue, or reject — so one tenant’s spike can’t exhaust the shared provider quota or the budget.
A single global limit lets the noisiest tenant consume it all — the quiet ones starve. Isolation has to be per-caller, not global.
Client-side limits are unenforceable and inconsistent; a bug ignores them entirely. Limits must be enforced where the resource is shared — the gateway.
Enforce per-key and per-tenant token buckets (requests and tokens/min) plus spend budgets at the gateway. Over the limit → throttle, queue, or reject with a clear 429 + Retry-After. Now a runaway tenant hits their ceiling, not everyone’s — and no key can silently overrun the bill.
Why this piece earns its place
The awkward fact underneath token limits is that you cannot meter what has not been generated yet. Input tokens are countable before the call; output tokens are known only when the stream ends. So the bucket has to reserve against the request's maximum output up front and reconcile to the true count afterwards, releasing the difference. Skip the reservation and a tenant running high concurrency with a generous ceiling sails past their limit long before the first debit lands, because every in-flight request is invisible to the accountant. The same lag makes a spend budget a soft ceiling by construction — you learn a cost after paying it — so decide how much overshoot you accept and state the number, instead of promising a cap the mechanism cannot deliver. Then there is the gateway being many replicas. A shared counter consulted on every request puts a store in the hot path and invents a new way for the gateway to die; local buckets reconciled periodically are faster and drift, which is fine when the drift is bounded and you sized it on purpose. Worth checking, too, that the limits you promise tenants sum to less than the quota the provider promises you.
- per-tenantbuckets + budgets
- tokens & requestsboth are limited
- 429 + Retry-Afterclear backpressure
What the new pieces do
- Rate Limiterservice
- Per-key and per-tenant token buckets and spend budgets. Rejects or queues requests over the limit so one noisy tenant can’t exhaust a shared provider quota or blow the bill.
Step 6 · Stop paying twice
Response caching cuts cost and latency
Many prompts repeat — the same FAQ, the same system-prompt+doc, near-identical questions. Each one is a full-price, full-latency model call. Where’s the cheapest token? The one you never send.
How do you avoid re-calling the model for repeated work?
With temperature 0 (or accept-a-close-answer) many calls are effectively repeatable, and exact-prompt hits are trivially safe. Skipping caching leaves the biggest cost/latency lever on the table.
An exact-match cache returns the stored completion for an identical request instantly. A semantic cache embeds the prompt and reuses an answer for a near-duplicate — with a similarity threshold. Both key on model+params so a cache entry is valid.
Prompts carry user/tenant context and private data — a global forever-cache leaks across tenants and serves stale answers. Cache must be scoped and TTL’d.
Put a Response Cache in the gateway: an exact-match layer (identical prompt + model + params → stored completion) and optionally a semantic layer (embed the prompt, reuse a near-duplicate above a similarity threshold). Scope keys per tenant, set a TTL, and skip the provider on a hit — often the single biggest win on both cost and p50 latency.
Why this piece earns its place
The parts of the cache key people get right are the ones written on the box. The ones that bite are the tool and function schemas offered on the call — change which tools a model may use and the same prompt legitimately produces a different answer — and the version of whatever template assembled the prompt. Templates are the one people miss: edit a system prompt in a config file and every stored entry built on the old wording is now wrong, with nothing in the system aware anything happened. Version the template, put that version in the key, or accept a flush you will forget to run. The other half is that a semantic hit fails in the direction your dashboard rewards. A threshold set slightly too loose raises the hit rate, cuts p50 and drops spend, so every number on the page improves while a fraction of users quietly receive the answer to a question they did not ask. No error, no latency signal, nothing to find it with. The defense is to spend a little of the saving back: sample some share of hits, run the real call behind them, and compare — which gives you a measured wrong-hit rate to set the threshold against, rather than a number someone picked once and nobody has revisited since.
- exact + semantictwo cache layers
- key on model+paramsso a hit is valid
- TTL + per-tenantfresh and isolated
What the new pieces do
- Response Cachecache
- Returns a stored completion for a repeated (or semantically-equivalent) prompt, skipping the provider call entirely — the biggest single lever on cost and latency for read-heavy workloads.
Back of the envelope
- temperature 0 caches best
- deterministic outputs are safely reusable
- semantic needs a threshold
- too loose = wrong answers served
- never cache across tenants
- prompts carry private context
- cache-hit is a metric
- track it — it’s money saved
Step 7 · See what it costs
Observability, cost, and attribution
You’re routing across providers, failing over, caching, and throttling. Now finance asks "what did each team spend?", an app owner asks "why did quality drop?", and you ask "is the router’s policy actually working?" You can’t answer any of it without data. What do you record?
What must the gateway log on every request to stay operable?
HTTP basics miss everything LLM-specific: which model ran, tokens in/out, dollar cost, cache-hit, fallback. You can’t bill, budget, or debug quality with just status codes.
Storing raw prompts/outputs wholesale is a privacy and cost liability (PII, huge volume). Log metadata by default; sample or redact content deliberately.
Structured records of model, token counts, cost, latency, cache-hit and any fallback, keyed to the tenant, power billing, budgets, alerting, and the quality investigations the router’s decisions demand.
Emit a structured record per request: tenant, model chosen (and why), tokens in/out, cost, latency, cache-hit, retries/fallback, and outcome — into a Cost & Traces store. That one ledger drives billing, budget enforcement, alerting, and quality debugging. Log metadata always; sample or redact prompt/response content deliberately.
Why this piece earns its place
Notice that cost here is computed, not reported: you multiply token counts by a number sitting in the registry, which makes the price table production data with the same care requirements as the routing policy. Write the price you used onto the record rather than joining to the current table at read time, or the next vendor price change silently rewrites last quarter. The zeroes want the same discipline. A cache hit legitimately costs nothing, and a request whose usage numbers never came back also lands as nothing, and on a chart those are the same bar — so record why a cost is zero, or the cheapest-looking tenant is really the one whose accounting is broken. Then there is the fact that this one store has two consumers with incompatible rules. Debugging wants everything, seconds old, and does not care that it disappears next month. Billing wants a narrow, complete, immutable record that survives for years and reconciles line by line against a vendor invoice. Sampling is a fine answer for the first and a wrong one for the second — you cannot bill from a sample, and the sampling rate is invisible at the moment someone sums the column. Split them early, because retrofitting completeness onto a sampled table leaves a quarter nobody can reconstruct.
- per requestmodel · tokens · cost · latency
- attributedto a tenant
- drivesbilling · budgets · debugging
What the new pieces do
- Cost & Tracesstore
- Structured logs of tokens in/out, cost, latency, model, and cache-hit per request — attributed to a tenant. The ledger for billing, budgets, and debugging quality.
Step 8 · Serve it at scale
Streaming, statelessness, and multi-region
Users expect tokens to stream as they’re generated, traffic is spiky, and the gateway now sits in the hot path of every LLM call — a tempting single point of failure. How do you make it fast and highly available?
How do you scale a gateway that’s in every request’s hot path?
Buffering throws away streaming — users stare at a spinner for seconds. The gateway should pass provider tokens straight through as they arrive.
The gateway proxies the provider’s SSE/streaming chunks to the client with minimal added latency, holds no per-request state (cache/registry/limits are external), so you scale horizontally and run multi-region for HA.
In-process state makes replicas non-interchangeable and lossy on restart. Keep the gateway stateless; put shared state in the cache/registry/limit stores.
Make the gateway stateless — cache, registry, limits and traces live in external stores — so you run many replicas behind a load balancer, multi-region for availability. Stream provider tokens straight through (SSE) so time-to-first-token stays low. The control plane must be more available than any single provider, because everything depends on it.
Why this piece earns its place
Streaming forces a decision an ordinary HTTP service never makes: the moment you relay the first chunk you have committed to a status code and to a provider, and everything after that has to be an in-band event. Where you place that commit point is the design. Hold the first chunk or two and an early provider failure is still a clean status code and an invisible switch; relay the instant bytes arrive and time-to-first-token is lower, but a later failure reaches the client as a truncated answer that a caller checking only the response code reads as a complete one. So the wire format needs a terminal event the client is written to check for. Operating it changes too. Total request duration now measures how long an answer was, not whether anything is healthy, so time-to-first-token is the number worth alerting on. Any intermediary can also buffer without telling you — a default in a hop you may not own — and your own metrics stay perfect throughout, because from the gateway's side the stream left on time. Drain windows want sizing against your longest generation, not a platform default written for request/response traffic. And stateless is true of the process, not the system: a region cut off from the shared stores has to choose between a cross-region hop on every request and enforcing its own copy of a budget that is now being spent twice.
- statelessshared state is external
- stream throughlow time-to-first-token
- multi-regionHA for the hot path
The payoff
You built an LLM gateway
From "every app hard-codes one vendor" to a control plane: one unified API, provider adapters, a capability-first router with cost/latency tie-breakers, retries + failover + circuit breaking, per-tenant limits and budgets, exact + semantic caching, per-request cost attribution, and a thin stateless streaming edge.
Now route by price alone — collapse the policy to "cheapest token wins" — and watch quality quietly fall apart: hard prompts sent to weak models return confident, wrong, un-parseable answers as clean 200s with no error. That’s why you route on capability first and let cost break ties only within a tier — and why the cost-and-quality ledger from step 7 exists to catch it.
Everything you assembled, in order
- Unified API — apps call one provider-agnostic endpoint; routing changes never touch them
- Adapters — one per provider — normalize the vendor mess at the edge
- Router — capability first, cost/latency as tie-breakers, policy in a hot-reloadable registry
- Fallback — bounded retries → failover to an equivalent model → circuit breaker
- Limits & budgets — per-tenant token buckets + spend caps isolate the shared quota
- Caching — exact + semantic — the cheapest call is the one you never send
- Observability — per-request cost/latency attributed to a tenant drives billing + debugging
- Scale — stateless, streaming, multi-region — the control plane outlives any one provider