Vibe Engines
YouTube
AI System Design

Design an LLM Gateway

Step 1 / 9

Learn AI system design by building an LLM gateway and model router step by step.

The numbers to beat1 adapterper providercanonical shapeone request/responseadd vendor= add adapter, no core change

The whole design, in writing

Learn AI system design by building an LLM gateway and model router step by step. An interactive guide to a multi-provider control plane — a unified API, provider adapters, policy-based routing, fallback and retries, per-tenant rate limits and budgets, response caching, and cost/latency observability — so one endpoint serves many models reliably and affordably.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

Why put a gateway in front of the models?

Your product calls an LLM. Easy — until you have a dozen services each hard-coding one vendor’s SDK, key, and quirks. Then that vendor has an outage, or triples its price, or a cheaper model ships, or one feature floods your shared rate limit at 2am. Every one of those is now a code change in a dozen places. How do you make "which model, from whom, under what limits" a runtime decision instead of a compile-time one?

Put an LLM Gateway in front of every provider: one stable, provider-agnostic API your apps code against, with a model router, fallback, rate limits, caching, and cost accounting behind it. Swap models, add vendors, and enforce budgets centrally — without touching a single app.

Step 1 · The skeleton

One API for every model

Apps need completions from many models over time. What should they actually call, so that changing the underlying model never changes their code?

Applicationone API callLLM Gatewayunified API
New in this step: Application, LLM Gateway.

What’s the right contract between your apps and the models?

  1. That scatters vendor keys, retry logic, and model choices across every service — the exact coupling a gateway removes. A model swap becomes a dozen deploys.

  2. A single unified API (often OpenAI-compatible) lets apps say "complete this" without naming a vendor. Routing, keys, limits and fallback all live behind it.

  3. A library still runs in every app’s process — you can’t change routing, revoke a key, or enforce a global budget without redeploying them all. Control has to be a network hop, not a dependency.

Apps call one LLM Gateway endpoint with a provider-agnostic request (messages, params, maybe a task hint). The gateway authenticates the caller and owns everything downstream. Because it’s a network boundary, you change routing, keys, limits and providers centrally — the apps never notice.

Why this piece earns its place

The unified API is a product you now own and have to version. Being OpenAI-compatible buys instant adoption — every SDK already speaks it — and quietly commits you to a schema you do not control, so the first time a vendor ships something the canonical shape has no field for, every app that wants it is blocked on a gateway release. The usual escape hatch is a passthrough bag of vendor-specific options the adapter forwards untouched. Take it deliberately: anything in that bag is a request the router can no longer freely move elsewhere, so the field that unblocks a team today is the field that breaks failover tomorrow. Mark those requests as pinned rather than pretending they are portable. The other thing that happens here and nowhere else is identity. The tenant has to be derived from the credential presented at the edge, never read out of a field the caller sets, because that derived tenant is the subject every later step keys on — buckets, budgets, cache scope, cost attribution. Get it from a header an app can lie about and all four inherit the lie at once.

What the new pieces do

Applicationclient
Any service or product surface that wants a completion. It codes against one stable, provider-agnostic API and never talks to a model vendor directly.
LLM Gatewaybackend
The single control-plane endpoint. Authenticates the caller, enforces limits, checks the cache, asks the router which model to use, calls the provider (with fallback), and records cost + latency — behind one OpenAI-compatible contract.

Step 2 · Speak every dialect

Provider adapters normalize the vendors

Provider A, Provider B and your self-hosted model each have different request shapes, auth, streaming formats, error codes, and token accounting. The gateway promises one contract. How does it talk to all of them?

Provider Afrontier modelProvider Bfast / cheapSelf-hostedOSS on your GPUs
New in this step: Provider A, Provider B, Self-hosted.

How does one API front three incompatible provider APIs?

  1. You don’t control vendors’ APIs — they’ll never conform. The translation has to live on your side, per provider.

  2. Then apps must special-case each vendor again — you’ve moved the coupling, not removed it. The gateway’s value is a single normalized shape.

  3. A thin per-provider adapter translates the gateway’s canonical request into the vendor’s format (and its response, streaming chunks, errors and token counts back). Add a vendor = add an adapter.

Each provider gets an adapter that maps the gateway’s canonical request ↔ that vendor’s native API — including auth, streaming format, error codes, and token accounting. The core stays vendor-neutral; onboarding a new model is writing one adapter, not rewriting the gateway.

Why this piece earns its place

Step 4's retry policy cannot branch on vendor error codes, so what the adapter has to hand inward is a class, not a message: retryable on this provider, retryable only somewhere else, never retryable, the caller's own fault. Only the adapter can decide which is which — one vendor signals an over-length context with a 400 while another buries it in a 200 carrying an error object in the body, and one status code can mean "slow down for a second" at one vendor and "you have no credit left" at the next, which are a wait and a dead end respectively. Token accounting is the other thing that stays invisible until finance asks. Usage counts arrive differently everywhere, and some vendors send them only in a final streaming frame — which never arrives if the client disconnects or the call times out, so the ledger loses precisely the requests that cost the most. The adapter owes an estimate when the true number is missing, flagged as an estimate. And because adapters are the one layer a vendor can break while you deploy nothing, they need contract tests run against the live provider on a schedule; recorded mocks keep passing happily for months after the thing they recorded changed.

  • 1 adapterper provider
  • canonical shapeone request/response
  • add vendor= add adapter, no core change

What the new pieces do

Provider Amodel
A hosted frontier model (e.g. a top-tier API). Strongest, priciest — the router sends hard tasks here.
Provider Bmodel
A second vendor’s smaller, faster, cheaper model. Good for easy tasks and a fallback when Provider A is down or rate-limited.
Self-hostedmodel
An open-weights model you run yourself. No per-token vendor cost and full data control — another lane for routing and failover.

Step 3 · Pick the right model

The router chooses per request

Now several models are reachable through one API. But which one should a given request use? "Always the biggest" is slow and expensive; "always the cheapest" is the trap you’ll feel later. How does the gateway decide, per request?

Model RouterpolicyModel Registrypolicy + prices
New in this step: Model Router, Model Registry.

What should the router optimize when choosing a model?

  1. Sending a one-line classification to a frontier model burns money and latency for no gain. Overkill is as wrong as underkill; the point is to match task to model.

  2. The router reads a policy: route by required capability (reasoning, tools, context length, JSON strictness), then let cost and latency choose among the models that can actually do the job.

  3. Cost-only routing sends hard tasks to weak models and silently tanks quality — no error, just wrong answers. (That’s literally the chaos button on this page.)

A Model Router chooses per request from a policy in the Model Registry: match the required capability first (reasoning depth, tool use, context length, structured-output reliability), then let cost and latency break ties within a capability tier. Pins and per-tenant overrides sit on top. The registry updates without a redeploy.

Why this piece earns its place

Hot-reloadable policy means the router's input is mutable, and that has a consequence you meet during an incident rather than at design time: with no version stamped on the decision, "why did this request go to that model" is unanswerable a week later. Record the registry version and the rule that matched beside the model id, so the ledger carries a decision you can replay instead of a name you can only guess at. Pin that snapshot for the life of a request, too — policy that reloads mid-flight can route a second attempt by different rules than the first, and a request explained by two policies is a request nobody can explain. The flip side of "no redeploy" is that a routing change also skips everything a deploy brings with it: review, a staging run, a rollback somebody has practised. Edit a tier late in the afternoon and it is live for every tenant in seconds with no artifact to revert to, so the registry wants versioned writes and a one-command rollback long before it wants more rules. And keep the decision cheap. The router sits in every request's critical path, so evaluate policy in memory against a cached snapshot with a background refresh; fetching config per request adds a network hop to the one component whose job is to add none, and turns the registry into a second thing that can take the gateway down.

  • capabilitythe primary key
  • cost / latencytie-breakers within a tier
  • registrypolicy, hot-reloaded

What the new pieces do

Model Routerservice
Chooses a target model per request from a policy: task/capability first, then cost and latency as tie-breakers, honoring pins and overrides. The brain of the gateway.
Model Registrystore
The config the router reads: available models, their capabilities, prices, context limits, health, and the routing rules — updatable without a redeploy.

Back of the envelope

classify the task
a hint or a tiny classifier picks the tier
pin when needed
some flows must use a specific model
A/B & canary
route a % to a new model, compare quality
policy as data
change routing in the registry, not in code

Step 4 · When a provider fails

Fallback, retries, and timeouts

Providers rate-limit you, time out, 500, or go fully down — regularly. Your gateway is now a single point every app depends on. How do you keep serving when the chosen model can’t?

ApplicationLLM GatewayModel RouterModel RegistryProvider AProvider BSelf-hosted
The system as it stands at this step. · swipe to pan the diagram

The chosen provider returns 429/503. What should the gateway do?

  1. That pushes retry and failover logic back into every app — the coupling the gateway exists to remove. Resilience belongs in the control plane.

  2. Hammering a rate-limited or dying provider makes it worse (a retry storm) and stalls the caller. Retries need backoff and a limit — and a Plan B.

  3. Bounded retries with jittered backoff handle blips; on repeated failure the router fails over to a capability-equivalent model elsewhere, and a circuit breaker stops sending to a provider that’s down.

On failure the gateway does bounded, jittered retries, then fails over to a capability-equivalent model on another provider. A circuit breaker stops routing to a provider that’s persistently failing (and probes for recovery). Hard timeouts and hedging cap tail latency. Multi-provider isn’t just for cost — it’s your redundancy.

Why this piece earns its place

The safety here comes from a budget shared across the whole request rather than per hop: one deadline and one total attempt count, set at the edge, decremented before each attempt, with any attempt that cannot finish inside what is left simply not started. Count attempts per hop instead and they compose by multiplication — the caller's own retries wrapping the gateway's, wrapping the adapter's — and the surge lands exactly when providers are least able to absorb it. Two operational consequences follow. A circuit breaker's counters are per-process unless you deliberately make them otherwise, so every replica has to learn independently that a provider is down, and the fleet keeps feeding it long after any single replica would have tripped; either share that state or accept that the real trip threshold is the configured one times your replica count. And a fallback lane that has never carried traffic is not a fallback: the spare provider's key usually still holds the quota you asked for while you were evaluating it, so the first genuine failover moves the entire load onto the one lane provisioned for none of it. Keep a slice of production flowing there continuously. Last, tell the caller which model actually served the request — a path that parses strict JSON has a real interest in knowing it was answered by the backup, and downgrading silently is how a provider incident becomes a data incident.

  • backoff + jitterabsorb transient errors
  • failoverto an equivalent model
  • circuit breakerisolate a down provider

Back of the envelope

timeouts everywhere
never wait forever on a provider
idempotency keys
so a retry doesn’t double-bill or double-act
hedge tail latency
a second request cancels the slow one
degrade, don’t 500
a smaller model beats no answer

Step 5 · Protect the shared resource

Rate limits, quotas, and budgets

Every app shares your provider quotas and your bill. One buggy loop or one heavy tenant can exhaust the rate limit for everyone — or spend a month’s budget by lunch. How do you isolate them?

LLM Gatewayroute · guardRate Limiterquotas / budgets
New in this step: Rate Limiter.

How do you stop one caller from starving the rest (or blowing the bill)?

  1. Each key gets its own rate (requests + tokens) and a spend budget. Over the limit → throttle, queue, or reject — so one tenant’s spike can’t exhaust the shared provider quota or the budget.

  2. A single global limit lets the noisiest tenant consume it all — the quiet ones starve. Isolation has to be per-caller, not global.

  3. Client-side limits are unenforceable and inconsistent; a bug ignores them entirely. Limits must be enforced where the resource is shared — the gateway.

Enforce per-key and per-tenant token buckets (requests and tokens/min) plus spend budgets at the gateway. Over the limit → throttle, queue, or reject with a clear 429 + Retry-After. Now a runaway tenant hits their ceiling, not everyone’s — and no key can silently overrun the bill.

Why this piece earns its place

The awkward fact underneath token limits is that you cannot meter what has not been generated yet. Input tokens are countable before the call; output tokens are known only when the stream ends. So the bucket has to reserve against the request's maximum output up front and reconcile to the true count afterwards, releasing the difference. Skip the reservation and a tenant running high concurrency with a generous ceiling sails past their limit long before the first debit lands, because every in-flight request is invisible to the accountant. The same lag makes a spend budget a soft ceiling by construction — you learn a cost after paying it — so decide how much overshoot you accept and state the number, instead of promising a cap the mechanism cannot deliver. Then there is the gateway being many replicas. A shared counter consulted on every request puts a store in the hot path and invents a new way for the gateway to die; local buckets reconciled periodically are faster and drift, which is fine when the drift is bounded and you sized it on purpose. Worth checking, too, that the limits you promise tenants sum to less than the quota the provider promises you.

  • per-tenantbuckets + budgets
  • tokens & requestsboth are limited
  • 429 + Retry-Afterclear backpressure

What the new pieces do

Rate Limiterservice
Per-key and per-tenant token buckets and spend budgets. Rejects or queues requests over the limit so one noisy tenant can’t exhaust a shared provider quota or blow the bill.

Step 6 · Stop paying twice

Response caching cuts cost and latency

Many prompts repeat — the same FAQ, the same system-prompt+doc, near-identical questions. Each one is a full-price, full-latency model call. Where’s the cheapest token? The one you never send.

LLM Gatewayroute · guardResponse Cacheexact + semantic
New in this step: Response Cache.

How do you avoid re-calling the model for repeated work?

  1. With temperature 0 (or accept-a-close-answer) many calls are effectively repeatable, and exact-prompt hits are trivially safe. Skipping caching leaves the biggest cost/latency lever on the table.

  2. An exact-match cache returns the stored completion for an identical request instantly. A semantic cache embeds the prompt and reuses an answer for a near-duplicate — with a similarity threshold. Both key on model+params so a cache entry is valid.

  3. Prompts carry user/tenant context and private data — a global forever-cache leaks across tenants and serves stale answers. Cache must be scoped and TTL’d.

Put a Response Cache in the gateway: an exact-match layer (identical prompt + model + params → stored completion) and optionally a semantic layer (embed the prompt, reuse a near-duplicate above a similarity threshold). Scope keys per tenant, set a TTL, and skip the provider on a hit — often the single biggest win on both cost and p50 latency.

Why this piece earns its place

The parts of the cache key people get right are the ones written on the box. The ones that bite are the tool and function schemas offered on the call — change which tools a model may use and the same prompt legitimately produces a different answer — and the version of whatever template assembled the prompt. Templates are the one people miss: edit a system prompt in a config file and every stored entry built on the old wording is now wrong, with nothing in the system aware anything happened. Version the template, put that version in the key, or accept a flush you will forget to run. The other half is that a semantic hit fails in the direction your dashboard rewards. A threshold set slightly too loose raises the hit rate, cuts p50 and drops spend, so every number on the page improves while a fraction of users quietly receive the answer to a question they did not ask. No error, no latency signal, nothing to find it with. The defense is to spend a little of the saving back: sample some share of hits, run the real call behind them, and compare — which gives you a measured wrong-hit rate to set the threshold against, rather than a number someone picked once and nobody has revisited since.

  • exact + semantictwo cache layers
  • key on model+paramsso a hit is valid
  • TTL + per-tenantfresh and isolated

What the new pieces do

Response Cachecache
Returns a stored completion for a repeated (or semantically-equivalent) prompt, skipping the provider call entirely — the biggest single lever on cost and latency for read-heavy workloads.

Back of the envelope

temperature 0 caches best
deterministic outputs are safely reusable
semantic needs a threshold
too loose = wrong answers served
never cache across tenants
prompts carry private context
cache-hit is a metric
track it — it’s money saved

Step 7 · See what it costs

Observability, cost, and attribution

You’re routing across providers, failing over, caching, and throttling. Now finance asks "what did each team spend?", an app owner asks "why did quality drop?", and you ask "is the router’s policy actually working?" You can’t answer any of it without data. What do you record?

LLM Gatewayroute · guard · meterResponse Cacheexact + semanticModel Routerpolicy + fallbackCost & Tracesper-tenant
New in this step: Cost & Traces.

What must the gateway log on every request to stay operable?

  1. HTTP basics miss everything LLM-specific: which model ran, tokens in/out, dollar cost, cache-hit, fallback. You can’t bill, budget, or debug quality with just status codes.

  2. Storing raw prompts/outputs wholesale is a privacy and cost liability (PII, huge volume). Log metadata by default; sample or redact content deliberately.

  3. Structured records of model, token counts, cost, latency, cache-hit and any fallback, keyed to the tenant, power billing, budgets, alerting, and the quality investigations the router’s decisions demand.

Emit a structured record per request: tenant, model chosen (and why), tokens in/out, cost, latency, cache-hit, retries/fallback, and outcome — into a Cost & Traces store. That one ledger drives billing, budget enforcement, alerting, and quality debugging. Log metadata always; sample or redact prompt/response content deliberately.

Why this piece earns its place

Notice that cost here is computed, not reported: you multiply token counts by a number sitting in the registry, which makes the price table production data with the same care requirements as the routing policy. Write the price you used onto the record rather than joining to the current table at read time, or the next vendor price change silently rewrites last quarter. The zeroes want the same discipline. A cache hit legitimately costs nothing, and a request whose usage numbers never came back also lands as nothing, and on a chart those are the same bar — so record why a cost is zero, or the cheapest-looking tenant is really the one whose accounting is broken. Then there is the fact that this one store has two consumers with incompatible rules. Debugging wants everything, seconds old, and does not care that it disappears next month. Billing wants a narrow, complete, immutable record that survives for years and reconciles line by line against a vendor invoice. Sampling is a fine answer for the first and a wrong one for the second — you cannot bill from a sample, and the sampling rate is invisible at the moment someone sums the column. Split them early, because retrofitting completeness onto a sampled table leaves a quarter nobody can reconstruct.

  • per requestmodel · tokens · cost · latency
  • attributedto a tenant
  • drivesbilling · budgets · debugging

What the new pieces do

Cost & Tracesstore
Structured logs of tokens in/out, cost, latency, model, and cache-hit per request — attributed to a tenant. The ledger for billing, budgets, and debugging quality.

Step 8 · Serve it at scale

Streaming, statelessness, and multi-region

Users expect tokens to stream as they’re generated, traffic is spiky, and the gateway now sits in the hot path of every LLM call — a tempting single point of failure. How do you make it fast and highly available?

ApplicationLLM GatewayRate LimiterResponse CacheModel RouterModel RegistryCost & TracesProvider AProvider BSelf-hosted
The system as it stands at this step. · swipe to pan the diagram

How do you scale a gateway that’s in every request’s hot path?

  1. Buffering throws away streaming — users stare at a spinner for seconds. The gateway should pass provider tokens straight through as they arrive.

  2. The gateway proxies the provider’s SSE/streaming chunks to the client with minimal added latency, holds no per-request state (cache/registry/limits are external), so you scale horizontally and run multi-region for HA.

  3. In-process state makes replicas non-interchangeable and lossy on restart. Keep the gateway stateless; put shared state in the cache/registry/limit stores.

Make the gateway stateless — cache, registry, limits and traces live in external stores — so you run many replicas behind a load balancer, multi-region for availability. Stream provider tokens straight through (SSE) so time-to-first-token stays low. The control plane must be more available than any single provider, because everything depends on it.

Why this piece earns its place

Streaming forces a decision an ordinary HTTP service never makes: the moment you relay the first chunk you have committed to a status code and to a provider, and everything after that has to be an in-band event. Where you place that commit point is the design. Hold the first chunk or two and an early provider failure is still a clean status code and an invisible switch; relay the instant bytes arrive and time-to-first-token is lower, but a later failure reaches the client as a truncated answer that a caller checking only the response code reads as a complete one. So the wire format needs a terminal event the client is written to check for. Operating it changes too. Total request duration now measures how long an answer was, not whether anything is healthy, so time-to-first-token is the number worth alerting on. Any intermediary can also buffer without telling you — a default in a hop you may not own — and your own metrics stay perfect throughout, because from the gateway's side the stream left on time. Drain windows want sizing against your longest generation, not a platform default written for request/response traffic. And stateless is true of the process, not the system: a region cut off from the shared stores has to choose between a cross-region hop on every request and enforcing its own copy of a budget that is now being spent twice.

  • statelessshared state is external
  • stream throughlow time-to-first-token
  • multi-regionHA for the hot path

The payoff

You built an LLM gateway

From "every app hard-codes one vendor" to a control plane: one unified API, provider adapters, a capability-first router with cost/latency tie-breakers, retries + failover + circuit breaking, per-tenant limits and budgets, exact + semantic caching, per-request cost attribution, and a thin stateless streaming edge.

ApplicationLLM GatewayRate LimiterResponse CacheModel RouterModel RegistryCost & TracesProvider AProvider BSelf-hosted
The finished design, end to end. · swipe to pan the diagram

Now route by price alone — collapse the policy to "cheapest token wins" — and watch quality quietly fall apart: hard prompts sent to weak models return confident, wrong, un-parseable answers as clean 200s with no error. That’s why you route on capability first and let cost break ties only within a tier — and why the cost-and-quality ledger from step 7 exists to catch it.

Everything you assembled, in order

  • Unified API — apps call one provider-agnostic endpoint; routing changes never touch them
  • Adapters — one per provider — normalize the vendor mess at the edge
  • Router — capability first, cost/latency as tie-breakers, policy in a hot-reloadable registry
  • Fallback — bounded retries → failover to an equivalent model → circuit breaker
  • Limits & budgets — per-tenant token buckets + spend caps isolate the shared quota
  • Caching — exact + semantic — the cheapest call is the one you never send
  • Observability — per-request cost/latency attributed to a tenant drives billing + debugging
  • Scale — stateless, streaming, multi-region — the control plane outlives any one provider

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. Who decides whether a task is "hard" or "easy" for capability routing — a fixed rule, or a classifier?

    Both, layered: a cheap heuristic (prompt length, presence of tool definitions, an explicit task-type hint from the calling app) handles the obvious majority instantly, and a small fast classifier model handles the ambiguous remainder — running a full LLM call just to decide which LLM to call would defeat the purpose. Some gateways also let the calling app pass an explicit capability hint, since the app often already knows more about the task than any classifier could infer from the prompt alone.

  2. Why do retries specifically risk double-billing or double-acting for LLM calls, and how does an idempotency key fix it?

    If a request times out on the client side but actually succeeded on the provider (the response just didn’t make it back), a naive retry re-runs the whole generation — double cost, and if the model called a tool (sent an email, charged a card), a double side effect, not just a double bill. An idempotency key lets the gateway recognize "I’ve already processed this exact request" and return the cached result instead of re-executing, the same pattern payment APIs use for exactly this reason.

  3. A circuit breaker trips on a failing provider. How does the gateway know it’s safe to send traffic back?

    It doesn’t wait for a fixed timer and hope — it sends a small trickle of probe traffic (a fraction of a percent) to the tripped provider while the breaker stays open for everyone else, and only closes the breaker once that probe traffic succeeds consistently. Flipping straight back to full traffic risks re-triggering whatever caused the failure the moment load returns.

  4. A tenant hits their spend budget mid-conversation. Hard-stop immediately, or degrade gracefully?

    Graceful degradation is usually right for anything other than a hard compliance limit: route remaining requests to a cheaper model tier so the tenant can finish their current task at reduced cost/quality rather than getting cut off mid-workflow, and reserve a true hard-stop for cases where continuing would mean uncontrolled overspend (no ceiling on the tier-down) or a contractual limit that can’t be exceeded at all.

  5. Provider A starts streaming a response, then fails mid-stream. Can you fail over to Provider B without the user seeing a broken or duplicated answer?

    Not cleanly mid-sentence — the practical approach is to buffer the first chunk or two before committing to a provider (so early failures fail over invisibly), but once substantial content has already streamed to the client, failing over means either restarting the visible response with a brief "retrying…" indicator or accepting the partial answer and appending a continuation, both of which are visibly worse than a clean failover before any tokens shipped. This is why hedging (racing two providers on hard-to-classify requests) is sometimes preferred over pure failover for latency-critical paths.

Check yourself — the answers, and why

Nine steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. An LLM gateway exists mainly to…

    It turns "which model, from whom, under what limits" into a runtime decision behind one API, instead of vendor SDKs hard-coded in every app.

  2. A per-provider adapter’s job is to…

    Adapters normalize each vendor’s request shape, streaming format, errors and token accounting so the gateway core works on one clean schema.

  3. The model router should optimize primarily for…

    Match task difficulty to a capable model, then let cost/latency choose among the models that can actually do the job.

  4. When a chosen provider returns 429/503, a good gateway…

    Bounded jittered retries handle blips; failover to a capability-equivalent model and a circuit breaker keep the gateway serving when a provider is down.

  5. The single biggest cost/latency lever in a gateway is usually…

    A cache hit costs ~nothing and returns in milliseconds; exact-match is trivially safe and semantic caching extends it to near-duplicates (behind a threshold).

  6. Routing purely by lowest price fails because it…

    Cost-only routing produces confident wrong answers and malformed output with no error; route on capability first and use the cost/quality ledger to catch drift.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Unified API: apps call one provider-agnostic endpoint; changing the model never changes their code.
  • Adapters: one per provider — map the canonical request/response to each vendor’s native API.
  • Route: pick a model per request by capability first, cost/latency as tie-breakers.
  • Resilience: retries with backoff, failover to an equivalent model, circuit-break a down provider.
  • Govern: per-tenant rate limits + budgets, response caching, and per-request cost attribution.

The qualities that shape everything

Each one names the mechanism that buys it.

Swap models without touching apps
One provider-agnostic gateway endpoint owns routing, keys, limits and providers behind a network boundary — apps never notice a change.
One contract over incompatible vendors
A per-provider adapter maps the gateway’s canonical request/response to each vendor’s native API, so the core stays vendor-neutral.
The right model per request
The router matches required capability first (reasoning, tools, context, JSON strictness), then lets cost and latency break ties within a tier.
Keep serving when a provider fails
Bounded jittered retries, failover to a capability-equivalent model, and a circuit breaker that isolates a down provider.
One tenant can’t starve the rest or blow the bill
Per-key and per-tenant token buckets plus spend budgets at the gateway — over the limit throttles that tenant, not everyone.
The biggest cut to cost and latency
A response cache — exact-match plus optional semantic — returns a stored completion and skips the provider on a hit.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

One gateway endpoint over each app importing vendor SDKs

Vendor SDKs scatter keys, retries and model choices across every service, so a model swap is a dozen deploys; a network boundary lets you change routing, keys and providers centrally.

Capability-first routing over cheapest token wins

Cost-only routing sends hard tasks to weak models and silently tanks quality; match task difficulty to a capable model, then let cost break ties within a tier.

Failover to an equivalent model over returning the error to the app

Pushing retry/failover into every app is the coupling the gateway removes; bounded retries plus failover to a capability-equivalent model keep it serving when a provider 429s or dies.

Per-tenant limits + budgets over one global rate limit

A single global limit lets the noisiest tenant consume it all while quiet ones starve; per-caller buckets and spend caps isolate a runaway tenant to their own ceiling.

Exact + semantic caching over caching nothing (output is non-deterministic)

At temperature 0 many calls repeat exactly and are trivially safe to reuse; skipping caching leaves the single biggest cost/latency lever untouched — a hit costs ~nothing.

The answer, out loud

What a strong answer to “Design an LLM Gateway” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Scope it before I draw anything

    Let me pin down which system this is, because "LLM gateway" means two different things. I am not designing model serving — no GPUs, nothing about how tokens get generated. I am designing the control plane between our applications and whoever runs the models, and what it delivers is the ability to change our mind: which vendor, which model, how much a team may spend. I would judge the design on one question — can I move a team onto a different model without anybody redeploying?

  2. 3–8 min

    One endpoint, and why not a library

    The skeleton is applications on one side, providers on the other, and one endpoint in the middle owning everything downstream. The alternative people propose is a shared client library, and my objection is organisational more than technical: a library ships inside somebody else's release train, so every default I want to change becomes an upgrade campaign across a dozen teams, and I run two policies for months while the stragglers catch up. Behind an endpoint, a change is live for everyone when I deploy it. It costs a hop of latency and one more thing that can be down, which I would rather name than pretend away.

    Built in step 1: One API for every model
  3. 8–13 min

    How wide the canonical request should be

    Every vendor gets an adapter, and the decision worth discussing is how wide the canonical request may be. Take the union of everything the vendors support and the adapters get thick, faking features the model behind them lacks. Take the intersection and the shape stays clean, but the first team wanting a vendor's good new feature waits on me. I would start near the intersection and treat vendor-specific things as named exceptions rather than quietly widening the contract, because the failure I am steering around is the gateway growing a second code path per vendor, until nobody can say what it does without reading all three.

    Built in step 2: Provider adapters normalize the vendors
  4. 13–20 min

    The router returns a ranking, not a winner

    Several models are now reachable through one API, so something has to choose per request. I would keep that policy in a registry, not in code: capability decides which models are eligible — reasoning depth, tool use, context length, structured-output strictness — and price only orders the ones that qualify. The detail I would push for is a ranked list back, not a single name. The next step needs a second choice anyway, and the moment the router returns one winner the fallback path invents its own notion of equivalence, which is how a request that needed tool calling lands on a model that cannot make one.

    Built in step 3: The router chooses per request
  5. 20–26 min

    Which failure gets which strategy

    Providers rate-limit, time out and go down often enough to be ordinary traffic, not an exception path. The part I would argue about is the ordering. The reflex is back off, retry, then move elsewhere — right for a flaky connection, wrong for a quota. If a provider has just told me I am over its limit, waiting politely is waiting for a wall that has not moved, so that case goes straight to another lane and the backoff applies to the lane, not the request. One sequence applied to every failure class is how a rate limit turns into a timeout.

    Built in step 4: Fallback, retries, and timeouts
  6. 26–31 min

    Two limits, two different clocks

    Tenants share one pool of provider quota and one bill, so isolation belongs here. What I would separate out is that two limits do different jobs on different clocks. A rate limit protects something instantaneous and shared: if one tenant is over it now, everyone else feels it now, and the answer is immediate backpressure. A budget protects money across a month, and going over does not hurt anybody else this second, so it should not trigger the same reflex. I would make both numbers per-tenant settings, not constants, because the first real incident this causes will be a legitimate launch hitting a ceiling nobody remembers setting.

    Built in step 5: Rate limits, quotas, and budgets
  7. 31–36 min

    Cache, but measure before you loosen it

    Caching is where I would ask a question rather than assert a win: the hit rate is a property of the traffic, not of the cache. Classification and extraction over a fixed set of templates repeat constantly. An open-ended chat product, where every request carries a conversation nobody else has had, barely repeats. So exact match ships first and I measure what it catches. If that comes back near zero, a semantic layer is a way of manufacturing hits by loosening what counts as the same question, and I want that chosen deliberately, not in a follow-up ticket.

    Built in step 6: Response caching cuts cost and latency
  8. 36–41 min

    One record per request, keyed to a human

    Every call leaves a record, and two things about it get missed. The identity on it has to be something the business recognises — a team, a product, a cost centre — not the API key that made the call, because a league table of top-spending key ids is something nobody can act on. And the caller's own trace id has to come in with the request and go out with the response, so when an app owner says something was slow or wrong, they hand me one identifier that lands on our record instead of both of us searching by timestamp.

    Built in step 7: Observability, cost, and attribution
  9. 41–45 min

    Close on the trade-off

    To close on the trade-off, because it is the real cost here: coupling that was spread thinly across a dozen services is now concentrated in one component all of them depend on, which has to be more available than any provider behind it. Much of it follows from that: it holds no per-request state, so any replica can serve any call. With more time I would go after two things this does not answer. Data residency as a routing constraint: some tenants will only accept the self-hosted lane, which turns a preference into a hard filter the router must honour even when it costs quality. And the migration, because a dozen live services do not move behind a new endpoint in one afternoon.

    Built in step 8: Streaming, statelessness, and multi-region

What this teaches

Learn AI system design by building an LLM gateway and model router step by step. An interactive guide to a multi-provider control plane — a unified API, provider adapters, policy-based routing, fallback and retries, per-tenant rate limits and budgets, response caching, and cost/latency observability — so one endpoint serves many models reliably and affordably.

Key takeaways

  • Unified API — apps call one provider-agnostic endpoint; routing changes never touch them
  • Adapters — one per provider — normalize the vendor mess at the edge
  • Router — capability first, cost/latency as tie-breakers, policy in a hot-reloadable registry
  • Fallback — bounded retries → failover to an equivalent model → circuit breaker
  • Limits & budgets — per-tenant token buckets + spend caps isolate the shared quota
  • Caching — exact + semantic — the cheapest call is the one you never send
  • Observability — per-request cost/latency attributed to a tenant drives billing + debugging
  • Scale — stateless, streaming, multi-region — the control plane outlives any one provider

Concepts covered

  • Why put a gateway in front of the models?
  • One API for every model
  • Provider adapters normalize the vendors
  • The router chooses per request
  • Fallback, retries, and timeouts
  • Rate limits, quotas, and budgets
  • Response caching cuts cost and latency
  • Observability, cost, and attribution
  • Streaming, statelessness, and multi-region
RUN IT YOURSELF

Every retry landed while the bucket was still empty

Step 5 puts a per-tenant token bucket at the gateway and step 4 puts retries in front of it; this runs both at once, because their interaction is where the outage is. Two clients hit the identical bucket with the identical attempt budget and deadline — only the spacing between attempts differs — and it prints served, attempts and 429s for each. Change BLIND_GAP from 1 to 3, hit Run, and the same four attempts serve every request.

HOW TO READ THE CODE — 5 IDEAS
  1. A burst of 40 leaves 20 callers over the limit, and at 3 tokens/tick the bucket needs 7 ticks to let them back in — the blind client is out of attempts after 3 (step 5).
  2. Both clients get 4 attempts and a 1 s deadline against the same bucket. Only the spacing changes, and success goes 72.5% → 98.3% (step 4).
  3. So the blind client sends 202 more attempts to get 155 fewer answers — more load, less work done, while offering 20/s against a 30/s ceiling it never approached in aggregate.
  4. The fixed-interval sweep is the control: the attempt budget is untouched and 300 ms flat already serves all 600. Backoff finds that interval without being told the refill rate; the jitter stops the rejected callers resynchronising.
  5. Raising the cap to 8 attempts reaches 600 too — for 1755 attempts against 900. That is the amplification the shared per-request budget in step 4 exists to bound.
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, route by price alone, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs