The whole design, in writing
Learn AI system design by building an LLM evaluation and observability pipeline step by step. An interactive guide to measuring AI quality — tracing every call, curating versioned golden datasets, scoring with programmatic/LLM-judge/human methods, a CI regression gate, online production monitoring, and a feedback loop — plus why trusting an unvalidated LLM judge silently corrupts every number you ship on.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
You can’t ship what you can’t measure
An LLM feature has no compiler and no unit test that says "correct." Outputs are open-ended, quality is subjective, and the same prompt can regress when you tweak a word, swap a model, or the provider updates theirs — silently, with no error. Demos are easy; knowing whether you’re actually getting better (or quietly getting worse) is the hard part that separates a toy from a product. How do you measure AI quality, continuously?
Build an eval + observability pipeline: trace every call (observability), score the system against versioned golden datasets with validated scorers (evaluation), gate changes in CI, and monitor production for drift — closing a loop where real failures become new test cases. Measurement is the system that lets you improve on purpose instead of by vibes.
Step 1 · Two loops
Offline evals and online observability
There are two questions: "is this change good before I ship it?" and "is it still good in production?" They need different machinery. What’s the overall shape?
What are the two complementary halves of measuring an LLM system?
A one-time benchmark misses regressions from later changes and real-world drift. Measurement is continuous, in two loops — pre-ship and in production.
Offline evals score changes against versioned golden sets to catch regressions before deploy; online observability traces and scores live traffic to catch drift and new failures after deploy. Both are needed.
Monitoring alone means every regression ships to users first. You want to catch what you can offline and watch production for what offline missed.
Two complementary loops: offline evaluation runs a change against fixed, versioned datasets before shipping (catch regressions early), and online observability traces and scores production traffic (catch drift and novel failures). The rest of this design builds both on a shared foundation — traces in, scores out.
Why this piece earns its place
The two loops look like one system on the diagram, and the expensive mistake is building them as two. The offline suite gets written first, as a script living beside the prompts; months later someone adds production scoring and writes a second implementation of the same checks in a different service. Now both loops report a number with the same name computed by different code, and when they disagree nobody can say whether quality drifted or the scorers did. Write the scoring once, as a library the runner calls, and make the source of inputs — a dataset, or a sample of traffic — the argument. Sharing the code does not make the two numbers interchangeable, though. An offline score is computed on a set you balanced on purpose; an online one on whatever mix of questions arrived that day. The same metric name therefore means two different things, and a gap between them is not evidence of anything — keep them as two series, each compared against its own history. Skip the online half only while nothing is live yet: a prototype has traces, a small offline set, and nothing to watch.
Step 2 · Capture everything
Tracing is the foundation
You can’t evaluate, debug, or mine examples from calls you didn’t record. Before any scoring, what has to be captured on every LLM interaction?
What must you log on every LLM call to make evaluation possible?
The output alone can’t explain a failure or reproduce it — you need the prompt, params, tool calls, cost, and latency too. Partial logging cripples both debugging and eval-set mining.
Ops metrics miss the content entirely; you can’t judge quality or build datasets from them. LLM observability is about the calls themselves, not just system health.
A Trace Collector records the complete interaction into a queryable store — the raw material for debugging, mining eval examples, monitoring, and cost attribution.
A Trace Collector captures the full interaction — prompt, output, tokens, latency, cost, tool calls, and any user feedback — into a durable Trace Store. This is the observability bedrock: you can’t evaluate, debug, attribute cost, or build datasets from calls you never recorded. Everything downstream reads from these traces.
Why this piece earns its place
The collector sits on the request path, and that one fact decides most of its design: it has to be fire-and-forget — buffer locally, ship asynchronously, drop on backpressure. A collector that blocks or throws takes the product down to protect a measurement, which is the worst trade available here. What it captures has to be structured rather than one concatenated string. Store the prompt template id and the resolved variables separately, store the model version the provider actually served rather than the alias that was configured, and store the parameters really sent. Without that you cannot ask how a given template version performed, and you cannot rebuild a case after the template moves on. One user interaction is usually several calls — retries, tool calls, a chain of steps — so they need a shared trace_id grouping them, or you evaluate one hop of an agent and call it the system. Retention is where the cost is decided: full text hot for a short window, sampled or summarised beyond it.
- full traceprompt · output · tools · cost
- durable storequeryable, replayable
- feeds alleval · monitor · debug
What the new pieces do
- LLM Appclient
- The live AI feature. Every call it makes is both a user interaction and a data point for evaluation — the source of traces and of real-world examples to test against.
- Trace Collectorbackend
- Instruments every LLM call: prompt, output, tokens, latency, cost, tool calls, and user feedback. The observability foundation — you can’t evaluate what you don’t capture.
- Trace Storestore
- Durable, queryable store of production traces. The raw material for debugging, for mining eval examples, and for online monitoring.
Step 3 · The yardstick
Versioned golden datasets
To say "better" or "worse" you need something fixed to measure against. Where does that yardstick come from, and why must it be versioned?
What do you evaluate a change against?
Build eval sets from production traces (representative), known failures, and edge cases, each with expected behavior; version them so a score is comparable across runs and a change to the data is explicit.
Ad-hoc examples that change every run make scores incomparable and easy to game. The dataset must be a stable, versioned artifact.
Public benchmarks don’t reflect your task, distribution, or failure modes. They’re a weak proxy; your own golden set from production is the real measure.
Curate versioned golden datasets: representative cases mined from production traces, plus known failures and edge cases, each labeled with expected behavior. Version them so scores are comparable across runs and dataset changes are explicit. Your own eval set — not a public benchmark — is the yardstick that matters, and it grows as production reveals new failures.
Why this piece earns its place
Start with what expected behavior actually means inside a row. A single gold answer pushes you toward exact match, which does not fit open-ended output, or invites a judge to grade phrasing instead of substance. Most cases are better stored as assertions — must state the refund window, must not promise a callback, must decline — because an assertion survives a rewrite of the wording and names what the failure actually was. Then guard the set against yourself. The cases you stare at while iterating are the ones you tune the prompt to satisfy, so their score stops predicting anything about unseen traffic; that is overfitting to a dev set, and the remedy is the split ML has always used — hold a slice back, never read it while iterating, report it separately. Finally, stamp the dataset version onto every score you keep, not only onto the dataset. A stored number with no version attached turns a long trend line into two yardsticks glued together, with the join invisible.
- golden setreal + edge cases
- versionedcomparable across runs
- mined from prodvia the traces
What the new pieces do
- Eval Setstore
- Curated, versioned datasets of inputs with expected behavior — mined from production and edge cases. The fixed yardstick every prompt/model change is measured against.
Step 4 · How to score
Programmatic, LLM-judge, human
For "did it return valid JSON?" a simple check works. For "was this answer helpful and correct?" there’s no regex. What scoring methods do you use, and which for what?
How do you score open-ended LLM outputs?
Exact match fails for open-ended generation where many phrasings are correct — it’s only right for constrained/structured outputs. Open-ended quality needs judgment.
Human grading is the gold standard but doesn’t scale to every eval run. You reserve humans for calibration and hard cases, and automate the rest.
Use cheap deterministic checks for structured/verifiable outputs, an LLM-as-judge for subjective quality at scale, and human labels as ground truth to validate the judge and handle the ambiguous.
Layer the scorers: programmatic checks for verifiable/structured outputs (JSON valid, contains fact, matches regex), an LLM-as-judge for open-ended quality at scale, and human labels as ground truth. Critically, the judge is only trustworthy if validated against humans — measured for agreement and bias — because an uncalibrated judge (the chaos button) poisons every number.
Why this piece earns its place
Ask the judge to compare, not to grade. An absolute score on a scale is unstable across judge prompts, judge models and time, so the same figure six months apart is not the same measurement, while a pairwise preference between two outputs on one input is steadier — and it is exactly the question you have when deciding whether a change helped. Pairwise brings its own discipline: run every pair in both orders, and treat a pair whose winner flips when you swap them as an abstention rather than a result, because the judge has just told you it read position instead of content. Split the rubric as well — grounded, complete and correctly formatted fail for three different reasons with three different fixes, so score them separately and derive the headline from them. And treat a programmatic check as an invitation to delete the problem: if valid JSON matters, constrained decoding makes that check pass by construction, and a failure mode you have made impossible beats one you measure well. A run costs one scorer pass per case, so the judge bill rises with a set that is built to grow — which is what decides when each scorer runs: the deterministic checks on every commit, the judge on merge and nightly.
- programmaticstructured/verifiable
- LLM-judgeopen-ended, at scale
- humanground truth + calibration
What the new pieces do
- Scorers / Judgemodel
- The scoring methods: programmatic checks, an LLM-as-judge for open-ended quality, and human review. The judge must be validated against humans or its numbers are noise.
- Human Labelsservice
- Expert judgments — the ground truth that calibrates the LLM judge, seeds golden sets, and adjudicates the cases automated scorers can’t.
Back of the envelope
- validate the judge
- measure agreement with human labels
- check judge bias
- length, style, self-preference, position
- pin judge version
- its prompt+model is part of the metric
- humans on a sample
- keep the judge honest continuously
Step 5 · Gate every change
Offline evals as a CI regression gate
A developer tweaks a prompt to fix one case. It quietly breaks five others. Without a gate, that ships. How do you stop regressions from reaching users?
How do you prevent a prompt/model change from silently regressing quality?
Spot-checks miss regressions in cases you didn’t look at and aren’t repeatable. The eval suite has to run automatically on every change.
The Eval Runner scores the change against the golden sets automatically; if key metrics regress past a threshold, the deploy is blocked — treating prompts/models like code under test.
Infrequent evals let regressions accumulate between releases and make it hard to attribute which change caused what. Gate every change, like unit tests.
Wire the Eval Runner into CI as a regression gate: every prompt, model, or config change runs the golden-set evals, and a regression past threshold blocks the deploy. Prompts and models become artifacts under test, just like code — so a "small tweak" can’t silently trade one fixed case for five broken ones.
Why this piece earns its place
The thing under test is nondeterministic, which is what breaks a naive gate. Run the same suite twice against an unchanged commit and the score moves. So measure that noise band before choosing a threshold: run the suite repeatedly against a system you have not touched, see how far it wanders, and put the block point outside it. Get this backwards and the gate flakes, people learn to re-run until it goes green, and you have the ceremony of a gate with none of the protection. Pin what you can — temperature, a seed where the provider offers one, the judge’s version — to shrink the band rather than widen the threshold. Report the per-case diff and not only the aggregate: a suite can land on the same total while ten cases flip to failing and ten others flip to passing, which is a real change hiding inside a flat number. And give it a documented override with a recorded reason, or the gate gets switched off the first time it blocks a launch, permanently.
- run on every changeprompt · model · config
- block regressionspast threshold
- prompts as codeunder test in CI
What the new pieces do
- Eval Runnerservice
- Runs the system under test over an eval set and applies scorers to each output, producing metrics. The engine behind both offline (pre-ship) and online (production) evaluation.
- CI Regression Gateservice
- Runs the eval suite on every prompt/model/config change and blocks the deploy on a regression — the safety net that stops a "small tweak" from silently degrading quality.
Step 6 · Watch production
Online evaluation and drift
Offline evals pass, you ship — and a week later quality drifts: real inputs differ from your golden set, the provider silently updated the model, a new failure mode appears. Offline can’t see any of it. How do you catch production decay?
What catches quality problems that offline evals never saw?
A production monitor samples real calls, scores them (programmatic + judge), and alerts on quality drift, novel failures, and cost/latency regressions — the real distribution offline evals can’t reproduce.
Production always drifts from a fixed dataset — new inputs, provider model updates, changing users. Assuming they match is how silent decay goes unnoticed.
Offline evals still only test the frozen golden set, not live traffic. You need to score production itself to catch what the dataset doesn’t contain.
A Prod Monitor continuously samples and scores live traffic (programmatic checks + judge), alerting on quality drift, new failure modes, and cost/latency regressions. It sees the real distribution — new inputs, silent provider model updates, changing user behavior — that a frozen offline set never will. Offline gates the change; online watches reality.
Why this piece earns its place
Online scoring works with one hand tied. A live request has no expected answer, so every scorer that needs a reference is simply unavailable, and what remains is reference-free: format validity, refusal rate, whether an answer is grounded in the context it was given, a judge asked about quality with nothing to compare against. That changes what an alert means. You are not watching a pass rate, you are watching a distribution move, so the alarm compares a window against a baseline window instead of firing on one bad answer, and it can tell you that something changed but not what — naming it is still the offline set’s job. Wire the cheap signals first, because they move before any quality score does and cost nothing to compute: output length distribution, refusal rate, tool-call failure rate, tokens per request. A provider serving a different model underneath you usually shows up in those well before a judge notices the answers got worse.
- sample livescore continuously
- drift alertsquality · cost · latency
- sees realitythe real distribution
What the new pieces do
- Prod Monitorservice
- Continuously samples live traffic and scores it, watching for quality drift, new failure modes, and cost/latency regressions that offline evals never saw.
Step 7 · Close the loop
Metrics, user signals, and feedback
Eval scores are proxies. The ground truth is whether users are actually served well. And every new production failure is a test case you don’t yet have. How do you tie it together and keep improving?
How do you keep the eval system honest and growing?
Eval scores are proxies that can diverge from real satisfaction. Correlating them with user signals (thumbs, edits, escalations) is how you validate the proxy.
A frozen dataset goes stale as new failure modes appear. The set must grow — production failures become new eval cases — or it stops reflecting reality.
Dashboards tie eval metrics to real user feedback (validating the proxy), and newly-discovered failures from monitoring flow back as fresh eval cases — a compounding measurement flywheel.
Metrics + Feedback correlates eval scores with real user signals (thumbs, edits, escalations) to validate that the proxy tracks reality, and feeds new production failures back into the golden set. The loop closes: observability finds failures → they become eval cases → the gate prevents their return → monitoring finds the next ones. Your eval suite compounds into a moat.
Why this piece earns its place
The loop needs a promotion rule or it quietly ruins the dataset. Not every production failure deserves a case. One furious user’s oddity, added as a blocking fixture, makes the suite a graveyard of one-offs that every future prompt must satisfy, and the cheapest way to satisfy it is a special-case instruction that helps nobody else. Ask three things: does the failure represent a class, can the expected behavior be agreed rather than argued, and which slice does it belong to — a case with no slice label is invisible to per-category reporting. Cases also have to be retired; a fixture for a feature that has left the product still runs, still costs a scorer call, and can still block a deploy over behavior nobody wants any more. Treat user signals as a biased sampler too: unhappy people click more than happy ones and most click nothing, so the correlation points at the slice worth reading rather than giving a number to push up.
- correlateeval scores ↔ user signals
- failures → datasetclose the loop
- compoundsthe suite is a moat
What the new pieces do
- Metrics + Feedbackstore
- Dashboards correlating eval scores with real user signals (thumbs, edits, escalations), and the pipe that turns new failures into fresh eval cases.
The payoff
You built LLM eval & observability
From "is this actually good?" to a measurement system: full tracing, versioned golden datasets mined from production, layered scorers (programmatic + validated judge + human), a CI regression gate, online production monitoring for drift, and a feedback loop where failures become fixtures.
Now trust the judge — wire in an unvalidated LLM-as-judge — and watch every number quietly go wrong: the judge rewards length and its own style, so you optimize the metric, ship a regression, and the gate stays green while real quality falls. That’s why the judge must be calibrated against humans, version-pinned, and spot-checked — an uncalibrated judge corrupts the whole pipeline it feeds.
Everything you assembled, in order
- Two loops — offline evals gate changes; online observability watches production
- Tracing — capture every call — the foundation for everything downstream
- Golden datasets — versioned, mined from production + edge cases — your real benchmark
- Scorers — programmatic + LLM-judge + human, each for the right output
- CI gate — evals run on every change and block regressions — prompts as code
- Prod monitor — sample and score live traffic for drift offline never sees
- Feedback loop — production failures become new eval cases — a compounding moat
- The meta-failure — an unvalidated judge silently corrupts every metric — calibrate it
