Vibe Engines
YouTube
AI System Design

Design LLM Eval & Observability

Learn AI system design by building an LLM evaluation and observability pipeline step by step.

The numbers to beatfull traceprompt · output · tools · costdurable storequeryable, replayablefeeds alleval · monitor · debug

The whole design, in writing

Learn AI system design by building an LLM evaluation and observability pipeline step by step. An interactive guide to measuring AI quality — tracing every call, curating versioned golden datasets, scoring with programmatic/LLM-judge/human methods, a CI regression gate, online production monitoring, and a feedback loop — plus why trusting an unvalidated LLM judge silently corrupts every number you ship on.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

You can’t ship what you can’t measure

An LLM feature has no compiler and no unit test that says "correct." Outputs are open-ended, quality is subjective, and the same prompt can regress when you tweak a word, swap a model, or the provider updates theirs — silently, with no error. Demos are easy; knowing whether you’re actually getting better (or quietly getting worse) is the hard part that separates a toy from a product. How do you measure AI quality, continuously?

Build an eval + observability pipeline: trace every call (observability), score the system against versioned golden datasets with validated scorers (evaluation), gate changes in CI, and monitor production for drift — closing a loop where real failures become new test cases. Measurement is the system that lets you improve on purpose instead of by vibes.

Step 1 · Two loops

Offline evals and online observability

There are two questions: "is this change good before I ship it?" and "is it still good in production?" They need different machinery. What’s the overall shape?

What are the two complementary halves of measuring an LLM system?

  1. A one-time benchmark misses regressions from later changes and real-world drift. Measurement is continuous, in two loops — pre-ship and in production.

  2. Offline evals score changes against versioned golden sets to catch regressions before deploy; online observability traces and scores live traffic to catch drift and new failures after deploy. Both are needed.

  3. Monitoring alone means every regression ships to users first. You want to catch what you can offline and watch production for what offline missed.

Two complementary loops: offline evaluation runs a change against fixed, versioned datasets before shipping (catch regressions early), and online observability traces and scores production traffic (catch drift and novel failures). The rest of this design builds both on a shared foundation — traces in, scores out.

Why this piece earns its place

The two loops look like one system on the diagram, and the expensive mistake is building them as two. The offline suite gets written first, as a script living beside the prompts; months later someone adds production scoring and writes a second implementation of the same checks in a different service. Now both loops report a number with the same name computed by different code, and when they disagree nobody can say whether quality drifted or the scorers did. Write the scoring once, as a library the runner calls, and make the source of inputs — a dataset, or a sample of traffic — the argument. Sharing the code does not make the two numbers interchangeable, though. An offline score is computed on a set you balanced on purpose; an online one on whatever mix of questions arrived that day. The same metric name therefore means two different things, and a gap between them is not evidence of anything — keep them as two series, each compared against its own history. Skip the online half only while nothing is live yet: a prototype has traces, a small offline set, and nothing to watch.

Step 2 · Capture everything

Tracing is the foundation

You can’t evaluate, debug, or mine examples from calls you didn’t record. Before any scoring, what has to be captured on every LLM interaction?

LLM Appprod trafficTrace Collectorcapture callsTrace Storeevery call
New in this step: LLM App, Trace Collector, Trace Store.

What must you log on every LLM call to make evaluation possible?

  1. The output alone can’t explain a failure or reproduce it — you need the prompt, params, tool calls, cost, and latency too. Partial logging cripples both debugging and eval-set mining.

  2. Ops metrics miss the content entirely; you can’t judge quality or build datasets from them. LLM observability is about the calls themselves, not just system health.

  3. A Trace Collector records the complete interaction into a queryable store — the raw material for debugging, mining eval examples, monitoring, and cost attribution.

A Trace Collector captures the full interaction — prompt, output, tokens, latency, cost, tool calls, and any user feedback — into a durable Trace Store. This is the observability bedrock: you can’t evaluate, debug, attribute cost, or build datasets from calls you never recorded. Everything downstream reads from these traces.

Why this piece earns its place

The collector sits on the request path, and that one fact decides most of its design: it has to be fire-and-forget — buffer locally, ship asynchronously, drop on backpressure. A collector that blocks or throws takes the product down to protect a measurement, which is the worst trade available here. What it captures has to be structured rather than one concatenated string. Store the prompt template id and the resolved variables separately, store the model version the provider actually served rather than the alias that was configured, and store the parameters really sent. Without that you cannot ask how a given template version performed, and you cannot rebuild a case after the template moves on. One user interaction is usually several calls — retries, tool calls, a chain of steps — so they need a shared trace_id grouping them, or you evaluate one hop of an agent and call it the system. Retention is where the cost is decided: full text hot for a short window, sampled or summarised beyond it.

  • full traceprompt · output · tools · cost
  • durable storequeryable, replayable
  • feeds alleval · monitor · debug

What the new pieces do

LLM Appclient
The live AI feature. Every call it makes is both a user interaction and a data point for evaluation — the source of traces and of real-world examples to test against.
Trace Collectorbackend
Instruments every LLM call: prompt, output, tokens, latency, cost, tool calls, and user feedback. The observability foundation — you can’t evaluate what you don’t capture.
Trace Storestore
Durable, queryable store of production traces. The raw material for debugging, for mining eval examples, and for online monitoring.

Step 3 · The yardstick

Versioned golden datasets

To say "better" or "worse" you need something fixed to measure against. Where does that yardstick come from, and why must it be versioned?

Trace Storeevery callEval Setgolden · versioned
New in this step: Eval Set.

What do you evaluate a change against?

  1. Build eval sets from production traces (representative), known failures, and edge cases, each with expected behavior; version them so a score is comparable across runs and a change to the data is explicit.

  2. Ad-hoc examples that change every run make scores incomparable and easy to game. The dataset must be a stable, versioned artifact.

  3. Public benchmarks don’t reflect your task, distribution, or failure modes. They’re a weak proxy; your own golden set from production is the real measure.

Curate versioned golden datasets: representative cases mined from production traces, plus known failures and edge cases, each labeled with expected behavior. Version them so scores are comparable across runs and dataset changes are explicit. Your own eval set — not a public benchmark — is the yardstick that matters, and it grows as production reveals new failures.

Why this piece earns its place

Start with what expected behavior actually means inside a row. A single gold answer pushes you toward exact match, which does not fit open-ended output, or invites a judge to grade phrasing instead of substance. Most cases are better stored as assertions — must state the refund window, must not promise a callback, must decline — because an assertion survives a rewrite of the wording and names what the failure actually was. Then guard the set against yourself. The cases you stare at while iterating are the ones you tune the prompt to satisfy, so their score stops predicting anything about unseen traffic; that is overfitting to a dev set, and the remedy is the split ML has always used — hold a slice back, never read it while iterating, report it separately. Finally, stamp the dataset version onto every score you keep, not only onto the dataset. A stored number with no version attached turns a long trend line into two yardsticks glued together, with the join invisible.

  • golden setreal + edge cases
  • versionedcomparable across runs
  • mined from prodvia the traces

What the new pieces do

Eval Setstore
Curated, versioned datasets of inputs with expected behavior — mined from production and edge cases. The fixed yardstick every prompt/model change is measured against.

Step 4 · How to score

Programmatic, LLM-judge, human

For "did it return valid JSON?" a simple check works. For "was this answer helpful and correct?" there’s no regex. What scoring methods do you use, and which for what?

Scorers / Judgemetric methodsHuman Labelsground truth
New in this step: Scorers / Judge, Human Labels.

How do you score open-ended LLM outputs?

  1. Exact match fails for open-ended generation where many phrasings are correct — it’s only right for constrained/structured outputs. Open-ended quality needs judgment.

  2. Human grading is the gold standard but doesn’t scale to every eval run. You reserve humans for calibration and hard cases, and automate the rest.

  3. Use cheap deterministic checks for structured/verifiable outputs, an LLM-as-judge for subjective quality at scale, and human labels as ground truth to validate the judge and handle the ambiguous.

Layer the scorers: programmatic checks for verifiable/structured outputs (JSON valid, contains fact, matches regex), an LLM-as-judge for open-ended quality at scale, and human labels as ground truth. Critically, the judge is only trustworthy if validated against humans — measured for agreement and bias — because an uncalibrated judge (the chaos button) poisons every number.

Why this piece earns its place

Ask the judge to compare, not to grade. An absolute score on a scale is unstable across judge prompts, judge models and time, so the same figure six months apart is not the same measurement, while a pairwise preference between two outputs on one input is steadier — and it is exactly the question you have when deciding whether a change helped. Pairwise brings its own discipline: run every pair in both orders, and treat a pair whose winner flips when you swap them as an abstention rather than a result, because the judge has just told you it read position instead of content. Split the rubric as well — grounded, complete and correctly formatted fail for three different reasons with three different fixes, so score them separately and derive the headline from them. And treat a programmatic check as an invitation to delete the problem: if valid JSON matters, constrained decoding makes that check pass by construction, and a failure mode you have made impossible beats one you measure well. A run costs one scorer pass per case, so the judge bill rises with a set that is built to grow — which is what decides when each scorer runs: the deterministic checks on every commit, the judge on merge and nightly.

  • programmaticstructured/verifiable
  • LLM-judgeopen-ended, at scale
  • humanground truth + calibration

What the new pieces do

Scorers / Judgemodel
The scoring methods: programmatic checks, an LLM-as-judge for open-ended quality, and human review. The judge must be validated against humans or its numbers are noise.
Human Labelsservice
Expert judgments — the ground truth that calibrates the LLM judge, seeds golden sets, and adjudicates the cases automated scorers can’t.

Back of the envelope

validate the judge
measure agreement with human labels
check judge bias
length, style, self-preference, position
pin judge version
its prompt+model is part of the metric
humans on a sample
keep the judge honest continuously

Step 5 · Gate every change

Offline evals as a CI regression gate

A developer tweaks a prompt to fix one case. It quietly breaks five others. Without a gate, that ships. How do you stop regressions from reaching users?

Eval Setgolden · versionedEval RunnerscoreCI Regression Gateblock regress
New in this step: Eval Runner, CI Regression Gate.

How do you prevent a prompt/model change from silently regressing quality?

  1. Spot-checks miss regressions in cases you didn’t look at and aren’t repeatable. The eval suite has to run automatically on every change.

  2. The Eval Runner scores the change against the golden sets automatically; if key metrics regress past a threshold, the deploy is blocked — treating prompts/models like code under test.

  3. Infrequent evals let regressions accumulate between releases and make it hard to attribute which change caused what. Gate every change, like unit tests.

Wire the Eval Runner into CI as a regression gate: every prompt, model, or config change runs the golden-set evals, and a regression past threshold blocks the deploy. Prompts and models become artifacts under test, just like code — so a "small tweak" can’t silently trade one fixed case for five broken ones.

Why this piece earns its place

The thing under test is nondeterministic, which is what breaks a naive gate. Run the same suite twice against an unchanged commit and the score moves. So measure that noise band before choosing a threshold: run the suite repeatedly against a system you have not touched, see how far it wanders, and put the block point outside it. Get this backwards and the gate flakes, people learn to re-run until it goes green, and you have the ceremony of a gate with none of the protection. Pin what you can — temperature, a seed where the provider offers one, the judge’s version — to shrink the band rather than widen the threshold. Report the per-case diff and not only the aggregate: a suite can land on the same total while ten cases flip to failing and ten others flip to passing, which is a real change hiding inside a flat number. And give it a documented override with a recorded reason, or the gate gets switched off the first time it blocks a launch, permanently.

  • run on every changeprompt · model · config
  • block regressionspast threshold
  • prompts as codeunder test in CI

What the new pieces do

Eval Runnerservice
Runs the system under test over an eval set and applies scorers to each output, producing metrics. The engine behind both offline (pre-ship) and online (production) evaluation.
CI Regression Gateservice
Runs the eval suite on every prompt/model/config change and blocks the deploy on a regression — the safety net that stops a "small tweak" from silently degrading quality.

Step 6 · Watch production

Online evaluation and drift

Offline evals pass, you ship — and a week later quality drifts: real inputs differ from your golden set, the provider silently updated the model, a new failure mode appears. Offline can’t see any of it. How do you catch production decay?

Trace CollectorEval RunnerProd Monitor
New in this step: Prod Monitor. · swipe to pan the diagram

What catches quality problems that offline evals never saw?

  1. A production monitor samples real calls, scores them (programmatic + judge), and alerts on quality drift, novel failures, and cost/latency regressions — the real distribution offline evals can’t reproduce.

  2. Production always drifts from a fixed dataset — new inputs, provider model updates, changing users. Assuming they match is how silent decay goes unnoticed.

  3. Offline evals still only test the frozen golden set, not live traffic. You need to score production itself to catch what the dataset doesn’t contain.

A Prod Monitor continuously samples and scores live traffic (programmatic checks + judge), alerting on quality drift, new failure modes, and cost/latency regressions. It sees the real distribution — new inputs, silent provider model updates, changing user behavior — that a frozen offline set never will. Offline gates the change; online watches reality.

Why this piece earns its place

Online scoring works with one hand tied. A live request has no expected answer, so every scorer that needs a reference is simply unavailable, and what remains is reference-free: format validity, refusal rate, whether an answer is grounded in the context it was given, a judge asked about quality with nothing to compare against. That changes what an alert means. You are not watching a pass rate, you are watching a distribution move, so the alarm compares a window against a baseline window instead of firing on one bad answer, and it can tell you that something changed but not what — naming it is still the offline set’s job. Wire the cheap signals first, because they move before any quality score does and cost nothing to compute: output length distribution, refusal rate, tool-call failure rate, tokens per request. A provider serving a different model underneath you usually shows up in those well before a judge notices the answers got worse.

  • sample livescore continuously
  • drift alertsquality · cost · latency
  • sees realitythe real distribution

What the new pieces do

Prod Monitorservice
Continuously samples live traffic and scores it, watching for quality drift, new failure modes, and cost/latency regressions that offline evals never saw.

Step 7 · Close the loop

Metrics, user signals, and feedback

Eval scores are proxies. The ground truth is whether users are actually served well. And every new production failure is a test case you don’t yet have. How do you tie it together and keep improving?

LLM AppTrace CollectorTrace StoreEval SetEval RunnerScorers / JudgeHuman LabelsCI Regression GateProd MonitorMetrics + Feedback
New in this step: Metrics + Feedback. · swipe to pan the diagram

How do you keep the eval system honest and growing?

  1. Eval scores are proxies that can diverge from real satisfaction. Correlating them with user signals (thumbs, edits, escalations) is how you validate the proxy.

  2. A frozen dataset goes stale as new failure modes appear. The set must grow — production failures become new eval cases — or it stops reflecting reality.

  3. Dashboards tie eval metrics to real user feedback (validating the proxy), and newly-discovered failures from monitoring flow back as fresh eval cases — a compounding measurement flywheel.

Metrics + Feedback correlates eval scores with real user signals (thumbs, edits, escalations) to validate that the proxy tracks reality, and feeds new production failures back into the golden set. The loop closes: observability finds failures → they become eval cases → the gate prevents their return → monitoring finds the next ones. Your eval suite compounds into a moat.

Why this piece earns its place

The loop needs a promotion rule or it quietly ruins the dataset. Not every production failure deserves a case. One furious user’s oddity, added as a blocking fixture, makes the suite a graveyard of one-offs that every future prompt must satisfy, and the cheapest way to satisfy it is a special-case instruction that helps nobody else. Ask three things: does the failure represent a class, can the expected behavior be agreed rather than argued, and which slice does it belong to — a case with no slice label is invisible to per-category reporting. Cases also have to be retired; a fixture for a feature that has left the product still runs, still costs a scorer call, and can still block a deploy over behavior nobody wants any more. Treat user signals as a biased sampler too: unhappy people click more than happy ones and most click nothing, so the correlation points at the slice worth reading rather than giving a number to push up.

  • correlateeval scores ↔ user signals
  • failures → datasetclose the loop
  • compoundsthe suite is a moat

What the new pieces do

Metrics + Feedbackstore
Dashboards correlating eval scores with real user signals (thumbs, edits, escalations), and the pipe that turns new failures into fresh eval cases.

The payoff

You built LLM eval & observability

From "is this actually good?" to a measurement system: full tracing, versioned golden datasets mined from production, layered scorers (programmatic + validated judge + human), a CI regression gate, online production monitoring for drift, and a feedback loop where failures become fixtures.

LLM AppTrace CollectorTrace StoreEval SetEval RunnerScorers / JudgeHuman LabelsCI Regression GateProd MonitorMetrics + Feedback
The finished design, end to end. · swipe to pan the diagram

Now trust the judge — wire in an unvalidated LLM-as-judge — and watch every number quietly go wrong: the judge rewards length and its own style, so you optimize the metric, ship a regression, and the gate stays green while real quality falls. That’s why the judge must be calibrated against humans, version-pinned, and spot-checked — an uncalibrated judge corrupts the whole pipeline it feeds.

Everything you assembled, in order

  • Two loops — offline evals gate changes; online observability watches production
  • Tracing — capture every call — the foundation for everything downstream
  • Golden datasets — versioned, mined from production + edge cases — your real benchmark
  • Scorers — programmatic + LLM-judge + human, each for the right output
  • CI gate — evals run on every change and block regressions — prompts as code
  • Prod monitor — sample and score live traffic for drift offline never sees
  • Feedback loop — production failures become new eval cases — a compounding moat
  • The meta-failure — an unvalidated judge silently corrupts every metric — calibrate it

Deep cut · 19:07

The score went up. The product got worse.

The interactive build above walks the pipeline: traces, versioned golden datasets, three kinds of scorer, the CI gate, production monitoring and the feedback loop. This film starts from a support bot whose eval score rose four tenths of a point while a customer complained it got wordier, then puts the judge model itself on the stand — checked against two people, it agreed only sixty-one percent of the time.

  • See the silent failure: nothing errors and every eval stays green, because the judge grading each answer prefers the longer one and picks a different winner when you swap the order.
  • Take it into the interview: defend full traces, a dataset sliced by kind so an average cannot hide a failing category, and a judge validated on three hundred labelled examples — then fix its bias and show agreement climb from 61% to 89%.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. The LLM judge and a human labeler disagree on a borderline case. What happens to that disagreement?

    It gets logged and counted toward the judge’s measured agreement rate, not silently discarded — a judge with 95% agreement is calibrated differently (trusted more broadly) than one at 70%, and the specific disagreement cases are exactly what you review to find the judge’s systematic biases (does it consistently favor one style, one length, one tone on the cases it gets wrong?). Disagreement is data, not noise to throw away.

  2. Should every metric in the CI gate hard-block a deploy, or should some just warn?

    Not uniformly — a hard-block belongs on metrics tied to correctness or safety (factual accuracy, format validity, safety-check pass rate) where a regression is unambiguous harm, while a softer warn-only tier fits noisier or more subjective metrics (a judge-scored "helpfulness" dipping half a point) where blocking every deploy on noise would grind velocity to a halt without proportional benefit. The threshold-per-metric decision mirrors content moderation’s per-category severity tiers.

  3. Production traces contain real user prompts — potentially names, order numbers, health questions. Does the observability pipeline need its own privacy handling?

    Yes, and it’s easy to miss because "it’s just logging for evals" makes it feel exempt from the privacy discipline the app itself follows. The trace store needs the same PII redaction/access-control posture as any other store holding user data, and golden sets MINED from those traces need to be scrubbed before humans or an LLM judge (a third party, from the data’s perspective) ever sees them.

  4. The prod monitor samples 1% of live traffic instead of scoring everything. What does that trade away?

    Rare failure modes — a bug that only triggers on 1-in-10,000 requests may simply never appear in a 1% sample until it’s already affected thousands of real users. The fix isn’t always "sample more" (cost scales with judge calls); it’s often stratified sampling that over-samples traffic segments more likely to contain edge cases (new features, flagged-uncertain outputs, users who previously gave negative feedback) rather than pure random sampling.

  5. You A/B test two model providers and eval scores differ. How do you know it’s the model and not just a shift in which users landed in which arm?

    By checking that the traffic assigned to each arm is actually comparable — same distribution of query types, user segments, and time-of-day — before attributing the score gap to the model itself; a naive A/B split can accidentally correlate with something else (one arm getting more complex queries) and the eval pipeline’s golden-set comparison (same fixed inputs to both models) is what isolates the model variable cleanly, versus the noisier live-traffic comparison.

Check yourself — the answers, and why

Eight steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. Measuring an LLM system needs which two loops?

    Offline evals catch regressions before shipping; online monitoring catches drift and novel failures after — neither alone is sufficient.

  2. Tracing comes before evaluation because…

    Full traces (prompt, output, tools, cost) are the raw material for golden sets, monitoring, and debugging — observability is the foundation.

  3. The right yardstick to evaluate against is…

    Public benchmarks don’t reflect your task; versioned golden sets from your traces make scores comparable and relevant.

  4. You should score open-ended outputs with…

    Deterministic checks for structured outputs, an LLM-judge for subjective quality at scale, humans to calibrate and adjudicate — right scorer per output.

  5. Making evals a CI regression gate means…

    Treating prompts/models like code under test stops a small tweak from silently trading one fixed case for several broken ones.

  6. An unvalidated LLM-as-judge is dangerous because it…

    If the judge isn’t calibrated against humans, you optimize its quirks and ship regressions while the gate stays green; calibrate, version-pin, and spot-check it.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Trace: capture every call — prompt, output, tokens, latency, cost, tool calls, feedback — into a durable store.
  • Curate: build versioned golden datasets from production traces plus edge cases.
  • Score: programmatic checks, a validated LLM-as-judge, and human labels — the right scorer per output.
  • Gate: run evals in CI on every prompt/model/config change and block a regression.
  • Monitor + loop: sample and score live traffic for drift, and feed new failures back as eval cases.

The qualities that shape everything

Each one names the mechanism that buys it.

Stop a bad change reaching users
Wire the eval runner into CI as a regression gate — every prompt/model/config change runs the golden-set evals and a regression past threshold blocks the deploy.
Be able to evaluate, debug, and mine examples at all
A trace collector captures the full interaction — prompt, output, tokens, latency, cost, tool calls, feedback — into a durable store; you can’t measure calls you never recorded.
A “better or worse” claim you can trust
Versioned golden datasets mined from production traces plus edge cases make scores comparable across runs and dataset changes explicit.
Score open-ended, subjective quality
Layer programmatic checks, an LLM-as-judge validated against humans, and human labels as ground truth — the right scorer per output.
Catch drift offline evals never see
A prod monitor continuously samples and scores live traffic, alerting on quality drift, new failure modes, and cost/latency regressions.
A measurement suite that compounds
The feedback loop pipes new production failures back into the versioned golden set, so the system provably never regresses on a known failure again.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

Offline evals + online observability over a one-time pre-launch benchmark

A benchmark run once misses every later regression and real-world drift; you need a pre-ship gate against fixed datasets and continuous production monitoring, not a single measurement.

Full traces over logging the output text alone

The output alone can’t explain or reproduce a failure or seed a dataset; capturing prompt, params, tool calls, cost and latency is what makes debugging and eval-mining possible.

Your own versioned golden set over a public benchmark

Public benchmarks don’t reflect your task, distribution, or failure modes; a versioned set mined from your production traces is the yardstick that matters — and comparable across runs.

A layered scorer mix over exact string match

Exact match only fits constrained outputs and fails open-ended generation where many phrasings are right; deterministic checks where you can, an LLM-judge for subjective quality, humans to calibrate.

A judge validated against humans over trusting the judge’s scores

An unvalidated LLM-judge rewards length and its own style, so you optimize its quirks and ship a regression while the gate stays green; calibrating against human labels is what makes its numbers mean anything.

The answer, out loud

What a strong answer to “Design LLM Eval & Observability” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Pin down what “better” means

    Before I draw anything I’d agree on what better means, because on a support bot that’s harder than the architecture. I’d ask who decides an answer was bad, and what bad looks like here: an invented promise, a missing caveat, a tone that turns a mild complaint into an escalation. Those are three different failures with three different fixes, and one blended score hides all of them. I’d also say this is infrastructure, not a pre-launch project — the prompt, the model and the provider all change under us, so it has to still answer “did this change help” a year from now.

  2. 3–8 min

    Two loops on one foundation

    The shape is two loops on one foundation. Offline evaluation asks whether a change is good before it ships, and answers with inputs I choose. Online observability asks whether it’s still good now, and answers with traffic I don’t. Neither question can be answered with the other’s evidence, which is why I want both. Underneath sits tracing — the whole record of one call: prompt, output, tokens, latency, cost, tool calls, any feedback. I’d build that first, before there’s anything to score, because it’s the one piece with a deadline. A scorer written next month can be run over last month’s traffic, but only if last month got written down.

    Built in step 2: Tracing is the foundation
  3. 8–14 min

    A yardstick that doesn’t move

    To say better I need something that doesn’t move: a golden set, inputs paired with the behavior we expect. I’d mine it from our own traces so it looks like the traffic we really get, then add the failures support already knows about. I’d be honest that this is the expensive part, and it’s expensive in people-time — somebody has to decide what the right answer was, and that somebody is usually support, not me. So I’d start small enough that a person has read every case. It lives in the repo and edits go through review, because changing a case changes what passing means.

    Built in step 3: Versioned golden datasets
  4. 14–22 min

    The right scorer, and a checked judge

    Then scoring, and I’d sort outputs by what can be checked without an opinion. Valid JSON, contains the policy number, inside the length cap — that’s code returning pass or fail, and I’d push as much as I can into that bucket, because everything left over costs money or people. “Was this helpful” has no such check: exact match fails when many phrasings are right, and humans grading every run doesn’t scale. So an LLM-as-judge — a second model prompted to grade answers — covers the open-ended middle, with human labels underneath as ground truth. Nobody quotes the judge until we know how often it agrees with a person. And it’s a dependency: when the judge changes, old numbers and new ones aren’t the same measurement, so I’d re-score the history rather than draw a line across the seam and hope.

    Built in step 4: Programmatic, LLM-judge, human
  5. 22–28 min

    Gate every change like code

    Then the eval runner goes into CI as a gate: every prompt, model or config change runs the set, and a regression past threshold stops the deploy. What I like about that isn’t only the catching — it’s who it lets ship. Once the suite is the thing that says yes, a support lead can edit a prompt and find out in minutes whether it broke anything, instead of an engineer eyeballing a few outputs and merging on a feeling. I’d run it report-only at first: it comments on every change without blocking, we learn what it flags and what that would have cost us, and only then does it get to say no.

    Built in step 5: Offline evals as a CI regression gate
  6. 28–36 min

    Watch reality, then close the loop

    After we ship, the set stops being the whole story — real inputs drift, the provider can update its model under us, and new failure modes turn up that nothing is looking for. So a monitor samples live traffic and scores what can be scored without a reference answer, since a live request doesn’t come with one. Then the part I’d push hardest: every scorer we own measures a failure someone already thought of. New ones get in exactly one way, and that is a person reading real answers. So I’d book a standing slot where somebody reads a sample, picks what is genuinely wrong, and writes those into the set. Thumbs and escalations make good bookmarks for that reading.

    Built in step 6: Online evaluation and drift
  7. 36–42 min

    What I’d watch, and how it fails

    On the dashboard: scores per slice, what the gate has blocked, live quality beside cost and latency, and the judge’s agreement with people. Both failures I’d plan for are the measurement breaking, not the product. One is the suite failing open — a judge call times out, a case errors, the runner scores what it can and prints a number that looks like all the others. I’d track how many cases actually got scored and treat a short run as red, not as a smaller run. The other is alerts nobody reads: set the monitor too sensitive, it fires on ordinary variation, the channel gets muted, and the one real drift lands in a muted channel. I’d start those thresholds embarrassingly quiet.

  8. 42–45 min

    Close on the trade-off

    To close, in one breath: trace every call, measure against a versioned golden set mined from production, score with checks, a validated judge and humans, gate every change in CI, watch live traffic for drift, and turn every production failure into a test. Every knob — judge calls, human labels, how much traffic we sample — trades the cost of measuring against how much of the truth we see. With more time I’d go at two things. What we do when two of our own labelers disagree about a case — that is a spec we never wrote, not a scoring problem. And letting support file cases themselves, so the set grows at the speed failures are found.

What this teaches

Learn AI system design by building an LLM evaluation and observability pipeline step by step. An interactive guide to measuring AI quality — tracing every call, curating versioned golden datasets, scoring with programmatic/LLM-judge/human methods, a CI regression gate, online production monitoring, and a feedback loop — plus why trusting an unvalidated LLM judge silently corrupts every number you ship on.

Key takeaways

  • Two loops — offline evals gate changes; online observability watches production
  • Tracing — capture every call — the foundation for everything downstream
  • Golden datasets — versioned, mined from production + edge cases — your real benchmark
  • Scorers — programmatic + LLM-judge + human, each for the right output
  • CI gate — evals run on every change and block regressions — prompts as code
  • Prod monitor — sample and score live traffic for drift offline never sees
  • Feedback loop — production failures become new eval cases — a compounding moat
  • The meta-failure — an unvalidated judge silently corrupts every metric — calibrate it

Concepts covered

  • You can’t ship what you can’t measure
  • Offline evals and online observability
  • Tracing is the foundation
  • Versioned golden datasets
  • Programmatic, LLM-judge, human
  • Offline evals as a CI regression gate
  • Online evaluation and drift
  • Metrics, user signals, and feedback
RUN IT YOURSELF

Two broken systems, the same 0.750, opposite repairs

One frozen golden set, scored on two metrics instead of one: retrieval recall (did we fetch the chunk holding the answer) and answer faithfulness (does the answer assert only what this run retrieved). It runs four systems for real in your browser — two break in opposite halves and land on the same blended number, and a third column prints what faithfulness turns into if you score against the whole policy instead. Edit STALE, or the value the model remembers for refund_card in ASK, then hit Run.

HOW TO READ THE CODE — 4 IDEAS
  1. The two metrics count different populations — cases retrieved vs claims made — so one sits at 1.00 while the other collapses (step 4: the right scorer per output).
  2. Rows 2 and 3 both blend to 0.750 and need opposite repairs, so a gate that blocks on one average cannot name the half to fix (steps 5–6).
  3. Faithfulness is scored against what that run fetched. The true column re-scores the same claims against the whole policy: row 3 reads 1.00 true and 0.50 faithful — every figure it volunteers is correct, and it retrieved none of them.
  4. Row 2 is the mirror: faithfulness 1.00 while a third of its answers are false, lifted intact from a chunk the reindex never replaced. Only recall sees that (step 3).
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, trust the judge, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs