Vibe Engines
YouTube
AI System Design

Design a Fine-Tuning Pipeline

Learn AI system design by building an LLM fine-tuning and training pipeline step by step.

The numbers to beatclean + dedupquality over volumeconsistent formatinstruction → responseheld-out setnever seen in training

The whole design, in writing

Learn AI system design by building an LLM fine-tuning and training pipeline step by step. An interactive guide to the offline path — dataset curation, base-model choice and parameter-efficient fine-tuning (LoRA/PEFT), the distributed training loop, a held-out evaluation gate, a versioned model registry, canary deployment, and the production data flywheel — plus the silent trap of eval-set contamination.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

When prompting isn’t enough, fine-tune

Prompting and RAG get you far: they steer a model with instructions and context at inference time. But sometimes you need the model to internalize a style, a format, a domain, or a skill — reliably, without a giant prompt every call. That means changing the weights. How do you turn "we have examples of the behavior we want" into a better model, safely and repeatably?

Build a fine-tuning pipeline: an offline, repeatable path from curated data through training and a hard evaluation gate to a versioned, deployable model — with a feedback loop that turns production into the next dataset. It’s an ML system, so it lives or dies on data quality and honest evaluation.

Step 1 · The skeleton

Data → train → eval → register → deploy

Fine-tuning isn’t "run a training script." It’s a pipeline with stages that must be reproducible and gated. What’s the minimal shape that keeps a model change safe?

What’s the right shape for a fine-tuning pipeline?

  1. No curation, no eval gate, no versioning — you can’t tell if the model improved or regressed, or reproduce it. That’s how a worse model reaches production.

  2. Each stage is reproducible and traceable: a versioned dataset trains a model, an eval gate decides promotion, the registry records lineage, and deployment is gradual and reversible.

  3. Online training on raw traffic is unstable and unsafe — no gate, no reproducibility, easy to poison. Training is an offline, gated pipeline; production only feeds it.

The pipeline is a chain of reproducible, gated stages: curate a versioned dataset → train → pass an evaluation gate → register the artifact with full lineage → deploy behind a canary. Every arrow is traceable, so you always know what a model learned from and can roll back to any prior version.

Why this piece earns its place

The stages matter less than the boundaries between them. Each boundary is a handoff of an artifact — a dataset version, a checkpoint, a score report — and if those artifacts are not addressable by a stable id, the pipeline can only ever be re-run whole. That is the expensive mistake, because training dominates what this whole page costs to operate, and re-running it to recover from something that broke downstream of it buys nothing at all. A gate that flakes, a failed registry write, a typo in a canary config: each should re-enter at its own stage against the checkpoint that already exists, not at the top. Make every stage take its inputs as explicit parameters — base id, dataset version, hyperparameters — instead of reading them from a file in whoever's branch triggered the run, and that re-entry is free. The shape buys an organisational contract too. Curation, training and serving usually sit with different people, and the artifact at each boundary is the thing they agree on rather than a meeting. Skip all of it for a one-off experiment; build it the moment you intend a second model, because the second one is where comparisons start to matter.

Step 2 · Garbage in, garbage model

Dataset curation is the real work

The instinct is "more data = better model." But a fine-tuned model mimics its training data exactly — including its noise, duplicates, and mistakes. What actually makes a good fine-tuning dataset?

Raw DataCurate & FormatTraining Set
New in this step: Raw Data, Curate & Format, Training Set. · swipe to pan the diagram

How do you build a fine-tuning dataset that helps?

  1. Raw dumps carry duplicates, junk, contradictions and PII; the model faithfully learns all of it. Volume without curation usually makes the model worse, not better.

  2. Too few examples underfit the target behavior. The goal isn’t minimal or maximal — it’s clean, diverse, correct examples that represent the task.

  3. Curation dominates outcomes: remove junk and near-duplicates, filter for correctness, format into consistent instruction→response pairs, and carve off a strictly separate held-out eval set.

Curate: clean and filter for correctness, deduplicate (including near-duplicates), format into consistent examples, strip PII, and — critically — split off a held-out eval set that never touches training. Version the result. Quality and diversity beat raw volume; the model will imitate exactly what you feed it.

Why this piece earns its place

Versioning a dataset is not the same as being able to rebuild it. A version defined as a query over production logs stops being reproducible the moment retention deletes the rows underneath it, and a deletion request removes them sooner than that. So curation materialises a snapshot — the actual examples, written once as an immutable artifact — and every run pins the snapshot, never the query. The uncomfortable corollary: a user asking to be forgotten is only half handled when the rows go, because their examples also sit inside the weights of every model trained on that snapshot, and the only real remedy there is retraining without them. Decide that policy while the registry still holds a handful of models, not once it holds a fleet of them. This stage quietly owns one more thing — the prompt template. Examples are formatted here, and the serving path has to format live requests identically: same system text, same delimiters, same role names. Drift by a newline and the model meets a distribution at inference it never trained on, degrading for reasons no metric explains. Treat that template as a shared versioned artifact, not a string that exists in two codebases.

  • clean + dedupquality over volume
  • consistent formatinstruction → response
  • held-out setnever seen in training

What the new pieces do

Raw Datastore
The raw material: production transcripts, labeled examples, documents, human demonstrations. Fine-tuning quality is capped by this — garbage in, garbage model.
Curate & Formatbus
Cleans, deduplicates, filters and formats raw data into training examples (e.g. instruction→response pairs). Also splits off a held-out eval set — and must keep it strictly separate.
Training Setstore
The finished, versioned training data. Every training run pins a dataset version so results are reproducible and traceable to exactly what the model learned from.

Back of the envelope

dedup near-duplicates
not just exact — fuzzy/semantic too
time-based split
eval on newer data to mimic the future
balance the mix
diversity prevents narrow overfitting
version datasets
pin exactly what a run trained on

Step 3 · Adapt, don’t retrain

Base model + parameter-efficient fine-tuning

You have data. Do you retrain a model from scratch? Update all of a pretrained model’s billions of weights? Both are expensive and risky. What’s the practical way to specialize a model?

Base Modelpretrained
New in this step: Base Model.

How do most teams actually fine-tune a large model?

  1. Pretraining costs millions of dollars and needs internet-scale data. Almost no one does it; you adapt an existing pretrained base instead.

  2. Parameter-efficient fine-tuning trains a tiny set of new parameters (low-rank adapters) on top of a frozen base — far cheaper, faster, less prone to catastrophic forgetting, and adapters can be swapped per task.

  3. Full fine-tuning works but is compute- and memory-heavy and can overwrite general capability (catastrophic forgetting). PEFT gets most of the benefit for a fraction of the cost.

Start from a pretrained base model and, in most cases, use parameter-efficient fine-tuning (LoRA/PEFT): freeze the base and train small adapters. It’s cheaper, faster, resists catastrophic forgetting, and lets you keep many task-specific adapters over one shared base — hot-swappable at serving time.

Why this piece earns its place

An adapter is meaningless without exactly the weights it was trained against. Base models move — a vendor re-uploads a checkpoint under the same tag, someone bumps a minor version — and a LoRA adapter trained on the old numbers now sits on slightly different ones and degrades with nothing in the logs to show for it. Pin the base by content hash and store that hash beside the adapter; it is the single most useful field in the registry entry later. Two notes from actually running these. The saving is in optimizer state and gradients, not the forward pass, since the frozen base has to be resident either way — and that saving is frequently the difference between a run that fits your hardware and one that does not, because full fine-tuning also carries gradients and optimizer state for every weight it updates. What PEFT cannot do is shrink the base itself: if the frozen weights alone overflow the device, the answer is loading them quantized, or splitting them across devices, not a smaller adapter. And at serving time you choose between merging the adapter into the base — one model, no per-request overhead, no swapping — and keeping it separate, which is what buys the hot swap. Reach for full fine-tuning only when the change is broad, a new language or a genuinely different domain, where a low-rank update has too little capacity to carry it.

  • pretrained baseadapt, don’t retrain
  • LoRA / PEFTtrain a small adapter
  • swap adaptersmany tasks, one base

What the new pieces do

Base Modelmodel
The pretrained foundation you adapt. Fine-tuning adjusts it toward your task/domain — cheaper and faster than training from scratch, which almost no one does.

Step 4 · The training loop

Batches, loss, checkpoints, at scale

Now you run the actual training. Data and base feed the trainer; out comes a model. What has to be right about the loop — and what breaks when the model is too big for one GPU?

Raw DataCurate & FormatTraining SetBase ModelTrainer
New in this step: Trainer. · swipe to pan the diagram

What does the training loop need to be reliable and scalable?

  1. Train in batches over epochs, watch the loss (and eval loss) for over/underfitting, checkpoint often so a crash doesn’t lose everything, and shard across GPUs when the model or data won’t fit on one.

  2. A single pass with no checkpointing loses all progress on any failure and gives you no signal to stop early. Long runs must checkpoint and monitor.

  3. Zero training loss is memorization/overfitting — eval loss will be rising. You stop based on held-out eval, not training loss.

The Trainer runs batched forward/backward passes over epochs, monitors training and eval loss (to catch over/underfitting), checkpoints regularly, and shards across GPUs with data/tensor parallelism when the model or batch won’t fit on one device. Hyperparameters (learning rate, epochs, batch size) are logged for reproducibility.

Why this piece earns its place

Checkpoint cadence is a straight trade: writes are large and stall the run, so checkpointing every step burns throughput and checkpointing rarely burns GPU-hours the next time a node dies. Pick the interval from what a crash costs, not from a default. What the checkpoint contains matters more. Resuming from weights alone restarts the optimizer cold and, worse, restarts the data loader at the beginning, so the resumed run silently trains on the first examples twice and the epoch count in your logs becomes a lie. Save optimizer state, scheduler state and the data position together. Reproducibility has a related trap: recorded hyperparameters do not reproduce a run on a different number of GPUs, because the cluster shape changes the effective batch size and therefore the result — so record the topology beside the learning rate. One refinement worth saying out loud in an interview: if you early-stop on held-out data, that data has become a training signal. Split a small dev slice for stopping and checkpoint selection, and leave the gate's set untouched by the loop.

  • batches × epochsmonitored loss
  • checkpointsurvive failures
  • data/tensor parallelscale past one GPU

What the new pieces do

Trainerservice
Runs the training loop on GPUs: forward/backward passes, loss, checkpoints — usually parameter-efficient (LoRA/PEFT) so you train a small adapter instead of all the weights, across data/tensor-parallel workers.

Step 5 · The gate that matters

Evaluate before you promote

Training finished; you have a checkpoint. Is it actually better than what’s in production? Ship it untested and you might deploy a regression. What decides whether a model is allowed out?

TrainerLoRA · distributedEval Gateheld-out
New in this step: Eval Gate.

What should gate a fine-tuned model’s promotion?

  1. Training loss measures fit to the training data, not real quality — it can be low precisely because the model memorized. The gate must use held-out data.

  2. Spot-checking misses regressions and is unrepeatable. You need a consistent eval suite scored against the base and a threshold.

  3. The eval gate runs the model on truly held-out data, compares to the current/base model, and blocks promotion on a regression or a missed bar. Objective, repeatable, and the last line of defense.

The Eval Gate scores the fine-tuned model on the held-out eval set, compares it to the base/current model, and blocks promotion on a regression or a missed threshold. This is the pipeline’s most important stage — and it’s only trustworthy if the eval set was never contaminated by training data (the failure you’ll trigger).

Why this piece earns its place

The hard part of gating generative output is that there is no accuracy column. Scores come from graded rubrics, exact-match on structured tasks, or another model acting as judge — and a judge is itself a version. Pin it, or this month's scores are not comparable with last month's and the bar re-baselines under you with nobody deciding it should. The eval suite deserves the same treatment as the dataset: an immutable artifact with a version id recorded in the run, because a suite people are free to edit during a bad week will eventually pass. Be equally explicit about what the candidate is measured against. Scoring it against the original base is the flattering version, and it flatters a little more with every release, since the incumbent has been improving while the base has not. Score against whatever is serving right now and the bar ratchets: a model that comfortably cleared the gate last year would not clear it today, which is the behaviour you want and also the reason a promotion can get harder for reasons nobody changed. Budget matters too: the gate is a full inference pass over the suite per candidate, so suite size sets how often you can afford to evaluate at all. Run a small subset on intermediate checkpoints, the full suite only on the one you mean to promote.

  • held-outthe only honest test
  • vs baseno regressions
  • thresholdobjective promotion

What the new pieces do

Eval Gateservice
Scores the fine-tuned model on a held-out eval set and against the base model. A regression or a failed threshold blocks promotion — the quality gate before anything ships.

Step 6 · Remember everything

Model registry, lineage, versioning

A model passed the gate. Six weeks later it misbehaves and you need to know exactly what produced it — or roll back. How do you make every model fully traceable and reversible?

Eval Gateheld-outModel Registryversioned
New in this step: Model Registry.

What must you record about every trained model?

  1. Weights alone don’t tell you which base, data, or settings made them — you can’t reproduce, audit, or safely roll back. Lineage is the point.

  2. A model registry stores the artifact with full provenance — which base, which dataset version, which hyperparameters, which eval results — so any model is reproducible, comparable, and rollback-able.

  3. Retraining isn’t reproducible without the recorded inputs, and it’s expensive. The registry makes past models first-class, versioned assets.

Store every promoted model in a Model Registry with full lineage: base model, dataset version, hyperparameters, training run, and eval scores. Now models are versioned, comparable, reproducible, and instantly rollback-able — the audit trail that makes fine-tuning safe to iterate on.

Why this piece earns its place

Two things live here and confusing them causes outages. A version record is immutable — written once, artifact and lineage together, never edited. Which version is serving is a separate, mutable pointer. Rollback is moving that pointer, which takes seconds, and it stays seconds only if nothing downstream needs a rebuild to honour it. Keep them apart and a rollback is boring; hang a mutable latest tag on the artifact itself and two deploys will race to define what production means. Lineage also has to outlive what it points at: a record naming a dataset version whose storage was garbage-collected is provenance in name only, so the retention policy for snapshots, eval suites and base hashes is set by the oldest model you might still be asked about, not by storage cost. Artifacts are large and nothing deletes itself, so decide deliberately — anything ever promoted is kept forever, experimental checkpoints expire on a clock. The audit case is what pays for all of it: when a customer asks why the model told them something, this is the only place the answer exists.

  • artifact + lineagebase · data · params · scores
  • versionedcompare and roll back
  • reproduciblerebuild any model

What the new pieces do

Model Registrystore
Stores model artifacts with lineage: which base, which dataset version, which hyperparameters, which eval scores. Reproducibility and rollback depend on it.

Step 7 · Ship it carefully

Canary deploy, rollback, adapter swap

The gate passed offline — but offline eval never perfectly predicts production. How do you deploy a new model so that if it’s worse in the real world, few users feel it and you recover instantly?

Servingcanary → prodModel RegistryversionedUsersreal traffic
New in this step: Serving, Users.

How do you roll out a fine-tuned model safely?

  1. Serve the new model to a slice of traffic, compare live quality/latency/cost to the incumbent, ramp up if it holds, and roll back instantly if not. For PEFT, hot-swap the adapter without redeploying the base.

  2. A big-bang swap exposes all users to any regression offline eval missed, with no gradual signal. Canary limits blast radius and gives you live evidence.

  3. A model that never ships delivers no value. The answer is gradual, reversible rollout — not avoidance.

Deploy behind a canary: route a small slice of traffic to the new model, compare live quality, latency and cost to the incumbent, then ramp — with instant rollback via the registry. With PEFT you hot-swap the adapter onto the shared base, making deploys and rollbacks cheap. Offline eval starts the decision; production finishes it.

Why this piece earns its place

Two details decide whether a canary tells you anything. Assignment has to be sticky per user, not random per request, or one person's conversation alternates between two models, and every comparison — including the ratings you will later train on — is confounded. And the signals arrive at different speeds. Latency and cost are readable within minutes; quality arrives as thumbs, edits and escalations over hours or days. A canary judged the same afternoon it launched has established only that the new model is not slower, which the gate already implied. Size the window to the slowest signal you actually intend to act on. Capacity is the quiet constraint underneath that: serving two versions means holding two in GPU memory, which for full models roughly doubles the footprint and makes a long canary genuinely expensive. This is where the adapter decision from step 3 pays out — two adapters over one shared base cost almost nothing to run side by side, which is exactly what lets you leave a canary up long enough for the quality signal to land.

  • canarysmall % first
  • rollbackinstant, via registry
  • adapter swapcheap PEFT deploys

What the new pieces do

Servingbackend
Where a promoted model is rolled out — behind a canary, with rollback. For PEFT, adapters can be hot-swapped onto a shared base without redeploying the whole model.
Usersclient
Production users of the deployed model. Their interactions are both the payoff and the richest source of the next round of training data.

Step 8 · Close the loop

The production data flywheel

The model is live. Its best feature is that it’s now generating exactly the data that would make the next version better — real prompts, real outcomes, real corrections. How do you turn production into fuel without poisoning the pipeline?

Curate & FormatTraining SetServingTrainerEval GateModel RegistryFeedback LoopUsers
New in this step: Feedback Loop. · swipe to pan the diagram

How does production improve the next model?

  1. Blindly training on raw output creates feedback loops and drift (the model learns from its own mistakes). Production data must be curated and gated like any other, not fed back raw.

  2. Your own production distribution is the most relevant data you have. Discarding it wastes the flywheel that makes each version better at the actual task.

  3. A feedback loop harvests production signals (thumbs, edits, escalations), curates them through the same cleaning/dedup/held-out discipline, and rolls them into the next training set — a compounding data advantage.

A Feedback Loop collects production interactions, ratings, and corrections, then routes them back through curation into the next Training Set. Each release generates the data for the next — a compounding data flywheel — as long as that data gets the same cleaning, dedup, and held-out discipline as everything else. Never train on raw output unchecked.

Why this piece earns its place

Feedback is only usable if it is attributable. Log the exact prompt, the model version and the adapter that produced each response beside the rating, or six weeks on you hold a pile of corrections and no way to say which version they indict — useless as training data and useless as a regression signal. That instrumentation belongs in the serving path, so it is something you build in step 7 and only discover you needed here. The subtler problem is selection. Production contains only prompts people thought to send this model, and users quietly stop asking for what it visibly fails at, so each turn of the flywheel narrows the distribution toward what the current version already does well. Left alone it converges on a model that is excellent at a shrinking task. Sample deliberately outside it — cold-start prompts, abandoned sessions, questions your competitors get — and keep that share explicit. The same drift eventually reaches the held-out set: one carved from last year's traffic stops representing production, so it gets refreshed on a schedule and scores across suite versions are treated as incomparable.

  • collectratings + corrections
  • curatesame discipline
  • flywheeleach release fuels the next

What the new pieces do

Feedback Loopbus
Collects production interactions, ratings, and corrections, feeding curated examples back into the next dataset — the data flywheel that keeps improving the model.

The payoff

You built a fine-tuning pipeline

From "we have examples of what we want" to a governed model factory: curated, versioned datasets; a pretrained base adapted with LoRA/PEFT; a monitored distributed training loop; a hard held-out eval gate; a registry with full lineage; canary deploys with rollback; and a production flywheel feeding the next round.

Raw DataCurate & FormatTraining SetServingBase ModelTrainerEval GateModel RegistryFeedback LoopUsers
The finished design, end to end. · swipe to pan the diagram

Now leak the eval set into training and watch every offline number light up green while the deployed model is no better than the base — because the metrics measured memorization, not generalization. That’s why the held-out split is sacred, why dedup must catch near-duplicates, and why time-based splits and re-verification protect the one gate that decides what ships.

Everything you assembled, in order

  • Pipeline shape — curate → train → eval gate → register → canary — reproducible, reversible
  • Data curation — clean, dedup, format, hold out — the model imitates exactly what you feed it
  • Base + PEFT — adapt a pretrained base with small LoRA adapters, not from scratch
  • Training loop — batched, checkpointed, distributed; watch eval loss, not train loss
  • Eval gate — held-out scores vs base — the honest test that gates promotion
  • Registry — artifact + lineage → reproducible, versioned, rollback-able
  • Canary — gradual rollout with instant rollback; hot-swap PEFT adapters
  • Flywheel — curated production data fuels the next model; never self-train raw
  • The failure — eval contamination inflates metrics silently — isolate the split

Deep cut · 22:26

Aced the test. Got worse.

The interactive build above lays out the fine-tuning pipeline: chat history curated into training examples, a LoRA adapter trained on a frozen base model, epochs and checkpoints, an eval gate against the live model, a model registry with lineage, a canary rollout and the data flywheel. This film is an autopsy. A cooking app’s AI told a user that cooked rice keeps out overnight, a day after its new version aced the app’s own test at 97 percent, and the old version had always said: into the fridge within two hours. The film builds the pipeline one reasonable fix at a time, with the release function written out line by line as pseudo code, then replays the Thursday card by card to find how a model aced its test and got worse.

  • See the wreck: the test’s questions came from the app’s own chats, and near-copies of them, in other words, sat in the practice pile, so 97 percent was the model remembering answers, not knowing them. The rice myth came from a forum answer with two hundred thumbs up, in forty wordings the letter-by-letter duplicate filter could not match, and the model practised it forty times. With the twins removed and the test split by time, the honest score is 86 percent: lower, but real.
  • Take it into the interview: curate chats into examples, scrubbing names and phone numbers and keeping only answers a cook has checked against a food safety guide; remove near-duplicates by meaning, not letters; lock a held-out eval set split by time, newer than anything the model practised on; train a LoRA adapter on the frozen base and keep the checkpoint that scores best on the eval set; gate every release against the live model and on safety questions; hang each version in a registry with its lineage; try it on a small slice of users first; and feed only checked answers back into the next version.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. Exact-hash dedup misses two examples that are 95% identical wording. How does "near-duplicate" dedup actually catch those?

    Via embedding similarity or minhash/locality-sensitive-hashing over the text — examples get vectorized and any pair above a similarity threshold (not just byte-identical) is flagged as a near-duplicate cluster, keeping one representative. This is the layer that catches an eval example that got lightly paraphrased and dropped back into a training-data scrape, which exact hashing would completely miss.

  2. PEFT freezes most of the base model — does that mean catastrophic forgetting is a solved problem?

    No, just reduced. A LoRA adapter can still push the model’s effective behavior hard enough in a narrow direction that general capability visibly degrades on tasks unrelated to the fine-tuning target, especially with a high learning rate or too many epochs on a narrow dataset. It’s why the eval gate compares against the base model on a BROAD suite, not just the fine-tuning task — a narrow win with a broad regression should still fail the gate.

  3. Why does a time-based split (eval on newer data) beat a random split for this pipeline specifically?

    A random split still lets the model’s training data span the SAME time window as its eval data — if your product or user behavior shifted mid-window, random splitting hides that shift by mixing old and new patterns into both sets. A time-based split (train on older data, eval on newer) forces the eval to mimic the actual future the model will face in production, which is a harder and more honest test than a random shuffle.

  4. The canary shows slightly worse quality than the incumbent, but on only 2% of traffic. Do you roll back?

    Not automatically — small canary slices have real statistical noise, and a premature rollback on noise means you never learn anything from canaries. The right move is checking whether the gap exceeds what’s explainable by sample size (a significance check, not a gut call) before ramping further or rolling back; canary decisions need the same rigor as the offline eval gate, just measured live.

  5. Production wants to serve five different customer-specific adapters off one shared base model. What changes?

    Serving now needs per-request adapter routing (which customer’s adapter loads for this call) instead of one fixed model, and the registry has to track lineage per adapter independently — each one has its own dataset version, eval scores, and promotion history, even though they share a base. The canary and rollback story also has to be per-adapter: one customer’s bad fine-tune shouldn’t block or affect another’s.

Check yourself — the answers, and why

Nine steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. You reach for fine-tuning (over prompting/RAG) when…

    RAG/prompting inject knowledge and instructions at inference; fine-tuning changes weights to bake in behavior. (Fresh facts are a RAG job.)

  2. The biggest lever on fine-tuning quality is usually…

    A fine-tuned model imitates its data exactly; curation (not the training loop) is where most quality is won or lost.

  3. LoRA / PEFT is the default because it…

    Parameter-efficient fine-tuning captures the task with a tiny adapter, preserving the base’s general capability and enabling per-task adapter swaps.

  4. A fine-tuned model should be promoted based on…

    Training loss can be low from memorization; the eval gate uses truly held-out data compared to the base to decide promotion.

  5. The model registry exists to…

    Weights alone aren’t reproducible or auditable; the registry records provenance so any model can be rebuilt, compared, or rolled back.

  6. Eval-set contamination is dangerous because it…

    If eval examples leak into training, the gate approves a model that just read the test; strict dedup and time-based splits keep the eval set honest.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Curate a dataset: clean, deduplicate and filter raw logs and docs for correctness, format them into consistent instruction→response pairs, and carve off a held-out eval set that never touches training.
  • Adapt a base model: start from a pretrained base and train a small LoRA/PEFT adapter rather than pretraining from scratch or updating every weight.
  • Gate on evaluation: score the fine-tuned model on the held-out set versus the base and block promotion on any regression or missed threshold.
  • Register with lineage: store every promoted artifact with its base, dataset version, hyperparameters and eval scores so any model is reproducible and rollback-able.
  • Deploy and learn: roll out behind a canary with instant rollback, then feed curated production interactions and corrections back into the next training set.

The qualities that shape everything

Each one names the mechanism that buys it.

Reproducible, auditable model changes
A chain of gated stages pins a versioned dataset, gates promotion on eval, and records full lineage — so you always know what a model learned from and can roll back to any prior version.
Quality not capped by noisy data
Curation cleans, deduplicates (including near-duplicates), filters for correctness and formats consistent examples — the model imitates its data exactly, so this is where quality is won or lost.
Specialize cheaply, without catastrophic forgetting
Parameter-efficient fine-tuning freezes the pretrained base and trains a small LoRA adapter — far cheaper and faster, and adapters hot-swap per task over one shared base.
Survive long, expensive training runs
The trainer checkpoints regularly, monitors eval loss (not train loss) to catch overfitting, and shards across GPUs with data/tensor parallelism when the model won’t fit on one.
Never promote a regression
A held-out eval gate scores the candidate against the base and blocks promotion on a regression or missed threshold — trustworthy only if the eval set was never contaminated by training data.
Improve each release without poisoning the model
A feedback loop routes production ratings and corrections back through the same curation, dedup and held-out discipline before they reach the next dataset — never raw self-training.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

Curated, deduped examples over dumping in every log and document

A fine-tuned model imitates its data exactly, so clean, diverse, correct examples decide the outcome. Raw dumps carry duplicates, junk, contradictions and PII the model faithfully learns — volume without curation usually makes it worse, not better.

LoRA/PEFT adapters on a frozen base over full fine-tuning of every weight

Training a tiny low-rank adapter over a frozen base is cheaper, faster and resists catastrophic forgetting, and adapters swap per task. Updating every weight is compute- and memory-heavy and can overwrite general capability.

A held-out eval gate vs the base over a low final training loss

Scoring on truly held-out data against the base catches regressions objectively and repeatably. Training loss measures fit to the training data — it can be low precisely because the model memorized.

Canary a small slice, then ramp over a big-bang swap for everyone

A gradual, reversible rollout limits blast radius and gives live evidence offline eval can’t model. Replacing the production model for everyone at once exposes all users to any regression the gate missed.

Curate production data before reuse over auto-training on all production traffic

Routing ratings and corrections through the same cleaning, dedup and held-out discipline builds a compounding advantage. Blindly training on raw output creates feedback loops and drift — the model learns from its own mistakes.

The answer, out loud

What a strong answer to “Design a Fine-Tuning Pipeline” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    What fine-tuning is actually for

    I’d start by narrowing the ask, because fine-tuning is the wrong answer to a lot of questions. Missing facts are a retrieval problem. A one-off behaviour is a prompt. It earns its place when the behaviour has to live in the weights, so the model answers in our shape and our domain by default rather than because somebody re-explained it in the request. Then the job is turning examples into a model we can show is better than what we serve today.

  2. 3–8 min

    The shape I’d draw first

    So: five boxes in a line. Curate, train, gate, register, deploy. Any one can be taken out and you’ll still get a model at the other end, which is exactly why they get taken out. Drop the gate and production is where we find out. Drop the registry and we can’t say which run produced what’s live, so we can’t undo it either. And the whole thing is offline, on a cadence — production feeds it, nothing here edits the serving model in flight.

    Built in step 1: Data → train → eval → register → deploy
  3. 8–13 min

    Most of the work is the dataset

    Then I’d slow right down on the data, because that’s where the result is decided. Whatever pattern sits in the examples is the pattern we’re buying, including ones nobody meant to teach, like a support agent’s habit of hedging. So: keep the answers that are actually correct, collapse near-copies so the model doesn’t rehearse one opinion over and over, one prompt-and-response shape throughout, personal data out. Two artifacts come out of it — a versioned training set, and an eval set carved off here and never trained on, split by time so we’re tested on newer questions than the model ever saw.

    Built in step 2: Dataset curation is the real work
  4. 13–18 min

    Adapt a base, don’t build one

    Nobody’s pretraining from scratch here — neither the data nor the budget is on the table, and a pretrained base already has the general ability we want to keep. The question is how much of it to disturb. Updating every weight works, and it’s the honest choice for a genuinely broad change, but it’s heavy and it can crowd out that general ability to make room for ours. My default is LoRA: freeze the base, train a small adapter over it. Cheaper, quicker to iterate on, and several task adapters can share one copy of the base.

    Built in step 3: Base model + parameter-efficient fine-tuning
  5. 18–23 min

    The run itself

    The run is batches, a few passes over the data, measuring the error and nudging the weights. It checkpoints often: a node dying deep into a long run shouldn’t cost us the start of it. And it watches loss on data it isn’t training on — if the training curve keeps falling while that one turns up, it’s memorising and we stop. I’d be specific about which data: a small dev slice kept for stopping and picking a checkpoint, not the gate’s set. Stop on the gate’s set and the gate is scoring a model tuned against it.

    Built in step 4: Batches, loss, checkpoints, at scale
  6. 23–29 min

    What decides that it ships

    Then the gate, the stage I’d defend hardest. A training number can’t make this call — it only says the model fits what it studied. Reading a few outputs can’t either; two people read differently and neither reads the same way twice. So: run the candidate over the held-out suite, score it beside the model we’re already serving, and make it win before it moves. I’d keep that suite wider than the task we fine-tuned for, because the real risk in teaching a model one job is what it quietly gives up on the others.

    Built in step 5: Evaluate before you promote
  7. 29–33 min

    Write down how it was made

    Anything that passes goes into a registry, with the provenance beside it rather than the weights alone: which base, which snapshot of the data, which settings, what it scored. Weights on their own are an artifact nobody can explain. My test is a customer complaining about an answer months from now — can we say which version said it, and what it was taught? If not, we can’t fix it and we can’t safely reverse it.

    Built in step 6: Model registry, lineage, versioning
  8. 33–37 min

    Ship it to a few people first

    Offline scores are evidence, not proof; the suite is our guess at what users will ask. So the new model meets a slice of real traffic first, beside the incumbent, and I’d want live quality read alongside latency and cost before anyone else is ramped onto it. If it’s worse, serving points back at the previous version, which is quick precisely because the registry is already holding it. Adapters make that cheap — same base in memory, swap the small piece on top.

    Built in step 7: Canary deploy, rollback, adapter swap
  9. 37–41 min

    Then production writes the next dataset

    Once it’s live it produces what we needed at the beginning: real questions, and people telling us when an answer was wrong. I’d feed that back, but through the front door — filtered, deduped and versioned like any other source, with the eval set out of its reach. Piping production output straight into training is the tempting shortcut, and what compounds then is the errors rather than the quality. It’s also a standing job: every release changes what the next dataset looks like, so curation is staffing, not a one-off cost.

    Built in step 8: The production data flywheel
  10. 41–45 min

    Where it breaks, and the trade-off

    What I’d watch: both loss curves during a run, the gate’s margin over the incumbent, and on canary, quality next to latency and cost. And I’d name the failure that worries me, because it never announces itself: if questions from the eval suite were seen in training, every number goes up and the deployed model doesn’t. Green dashboards, same answers. The trade-off underneath is speed against certainty — every snapshot, gate and canary window slows how fast we can change the model and buys confidence the change is real. Given more time I’d settle what triggers the next run at all: enough new curated examples, or a measured fall in production quality, rather than the calendar. And who can say yes when the gate says no, since that conversation happens under pressure and should be on paper first.

What this teaches

Learn AI system design by building an LLM fine-tuning and training pipeline step by step. An interactive guide to the offline path — dataset curation, base-model choice and parameter-efficient fine-tuning (LoRA/PEFT), the distributed training loop, a held-out evaluation gate, a versioned model registry, canary deployment, and the production data flywheel — plus the silent trap of eval-set contamination.

Key takeaways

  • Pipeline shape — curate → train → eval gate → register → canary — reproducible, reversible
  • Data curation — clean, dedup, format, hold out — the model imitates exactly what you feed it
  • Base + PEFT — adapt a pretrained base with small LoRA adapters, not from scratch
  • Training loop — batched, checkpointed, distributed; watch eval loss, not train loss
  • Eval gate — held-out scores vs base — the honest test that gates promotion
  • Registry — artifact + lineage → reproducible, versioned, rollback-able
  • Canary — gradual rollout with instant rollback; hot-swap PEFT adapters
  • Flywheel — curated production data fuels the next model; never self-train raw
  • The failure — eval contamination inflates metrics silently — isolate the split

Concepts covered

  • When prompting isn’t enough, fine-tune
  • Data → train → eval → register → deploy
  • Dataset curation is the real work
  • Base model + parameter-efficient fine-tuning
  • Batches, loss, checkpoints, at scale
  • Evaluate before you promote
  • Model registry, lineage, versioning
  • Canary deploy, rollback, adapter swap
  • The production data flywheel
RUN IT YOURSELF

The eval score your split already knew the answer to

One fine-tuning corpus, cut into train and eval two ways — by row, and by a dedupe key computed from normalized content — then both gates scored against questions that were never logged at all. It prints what each gate claims and what the same model actually does. Change NEAR_DUP_RATE, hit Run, and watch one of the two numbers walk away from the truth.

HOW TO READ THE CODE — 5 IDEAS
  1. Curation, not the training loop, is where this is won or lost (step 2): the copies here are rewordings, so exact-string dedup finds 0 of them.
  2. The row split lets a question sit on both sides of the line. Its gate reads 0.810; the same model scores 0.582 on questions it never saw.
  3. The dedupe-key split moves whole groups of rewordings together and reads 0.604 against a real 0.602 — lower, and true. That is the only gate step 5 can trust.
  4. One model, one training run, its own eval set cut by what it had read: 0.949 on leaked rows, 0.589 on clean ones. Nothing else differs between them.
  5. At NEAR_DUP_RATE = 0 both splits agree, so the gap is bought by duplicates and not by the method — the page’s leak the eval set failure, as a number.
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, leak the eval set, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs