The whole design, in writing
Learn AI system design by building an LLM fine-tuning and training pipeline step by step. An interactive guide to the offline path — dataset curation, base-model choice and parameter-efficient fine-tuning (LoRA/PEFT), the distributed training loop, a held-out evaluation gate, a versioned model registry, canary deployment, and the production data flywheel — plus the silent trap of eval-set contamination.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
When prompting isn’t enough, fine-tune
Prompting and RAG get you far: they steer a model with instructions and context at inference time. But sometimes you need the model to internalize a style, a format, a domain, or a skill — reliably, without a giant prompt every call. That means changing the weights. How do you turn "we have examples of the behavior we want" into a better model, safely and repeatably?
Build a fine-tuning pipeline: an offline, repeatable path from curated data through training and a hard evaluation gate to a versioned, deployable model — with a feedback loop that turns production into the next dataset. It’s an ML system, so it lives or dies on data quality and honest evaluation.
Step 1 · The skeleton
Data → train → eval → register → deploy
Fine-tuning isn’t "run a training script." It’s a pipeline with stages that must be reproducible and gated. What’s the minimal shape that keeps a model change safe?
What’s the right shape for a fine-tuning pipeline?
No curation, no eval gate, no versioning — you can’t tell if the model improved or regressed, or reproduce it. That’s how a worse model reaches production.
Each stage is reproducible and traceable: a versioned dataset trains a model, an eval gate decides promotion, the registry records lineage, and deployment is gradual and reversible.
Online training on raw traffic is unstable and unsafe — no gate, no reproducibility, easy to poison. Training is an offline, gated pipeline; production only feeds it.
The pipeline is a chain of reproducible, gated stages: curate a versioned dataset → train → pass an evaluation gate → register the artifact with full lineage → deploy behind a canary. Every arrow is traceable, so you always know what a model learned from and can roll back to any prior version.
Why this piece earns its place
The stages matter less than the boundaries between them. Each boundary is a handoff of an artifact — a dataset version, a checkpoint, a score report — and if those artifacts are not addressable by a stable id, the pipeline can only ever be re-run whole. That is the expensive mistake, because training dominates what this whole page costs to operate, and re-running it to recover from something that broke downstream of it buys nothing at all. A gate that flakes, a failed registry write, a typo in a canary config: each should re-enter at its own stage against the checkpoint that already exists, not at the top. Make every stage take its inputs as explicit parameters — base id, dataset version, hyperparameters — instead of reading them from a file in whoever's branch triggered the run, and that re-entry is free. The shape buys an organisational contract too. Curation, training and serving usually sit with different people, and the artifact at each boundary is the thing they agree on rather than a meeting. Skip all of it for a one-off experiment; build it the moment you intend a second model, because the second one is where comparisons start to matter.
Step 2 · Garbage in, garbage model
Dataset curation is the real work
The instinct is "more data = better model." But a fine-tuned model mimics its training data exactly — including its noise, duplicates, and mistakes. What actually makes a good fine-tuning dataset?
How do you build a fine-tuning dataset that helps?
Raw dumps carry duplicates, junk, contradictions and PII; the model faithfully learns all of it. Volume without curation usually makes the model worse, not better.
Too few examples underfit the target behavior. The goal isn’t minimal or maximal — it’s clean, diverse, correct examples that represent the task.
Curation dominates outcomes: remove junk and near-duplicates, filter for correctness, format into consistent instruction→response pairs, and carve off a strictly separate held-out eval set.
Curate: clean and filter for correctness, deduplicate (including near-duplicates), format into consistent examples, strip PII, and — critically — split off a held-out eval set that never touches training. Version the result. Quality and diversity beat raw volume; the model will imitate exactly what you feed it.
Why this piece earns its place
Versioning a dataset is not the same as being able to rebuild it. A version defined as a query over production logs stops being reproducible the moment retention deletes the rows underneath it, and a deletion request removes them sooner than that. So curation materialises a snapshot — the actual examples, written once as an immutable artifact — and every run pins the snapshot, never the query. The uncomfortable corollary: a user asking to be forgotten is only half handled when the rows go, because their examples also sit inside the weights of every model trained on that snapshot, and the only real remedy there is retraining without them. Decide that policy while the registry still holds a handful of models, not once it holds a fleet of them. This stage quietly owns one more thing — the prompt template. Examples are formatted here, and the serving path has to format live requests identically: same system text, same delimiters, same role names. Drift by a newline and the model meets a distribution at inference it never trained on, degrading for reasons no metric explains. Treat that template as a shared versioned artifact, not a string that exists in two codebases.
- clean + dedupquality over volume
- consistent formatinstruction → response
- held-out setnever seen in training
What the new pieces do
- Raw Datastore
- The raw material: production transcripts, labeled examples, documents, human demonstrations. Fine-tuning quality is capped by this — garbage in, garbage model.
- Curate & Formatbus
- Cleans, deduplicates, filters and formats raw data into training examples (e.g. instruction→response pairs). Also splits off a held-out eval set — and must keep it strictly separate.
- Training Setstore
- The finished, versioned training data. Every training run pins a dataset version so results are reproducible and traceable to exactly what the model learned from.
Back of the envelope
- dedup near-duplicates
- not just exact — fuzzy/semantic too
- time-based split
- eval on newer data to mimic the future
- balance the mix
- diversity prevents narrow overfitting
- version datasets
- pin exactly what a run trained on
Step 3 · Adapt, don’t retrain
Base model + parameter-efficient fine-tuning
You have data. Do you retrain a model from scratch? Update all of a pretrained model’s billions of weights? Both are expensive and risky. What’s the practical way to specialize a model?
How do most teams actually fine-tune a large model?
Pretraining costs millions of dollars and needs internet-scale data. Almost no one does it; you adapt an existing pretrained base instead.
Parameter-efficient fine-tuning trains a tiny set of new parameters (low-rank adapters) on top of a frozen base — far cheaper, faster, less prone to catastrophic forgetting, and adapters can be swapped per task.
Full fine-tuning works but is compute- and memory-heavy and can overwrite general capability (catastrophic forgetting). PEFT gets most of the benefit for a fraction of the cost.
Start from a pretrained base model and, in most cases, use parameter-efficient fine-tuning (LoRA/PEFT): freeze the base and train small adapters. It’s cheaper, faster, resists catastrophic forgetting, and lets you keep many task-specific adapters over one shared base — hot-swappable at serving time.
Why this piece earns its place
An adapter is meaningless without exactly the weights it was trained against. Base models move — a vendor re-uploads a checkpoint under the same tag, someone bumps a minor version — and a LoRA adapter trained on the old numbers now sits on slightly different ones and degrades with nothing in the logs to show for it. Pin the base by content hash and store that hash beside the adapter; it is the single most useful field in the registry entry later. Two notes from actually running these. The saving is in optimizer state and gradients, not the forward pass, since the frozen base has to be resident either way — and that saving is frequently the difference between a run that fits your hardware and one that does not, because full fine-tuning also carries gradients and optimizer state for every weight it updates. What PEFT cannot do is shrink the base itself: if the frozen weights alone overflow the device, the answer is loading them quantized, or splitting them across devices, not a smaller adapter. And at serving time you choose between merging the adapter into the base — one model, no per-request overhead, no swapping — and keeping it separate, which is what buys the hot swap. Reach for full fine-tuning only when the change is broad, a new language or a genuinely different domain, where a low-rank update has too little capacity to carry it.
- pretrained baseadapt, don’t retrain
- LoRA / PEFTtrain a small adapter
- swap adaptersmany tasks, one base
What the new pieces do
- Base Modelmodel
- The pretrained foundation you adapt. Fine-tuning adjusts it toward your task/domain — cheaper and faster than training from scratch, which almost no one does.
Step 4 · The training loop
Batches, loss, checkpoints, at scale
Now you run the actual training. Data and base feed the trainer; out comes a model. What has to be right about the loop — and what breaks when the model is too big for one GPU?
What does the training loop need to be reliable and scalable?
Train in batches over epochs, watch the loss (and eval loss) for over/underfitting, checkpoint often so a crash doesn’t lose everything, and shard across GPUs when the model or data won’t fit on one.
A single pass with no checkpointing loses all progress on any failure and gives you no signal to stop early. Long runs must checkpoint and monitor.
Zero training loss is memorization/overfitting — eval loss will be rising. You stop based on held-out eval, not training loss.
The Trainer runs batched forward/backward passes over epochs, monitors training and eval loss (to catch over/underfitting), checkpoints regularly, and shards across GPUs with data/tensor parallelism when the model or batch won’t fit on one device. Hyperparameters (learning rate, epochs, batch size) are logged for reproducibility.
Why this piece earns its place
Checkpoint cadence is a straight trade: writes are large and stall the run, so checkpointing every step burns throughput and checkpointing rarely burns GPU-hours the next time a node dies. Pick the interval from what a crash costs, not from a default. What the checkpoint contains matters more. Resuming from weights alone restarts the optimizer cold and, worse, restarts the data loader at the beginning, so the resumed run silently trains on the first examples twice and the epoch count in your logs becomes a lie. Save optimizer state, scheduler state and the data position together. Reproducibility has a related trap: recorded hyperparameters do not reproduce a run on a different number of GPUs, because the cluster shape changes the effective batch size and therefore the result — so record the topology beside the learning rate. One refinement worth saying out loud in an interview: if you early-stop on held-out data, that data has become a training signal. Split a small dev slice for stopping and checkpoint selection, and leave the gate's set untouched by the loop.
- batches × epochsmonitored loss
- checkpointsurvive failures
- data/tensor parallelscale past one GPU
What the new pieces do
- Trainerservice
- Runs the training loop on GPUs: forward/backward passes, loss, checkpoints — usually parameter-efficient (LoRA/PEFT) so you train a small adapter instead of all the weights, across data/tensor-parallel workers.
Step 5 · The gate that matters
Evaluate before you promote
Training finished; you have a checkpoint. Is it actually better than what’s in production? Ship it untested and you might deploy a regression. What decides whether a model is allowed out?
What should gate a fine-tuned model’s promotion?
Training loss measures fit to the training data, not real quality — it can be low precisely because the model memorized. The gate must use held-out data.
Spot-checking misses regressions and is unrepeatable. You need a consistent eval suite scored against the base and a threshold.
The eval gate runs the model on truly held-out data, compares to the current/base model, and blocks promotion on a regression or a missed bar. Objective, repeatable, and the last line of defense.
The Eval Gate scores the fine-tuned model on the held-out eval set, compares it to the base/current model, and blocks promotion on a regression or a missed threshold. This is the pipeline’s most important stage — and it’s only trustworthy if the eval set was never contaminated by training data (the failure you’ll trigger).
Why this piece earns its place
The hard part of gating generative output is that there is no accuracy column. Scores come from graded rubrics, exact-match on structured tasks, or another model acting as judge — and a judge is itself a version. Pin it, or this month's scores are not comparable with last month's and the bar re-baselines under you with nobody deciding it should. The eval suite deserves the same treatment as the dataset: an immutable artifact with a version id recorded in the run, because a suite people are free to edit during a bad week will eventually pass. Be equally explicit about what the candidate is measured against. Scoring it against the original base is the flattering version, and it flatters a little more with every release, since the incumbent has been improving while the base has not. Score against whatever is serving right now and the bar ratchets: a model that comfortably cleared the gate last year would not clear it today, which is the behaviour you want and also the reason a promotion can get harder for reasons nobody changed. Budget matters too: the gate is a full inference pass over the suite per candidate, so suite size sets how often you can afford to evaluate at all. Run a small subset on intermediate checkpoints, the full suite only on the one you mean to promote.
- held-outthe only honest test
- vs baseno regressions
- thresholdobjective promotion
What the new pieces do
- Eval Gateservice
- Scores the fine-tuned model on a held-out eval set and against the base model. A regression or a failed threshold blocks promotion — the quality gate before anything ships.
Step 6 · Remember everything
Model registry, lineage, versioning
A model passed the gate. Six weeks later it misbehaves and you need to know exactly what produced it — or roll back. How do you make every model fully traceable and reversible?
What must you record about every trained model?
Weights alone don’t tell you which base, data, or settings made them — you can’t reproduce, audit, or safely roll back. Lineage is the point.
A model registry stores the artifact with full provenance — which base, which dataset version, which hyperparameters, which eval results — so any model is reproducible, comparable, and rollback-able.
Retraining isn’t reproducible without the recorded inputs, and it’s expensive. The registry makes past models first-class, versioned assets.
Store every promoted model in a Model Registry with full lineage: base model, dataset version, hyperparameters, training run, and eval scores. Now models are versioned, comparable, reproducible, and instantly rollback-able — the audit trail that makes fine-tuning safe to iterate on.
Why this piece earns its place
Two things live here and confusing them causes outages. A version record is immutable — written once, artifact and lineage together, never edited. Which version is serving is a separate, mutable pointer. Rollback is moving that pointer, which takes seconds, and it stays seconds only if nothing downstream needs a rebuild to honour it. Keep them apart and a rollback is boring; hang a mutable latest tag on the artifact itself and two deploys will race to define what production means. Lineage also has to outlive what it points at: a record naming a dataset version whose storage was garbage-collected is provenance in name only, so the retention policy for snapshots, eval suites and base hashes is set by the oldest model you might still be asked about, not by storage cost. Artifacts are large and nothing deletes itself, so decide deliberately — anything ever promoted is kept forever, experimental checkpoints expire on a clock. The audit case is what pays for all of it: when a customer asks why the model told them something, this is the only place the answer exists.
- artifact + lineagebase · data · params · scores
- versionedcompare and roll back
- reproduciblerebuild any model
What the new pieces do
- Model Registrystore
- Stores model artifacts with lineage: which base, which dataset version, which hyperparameters, which eval scores. Reproducibility and rollback depend on it.
Step 7 · Ship it carefully
Canary deploy, rollback, adapter swap
The gate passed offline — but offline eval never perfectly predicts production. How do you deploy a new model so that if it’s worse in the real world, few users feel it and you recover instantly?
How do you roll out a fine-tuned model safely?
Serve the new model to a slice of traffic, compare live quality/latency/cost to the incumbent, ramp up if it holds, and roll back instantly if not. For PEFT, hot-swap the adapter without redeploying the base.
A big-bang swap exposes all users to any regression offline eval missed, with no gradual signal. Canary limits blast radius and gives you live evidence.
A model that never ships delivers no value. The answer is gradual, reversible rollout — not avoidance.
Deploy behind a canary: route a small slice of traffic to the new model, compare live quality, latency and cost to the incumbent, then ramp — with instant rollback via the registry. With PEFT you hot-swap the adapter onto the shared base, making deploys and rollbacks cheap. Offline eval starts the decision; production finishes it.
Why this piece earns its place
Two details decide whether a canary tells you anything. Assignment has to be sticky per user, not random per request, or one person's conversation alternates between two models, and every comparison — including the ratings you will later train on — is confounded. And the signals arrive at different speeds. Latency and cost are readable within minutes; quality arrives as thumbs, edits and escalations over hours or days. A canary judged the same afternoon it launched has established only that the new model is not slower, which the gate already implied. Size the window to the slowest signal you actually intend to act on. Capacity is the quiet constraint underneath that: serving two versions means holding two in GPU memory, which for full models roughly doubles the footprint and makes a long canary genuinely expensive. This is where the adapter decision from step 3 pays out — two adapters over one shared base cost almost nothing to run side by side, which is exactly what lets you leave a canary up long enough for the quality signal to land.
- canarysmall % first
- rollbackinstant, via registry
- adapter swapcheap PEFT deploys
What the new pieces do
- Servingbackend
- Where a promoted model is rolled out — behind a canary, with rollback. For PEFT, adapters can be hot-swapped onto a shared base without redeploying the whole model.
- Usersclient
- Production users of the deployed model. Their interactions are both the payoff and the richest source of the next round of training data.
Step 8 · Close the loop
The production data flywheel
The model is live. Its best feature is that it’s now generating exactly the data that would make the next version better — real prompts, real outcomes, real corrections. How do you turn production into fuel without poisoning the pipeline?
How does production improve the next model?
Blindly training on raw output creates feedback loops and drift (the model learns from its own mistakes). Production data must be curated and gated like any other, not fed back raw.
Your own production distribution is the most relevant data you have. Discarding it wastes the flywheel that makes each version better at the actual task.
A feedback loop harvests production signals (thumbs, edits, escalations), curates them through the same cleaning/dedup/held-out discipline, and rolls them into the next training set — a compounding data advantage.
A Feedback Loop collects production interactions, ratings, and corrections, then routes them back through curation into the next Training Set. Each release generates the data for the next — a compounding data flywheel — as long as that data gets the same cleaning, dedup, and held-out discipline as everything else. Never train on raw output unchecked.
Why this piece earns its place
Feedback is only usable if it is attributable. Log the exact prompt, the model version and the adapter that produced each response beside the rating, or six weeks on you hold a pile of corrections and no way to say which version they indict — useless as training data and useless as a regression signal. That instrumentation belongs in the serving path, so it is something you build in step 7 and only discover you needed here. The subtler problem is selection. Production contains only prompts people thought to send this model, and users quietly stop asking for what it visibly fails at, so each turn of the flywheel narrows the distribution toward what the current version already does well. Left alone it converges on a model that is excellent at a shrinking task. Sample deliberately outside it — cold-start prompts, abandoned sessions, questions your competitors get — and keep that share explicit. The same drift eventually reaches the held-out set: one carved from last year's traffic stops representing production, so it gets refreshed on a schedule and scores across suite versions are treated as incomparable.
- collectratings + corrections
- curatesame discipline
- flywheeleach release fuels the next
What the new pieces do
- Feedback Loopbus
- Collects production interactions, ratings, and corrections, feeding curated examples back into the next dataset — the data flywheel that keeps improving the model.
The payoff
You built a fine-tuning pipeline
From "we have examples of what we want" to a governed model factory: curated, versioned datasets; a pretrained base adapted with LoRA/PEFT; a monitored distributed training loop; a hard held-out eval gate; a registry with full lineage; canary deploys with rollback; and a production flywheel feeding the next round.
Now leak the eval set into training and watch every offline number light up green while the deployed model is no better than the base — because the metrics measured memorization, not generalization. That’s why the held-out split is sacred, why dedup must catch near-duplicates, and why time-based splits and re-verification protect the one gate that decides what ships.
Everything you assembled, in order
- Pipeline shape — curate → train → eval gate → register → canary — reproducible, reversible
- Data curation — clean, dedup, format, hold out — the model imitates exactly what you feed it
- Base + PEFT — adapt a pretrained base with small LoRA adapters, not from scratch
- Training loop — batched, checkpointed, distributed; watch eval loss, not train loss
- Eval gate — held-out scores vs base — the honest test that gates promotion
- Registry — artifact + lineage → reproducible, versioned, rollback-able
- Canary — gradual rollout with instant rollback; hot-swap PEFT adapters
- Flywheel — curated production data fuels the next model; never self-train raw
- The failure — eval contamination inflates metrics silently — isolate the split
