The whole design, in writing
Learn AI system design by building a feature store for ML step by step. An interactive guide to the system that kills training/serving skew — a shared feature registry, an offline store for point-in-time-correct training data, a low-latency online store for serving, batch and streaming materialization, and drift monitoring — so a model sees the same features in production that it trained on.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
Why models need a feature store
A model is only as good as the features it’s fed. In training, a data scientist computes features in a notebook over historical data. In production, an engineer recomputes "the same" features in low-latency serving code. These are two different code paths — and the moment they diverge, the model sees inputs it never trained on and quietly gets worse. Every ML team reinvents feature plumbing, and every one hits this. How do you make features consistent, reusable, and fresh?
A feature store is the shared system for defining, computing, storing, and serving features. One definition produces both the historical values that train the model and the current values that serve it — so training and serving can’t drift apart, and features become reusable assets across models.
Step 1 · The core problem
Training/serving skew, and reuse
Two teams compute "user’s 30-day average spend": one in a training notebook, one in serving code. What goes wrong, and what does the right structure prevent?
What’s the fundamental thing a feature store must guarantee?
Latency matters for serving, but the primary guarantee is consistency: the value the model trains on must equal the value it serves on. Fast-but-inconsistent is the failure.
A single definition, materialized to both an offline (training) and online (serving) store, means the model can’t receive a differently-computed feature in production than it learned from — no skew — and features become reusable across models.
Independent implementations are exactly what causes skew and duplicated work. The store exists to make one definition the source of truth.
The feature store’s job is consistency and reuse: define a feature once, and serve it identically to training (historical values) and serving (current values). That single-definition rule eliminates training/serving skew — the silent killer — and turns features into shared assets many models can reuse.
Why this piece earns its place
The thing that decides whether this actually works is where you draw the line around a feature. The store guarantees consistency for whatever sits inside a definition; anything a model does to a value afterwards — scaling it against training-set statistics, bucketing it, filling a missing value with a default — sits outside that guarantee, and the skew you just eliminated quietly reappears in the model’s own preprocessing code. Either pull those transforms in as registered features, or ship their constants inside the model artifact and version them with it. The other half is the key. A definition is only shared if the thing it is keyed by is shared, and entity keys are where that fails without anyone noticing: one team keys a user by account id, another by a hashed login, and both read a feature of the same name and get different rows. Agreeing that namespace is unglamorous prerequisite work this design does not survive without. Grain is the same decision one level up — a feature about a user costs a row per user, the same feature about a user and a merchant costs a row per pair, which is where an online store stops being cheap. Settle grain at registration, because changing it later is a different feature under a new name, not an edit.
Step 2 · Compute the features
The pipeline + raw data
Features come from raw events and tables that a model can’t use directly. Something has to transform "a stream of orders" into "30-day average order value per user." Where does that logic live?
Where should feature transformation logic live?
Embedding transforms in training code alone means serving must reimplement them — the origin of skew. The logic has to be shared, not per-consumer.
Then training can’t reproduce it and other models can’t reuse it. Feature logic belongs in a shared pipeline, defined once.
A single pipeline reads raw events and applies the registered transformation to produce feature values — the one place the logic lives, feeding both stores.
A shared Feature Pipeline reads Raw Events and applies transformations to produce feature values. This is the one place feature logic lives, so both stores (and every model) get identically-computed values. The pipeline is the engine; the registry (soon) holds the definitions it runs.
Why this piece earns its place
One pipeline only produces one answer if its inputs can be read the same way twice. Most feature logic reaches past the event stream into dimension tables — a user’s plan tier, a merchant’s category — and those tables are usually updated in place. Recompute last Tuesday’s value today and you get a number last Tuesday could not have produced, so the pipeline is not reproducible even though nobody touched the definition. Anything feeding a feature wants to be append-only or snapshotted, and that is a constraint you have to push upstream onto teams who do not report to you. The other thing this box owes you is a backfill mode, which is usually the one nobody builds. Steady state handles the last window; registering a new feature, or correcting a definition, means recomputing years of it — each value as steady state would have produced it at that moment, same code, different clock. Write the backfill as its own script and you have rebuilt the two-implementations problem inside the platform, one layer down. It is also where the bill lands, because it reads the whole history of the raw events, so no model can train on more history of a feature than those raw events still cover.
- shared pipelineone transformation path
- raw → featuresfrom events/tables
- feeds both storesoffline + online
What the new pieces do
- Raw Eventsstore
- The raw inputs features are computed from — event streams, transactional tables, clickstreams. Raw and unusable by a model until transformed into features.
- Feature Pipelinebus
- Computes features from raw data using the definitions in the registry — the single transformation logic that produces values for both the offline and online stores.
Step 3 · History for training
The offline store + point-in-time joins
To train, you need each example’s features as they were at that moment — not today’s values. Join naively and you leak the future into the past. How do you build a correct training set?
How do you assemble features for a historical training set correctly?
The offline store keeps feature values over time; a point-in-time join attaches, to each training example, only the feature values that existed at that example’s moment — no leakage from the future.
Using today’s values for a year-old example leaks future information the model won’t have at serving — inflated offline metrics, worse production performance. Classic label/feature leakage.
Recomputing per run is slow, non-reproducible, and easy to get subtly wrong. The offline store materializes historical values once, correctly, for reuse.
The Offline Store holds historical feature values over time; Training builds datasets with point-in-time-correct joins — each example gets only the feature values known at its own timestamp. That prevents feature leakage (using future information) and makes training sets reproducible. It’s throughput-optimized (a warehouse), not latency-optimized.
Why this piece earns its place
The detail that decides whether this is correct is which timestamp you join at. It is not when the row landed in the warehouse and it is not when the outcome became known — it is the moment the model would have been called. Join on ingestion time and you hand training the values that existed only because your own pipeline had already caught up, so the model learns from a world a few minutes ahead of the one it serves in. Nothing flags that; the eval is simply better than production, and every explanation you reach for will be the wrong one. The join itself is ordinary work done carefully: partition both sides on the entity key, sort each partition by time, and walk them together taking the last value at or before each label. The cost sits in that sort, not in the comparison — what explodes is the naive version that pairs every label with every value and filters afterwards. Keep the joined result as a dated artifact beside the model, because the join is deterministic while the data underneath it is not: a retention rule written to save storage will quietly make every model older than it unrebuildable, and nobody who set that rule was thinking about models.
- historicalvalues over time
- point-in-time joinno future leakage
- throughputwarehouse, not KV
What the new pieces do
- Offline Storestore
- A columnar warehouse of historical feature values over time. Serves large point-in-time-correct joins to build training datasets. Optimized for throughput, not latency.
- Trainingservice
- Builds training datasets by joining labels with historical features from the offline store — using point-in-time joins so it only sees values known at each example’s timestamp.
Step 4 · Freshness for serving
The online store, at millisecond latency
At serving time a prediction must return in milliseconds, and re-running the batch pipeline per request is impossible. But the features must be the same ones training used. How do you serve them fast without recomputing?
How do you serve features in milliseconds without skew?
Recomputing per request is too slow for online serving and risks diverging from the training computation — the skew you’re avoiding. Precompute and look up instead.
The same pipeline writes the latest feature values into a key-value online store; serving does a fast lookup by entity id — same definition as offline, millisecond latency, no recompute.
Caching serving-side computation still starts from a separate implementation — it caches the skew. The values must come from the shared pipeline, not re-derived at the edge.
The Online Store is a low-latency key-value store holding the latest feature values per entity, materialized by the same pipeline that fills the offline store. Model Serving looks up features by entity key in milliseconds and scores. Because both stores come from one definition, the serving features match the training features exactly — no recompute, no skew.
Why this piece earns its place
The online store holds the latest value and nothing else, which means that the instant a prediction is served, the exact input that produced it is gone. Log the feature vector alongside the prediction, keyed to the request. That costs one write per prediction and buys two things nothing else can: a decision you can explain months later, and a direct measurement of skew — replay a day of logged vectors against the offline store’s reconstruction for the same moments and count the columns that disagree. Without it, "no skew" is an architecture claim rather than a number, which is a bad position to be in when accuracy is down and the diagram says this cannot happen. Be deliberate about a lookup that finds nothing, too, because a row that is absent cannot tell you why. A genuinely new entity and one materialization has not reached yet arrive as the same empty response and want opposite handling, so the signal has to come from outside the row: publish a per-feature watermark — how far materialization has got — that serving can read, or keep the registered entity set explicit. And plan the read as a fan-out: one prediction wants many features, and the latency you promise is the slowest lookup, not the average.
- latest valueper entity key
- ms lookupno recompute
- same as offlineby construction
What the new pieces do
- Online Storecache
- A key-value store holding the latest feature values per entity, for millisecond lookups at serving time. Same features as offline — that sameness is the whole point.
- Model Servingbackend
- At inference, looks up the entity’s current features from the online store and feeds them to the model — the same features, by definition, that training used.
- Applicationclient
- Sends a prediction request (a user id, a transaction) and gets a scored result. It supplies the entity key; the store supplies the features.
Step 5 · One source of truth
The feature registry + versioning
If the pipeline holds the logic but each run hard-codes it, definitions drift, nobody knows who owns a feature, and changing one silently breaks a model. What makes a feature a governed, discoverable asset?
What turns a computed value into a reusable, governed feature?
A comment isn’t discoverable, versioned, or enforceable. Features need a registered definition other teams can find and depend on.
Without versioning, changing a feature silently alters what every dependent model sees. Definitions must be versioned so changes are explicit.
The Feature Registry holds each feature’s definition, data type, owner, and version. The pipeline runs these definitions, and both stores derive from them — so a feature is discoverable, reusable, and changes are versioned.
The Feature Registry is the single source of truth: each feature’s transformation logic, type, owner, and version. The pipeline executes these definitions and both stores derive from them, so there is exactly one meaning of a feature. Versioning makes changes explicit — a new version instead of a silent redefinition that breaks dependents.
Why this piece earns its place
Versions are cheap to create and expensive to keep, so the registry needs the other half: knowing who reads each one. Every live version is storage in two places plus a share of every materialization run, and nothing retires it on its own — without usage tracking you accumulate versions nobody will delete because nobody can prove they are unread. Record which model read which version and when it last did, and deprecation becomes a decision instead of a standoff. The registry is also the only thing that can answer the question that always arrives at the worst moment: an upstream table is being dropped or a column renamed — what breaks? Because definitions name their inputs, the registry holds a dependency graph from raw source through feature to model, and that graph turns an upstream migration from an outage into a list of owners to warn. Which is the other job the owner field does: when something fires on a feature later, ownership decides who gets paged, and a feature with no owner is one nobody fixes. Keep the definitions in version control beside the pipeline rather than editable in a console — a feature change should arrive through review, like the code it is.
- one definitionlogic · type · owner
- versionedchanges are explicit
- discoverablereuse, don’t reinvent
What the new pieces do
- Feature Registrystore
- The single source of truth for feature definitions: transformation logic, data types, owners, versions. Both training and serving derive from these definitions — never their own copies.
Step 6 · Keep it fresh
Batch + streaming materialization
Some features are slow-moving (lifetime order count); some are fast (activity in the last 5 minutes). A nightly batch keeps the first fresh but leaves the second hours stale. How do you keep the online store current for both?
How do you keep fast-moving features fresh in the online store?
Use both: scheduled batch jobs for features that change slowly, and streaming ingestion from event streams for features that must reflect the last minutes — each store stays as fresh as its features need.
Minute-batching heavy features wastes compute and still lags real-time; streaming is the right tool for fast features, batch for slow ones.
Some features (fraud signals, session activity) are worthless stale. Streaming materialization exists precisely for them.
Materialize with the right cadence per feature: batch jobs for slow-moving features and streaming ingestion for fast-moving ones, both writing the online store. Freshness becomes a per-feature property (its TTL/SLA), so the online store reflects reality at the speed each feature demands — without recomputing everything constantly.
Why this piece earns its place
Two cadences means one definition now runs in two engines, and that is where "defined once" gets tested. A thirty-day sum computed by a batch job over a settled day and the same sum maintained incrementally from a stream will not agree at the edges, because events arrive late and out of order and the streaming job has to decide how long to wait before it calls a window closed. Which of the two writes lands is a write-ordering question with a clean answer. The quieter problem is which number the model met: the offline store is filled by the batch path, so training almost always learns from the settled value while serving may be reading the streamed approximation of it — skew again, arriving inside the machine built to prevent it. Close that by making lateness a property of the definition, the window an event may still be counted in, and holding both engines to the same one, so the streamed value converges on the batch value before training ever reads that day. And a streamed feature is not free per feature: it is a job running forever whether or not any model reads that value today, so a feature whose freshness target is longer than the batch period should stay in batch.
- batchslow-moving features
- streamingfresh in near-real-time
- per-feature SLAfreshness as a property
What the new pieces do
- Streaming Ingestbus
- Updates the online store in near-real-time from event streams, so fast-moving features (last-5-min activity) are fresh at serving without waiting for the next batch.
Step 7 · Watch for rot
Drift and staleness monitoring
Everything works on launch day. Months later the model quietly underperforms — the world shifted, or a materialization job silently stalled and features went stale. Nothing errored. How do you catch decay before users do?
What tells you features are silently degrading the model?
By the time aggregate accuracy visibly drops, users have felt it for a while — and you still don’t know why. Monitoring features directly catches the cause earlier.
Track each feature’s distribution over time to catch drift (the world changed) and its update recency to catch staleness (materialization stalled) — the two silent ways features degrade a model.
The definition is stable, but the data flowing through it drifts, and pipelines stall. Both silently hurt the model and both need monitoring.
A Drift Monitor watches feature distributions (has the data shifted from what the model trained on?) and freshness (did a materialization job stall?), alerting on both. These are the two silent ways a healthy-looking system decays — catching them at the feature level surfaces the cause before aggregate model metrics even move.
Why this piece earns its place
Two practical problems decide whether anyone still trusts this in six months. The first is what you compare against. A distribution test needs a reference, and the only defensible one is the training window of the model version currently serving, pinned alongside that version — not a rolling window of recent production, or you are measuring today against yesterday and a slow shift never trips anything. That pin belongs next to the version in the registry, which is a quiet second job the registry does. The second is that you are not comparing data, you are comparing summaries. Nobody keeps every value, so the monitor holds a histogram or a sketch per feature per window, and if its bucket edges are recomputed each window the test partly measures your bucketing rather than the world. Freeze the edges with the reference, and choose them by quantile for anything long-tailed, or one heavy bucket absorbs the shift you cared about. Then compute that summary per slice as well as overall: a feature’s total distribution can sit perfectly still while one country’s or one platform’s values have collapsed to a constant, and the aggregate is exactly what hides it. Slices multiply the work, so tie the set to the cuts the product is already reported on.
- driftdistribution shift
- stalenessstalled materialization
- alert earlybefore accuracy visibly drops
What the new pieces do
- Drift Monitorservice
- Watches feature distributions and freshness over time, alerting on drift (the world changed) or staleness (materialization stalled) before they quietly degrade the model.
The payoff
You built a feature store
From "two teams compute the same feature two ways" to a shared platform: one versioned registry of definitions, a pipeline that computes once, an offline store with point-in-time-correct history for training, a low-latency online store for serving, batch + streaming materialization for freshness, and drift/staleness monitoring.
Now skew train vs serve — recompute a feature at serving instead of reading the store — and watch production accuracy fall while every offline metric stays green, because the model is now fed a distribution it never trained on. That’s why a feature is defined once and read everywhere, and why training/serving skew is the silent failure a feature store exists to kill.
Everything you assembled, in order
- One definition — define a feature once; training and serving read the same values — no skew
- Feature pipeline — the single place transformation logic lives; computes once, feeds both stores
- Offline store — historical values + point-in-time joins → correct, leak-free training sets
- Online store — latest values, millisecond lookup → fast serving, same features as training
- Registry — versioned, owned, discoverable definitions — features as reusable assets
- Materialization — batch for slow features, streaming for fast — freshness per feature
- Monitoring — drift + staleness alerts catch silent decay before users do
- The failure — training/serving skew degrades the model silently — read the store, never re-derive
