Vibe Engines
YouTube
AI System Design

Design a Feature Store

Learn AI system design by building a feature store for ML step by step.

The numbers to beatshared pipelineone transformation pathraw → featuresfrom events/tablesfeeds both storesoffline + online

The whole design, in writing

Learn AI system design by building a feature store for ML step by step. An interactive guide to the system that kills training/serving skew — a shared feature registry, an offline store for point-in-time-correct training data, a low-latency online store for serving, batch and streaming materialization, and drift monitoring — so a model sees the same features in production that it trained on.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

Why models need a feature store

A model is only as good as the features it’s fed. In training, a data scientist computes features in a notebook over historical data. In production, an engineer recomputes "the same" features in low-latency serving code. These are two different code paths — and the moment they diverge, the model sees inputs it never trained on and quietly gets worse. Every ML team reinvents feature plumbing, and every one hits this. How do you make features consistent, reusable, and fresh?

A feature store is the shared system for defining, computing, storing, and serving features. One definition produces both the historical values that train the model and the current values that serve it — so training and serving can’t drift apart, and features become reusable assets across models.

Step 1 · The core problem

Training/serving skew, and reuse

Two teams compute "user’s 30-day average spend": one in a training notebook, one in serving code. What goes wrong, and what does the right structure prevent?

What’s the fundamental thing a feature store must guarantee?

  1. Latency matters for serving, but the primary guarantee is consistency: the value the model trains on must equal the value it serves on. Fast-but-inconsistent is the failure.

  2. A single definition, materialized to both an offline (training) and online (serving) store, means the model can’t receive a differently-computed feature in production than it learned from — no skew — and features become reusable across models.

  3. Independent implementations are exactly what causes skew and duplicated work. The store exists to make one definition the source of truth.

The feature store’s job is consistency and reuse: define a feature once, and serve it identically to training (historical values) and serving (current values). That single-definition rule eliminates training/serving skew — the silent killer — and turns features into shared assets many models can reuse.

Why this piece earns its place

The thing that decides whether this actually works is where you draw the line around a feature. The store guarantees consistency for whatever sits inside a definition; anything a model does to a value afterwards — scaling it against training-set statistics, bucketing it, filling a missing value with a default — sits outside that guarantee, and the skew you just eliminated quietly reappears in the model’s own preprocessing code. Either pull those transforms in as registered features, or ship their constants inside the model artifact and version them with it. The other half is the key. A definition is only shared if the thing it is keyed by is shared, and entity keys are where that fails without anyone noticing: one team keys a user by account id, another by a hashed login, and both read a feature of the same name and get different rows. Agreeing that namespace is unglamorous prerequisite work this design does not survive without. Grain is the same decision one level up — a feature about a user costs a row per user, the same feature about a user and a merchant costs a row per pair, which is where an online store stops being cheap. Settle grain at registration, because changing it later is a different feature under a new name, not an edit.

Step 2 · Compute the features

The pipeline + raw data

Features come from raw events and tables that a model can’t use directly. Something has to transform "a stream of orders" into "30-day average order value per user." Where does that logic live?

Raw Eventslogs · tablesFeature Pipelinetransform
New in this step: Raw Events, Feature Pipeline.

Where should feature transformation logic live?

  1. Embedding transforms in training code alone means serving must reimplement them — the origin of skew. The logic has to be shared, not per-consumer.

  2. Then training can’t reproduce it and other models can’t reuse it. Feature logic belongs in a shared pipeline, defined once.

  3. A single pipeline reads raw events and applies the registered transformation to produce feature values — the one place the logic lives, feeding both stores.

A shared Feature Pipeline reads Raw Events and applies transformations to produce feature values. This is the one place feature logic lives, so both stores (and every model) get identically-computed values. The pipeline is the engine; the registry (soon) holds the definitions it runs.

Why this piece earns its place

One pipeline only produces one answer if its inputs can be read the same way twice. Most feature logic reaches past the event stream into dimension tables — a user’s plan tier, a merchant’s category — and those tables are usually updated in place. Recompute last Tuesday’s value today and you get a number last Tuesday could not have produced, so the pipeline is not reproducible even though nobody touched the definition. Anything feeding a feature wants to be append-only or snapshotted, and that is a constraint you have to push upstream onto teams who do not report to you. The other thing this box owes you is a backfill mode, which is usually the one nobody builds. Steady state handles the last window; registering a new feature, or correcting a definition, means recomputing years of it — each value as steady state would have produced it at that moment, same code, different clock. Write the backfill as its own script and you have rebuilt the two-implementations problem inside the platform, one layer down. It is also where the bill lands, because it reads the whole history of the raw events, so no model can train on more history of a feature than those raw events still cover.

  • shared pipelineone transformation path
  • raw → featuresfrom events/tables
  • feeds both storesoffline + online

What the new pieces do

Raw Eventsstore
The raw inputs features are computed from — event streams, transactional tables, clickstreams. Raw and unusable by a model until transformed into features.
Feature Pipelinebus
Computes features from raw data using the definitions in the registry — the single transformation logic that produces values for both the offline and online stores.

Step 3 · History for training

The offline store + point-in-time joins

To train, you need each example’s features as they were at that moment — not today’s values. Join naively and you leak the future into the past. How do you build a correct training set?

Offline StorehistoricalTrainingreads offline
New in this step: Offline Store, Training.

How do you assemble features for a historical training set correctly?

  1. The offline store keeps feature values over time; a point-in-time join attaches, to each training example, only the feature values that existed at that example’s moment — no leakage from the future.

  2. Using today’s values for a year-old example leaks future information the model won’t have at serving — inflated offline metrics, worse production performance. Classic label/feature leakage.

  3. Recomputing per run is slow, non-reproducible, and easy to get subtly wrong. The offline store materializes historical values once, correctly, for reuse.

The Offline Store holds historical feature values over time; Training builds datasets with point-in-time-correct joins — each example gets only the feature values known at its own timestamp. That prevents feature leakage (using future information) and makes training sets reproducible. It’s throughput-optimized (a warehouse), not latency-optimized.

Why this piece earns its place

The detail that decides whether this is correct is which timestamp you join at. It is not when the row landed in the warehouse and it is not when the outcome became known — it is the moment the model would have been called. Join on ingestion time and you hand training the values that existed only because your own pipeline had already caught up, so the model learns from a world a few minutes ahead of the one it serves in. Nothing flags that; the eval is simply better than production, and every explanation you reach for will be the wrong one. The join itself is ordinary work done carefully: partition both sides on the entity key, sort each partition by time, and walk them together taking the last value at or before each label. The cost sits in that sort, not in the comparison — what explodes is the naive version that pairs every label with every value and filters afterwards. Keep the joined result as a dated artifact beside the model, because the join is deterministic while the data underneath it is not: a retention rule written to save storage will quietly make every model older than it unrebuildable, and nobody who set that rule was thinking about models.

  • historicalvalues over time
  • point-in-time joinno future leakage
  • throughputwarehouse, not KV

What the new pieces do

Offline Storestore
A columnar warehouse of historical feature values over time. Serves large point-in-time-correct joins to build training datasets. Optimized for throughput, not latency.
Trainingservice
Builds training datasets by joining labels with historical features from the offline store — using point-in-time joins so it only sees values known at each example’s timestamp.

Step 4 · Freshness for serving

The online store, at millisecond latency

At serving time a prediction must return in milliseconds, and re-running the batch pipeline per request is impossible. But the features must be the same ones training used. How do you serve them fast without recomputing?

Online Storelow latencyModel Servingreads onlineApplicationprediction req
New in this step: Online Store, Model Serving, Application.

How do you serve features in milliseconds without skew?

  1. Recomputing per request is too slow for online serving and risks diverging from the training computation — the skew you’re avoiding. Precompute and look up instead.

  2. The same pipeline writes the latest feature values into a key-value online store; serving does a fast lookup by entity id — same definition as offline, millisecond latency, no recompute.

  3. Caching serving-side computation still starts from a separate implementation — it caches the skew. The values must come from the shared pipeline, not re-derived at the edge.

The Online Store is a low-latency key-value store holding the latest feature values per entity, materialized by the same pipeline that fills the offline store. Model Serving looks up features by entity key in milliseconds and scores. Because both stores come from one definition, the serving features match the training features exactly — no recompute, no skew.

Why this piece earns its place

The online store holds the latest value and nothing else, which means that the instant a prediction is served, the exact input that produced it is gone. Log the feature vector alongside the prediction, keyed to the request. That costs one write per prediction and buys two things nothing else can: a decision you can explain months later, and a direct measurement of skew — replay a day of logged vectors against the offline store’s reconstruction for the same moments and count the columns that disagree. Without it, "no skew" is an architecture claim rather than a number, which is a bad position to be in when accuracy is down and the diagram says this cannot happen. Be deliberate about a lookup that finds nothing, too, because a row that is absent cannot tell you why. A genuinely new entity and one materialization has not reached yet arrive as the same empty response and want opposite handling, so the signal has to come from outside the row: publish a per-feature watermark — how far materialization has got — that serving can read, or keep the registered entity set explicit. And plan the read as a fan-out: one prediction wants many features, and the latency you promise is the slowest lookup, not the average.

  • latest valueper entity key
  • ms lookupno recompute
  • same as offlineby construction

What the new pieces do

Online Storecache
A key-value store holding the latest feature values per entity, for millisecond lookups at serving time. Same features as offline — that sameness is the whole point.
Model Servingbackend
At inference, looks up the entity’s current features from the online store and feeds them to the model — the same features, by definition, that training used.
Applicationclient
Sends a prediction request (a user id, a transaction) and gets a scored result. It supplies the entity key; the store supplies the features.

Step 5 · One source of truth

The feature registry + versioning

If the pipeline holds the logic but each run hard-codes it, definitions drift, nobody knows who owns a feature, and changing one silently breaks a model. What makes a feature a governed, discoverable asset?

Feature PipelinetransformOffline StorehistoricalOnline Storelow latencyFeature Registrydefinitions
New in this step: Feature Registry.

What turns a computed value into a reusable, governed feature?

  1. A comment isn’t discoverable, versioned, or enforceable. Features need a registered definition other teams can find and depend on.

  2. Without versioning, changing a feature silently alters what every dependent model sees. Definitions must be versioned so changes are explicit.

  3. The Feature Registry holds each feature’s definition, data type, owner, and version. The pipeline runs these definitions, and both stores derive from them — so a feature is discoverable, reusable, and changes are versioned.

The Feature Registry is the single source of truth: each feature’s transformation logic, type, owner, and version. The pipeline executes these definitions and both stores derive from them, so there is exactly one meaning of a feature. Versioning makes changes explicit — a new version instead of a silent redefinition that breaks dependents.

Why this piece earns its place

Versions are cheap to create and expensive to keep, so the registry needs the other half: knowing who reads each one. Every live version is storage in two places plus a share of every materialization run, and nothing retires it on its own — without usage tracking you accumulate versions nobody will delete because nobody can prove they are unread. Record which model read which version and when it last did, and deprecation becomes a decision instead of a standoff. The registry is also the only thing that can answer the question that always arrives at the worst moment: an upstream table is being dropped or a column renamed — what breaks? Because definitions name their inputs, the registry holds a dependency graph from raw source through feature to model, and that graph turns an upstream migration from an outage into a list of owners to warn. Which is the other job the owner field does: when something fires on a feature later, ownership decides who gets paged, and a feature with no owner is one nobody fixes. Keep the definitions in version control beside the pipeline rather than editable in a console — a feature change should arrive through review, like the code it is.

  • one definitionlogic · type · owner
  • versionedchanges are explicit
  • discoverablereuse, don’t reinvent

What the new pieces do

Feature Registrystore
The single source of truth for feature definitions: transformation logic, data types, owners, versions. Both training and serving derive from these definitions — never their own copies.

Step 6 · Keep it fresh

Batch + streaming materialization

Some features are slow-moving (lifetime order count); some are fast (activity in the last 5 minutes). A nightly batch keeps the first fresh but leaves the second hours stale. How do you keep the online store current for both?

Online StoreFeature RegistryStreaming Ingest
New in this step: Streaming Ingest. · swipe to pan the diagram

How do you keep fast-moving features fresh in the online store?

  1. Use both: scheduled batch jobs for features that change slowly, and streaming ingestion from event streams for features that must reflect the last minutes — each store stays as fresh as its features need.

  2. Minute-batching heavy features wastes compute and still lags real-time; streaming is the right tool for fast features, batch for slow ones.

  3. Some features (fraud signals, session activity) are worthless stale. Streaming materialization exists precisely for them.

Materialize with the right cadence per feature: batch jobs for slow-moving features and streaming ingestion for fast-moving ones, both writing the online store. Freshness becomes a per-feature property (its TTL/SLA), so the online store reflects reality at the speed each feature demands — without recomputing everything constantly.

Why this piece earns its place

Two cadences means one definition now runs in two engines, and that is where "defined once" gets tested. A thirty-day sum computed by a batch job over a settled day and the same sum maintained incrementally from a stream will not agree at the edges, because events arrive late and out of order and the streaming job has to decide how long to wait before it calls a window closed. Which of the two writes lands is a write-ordering question with a clean answer. The quieter problem is which number the model met: the offline store is filled by the batch path, so training almost always learns from the settled value while serving may be reading the streamed approximation of it — skew again, arriving inside the machine built to prevent it. Close that by making lateness a property of the definition, the window an event may still be counted in, and holding both engines to the same one, so the streamed value converges on the batch value before training ever reads that day. And a streamed feature is not free per feature: it is a job running forever whether or not any model reads that value today, so a feature whose freshness target is longer than the batch period should stay in batch.

  • batchslow-moving features
  • streamingfresh in near-real-time
  • per-feature SLAfreshness as a property

What the new pieces do

Streaming Ingestbus
Updates the online store in near-real-time from event streams, so fast-moving features (last-5-min activity) are fresh at serving without waiting for the next batch.

Step 7 · Watch for rot

Drift and staleness monitoring

Everything works on launch day. Months later the model quietly underperforms — the world shifted, or a materialization job silently stalled and features went stale. Nothing errored. How do you catch decay before users do?

Online Storelow latencyDrift Monitorskew · staleness
New in this step: Drift Monitor.

What tells you features are silently degrading the model?

  1. By the time aggregate accuracy visibly drops, users have felt it for a while — and you still don’t know why. Monitoring features directly catches the cause earlier.

  2. Track each feature’s distribution over time to catch drift (the world changed) and its update recency to catch staleness (materialization stalled) — the two silent ways features degrade a model.

  3. The definition is stable, but the data flowing through it drifts, and pipelines stall. Both silently hurt the model and both need monitoring.

A Drift Monitor watches feature distributions (has the data shifted from what the model trained on?) and freshness (did a materialization job stall?), alerting on both. These are the two silent ways a healthy-looking system decays — catching them at the feature level surfaces the cause before aggregate model metrics even move.

Why this piece earns its place

Two practical problems decide whether anyone still trusts this in six months. The first is what you compare against. A distribution test needs a reference, and the only defensible one is the training window of the model version currently serving, pinned alongside that version — not a rolling window of recent production, or you are measuring today against yesterday and a slow shift never trips anything. That pin belongs next to the version in the registry, which is a quiet second job the registry does. The second is that you are not comparing data, you are comparing summaries. Nobody keeps every value, so the monitor holds a histogram or a sketch per feature per window, and if its bucket edges are recomputed each window the test partly measures your bucketing rather than the world. Freeze the edges with the reference, and choose them by quantile for anything long-tailed, or one heavy bucket absorbs the shift you cared about. Then compute that summary per slice as well as overall: a feature’s total distribution can sit perfectly still while one country’s or one platform’s values have collapsed to a constant, and the aggregate is exactly what hides it. Slices multiply the work, so tie the set to the cuts the product is already reported on.

  • driftdistribution shift
  • stalenessstalled materialization
  • alert earlybefore accuracy visibly drops

What the new pieces do

Drift Monitorservice
Watches feature distributions and freshness over time, alerting on drift (the world changed) or staleness (materialization stalled) before they quietly degrade the model.

The payoff

You built a feature store

From "two teams compute the same feature two ways" to a shared platform: one versioned registry of definitions, a pipeline that computes once, an offline store with point-in-time-correct history for training, a low-latency online store for serving, batch + streaming materialization for freshness, and drift/staleness monitoring.

Raw EventsFeature PipelineOffline StoreOnline StoreFeature RegistryTrainingModel ServingApplicationStreaming IngestDrift Monitor
The finished design, end to end. · swipe to pan the diagram

Now skew train vs serve — recompute a feature at serving instead of reading the store — and watch production accuracy fall while every offline metric stays green, because the model is now fed a distribution it never trained on. That’s why a feature is defined once and read everywhere, and why training/serving skew is the silent failure a feature store exists to kill.

Everything you assembled, in order

  • One definition — define a feature once; training and serving read the same values — no skew
  • Feature pipeline — the single place transformation logic lives; computes once, feeds both stores
  • Offline store — historical values + point-in-time joins → correct, leak-free training sets
  • Online store — latest values, millisecond lookup → fast serving, same features as training
  • Registry — versioned, owned, discoverable definitions — features as reusable assets
  • Materialization — batch for slow features, streaming for fast — freshness per feature
  • Monitoring — drift + staleness alerts catch silent decay before users do
  • The failure — training/serving skew degrades the model silently — read the store, never re-derive

Deep cut · 11:07

One number. Two answers.

The interactive build above lays out a feature store: one registry of definitions, a pipeline that computes each feature once, an offline store for training and an online store for serving, streaming ingest to keep the fast features fresh, and a drift monitor watching both. The film builds the same system in a pub. Wick’s discount rule runs on one number — what a regular has spent this month — and that number is read twice: at the door, off a chalked board written up to now, and on a Sunday, off the finished till roll. Then it asks why a rule that scores right nine times in ten lost two hundred pounds.

  • See the wreck: Denny was refused a discount on the ninth with thirty-eight pounds on the board. Sunday marked that call against a hundred and fifteen — seventy-seven of which he spent on the twenty-ninth, three weeks after the decision being graded. Nothing errored and nobody lied; two copies of one number drifted apart under the same name, and the only place it showed was the takings.
  • Take it into the interview: training-serving skew is one definition computed twice. Write the value once with the time it became true, serve the fast copy at the door, and build training sets with point-in-time joins so each old decision is scored against only what was known when it was made — then backfill the months that came before the ledger existed.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. A point-in-time join needs a feature’s value at a timestamp before that feature even existed. What happens?

    It resolves to null (or a documented default), not an error and not the nearest available value — using the nearest-available value would itself be a subtle leak (borrowing information from closer to "now" than the example is allowed to see). A model has to be trained to handle that null explicitly, the same way it would in true production cold-start, which is exactly the scenario point-in-time correctness is protecting the model from being blind to.

  2. Both a nightly batch job and streaming ingestion write to the online store for the same feature. Won’t the batch job overwrite a fresher streamed value with a stale one?

    Yes, unless writes are timestamped and the store keeps the most-recent-by-timestamp rather than most-recent-by-write-order — the batch job’s payload carries the "as of" time its data represents, and a write with an older as-of time than what’s already stored is a no-op. This is why materialization isn’t just "run two jobs" — it’s two jobs cooperating under one freshness contract per feature.

  3. The drift monitor flags a feature distribution has shifted. Does that halt serving?

    No — halting on drift alone would take the model offline for a change that might be legitimate (a real shift in user behavior, not a bug), so drift triggers an alert and an investigation, not an automatic kill switch. Staleness (a stalled materialization job) is the one that more often warrants an automatic fallback, because a stalled pipeline has no legitimate interpretation — it’s just broken.

  4. How do you migrate a model from feature-v1 to feature-v2 without a retrain outage?

    Both versions coexist in the registry and both stores materialize both versions for an overlap window — the new model trains and validates against v2 while the old model keeps serving on v1, and the cutover happens only after the v2-trained model passes evaluation. Versioning is what makes this a planned migration instead of a forced simultaneous flip.

  5. The app now serves users on three continents. What changes about the online store?

    A single global key-value store adds cross-region latency that defeats the point of "millisecond lookup," so the online store is typically replicated per region with the feature pipeline writing to all regions (or a region-local pipeline reading from region-local streams) — the offline store can usually stay centralized since training isn’t latency-sensitive, but serving-time freshness now has to survive replication lag, not just pipeline lag.

Check yourself — the answers, and why

Eight steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. The primary problem a feature store solves is…

    Define a feature once and serve it identically to training and serving so the model never sees a differently-computed input in production.

  2. The offline store uses point-in-time joins to…

    Joining to current values leaks the future into the past and inflates offline metrics; point-in-time correctness reconstructs history as it actually was.

  3. The online store exists to…

    A key-value online store gives serving fast lookups of current features — materialized from the same pipeline as offline, so no skew.

  4. The feature registry’s job is to…

    A single versioned definition makes features reusable and governable, and makes changes explicit instead of silent redefinitions.

  5. Fast-moving features stay fresh via…

    Batch handles slow features; streaming updates fast ones (last-5-min activity) so freshness matches each feature’s SLA.

  6. Training/serving skew is dangerous because it…

    Recomputing a feature differently at serving silently degrades the model while offline metrics stay perfect; route both through the same definition.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Define once: register a feature’s transformation logic in one place and serve it identically to training and serving.
  • Train correctly: build training sets with point-in-time joins so each example sees only the feature values known at its own timestamp.
  • Serve fast: look up an entity’s latest feature values in milliseconds at inference, with no recompute.
  • Keep fresh: materialize features on the right cadence — batch for slow-moving, streaming for fast-moving.
  • Catch decay: monitor feature distributions and freshness, alerting on drift and staleness before accuracy visibly drops.

The qualities that shape everything

Each one names the mechanism that buys it.

The model never sees a differently-computed feature
One registered definition is materialized to both the offline (training) and online (serving) store, so training and serving read the same values by construction — no skew.
Leak-free, reproducible training sets
The offline store keeps historical values over time and Training joins with point-in-time correctness — each example gets only the feature values known at its own timestamp, never future ones.
Millisecond serving with no recompute
A low-latency key-value online store holds the latest feature values per entity, materialized by the same pipeline; serving looks up by entity key instead of re-deriving anything.
A change can’t silently poison a live model
The registry versions each definition (logic, type, owner), so a change is a new version rather than an in-place edit and dependent models stay pinned to the version they validated against.
Fast-moving features stay fresh
Batch jobs materialize slow-moving features while streaming ingestion updates fast-moving ones into the online store in near-real-time, making freshness a per-feature SLA.
Silent decay is caught before users feel it
A drift monitor watches feature distributions (has the data shifted?) and freshness (did materialization stall?) and alerts on both at the feature level, before aggregate accuracy visibly drops.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

One shared definition over each team writing its own feature code

Independent implementations are exactly what cause training/serving skew and duplicated work; a single definition materialized to both stores means the model can’t be fed a differently-computed feature in production.

Point-in-time joins over joining every example to today’s values

Using current values for a year-old example leaks future information the model won’t have at serving — inflated offline metrics that collapse in production; point-in-time joins reconstruct only what was known at each example’s timestamp.

Precompute into a low-latency online store over recomputing features per request

Recomputing per request is too slow for online serving and risks diverging from the training computation — the very skew you’re avoiding; materializing the same pipeline’s output and looking up by entity key is fast and consistent.

Versioned registry definitions over trusting whatever the latest pipeline run produced

Without versioning, changing a feature silently alters what every dependent model sees; a versioned, owned definition makes a change a new version dependents opt into rather than a silent redefinition.

Batch slow features, stream fast ones over batching everything on one schedule

Minute-batching heavy features wastes compute and still lags real-time, while nightly-batching fast features leaves them hours stale; matching cadence per feature meets each one’s freshness SLA cheaply.

The answer, out loud

What a strong answer to “Design a Feature Store” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Name the failure before I draw anything

    I’d start by naming what breaks, because the words "feature store" hide it. One feature ends up with two implementations — one running over history to train the model, one running per request to serve it — and they can drift apart without either being broken, since each is correct on its own terms. Nothing throws an error, so this gets found late and from the wrong direction. My first requirement isn’t latency or scale: the number the model learned from has to be the number it’s handed in production. The second is reuse, and that is what pays for the platform.

  2. 3–8 min

    One definition, two consumers

    The shape I’d propose is one definition materialized twice, and the second copy exists because the readers want opposite things. Training reads every entity at once, across months, and can wait. Serving reads one entity, now, with somebody waiting on a page. No single store is good at both, so I’d stop trying: the same values land in a columnar store for one and a key-value store for the other. What keeps that from becoming two implementations again is that neither store is written by its reader — one job fills both from one definition.

    Built in step 1: Training/serving skew, and reuse
  3. 8–13 min

    Put the logic in a shared pipeline

    The first concrete box is the pipeline that turns raw events and tables into feature values, and I’d refuse to let either end keep a copy of that logic. My reason is about time more than correctness: a copy is never wrong the day it’s written, it goes wrong months later when one side gets patched and the other doesn’t, and by then nobody remembers there were two. I’d also pin down what the pipeline emits — not just a number, but the entity it describes and the moment it became true. Both of the next two boxes stand on that timestamp.

    Built in step 2: The pipeline + raw data
  4. 13–20 min

    History, as it actually was

    Training reads from an offline store: a warehouse of historical values, optimized for throughput rather than latency. The interesting part isn’t the storage, it’s the join. Each labeled example needs the feature values known at its own timestamp, not today’s. Join everything to current values and I leak the future into the past — a year-old example gets information the model will never have at serving time, the offline numbers look wonderful, and production doesn’t match them. That’s a point-in-time join, and I’d say the words out loud, because it’s the thing that separates having built one of these from having read about it.

    Built in step 3: The offline store + point-in-time joins
  5. 20–26 min

    Serve it without recomputing

    Serving is the mirror image. The answer has to come back while someone waits, so nothing is computed inside the request: the same pipeline writes the current value per entity into a key-value store, and serving looks it up by entity key and scores. The shortcut I’d argue against out loud is letting the serving code compute the feature behind a cache — a cache makes a second implementation fast, it doesn’t make it agree with the first. And this is the box with a real availability requirement: if it is down, the model isn’t slow, it is guessing.

    Built in step 4: The online store, at millisecond latency
  6. 26–32 min

    Make the definition a governed asset

    Then the registry, which is what makes this a shared system rather than one team’s pipeline. Everything so far works without it — right until a second team wants a feature that already exists. They need to find it, see who owns it, see what it is computed from, and trust it won’t move under them. Take any of those away and the rational move is to copy the definition and rename it, and now the same feature exists under a second name, and the reuse that justified the platform is gone. So the definition is the registered object, and the pipeline runs registered definitions rather than whatever code shipped last.

    Built in step 5: The feature registry + versioning
  7. 32–38 min

    Freshness is per feature

    Freshness I’d attach to the feature rather than the schedule. Each definition states how stale its value may be, and that target picks the machinery: a slow-moving count is happy on a batch job, while anything about the last few minutes comes off the stream into the online store as events land. I’d put the target on the feature because the model owner is the one who knows it and the platform is the one who pays for it — if the platform decides, it either overpays everywhere or under-serves the one feature that mattered. I’d ask here what the tightest feature really is.

    Built in step 6: Batch + streaming materialization
  8. 38–43 min

    Catch the decay

    Last is monitoring, and I’d watch the features rather than the model, because in most of these systems the label arrives late or never. Scoring risk, the truth about today’s decision turns up whenever a dispute does; until then there is no accuracy to plot. The inputs are observable immediately. So per feature: does its distribution still look like what the model learned on, and is it being updated at all. A stalled job keeps serving the last value it wrote, which looks exactly like a feature that hasn’t changed.

    Built in step 7: Drift and staleness monitoring
  9. 43–45 min

    The trade-off, and what I’d do next

    I’d close on the trade-off, because that is the real question. This is a platform: storage in two shapes, jobs that run whether or not anyone reads them, and someone carrying a pager — and one team shipping one model is faster without it. It starts earning at the second consumer of a feature, and the case for building it earlier is that retrofitting means asking a team with something that works to give it up. With more time, two pieces I’d design properly. Features computable only from what is in the request itself — the amount of the transaction being scored — which this design has nowhere to put. And deleting a person, whose values sit in both stores and in every training set built off the offline one; compute-once-read-everywhere makes that harder, not easier.

What this teaches

Learn AI system design by building a feature store for ML step by step. An interactive guide to the system that kills training/serving skew — a shared feature registry, an offline store for point-in-time-correct training data, a low-latency online store for serving, batch and streaming materialization, and drift monitoring — so a model sees the same features in production that it trained on.

Key takeaways

  • One definition — define a feature once; training and serving read the same values — no skew
  • Feature pipeline — the single place transformation logic lives; computes once, feeds both stores
  • Offline store — historical values + point-in-time joins → correct, leak-free training sets
  • Online store — latest values, millisecond lookup → fast serving, same features as training
  • Registry — versioned, owned, discoverable definitions — features as reusable assets
  • Materialization — batch for slow features, streaming for fast — freshness per feature
  • Monitoring — drift + staleness alerts catch silent decay before users do
  • The failure — training/serving skew degrades the model silently — read the store, never re-derive

Concepts covered

  • Why models need a feature store
  • Training/serving skew, and reuse
  • The pipeline + raw data
  • The offline store + point-in-time joins
  • The online store, at millisecond latency
  • The feature registry + versioning
  • Batch + streaming materialization
  • Drift and staleness monitoring
RUN IT YOURSELF

The training set that already knows the answer

One event log, one shared feature pipeline, one offline store — and the same training set assembled three ways, differing only in which timestamp the join reads at. It trains a real model on each and prints offline accuracy beside the accuracy that same model gets in production. Change DISPUTE_LAG_DAYS to 300 so the outcome falls off the end of the log, hit Run, and watch the gap collapse to +0.000.

HOW TO READ THE CODE — 6 IDEAS
  1. The naive join takes each entity’s latest value. Chargebacks land ~21 days after the decision they grade, so disputes_total arrives back in the table carrying the label (step 3).
  2. That join wins the offline bake-off — 0.906 against 0.728 — and loses production, 0.606 against 0.728. A leak always makes the eval better, which is why nobody goes looking for it.
  3. Push DISPUTE_LAG_DAYS to 90 or 150 and every number is byte-identical: the gap is not the distance to the future, it is the fact that the future is in the table at all. Only at 300, past the end of the log, does it collapse.
  4. Read the w[disputes] column: the leaky model puts 1.82 on a feature that is 0.05 at decision time and 0.48 afterwards. The as-of model puts 0.00 on it.
  5. The as-of row’s gap is 0.000 by identity, not by luck — the join replays what the online store would have returned, so there is nothing for serving to diverge from (steps 1 and 4).
  6. The third row is the subtler bug the page names: joining on the event clock instead of the materialized one returns 139 of 540 lookups the store did not hold yet (step 3).
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, skew train vs serve, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs