Vibe Engines
YouTube
AI System Design

Design a Recommendation System

Step 1 / 9

Learn AI system design by building a large-scale recommendation system step by step.

The numbers to beatmillionsitems, each a vector1 vectorper userdistance= predicted affinity

The whole design, in writing

Learn AI system design by building a large-scale recommendation system step by step. An interactive guide covering the two-stage retrieve-and-rank funnel, two-tower embedding retrieval with ANN, a heavy ranking model, the feature store and its offline/online split, filtering business rules, and the feedback loop that keeps recommendations fresh.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

Pick the best few from millions

A user opens the app. Somewhere in a catalog of millions of items are the ten they’d most want to see right now. You have tens of milliseconds to find them. You can’t score millions of items per request — but you can’t guess, either. How do you do this fast and well?

Useropens the app
New in this step: User.

Use a funnel. A cheap, fast stage narrows millions to a few hundred plausible candidates; a heavy, accurate stage ranks those few hundred precisely. Cheap-and-wide, then expensive-and-narrow — the pattern behind every large recommender.

What the new pieces do

Userclient
A person who opens the app expecting a feed of items they’ll actually like — picked from millions.

Step 1 · The skeleton

A request and a funnel

The user opens the app and needs a feed. We can’t run a heavy model over millions of items in the time budget. What overall structure makes "best ten from millions" feasible?

Useropens the appRecs Gatewayorchestrates
New in this step: Recs Gateway.

You must return the best 10 of millions in ~tens of ms. What’s the shape?

  1. Running a heavy model over millions of items per request is impossible in the latency budget. You must shrink the set first.

  2. Popularity is a fine fallback but ignores the individual — it’s not personalization, and engagement suffers. You need per-user relevance.

  3. Stage 1 narrows millions to hundreds cheaply; stage 2 ranks those hundreds with a heavy model. Each stage is sized for its job.

A Recs Gateway runs a two-stage funnel for every request: candidate generation (millions → hundreds, fast) then ranking (hundreds → ordered, accurate), with filtering before the final list ships to the user. The shape is the whole idea.

Why this piece earns its place

The funnel sets a ceiling nothing downstream can lift: an item candidate generation never returns cannot be recommended, however good the ranker becomes. That matters because ranking metrics are computed over the candidate set, so they keep looking healthy while the item a user would have loved was never in the room — a retrieval regression surfaces as a ranking plateau, and a team can spend a quarter tuning the wrong stage. The stages are also less sequential than the drawing suggests. Nothing the ranker needs about the user depends on which candidates come back, so that fetch can start the instant the request arrives and overlap retrieval entirely, while anything about the items cannot begin until there is a candidate list. Which half of the work is blocked on what is usually how the budget gets met, rather than by making any one stage faster. The gateway is also where the per-stage deadlines live, and every stage needs an answer for blowing one. Ranking times out and you can ship the candidate order; filters time out and you ship a shorter feed. What you cannot do is return nothing — unlike a search box, where no results is an honest answer, an empty home screen reads as a broken app.

What the new pieces do

Recs Gatewaybackend
Runs the funnel for each request: fetch candidates, rank them, apply rules, return the final list.

Step 2 · Represent taste

Embeddings for users and items

To find items a user will like, you need to compare a user against items numerically. "Action movies" and "this thriller" need to be close somehow. How do you represent taste so similarity is computable?

Useropens the appRecs Gatewayorchestrates
The system as it stands at this step.

How do you make "this user" and "this item" comparable for similarity?

  1. Tags are coarse and miss latent taste — two thrillers can feel totally different. Learned embeddings capture nuance hand-tags can’t.

  2. Train embeddings from engagement so a user vector lands near the item vectors they’d enjoy. Similarity becomes distance in that space.

  3. IDs carry no notion of similarity — item 5012 isn’t "near" item 5013. You need a learned space where proximity means relevance.

The system learns embeddings: users and items map to vectors in a shared space where a user sits near the items they’d engage with. These come from training on past engagement, so proximity encodes learned taste — not hand-typed tags.

Why this piece earns its place

Treat the space as a versioned artifact, not a property of a model. Both towers are trained together and have to go live together, which makes this a cutover rather than a sequence: build the new index from the new item tower beside the one already serving, check it, then move queries and the new user tower across in one switch. The new user tower must never address the old index, and the old one must never address the new. Either mismatch has queries from one space searching an index built in another, and nearest-neighbour across two spaces returns real-looking distances over essentially arbitrary items, with nothing anywhere throwing an error. Stamp every index with the model version that produced it so a mismatch is a startup assertion rather than an incident, and keep the previous index warm afterwards, because a rollback has to take both halves back together. The second thing worth knowing is what this space buys you elsewhere. Item vectors are the cheapest similarity you will ever have, and the near-duplicate rules further down the funnel run on those same distances — one artifact, several consumers. That makes its refresh cadence a shared dependency: bump the embeddings to improve retrieval and you have quietly changed what the feed considers repetitive, which is not a connection anyone remembers when the complaint arrives.

  • millionsitems, each a vector
  • 1 vectorper user
  • distance= predicted affinity

Step 3 · Narrow the field

Two-tower retrieval with ANN

You have a user vector and millions of item vectors. You need the few hundred closest — in milliseconds, per request. Comparing against every item exactly is too slow. How do you retrieve candidates fast?

Candidate Gentwo-tower ANNItem Indexembeddings + ANN
New in this step: Candidate Gen, Item Index.

Find the few hundred closest items to the user vector, fast. How?

  1. Exact scoring over millions per request blows the latency budget. At this scale you use an approximate index.

  2. A "two-tower" model produces the user vector live and item vectors offline; an ANN index returns the nearest few hundred in milliseconds.

  3. Category filters are a blunt pre-filter that miss cross-category gems and still leave too many to score exactly. ANN over embeddings is the scalable answer.

A two-tower model has a user tower (computed live from the request) and an item tower (item vectors precomputed offline and stored in the Item Index). Candidate Generation embeds the user, then runs ANN search to pull the nearest few hundred items in milliseconds. Recall stage: don’t miss the good ones.

Why this piece earns its place

The reason this is two towers and not one model is a hard constraint, and naming it is most of the interview. The towers never see each other — the item side was computed hours ago, the user side is computed from the request, and they meet only at a distance. So retrieval structurally cannot use any user-item cross feature: whether this person has watched this creator before, how people in this city reacted to this item in the last hour. Those only become computable once the set is small enough to score pairs, which is why the next stage exists and why you cannot simply move it earlier. Production systems buy back some of that lost expressiveness by running several retrievers rather than one — a recent-interest source, a social-graph source, a trending source — and unioning the results before ranking. That is a deliberately cheap change: each source is an independent index with its own quota of the candidate budget, and adding one never touches the ranker. It also means the candidate count stops being a single number you tune and becomes an allocation you argue about.

What the new pieces do

Candidate Genservice
Narrows millions of items to a few hundred plausible candidates fast, using embedding similarity.
Item Indexindex
Vector index of every item’s embedding, so candidate generation is a nearest-neighbour lookup.

Step 4 · Rank with care

The heavy ranking model

Retrieval handed back ~hundreds of candidates, ordered only by rough embedding similarity. Similarity isn’t the same as "will this specific user engage right now." How do you get the order right?

Recs GatewayorchestratesCandidate Gentop few hundredRanking Modelpredict engage
New in this step: Ranking Model.

You have ~hundreds of candidates. How do you order them precisely?

  1. Embedding distance is a coarse recall signal, not a precise engagement prediction. The top few hundred need a real scoring model.

  2. Popularity ignores the individual and context. The whole point of ranking is per-user, per-moment precision.

  3. A ranking model scores P(engage) for each candidate using many user/item/context features — affordable now because there are only hundreds, not millions.

A heavy Ranking Model scores each candidate’s probability of engagement (click, watch, purchase) using rich features — user history, item attributes, context (time, device), cross features. It’s far too expensive to run over millions, but perfect over a few hundred. Precision stage: order them right.

Why this piece earns its place

Two things about this model only show up once it is running. First, ranking well and predicting well are not the same. A model can order candidates correctly and still be badly calibrated — the number it emits reads as a probability and is not one. Ordering is all you need while the score’s only job is sorting, but the moment anything does arithmetic with it — blending several predictions, comparing against a fixed cutoff, deciding whether a slot is worth filling at all — an uncalibrated score corrupts that arithmetic quietly, and rank-based offline metrics will not notice, because any transformation that preserves order leaves them unchanged. Calibrate against held-out outcomes and re-check after every retrain, since calibration drifts even when ranking quality does not. Second, this is where the compute bill lives. The ranker runs over every candidate on every request, so its cost scales with candidates times traffic, not traffic — far and away the largest inference spend in the design. When it gets too expensive, reach for the size of the candidate set before a smaller model. Both costs are measurable; the difference is reversibility. The candidate count is a number you can turn down on one surface this afternoon and put back tomorrow, while a smaller ranker is a retrain, an evaluation and a release before you know what you gave up.

What the new pieces do

Ranking Modelservice
A heavy model that scores each candidate’s probability of engagement using rich user/item features.

Step 5 · Feed the model

The feature store

Ranking needs features — "how many cooking videos did this user watch today?", "this item’s 1-hour click-rate". Some are slow to compute; all must be fresh and identical between training and serving. Where do features come from?

Ranking Modelpredict engageFeature Storeonline + offline
New in this step: Feature Store.

Ranking needs fresh features, and training must see the SAME values. How?

  1. Heavy aggregations can’t be computed within the latency budget per request. And ad-hoc computation drifts from what training saw.

  2. Compute features in an offline pipeline AND serve them online with low latency, from one definition — so training and serving see identical values (no skew).

  3. Features are dynamic (today’s watch count); they can’t be frozen into weights. They must be served live and refreshed.

A Feature Store serves ranking its features and has two halves that must agree: an offline pipeline computes features in batch (for training), and an online store serves the freshest values at request time (for serving) — from a single definition, so both see the same numbers.

Why this piece earns its place

The half of this that nobody mentions until it bites is point-in-time correctness. Training rows are built by joining logged impressions to features, and the value you have to join is the one the online store would have served at that instant — not the one sitting there now. Join current numbers onto last month’s impressions and you have handed the model a feature that already knows the outcome: a user’s watch count for a creator includes the watch you are trying to predict. Offline metrics improve, the online test does not move, and nothing errors. Getting it right means every feature is written with its own event time and read as of a timestamp, which is a far heavier requirement than serving the latest value, and the real reason this is infrastructure rather than a cache. The same requirement decides what happens when a value simply is not there at request time, which the online half will occasionally do. The ranker scores the row regardless, so whatever gets substituted has to be what training substituted — a default picked quietly in serving code is the one skew nobody documents, and it shows up as a model that ranks a small, unlucky slice of users badly while every aggregate metric looks fine.

What the new pieces do

Feature Storestore
Serves the user/item/context features ranking needs, with offline (batch) and online (low-latency) halves.

Back of the envelope

offline half
batch-compute features → training data
online half
serve freshest values at request, low latency
one definition
same logic both sides → no train/serve skew
freshness matters
today’s behavior must reach ranking today

Step 6 · The last mile

Filters, rules and diversity

Ranking gives a perfectly-ordered list — but the top items might be things the user already saw, items that violate policy, or ten near-identical videos. A great score isn’t a shippable feed. What sits between ranking and the screen?

Recs GatewayorchestratesRanking Modelpredict engageFilters & Rulesdedupe · policy
New in this step: Filters & Rules.

The ranked list is perfect by score. Why isn’t it ready to ship?

  1. Top-by-score can repeat already-seen items, surface blocked content, and stack near-duplicates. Real feeds need post-ranking rules.

  2. Diversity matters, but blindly maximizing it tanks relevance. It’s one constraint among several applied after ranking, not a replacement for it.

  3. A filtering stage removes already-seen/blocked items, enforces policy and freshness, and spreads diversity — turning a scored list into a shippable feed.

A Filters & Rules stage takes the ranked list and makes it shippable: drop already-seen and blocked/policy-violating items, enforce diversity (don’t show ten near-identical items), and apply business rules (freshness, sponsored slots, fairness). Then the final feed goes to the user.

Why this piece earns its place

The rules here are not all of one kind, and the difference decides what the stage above owes this one. Already-seen and blocked are subtractions: they remove candidates outright, so the candidate budget has to be set from the observed survival rate rather than from the length of the feed. A heavy user has seen most of what retrieval returns, so a few hundred candidates can collapse to a handful and the feed ships short — and the user whose survival rate is unusual needs more candidates, not a thinner screen. Diversity and sponsored slots subtract nothing; they re-order and they reserve, changing which items reach the screen without changing how many you had to retrieve for it. The cheapest structural fix is to push the hard exclusions up into retrieval so the index never returns them, leaving only the order-dependent rules down here. The already-seen set is the awkward state on this path: it has to reflect an impression from seconds ago, because page two of a scroll must not repeat page one, so the impression write sits on the hot path, and it grows per user forever, which in practice means a TTL and an approximate structure. Approximate means it will sometimes suppress an item the user never saw. That is the right trade — just know you made it.

What the new pieces do

Filters & Rulesservice
Removes already-seen, blocked, or policy-violating items and applies diversity/business constraints.

Step 7 · Close the loop

Log engagement, retrain, repeat

The feed shipped. The user clicks some things, ignores others. That reaction is the single most valuable signal you have — it’s both the truth about what’s good and the fuel for tomorrow’s models. How do you use it?

Engagement LogTraining Pipeline
New in this step: Engagement Log, Training Pipeline. · swipe to pan the diagram

The user reacts to the feed. What do you do with that signal?

  1. Ignoring engagement freezes the system — it never learns the user’s shifting taste. The feedback loop is what keeps recs alive.

  2. Engagement events become training data and freshness signals: retrain the models periodically and refresh features so the system keeps improving.

  3. Without impressions you can’t tell "shown and ignored" from "never shown" — you lose the negatives the model needs to learn from.

Every impression and interaction is logged to the Engagement Log. A Training Pipeline periodically retrains the retrieval and ranking models on that data and refreshes the feature store — so the system learns from what worked and adapts as taste shifts. The loop is the product.

Why this piece earns its place

What you log is a schema decision you get to make roughly once. Every model, every feature and every offline evaluation downstream reads this stream, and anything not written at serving time cannot be recovered afterwards: which model versions produced this feed, which retriever the item came from, what position it occupied, what the ranker scored it. Without those columns you can attribute nothing — a metric moves and there is no way to say which of the three things that shipped that week moved it. Adding a field only helps from the day you add it, so an omission costs weeks of history. The second thing is that labels arrive late, and at different speeds. A click is immediate, a completed watch is minutes, a purchase that survives a return window is days. Join the log too eagerly and every slow positive is recorded as a negative, and the model learns to prefer exactly the fast, shallow engagement you were trying to move away from. Each label needs its own waiting period before its row is trainable, and that period — not the retrain schedule — sets how quickly the system can actually learn.

What the new pieces do

Engagement Logbus
Every impression and interaction, logged — the training data and freshness signal for the whole system.
Training Pipelinebus
Periodically retrains the retrieval and ranking models on logged engagement so they keep improving.

The payoff

You built a recommender

From "best ten of millions" to a self-improving funnel: two-tower retrieval, a heavy ranker, a consistent feature store, post-ranking rules, and a feedback loop that retrains on engagement.

UserRecs GatewayCandidate GenRanking ModelFilters & RulesItem IndexFeature StoreEngagement LogTraining Pipeline
The finished design, end to end. · swipe to pan the diagram

Now stale the feature store and watch relevance quietly rot — recommendations built on yesterday’s signals — and see why offline/online feature consistency is the system’s spine, not a detail.

Everything you assembled, in order

  • Retrieve → rank — cheap-wide recall, then expensive-narrow precision
  • Embeddings — taste as geometry — similarity becomes distance
  • Two-tower + ANN — live user vector × offline item index, in ms
  • Ranking Model — score P(engage) over hundreds with rich features
  • Feature Store — offline + online, consistent — no train/serve skew
  • Filters & Rules — dedupe, policy, diversity → a shippable feed
  • Feedback loop — log engagement, retrain, refresh — the flywheel
  • Freshness — stale features rot relevance with no error at all

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. A brand-new user opens the app with zero engagement history. What does two-tower retrieval actually do for them?

    There’s no learned user vector yet, so the system falls back to onboarding signals (explicit taste picks at signup, demographic/context priors) or simply popularity/trending candidates for the first session — the two-tower model only becomes meaningfully personalized once there’s enough engagement to compute a real user embedding. Cold-start users are usually served a DIFFERENT candidate-generation path entirely, not a degraded version of the personalized one.

  2. A brand-new item is uploaded with zero engagement. How does it ever get recommended if ranking is trained on historical engagement?

    Content-based signals (the item’s own embedding from its title/description/category, computed without any engagement data) let it enter candidate generation via similarity to items users already liked, and the exploration mechanism from the chaos scenario deliberately gives it some impressions despite having no track record — without either, a new item is invisible by construction, since a P(engage) ranking model trained on engagement history has nothing to score a zero-history item against.

  3. "Add some diversity" — concretely, how does a diversity rule get compared against a relevance score to decide the final order?

    A common approach is a penalty term subtracted from an item’s score based on similarity to items already placed higher in the list (maximal marginal relevance) — so the 3rd cooking video in a row takes an increasing diversity penalty even if its raw engagement score is high, until a different category’s next-best item overtakes it. It’s a tunable tradeoff (how much penalty per repeat), not a hard cap on category count.

  4. "Today’s watch count" needs to be fresh — how fresh is fresh in practice? Hourly batch, or true streaming?

    It depends on the feature’s decay rate: a feature like "watched in the last 5 minutes" (signals someone is mid-session on a topic right now) needs streaming updates to be useful at all, while "total watches this month" barely changes hour to hour and is fine on a batch cadence. The feature store’s batch+streaming split (mirrored in the feature-store system design) lets each feature pick the cadence its own volatility actually demands, rather than forcing one global freshness SLA.

  5. The ranking model predicts P(engage). How does a single score balance engagement against revenue and long-term retention, which can conflict?

    It usually doesn’t stay single — production rankers often predict multiple objectives (P(click), P(watch-to-completion), predicted revenue) and combine them with tunable weights into one final score, rather than training one model to secretly balance everything. This makes the tradeoff a product decision exposed as a weight you can dial (favor engagement this quarter, favor retention next), not a fixed thing baked irreversibly into training.

Check yourself — the answers, and why

Eight steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. Recommenders use a two-stage funnel because…

    Cheap retrieval narrows millions to hundreds; a heavy ranker orders those hundreds precisely. Each stage is sized for its job.

  2. In a two-tower model, item vectors are…

    The item tower runs offline over millions and is cached in an ANN index; only the user vector is computed at request time.

  3. The feature store’s offline and online halves must agree to avoid…

    If serving computes a feature differently than training did, the model scores on unfamiliar values and quality silently degrades.

  4. Why log impressions, not just clicks?

    Negatives (shown-and-ignored) are essential training signal; impressions provide them.

  5. Training only on engagement with items the system already recommended causes…

    Without deliberate exploration and off-policy correction, the model only ever learns from its own biased serving policy, and the effective catalog it recommends from quietly narrows.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Best few of millions: return a personalized feed from a catalog of millions in tens of milliseconds.
  • Two-stage funnel: cheap candidate generation narrows millions to hundreds, a heavy ranker orders those precisely.
  • Represent taste: learn user and item embeddings so similarity becomes distance in a shared space.
  • Score & filter: a heavy ranking model predicts engagement, then filters dedupe seen/blocked and add diversity.
  • Close the loop: log every impression and interaction, retrain the models, and refresh features.

The qualities that shape everything

Each one names the mechanism that buys it.

Both fast over millions and precise per item
A two-stage funnel: cheap retrieval narrows millions to hundreds, then a heavy ranker orders those hundreds — each stage sized for its job.
Retrieve candidates in milliseconds
A two-tower model: item vectors precomputed offline and indexed, the user vector computed live, meeting via an ANN search at request time.
Order candidates by real engagement
A heavy ranking model scores each candidate’s P(engage) using rich user/item/context features — affordable over hundreds, not millions.
No train/serve skew
A feature store with offline (batch) and online (low-latency) halves computed from one definition, so training and serving see identical values.
A scored list becomes a shippable feed
A filters & rules stage drops already-seen and blocked items, enforces diversity, and applies business rules after ranking.
Keep improving as taste shifts
Log every impression and interaction and retrain retrieval + ranking on it, refreshing the feature store — the flywheel.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

A two-stage funnel over scoring every item with the best model

Running a heavy model over millions per request is impossible in the latency budget — you shrink the set cheaply first, then rank precisely.

Learned embeddings over matching on shared tags/categories

Tags are coarse and miss latent taste (two thrillers can feel totally different); learned embeddings place similar users and items near each other so similarity is distance.

Two-tower + ANN over a dot-product against every item vector

Exact scoring over millions per request blows the budget; precompute item vectors offline and run an ANN index to pull the nearest few hundred in milliseconds.

A feature store (offline + online) over recomputing features inline per request

Heavy aggregations can’t be computed in the latency budget and ad-hoc computation drifts from what training saw — one definition serving both sides prevents train/serve skew.

Post-ranking filters and rules over returning the top-by-score list

A top-by-score list can repeat already-seen items, surface blocked content, and stack near-duplicates; post-ranking rules encode the product judgment a score can’t.

The answer, out loud

What a strong answer to “Design a Recommendation System” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Frame it as a budget problem

    Before I draw anything I want to state the problem as a budget. The catalog is millions of items, the user wants the ten they’d most like to see right now, and I have tens of milliseconds. Those three pull against each other — a model good enough to pick the ten is far too slow to look at the millions. So I’m not designing one system, I’m designing a shape that spends compute unevenly, and almost everything after this follows from that shape.

  2. 3–8 min

    Two stages, each sized for its job

    So: a gateway running a funnel on every request. Cheap retrieval takes millions down to a few hundred candidates; a heavy ranker orders those; the top handful ships. I’d put those three numbers up before naming a single model: the shape is the part I can’t change later, while the models inside it get swapped. The obvious alternatives, one model over the whole catalog or one popular list for everybody, I’d name and set aside in a sentence each rather than argue. I’d spend the time instead on the two stages answering different questions: was the right item in the set at all, and did it come out near the top.

    Built in step 1: A request and a funnel
  3. 8–14 min

    Make taste a distance

    For retrieval to be cheap, similarity has to be something I can compute geometrically, so I’d learn embeddings: users and items as vectors in one shared space, trained on past engagement, with a user landing near the items they’d actually engage with. The alternative I’d expect to be offered is matching on tags, and the trouble isn’t that tags are wrong — it’s that they describe the item and not its audience. A learned space can put two items beside each other because the same people like both, even when they share no words and no genre. That relationship is what I want to retrieve on, and nobody types it in by hand.

    Built in step 2: Embeddings for users and items
  4. 14–20 min

    Two towers meeting at an index

    Retrieval is a two-tower model. The item tower runs offline across the catalog and its vectors go into an ANN index; the user tower runs live, and at query time I embed the user and ask the index for the nearest few hundred items. The asymmetry is the trick — millions of vectors computed hours ago, one vector computed now. I’d be upfront that the search is approximate on purpose: it will occasionally miss a genuinely close item, and I’m fine with it, because I’m not asking this stage for an order, only for a set. Putting the set in the right order is the next thing I’d build.

    Built in step 3: Two-tower retrieval with ANN
  5. 20–26 min

    Rank the survivors properly

    Then the ranker, the stage that does the actual personalizing. Every surviving candidate gets scored by a heavy model predicting how likely this person is to engage — their history, the item’s attributes, the context of the request like time and device, and features built from the user and the item together. That model is unaffordable over millions and comfortable over a few hundred, which is the whole justification for the stage before it. The line I’d want heard is that the same person gets a genuinely different order on a phone at breakfast and on a television at night: personalization is a property of the request, not of the user.

    Built in step 4: The heavy ranking model
  6. 26–32 min

    Where the features come from

    Ranking is only as good as its features, and features are where this design quietly breaks. Something like how many cooking videos this user watched today can’t be recomputed inline inside the latency budget, and if I compute it one way for training and another way at serving, the model scores on values it never learned from. So: a feature store with an offline batch half that builds training data and an online half serving the freshest values at request time, both from a single definition. Train/serve skew is the failure I’d name unprompted, because nothing errors when it happens.

    Built in step 5: The feature store
  7. 32–38 min

    A score isn’t a feed

    The ranked list still isn’t shippable. The top items might be things the user already saw yesterday, or blocked content, or ten near-identical videos in a row. So there’s a filters and rules stage after ranking: drop already-seen and policy-violating items, enforce diversity, apply business rules like freshness or sponsored slots. I’d keep those deliberately outside the model. They’re product judgments that change weekly, and I don’t want them smuggled into a training objective where nobody can read them or switch one off.

    Built in step 6: Filters, rules and diversity
  8. 38–44 min

    Close the loop, and where it bites

    Then the loop, which I’d argue is the actual product. Every impression and every interaction is logged, a training pipeline retrains retrieval and ranking on that log, and the feature store gets refreshed. Impressions earn their place there as much as clicks — without them I can’t separate shown-and-ignored from never-shown, and the ignored ones are my negatives. It runs on two clocks, too: features refresh continuously, so this morning’s behaviour reaches ranking this morning, while models retrain offline on a schedule. Both get called freshness and they rot differently. And the data this system trains on is produced by this system, so it never gets evaluated against a fixed dataset the way the other models here would be.

    Built in step 7: Log engagement, retrain, repeat
  9. 44–45 min

    The trade-off, and what’s next

    To close, in one breath: a funnel, two-tower retrieval into an ANN index, a heavy ranker over the survivors, one feature store feeding training and serving, post-ranking rules, and a loop that retrains on logged engagement. The trade-off running through all of it is precision against the clock: every stage buys speed with an approximation, and it holds up only because each approximation gets corrected by something after it. With more time I’d go at position bias, since an item near the top earns clicks partly for being near the top, and at measurement, because deciding that a change to one stage improved the feed as a whole takes more machinery than I’ve described here.

What this teaches

Learn AI system design by building a large-scale recommendation system step by step. An interactive guide covering the two-stage retrieve-and-rank funnel, two-tower embedding retrieval with ANN, a heavy ranking model, the feature store and its offline/online split, filtering business rules, and the feedback loop that keeps recommendations fresh.

Key takeaways

  • Retrieve → rank — cheap-wide recall, then expensive-narrow precision
  • Embeddings — taste as geometry — similarity becomes distance
  • Two-tower + ANN — live user vector × offline item index, in ms
  • Ranking Model — score P(engage) over hundreds with rich features
  • Feature Store — offline + online, consistent — no train/serve skew
  • Filters & Rules — dedupe, policy, diversity → a shippable feed
  • Feedback loop — log engagement, retrain, refresh — the flywheel
  • Freshness — stale features rot relevance with no error at all

Concepts covered

  • Pick the best few from millions
  • A request and a funnel
  • Embeddings for users and items
  • Two-tower retrieval with ANN
  • The heavy ranking model
  • The feature store
  • Filters, rules and diversity
  • Log engagement, retrain, repeat
RUN IT YOURSELF

The click-through rate climbs while the catalog it draws from collapses

This runs the flywheel of step 7 for real — recommend, log impressions and clicks, retrain, twelve times — on a catalog that grows the way a real one does. It runs the loop twice on the identical world with the identical coin on every impression, once serving purely what the model already believes and once giving up one slot in six to items the log cannot yet judge, then prints what each arm’s coverage and true match rate did. Change EXPLORE_SLOTS, hit Run, and watch the 19 genuinely best items exploitation never showed come back into view.

HOW TO READ THE CODE — 5 IDEAS
  1. Pure exploitation’s measured CTR climbs 0.167 → 0.525 while its true match rate peaks at 51% in round 5 and falls back to 34%. The dashboard is not lying — it only ever measures the items you chose to show (step 7).
  2. It serves 41 of 170 items across all twelve rounds against exploration’s 141. Split by arrival date the bias is plain: it shows 26 of the 60 launch items but only 15 of the 110 uploaded later, and 19 of the 30 items that belong in somebody’s true top six are never shown to anybody. Exploration misses none of them.
  3. The cause is one line: an item with zero impressions sits exactly on its content prior for ever, capped at a lift of 1.5, while a proven item’s learned lift passes 3.0. No constant punishes the newcomer; a neutral prior is enough (step 2).
  4. Reserving slots is not enough on its own. Set BONUS = 0 and the exploring arm becomes byte-identical to the greedy one — spread evenly over 170 items, exploration never gives any single item enough impressions to overturn its prior. It has to rank by an upper bound (steps 4 and 5).
  5. Change SEED to re-roll the whole world without touching the mechanism. The coverage collapse holds in all 13 seeds tried (29–45 of 170, against 114–149); the size of the true-match fall does not — it ranges from 0 to 18 points. The collapse is the number to trust.
CPython · WebAssembly
RUN IT YOURSELF

Item-based collaborative filtering

"Because you liked X" works by finding items similar to ones you already liked — item-item cosine similarity over a rating/feature matrix. Here it is in real Python, running live. Read the comments, edit the catalog, and hit Run.

HOW TO READ THE CODE — 4 IDEAS
  1. Each item is a feature vector (here: how action / comedy / romance it is).
  2. Two items are similar if their vectors point the same way — cosine similarity.
  3. Score each unseen item by its best similarity to something you liked (steps 1–2).
  4. Recommend the top-k, never re-suggesting what you have seen (step 3).
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, stale the features, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs