The whole design, in writing
Learn system design by building a social news feed step by step. An interactive guide covering the social graph, pull vs push fan-out, a precomputed feed cache, ranking, the hybrid model for celebrities, mixing content sources, and pagination.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
What is a news feed?
Open the app and you see a personalized stream of posts from everyone you follow, freshest and most relevant on top. Simple to use — but building one reader’s feed means merging posts from hundreds or thousands of authors, ranking them, in a few hundred milliseconds, for billions of readers.
The whole game is when you do the work: merge everyone’s posts at read time (simple, slow) or pre-build each feed at write time (fast reads, heavy writes). We’ll start naive and evolve into the hybrid that real feeds use.
What the new pieces do
- Authorclient
- Anyone who publishes. A normal user has hundreds of followers; a celebrity has millions — and that difference reshapes the whole design.
- Readerclient
- Pulls down to refresh and expects a fresh, ranked feed in a few hundred milliseconds. Reads vastly outnumber posts.
Step 1 · The skeleton
Post and read
Two operations: an author publishes a post, and a reader asks for their feed. Both go through one place, and the post content has to live somewhere durable. How should feeds refer to posts?
A post can appear in millions of feeds. How should those feeds store it?
Now an edit or delete must hunt down millions of copies, and storage explodes. Duplicated content turns every change into a distributed cleanup problem.
One durable copy in the Post Store; feeds hold only IDs and hydrate bodies at read time. Edits, deletes and privacy changes resolve correctly everywhere, for free.
That’s pure pull — every reader re-queries every author they follow on each refresh. Simple, but it makes the hot read path do all the work (the next step’s problem).
Stand up a Feed API with a Post Store behind it. Writing a post saves a row; reading a feed will (for now) go figure out what to show. This is the spine everything else hangs off.
Why this piece earns its place
Storing IDs is the right call, but notice that it moves the join rather than removing it. Every feed read now ends in a multi-get against the Post Store, so the store that looks like cold durable storage is really the hottest read path in the system — touched once per post per reader. That means hydration has to be one batched multi-get per page, never a lookup per ID, and it wants its own read-through cache of post bodies in front of it. That cache has the opposite shape to the feed cache: a feed is one entry per reader, a post body is one entry shared by everyone who sees it, so the wildly popular content this design struggles hardest to fan out is exactly the content a body cache serves best. And build the read for partial resolution from day one — a page where a few IDs come back with nothing should render the rest, because a missing body is the normal expected outcome of a delete, not an error.
What the new pieces do
- Feed APIbackend
- The entry point for both posting and reading. It assembles a reader’s feed and accepts an author’s new posts.
- Post Storestore
- The durable home of post bodies, media references and metadata. Feeds store only IDs; this is where the actual content is hydrated from.
Step 2 · The naive read
Pull (fan-out on read)
To build a feed on demand, you’d look up everyone the reader follows in the Social Graph, fetch each one’s recent posts, merge and sort them — every single time they refresh.
This pull model is dead simple and always fresh, and writes are trivial (just save the post). But a reader following 2,000 accounts triggers a 2,000-way query and merge on every open — far too slow at scale.
Why this piece earns its place
The number that kills pull is not the two thousand fetches, it is which one of them is slowest. A merged read cannot return until its last fetch does, so the reader waits on the maximum across two thousand requests, not the average — the single author whose shard happens to be compacting sets the response time for that refresh. Averages look healthy right up to the moment the p99 of one dependency becomes the p50 of the feed. If you had to ship pull, you would cap it: a hard per-author deadline, hedged requests for stragglers, and return the merge you have rather than wait for completeness. That is worth building properly even though the next step replaces it, because pull does not die here. Step 5 hands it the celebrities, and it is what the system falls back on when the feed cache is gone. A fallback path nobody maintains is the one that fails during the incident it exists for.
- 2,000follows to merge
- everyrefresh re-queries
- freshbut slow
What the new pieces do
- Social Graphservice
- Stores the follow edges. To build a feed you first need to know whose posts a reader should even see.
Back of the envelope
- follows 2,000 × recent posts
- a 2,000-way fetch + merge on every single refresh
- reads ≫ writes (~100:1)
- paying the big cost on the common operation is backwards
- writes are O(1)
- posting is just one insert — pull’s one redeeming virtue
Step 3 · Flip the work
Push (fan-out on write)
Reads are the hot path, yet pull does the most work there. We want reading a feed to be a single cheap lookup — which means the answer must already exist before the reader asks. When do you build the feed?
Reads dominate but pull does its heaviest work on reads. How do you make a feed read O(1)?
It helps a little, but the first reader still pays the full 2,000-way merge, and the cache is cold or stale exactly when new posts arrive. You’re patching pull, not fixing it.
Replicas scale raw query throughput but every read still does the expensive multi-author merge. You’ve made a slow operation parallel, not cheap.
When an author posts, push the ID into every follower’s feed cache. Reading is then O(1) — grab your precomputed list. Work moves from many reads to one write.
When an author posts, Fan-out Workers push the post ID into the Feed Cache of every follower. Reading is now O(1): grab your precomputed list of IDs and hydrate the bodies. Work moved from many reads to one write.
Why this piece earns its place
Push needs a bound or it stops being a cache and becomes a second copy of the corpus: a feed that grows forever costs memory per follower per post, multiplied by every reader who is still active. So each cached feed is trimmed to a fixed window of recent IDs, and that window is the real knob here. Keep it short and memory is cheap, but a reader who scrolls past the end falls off the precomputed list and has to be served the slow way — you have handed a rare but real request back to the pull path you just tried to retire. Make it deep and every extra ID is paid for once per follower, which is the same amplification this step already charges against write throughput, now charged against RAM as well. I would set it from how far readers actually scroll in a session and accept that the tail goes to pull. The same bound decides who gets a feed maintained at all: pushing into the cache of an account that has not opened the app in a year is pure waste, so dormant readers are better rebuilt on their next visit.
- O(1)feed read
- precomputedper follower
- ×followerswrite cost
What the new pieces do
- Feed Cachecache
- A ready-made list of post IDs per user. Reading the feed becomes a single fast lookup instead of a live query across everyone they follow.
- Fan-out Workersworker
- When someone posts, these workers push the post ID into each follower’s feed cache — doing the join once at write time, not on every read.
Back of the envelope
- 1 post × F followers
- F cache inserts at write time — the amplification
- typical F ≈ hundreds
- cheap: a few hundred tiny ID inserts per post
- celebrity F ≈ 50M
- 50M inserts for one post — push collapses here (step 5)
Step 4 · Newest isn’t best
Ranking
A purely chronological feed buries the post you’d most want to see under noise. Engagement craters when relevance is left to luck and timestamps. How do you order the feed?
A chronological feed buries the best post. How do you order what the reader sees?
Simple and predictable, but a single chatty account drowns the post you actually care about. Recency alone is a weak proxy for relevance.
That hands ranking to the people with the most incentive to game it — every post becomes "top priority". Ordering must be decided by the platform, from signals, not by the author.
Two stages: cheaply gather candidates, then a model orders just those by how close you are to the author, freshness and likely engagement. The same retrieve-then-rank pattern powers search.
Run candidates through a Ranking stage that scores each by affinity (how close you are to the author), recency, and predicted engagement. The feed cache holds candidates; ranking decides their order at (or near) read time.
Why this piece earns its place
This step quietly assumes two things: that engagement is a decent proxy for value, and that the model sees a fair sample. Neither is free. The training data is the feed the ranker itself produced, so it only ever observes outcomes for posts it chose to show — the candidates it buries generate no evidence that burying them was wrong, and the model grows confident in its own habits. It shows up as a feed that narrows: the same few authors, the same formats, engagement flat or even up while the thing gets duller. The defenses all cost you slots on purpose. Reserve part of the page for candidates the model is unsure about, cap how many posts one author may hold so no score can monopolize a page, and keep a held-out chronological group so ranking is measured against something that is not itself. Worth saying out loud too: the ranker can only order what retrieval handed it, so when a feed feels empty the fault is usually upstream in the candidate set, not in the scoring.
What the new pieces do
- Rankingservice
- Scores candidate posts by relevance — affinity, recency, engagement — so the feed shows what matters, not just the newest thing.
Step 5 · The celebrity problem
Go hybrid
Push breaks for the mega-popular: one post by a celebrity with 50M followers means 50M cache writes — a thundering, slow, expensive fan-out that also wastes effort on inactive followers. What do you do for them?
A celebrity with 50M followers posts. Push to all 50M feeds, or something else?
50M cache writes per post, per celebrity, including for followers who’ll never open the app. You can parallelize it, but it’s enormous wasted work on the write path.
Now every ordinary reader pays the slow multi-author merge again — you’ve thrown away the O(1) reads that push bought you, to fix a problem only celebrities have.
Skip fan-out for the few mega-accounts and merge their recent posts into the precomputed feed on read. The merge set is tiny (a handful of celebs) so reads stay fast and writes stay sane.
Go hybrid. Push for ordinary authors; for celebrities, skip fan-out and let readers pull their recent posts at read time, merging them into the precomputed feed. Route new posts through an event stream so fan-out can absorb spikes and run async.
Why this piece earns its place
The hybrid sorts authors into two classes, and the sharp edge is that the class is not stable — follower counts drift across the cutoff, sometimes overnight. If the read path decides how to treat an author by checking their status now, a post that was published while they were ordinary and pushed into feeds gets pulled as well the moment they cross the line, and the same post lands twice in one feed. Flip the timing and you get the opposite: a post that skipped fan-out because the author was above the line, then was never pulled because the count slipped back below it, is simply absent, and nothing logs its absence. So the routing decision belongs on the post, stamped at publish time, rather than re-derived at read time from a number that moves. The merge should be idempotent on post ID regardless, because a reclassification part-way through fan-out leaves feeds where some followers got the push and some did not — and ID-level dedup is something the scroll needs a couple of steps from here anyway.
- pushnormal authors
- pullcelebrities
- mergeat read time
What the new pieces do
- Post Eventsbus
- A stream of new posts. Fan-out, search indexing and analytics all consume it independently — and it absorbs celebrity-sized spikes.
Back of the envelope
- a 10k-follower author ⇒ 10k inbox writes per post
- expensive, but bounded — push is still the right call
- a 50M-follower account ⇒ 50M writes for that same single post
- 5,000× the cost of the author above — this one number is why the cutoff exists
- so push below ~10k followers, pull above it
- the threshold is not a convention; it is where write cost stops being survivable
- pull the few thousand mega-accounts
- their posts are fetched and merged on read
- read = precomputed feed + N celeb pulls
- N is single digits, so reads stay fast
Step 6 · More than friends
Mixing sources
A modern feed isn’t only people you follow — it’s followed pages, recommended posts you don’t follow yet, and ads. These come from different systems but must feel like one coherent stream. How do you combine them?
Friends, pages, recommendations and ads all need to share one feed. How?
That’s not a feed, it’s a dashboard — and it ignores relevance across sources. A great recommended post should be able to outrank a dull friend post, which fixed sections forbid.
Friends, pages, recs and ads each emit candidates; a shared ranker scores and interleaves them under one policy (with ad-spacing and diversity rules). Adding a new source is a new generator, not a new feed.
Phones can’t see global ranking signals, the merge logic forks across platforms, and you ship ranking secrets to the client. Blending belongs on the server, in one place.
Gather candidates from each source — friends (feed cache), pages, recommendations, ads — then rank and interleave them under one policy, with rules for ad spacing and diversity so the feed doesn’t clump.
Why this piece earns its place
One shared ranker only works if the candidate generators speak the same language, and out of the box they do not. A friend post arrives carrying a predicted engagement probability, a recommendation carries a relevance score from a different model trained on different data, an ad carries a bid in money. Ordering those by raw score is not ranking, it is whichever team’s numbers happen to come out largest. What makes the blend honest is a common currency: every generator emits a calibrated probability, meaning that when it claims something is likely it happens about that often, and each candidate is then scored as that probability times what the outcome is worth — which is how an ad’s bid enters the same arithmetic as a friend post without a thumb on the scale. Keep the policy out of the score, too. Ad spacing and per-source diversity belong as constraints applied to the final ordering, not as bonuses added to candidates, because a bonus big enough to guarantee a slot has no ceiling and will eventually outrank the things it was only meant to sit beside.
What the new pieces do
- Other Sourcesstore
- Followed pages, recommended posts and ads. The final feed blends these with friends’ posts under one ranking.
Step 7 · The sharp edges
Pagination, dedup & freshness
Readers scroll for ages, so feeds must page without showing the same post twice or skipping new ones that arrive mid-scroll. And a post deleted after fan-out shouldn’t haunt a million feeds. How do you page an ever-changing feed?
Infinite scroll, with posts constantly inserted and deleted. How do you paginate?
When new posts arrive at the top, every offset shifts — page 2 now repeats items from page 1, or skips some. Offset pagination and a live feed don’t mix.
A stable cursor marks your position so inserts can’t cause dupes or gaps; dedup drops already-seen IDs; and because feeds store IDs, hydration re-checks the Post Store so deleted/edited posts resolve correctly at read time.
Refetching everything is wasteful, janky, and still doesn’t give a stable scroll position — the reader keeps losing their place as the top churns.
Use cursor-based pagination (a stable position, not page numbers) so inserts don’t cause dupes or gaps. De-dupe already-seen IDs per session. Because feeds store IDs, hydration re-checks the Post Store — so deleted or edited posts resolve correctly at read time.
Why this piece earns its place
Per-session dedup is the one piece of this design that has to remember something, and memory on an otherwise stateless read path is where it gets expensive. The seen set grows with the session, and a determined scroller’s session is long. Hold it server-side and you have added per-session state to every feed request plus a cleanup job for the sessions that simply stop. Carry it in the cursor and the cursor grows with every page, until a fat token is travelling up and back on each request. So bound it deliberately — a fixed window of recent IDs, which forgets the distant past on purpose, or a compact probabilistic membership filter, which bounds the space instead and pays for it in accuracy as it fills. Then choose which way it is allowed to be wrong, because a compact filter errs by claiming it has seen something it has not, and that post is silently dropped. A repeat the reader can see and shrug at is a far cheaper mistake than a post nobody ever learns was suppressed, so tune it to err toward showing the dupe.
You did it
You just designed a news feed.
Everything you assembled, in order
- A Feed API over a Post Store; feeds reference posts by ID, never copy bodies.
- Pull (fan-out on read) is simple and fresh but slow for big follow counts.
- Push (fan-out on write) precomputes feeds for O(1) reads at write-time cost.
- A ranking stage orders candidates by affinity, recency and engagement.
- A hybrid model pushes for normal authors and pulls celebrities to dodge mega-fan-out.
- Multiple candidate sources (pages, recs, ads) blended under one ranker.
- Cursor pagination, per-session dedup and ID hydration keep an infinite scroll correct.