The whole design, in writing
Learn AI system design by building a retrieval-augmented generation (RAG) pipeline step by step. An interactive guide covering document chunking, embeddings, the vector store, top-k retrieval, reranking, grounded prompt assembly, and keeping the index fresh — so an LLM answers from your knowledge, not its training data.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
Why retrieval at all?
An LLM only knows what it was trained on. Ask it about your internal docs, last week’s release, or a private wiki and it confidently makes something up. How do you make it answer from knowledge it never saw?
Retrieval-Augmented Generation: before generating, fetch the most relevant passages from your knowledge and put them in the prompt. The model reasons over real, current text instead of fuzzy memory — and can cite its sources.
What the new pieces do
- Userclient
- A person asking a question that should be answered from your documents, not the model’s training data.
Step 1 · The skeleton
Retrieve, then generate
A user asks a question. We can’t just forward it to the LLM — it would answer from training data. What has to happen between the question and the answer?
The user asks a question about your private docs. What’s the flow?
It has never seen your documents, so it answers from training data — confidently and often wrong. This is exactly what RAG exists to fix.
A gateway orchestrates retrieve-then-generate: find the chunks that matter, hand them to the model as context, and let it answer from real text.
Fine-tuning bakes knowledge in slowly and expensively, can’t cite sources, and goes stale the moment a doc changes. Retrieval keeps knowledge live and external.
A RAG Gateway orchestrates two phases for every question: retrieve the most relevant passages, then generate an answer grounded in them. Knowledge lives outside the model, so it’s always current and citable.
Why this piece earns its place
The gateway looks like plumbing, and skipping it is the most common shortcut: retrieval is three function calls, so the app makes them itself. What you lose is the only place that can degrade. Retrieval has parts that fail independently — the vector store times out, the embedding API rate-limits, the reranker is slow today — and each failure has a different right answer: serve a cached response, answer ungrounded with an explicit warning, or refuse outright. That policy has to live somewhere, and if it does not live in one component then every caller invents its own version of it, inconsistently. The gateway is also where the per-request budget lives: how many milliseconds retrieval gets before generation starts without it, how many tokens the context may consume, which tenant's documents this request is allowed to see. None of those are model decisions and none of them are UI decisions. Build it as a seam from the start, even when it is thin, because the second consumer arrives sooner than you expect.
What the new pieces do
- RAG Gatewaybackend
- Coordinates the retrieve-then-generate flow: embeds the query, fetches chunks, builds the grounded prompt, calls the model.
Step 2 · Prepare the knowledge
Chunk and embed your documents
Your knowledge is a pile of long documents. You can’t hand a whole 80-page PDF to the model per query. How do you make documents searchable by meaning?
How do you store documents so you can find the relevant bits by meaning?
Keyword search misses paraphrase — "how do I cancel" won’t match "terminating your subscription." Meaning needs semantic vectors, not exact words.
Whole docs are too big for the context window and too coarse to pinpoint the relevant passage. You need to split first.
An offline pipeline chunks documents, embeds each chunk into a vector capturing its meaning, and upserts vectors + text into the stores. Now "find by meaning" is a nearest-neighbour search.
An offline Ingestion pipeline loads each document, splits it into chunks, runs every chunk through an Embedding Model to get a vector, and upserts the vectors into a Vector Store (plus the raw text into a Document Store). This is the prep work that makes retrieval possible.
Why this piece earns its place
Ingestion reads like a script you run once, and that framing is what breaks RAG three months in. It is a stateful, versioned system, and the state is the hard part. Chunk ids have to be deterministic — derived from source id plus position, not a fresh UUID per run — or re-ingesting a lightly edited document appends a second copy instead of updating the first, and retrieval starts returning both the old and the new wording of the same policy. Deletes are worse: remove a document at the source and nothing removes its vectors, so the pipeline keeps confidently citing a page that no longer exists. Tombstones or a reconciliation sweep are not optional. This is also where the money is. Embedding is charged per token across the whole corpus, not per query, so the expensive operation is the backfill — a re-chunk, an embedder upgrade, a new document set — while steady-state ingestion is cheap. The upside: it is offline work, so unlike query latency you can schedule it, batch it, and run it on spare capacity.
- ~300–800tokens / chunk
- 1 vectorper chunk
- offlineruns ahead of time
What the new pieces do
- Vector Storeindex
- Holds chunk embeddings and answers nearest-neighbour search in milliseconds via an ANN index.
- Document Storestore
- The original chunk text and metadata, returned alongside vectors so the prompt and citations use real content.
- Ingestionbus
- The offline pipeline: load documents, split into chunks, embed each, and upsert into the vector + document stores.
- Knowledgestore
- The source of truth: PDFs, docs, wikis, tickets — whatever the system must answer from.
- Embedding Modelbus
- Maps text to vectors. The SAME model must embed both stored chunks and incoming queries so they share a space.
Back of the envelope
- split on structure
- paragraphs/sections, not arbitrary character counts
- add overlap
- ~10–15% so ideas spanning a boundary aren’t cut in half
- same embedding model
- for chunks AND queries — they must share a space
- store metadata
- source, title, URL — for filtering and citations
Step 3 · Find the right chunks
Embed the query, search by similarity
A question comes in. The chunks are vectors in a store. How do you find the handful that actually answer this question — out of possibly millions?
You have millions of chunk vectors. How do you find the closest to the query?
Exact nearest-neighbour over millions of high-dim vectors is too slow per query. At scale you trade a little accuracy for huge speed.
Embed the query with the SAME model, then let an ANN index (HNSW/IVF) return the closest chunks in milliseconds — near-exact recall, sub-linear cost.
That’s back to lexical search and its paraphrase blind spot. Semantic retrieval is the whole point. (Hybrid keyword+vector is a refinement, not a replacement.)
The Retriever embeds the query with the same model, then asks the Vector Store for the top-k nearest chunks via an ANN index. The Document Store returns their text. Semantic match means "cancel my plan" finds "terminating your subscription."
Why this piece earns its place
Approximate search is the one component whose failure you cannot see from production traffic. A missing chunk raises no error, adds no latency and fills no dashboard — the answer just quietly draws on the second-best passage. So recall has to be measured off traffic, against a held-out set of queries with known-correct chunks, re-run whenever the index, the embedder or the ingest rules change. Without that set, "we tuned the index" is a claim nobody can check. The knobs are worth knowing by name. HNSW holds its graph in memory, so the index is sized in RAM per million vectors rather than on disk, and that is usually the real cost ceiling. Build and query cost are asymmetric: the graph is expensive to construct and cheap to search, which pushes you toward rebuilding on a schedule rather than continuously. And the search-effort knob (ef_search in HNSW, nprobe in IVF) is set per query, not per index — so a request that fails a confidence check can be retried wider before you give up on it.
What the new pieces do
- Retrieverservice
- Turns the query into a vector and searches the store for the nearest chunks (semantic, not keyword, match).
Step 4 · How much to fetch
Top-k and the recall/precision dial
Retrieve too few chunks and you miss the answer. Retrieve too many and you bury the model in noise, blow the token budget, and slow it down. Where’s the line?
How many chunks should you stuff into the prompt?
More context isn’t better — irrelevant chunks distract the model ("lost in the middle"), raise cost, and slow generation. Precision beats volume.
Too brittle: the answer often spans two or three passages, and the very top hit isn’t always the right one. You need a small set, not a single bet.
Retrieve a modest candidate set (e.g. k≈20), then narrow to the best few. Enough recall to catch the answer, enough precision to keep the prompt clean.
Retrieve a modest top-k candidate set from the Vector + Document stores — wide enough that the answer is almost certainly in there (recall), but not so wide it drowns the prompt. The next step tightens it to the best few (precision).
Why this piece earns its place
k is the only dial on this page that moves latency, cost and answer quality at the same time, which is exactly why it belongs in the request as a parameter, not a constant. Query classes differ: a lookup with one right answer ("what is the refund window") is served by a handful of chunks, while a synthesis question ("how has this policy changed") genuinely needs a wide net, and one hardcoded k serves one of them badly. Route by query type, or start narrow and widen when confidence comes back low. The other thing worth saying out loud is that there are two budgets here and they get confused constantly. The candidate set — k≈20 — is read only by the reranker, a small model that does not much care about volume. The prompt budget is what the generator reads, and it is the one that costs real money per request and degrades with noise. Growing k spends reranker time; growing the post-rerank shortlist spends generation quality. Naming which budget you are spending is most of the argument.
Back of the envelope
- k ≈ 20 candidates × ~500 tok/chunk (mid-range of Step 2’s ~300–800) ≈ 10,000 tok
- the retrieval-stage set — only the reranker reads this, not the generator
- → rerank → top 3–5 × ~500 tok ≈ 1,500–2,500 tok
- what actually lands in the prompt
- on an 8K-context model, worst case that leaves ≈ 5,500+ tok
- room for the system prompt, conversation history, and the answer
- optional metadata filter
- restrict by source/recency before ranking
Step 5 · Sharpen the shortlist
Rerank for precision
The vector search is fast but coarse — it ranks by embedding similarity, which isn’t the same as "actually answers the question." The true best chunk might sit at position 8. How do you fix the order?
ANN similarity ≠ true relevance. How do you get the best chunks to the top?
Bi-encoder similarity is a fast approximation. The genuinely most relevant passage often isn’t the nearest vector — order needs a second, sharper pass.
A cross-encoder reads the query and each chunk together and scores true relevance — far more accurate. Run it on the ~20 candidates (cheap), keep the best 3–5.
More candidates raises recall but not precision — you still feed the model noise. Reranking is what turns a long candidate list into a clean shortlist.
A Reranker (a cross-encoder) re-scores the top-k candidates by reading the query and each chunk together, then keeps only the best 3–5. It’s too slow to run over the whole store — but perfect over a 20-candidate shortlist.
Why this piece earns its place
The reranker buys two things and the second one is the one people miss. The obvious one is precision. The quiet one is a usable score. Bi-encoder cosine similarity is not calibrated and not comparable across queries — 0.82 is high for one question and mediocre for another, because it depends on where that query happens to land in the embedding space. You cannot threshold on it. A cross-encoder reads the pair together and produces a relevance score you can calibrate against labelled data, which is what makes Step 7's "say I don't know" mechanically possible at all; without it the decline rule has nothing trustworthy to fire on. The reranker is also the cheapest place to buy quality. Swapping it is a config change that takes effect on the next request, whereas swapping the embedding model invalidates every stored vector and means a full reindex of the corpus. So when retrieval quality is not good enough, reach for the reranker first and treat the embedder as the change you make deliberately and rarely.
What the new pieces do
- Rerankerservice
- Re-scores the top candidates with a heavier cross-encoder so the few chunks that reach the prompt are the most relevant.
Step 6 · Ground the answer
Build the prompt and generate
You have the 3–5 best chunks. Now the model has to answer — but you need it to use those chunks and not slip back into making things up. How do you assemble the prompt?
You have the best chunks. How do you get a grounded, citable answer?
If the chunks aren’t in the prompt, the model can’t use them. Grounding only works when the retrieved text is actually in the context.
The gateway builds a grounded prompt: "answer ONLY from these passages, cite them, say you don’t know if they don’t cover it" — then generates.
Back to noise and cost. You reranked for a reason — feed the clean shortlist, not the raw candidate pile.
The Gateway assembles a grounded prompt — system instructions + the reranked chunks + the question — and tells the LLM to answer only from the passages, cite them, and admit when they don’t cover the question. The chunks carry their source metadata, so citations point at real documents.
Why this piece earns its place
Prompt assembly is the step most likely to be written as string concatenation and never revisited, and it carries a real contract. Chunks need unambiguous delimiters and a stable id printed beside each one, because "cite your sources" only works if the model's citation can be parsed back to a chunk and verified — an answer naming a plausible document title is not a citation, it is a hallucination wearing one. Order matters too: models attend unevenly across a long context, so put the highest-ranked chunk where it will actually be read, and decide deliberately whether that is first or last for your model. Truncation needs a policy — drop whole chunks by rank, never tail-cut the last one, or you ship half a sentence as evidence. And log the assembled prompt alongside the answer. Without it a bad answer is undebuggable: you cannot tell whether retrieval never found the right passage or the model had it and ignored it, and those two have completely different fixes.
What the new pieces do
- LLMservice
- Generates the answer using only the retrieved chunks as context, with instructions to cite and not invent.
Step 7 · Trust the answer
Citations, "I don’t know", and evals
Even grounded, the model can over-claim or cite the wrong chunk. Users need to trust the answer — and you need to know when retrieval is failing. How do you keep RAG honest?
How do you keep a RAG system trustworthy over time?
Grounding reduces hallucination but doesn’t eliminate it — the model can still misread or over-generalize from a chunk. Trust must be verifiable.
Citations are the trust mechanism — they let a user verify the claim against the source. Hiding them removes the one thing that makes RAG auditable.
Surface the source chunks as citations, let the model decline when context is thin, and run evals (retrieval recall, answer faithfulness) to catch drift.
The answer ships with citations back to the source chunks, the model is allowed to say "I don’t know" when retrieval comes up thin, and an eval harness tracks retrieval quality (did we fetch the right chunks?) and answer faithfulness (did the answer stick to them?). That’s how you catch a silently drifting index before users do.
Why this piece earns its place
Two metrics, not one, and the reason is that the halves fail independently. A single end-to-end "was that a good answer" score tells you something regressed and nothing about what to do, because retrieval recall and answer faithfulness point at opposite repairs. Recall dropped: reindex, revisit chunking, widen k. Faithfulness dropped while recall held: the model had the right passage and drifted, so the fix is in the prompt or the model, and touching the index would waste a week. The eval set itself has to be frozen and versioned separately from the corpus, or you cannot distinguish a real regression from the ground truth moving underneath you — which is the failure mode of golden sets that get "refreshed" whenever they start failing. Offline evals are slow, so pair them with online signals that move daily: the rate at which the system declines to answer, whether users open the citations, thumbs-down clustering by source. None of those are ground truth. All of them tell you where to point the offline set next.
The payoff
You built a RAG pipeline
From a forgetful model to a grounded, citable system: an ingestion pipeline that chunks and embeds, a vector store for semantic search, two-stage retrieval with reranking, grounded generation, and evals to keep it honest.
Now stale the index and watch RAG’s signature failure — confident answers built on outdated chunks — and see why ingestion is a first-class, ongoing system, not a one-time script.
Everything you assembled, in order
- RAG Gateway — orchestrates retrieve → augment → generate
- Ingestion — chunk + embed documents, offline and ongoing
- Embeddings — meaning as coordinates — same model for chunks + queries
- Vector Store — ANN search returns top-k by similarity in ms
- Top-k — wide for recall, then narrowed for precision
- Reranker — cross-encoder sharpens the shortlist to the best few
- Grounded prompt — answer only from context, cite, decline if thin
- Evals + citations — make a quiet failure mode visible and auditable
