Vibe Engines
YouTube
AI System Design

Design an AI Coding Assistant

Learn AI system design by building an AI coding assistant (like Copilot or Cursor) step by step.

The numbers to beatfill-in-the-middleprefix + suffixopen tabs + importslocal contexttoken budgetrank by relevance

The whole design, in writing

Learn AI system design by building an AI coding assistant (like Copilot or Cursor) step by step. An interactive guide to sub-second inline completion / tab prediction — context assembly with fill-in-the-middle, repo-aware retrieval, a specialized code model, latency tricks (debounce, cancel, cache), acceptance-rate telemetry, an agentic chat mode, and privacy/secret filtering — plus why starving the context produces confident, wrong code.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

Autocomplete that understands your code

A coding assistant suggests the next lines as you type. To feel magical it must be three things at once: fast (a suggestion in a few hundred milliseconds, or the developer has moved on), context-aware (it knows this file, this repo, these imports), and right enough that accepting saves time instead of introducing bugs. Those three fight each other — more context and a bigger model mean more latency. How do you build inline completion that’s fast and grounded?

An AI coding assistant is a latency-obsessed retrieval + inference system: assemble the tightest useful context (cursor prefix/suffix, open tabs, retrieved repo snippets), feed a code-specialized model doing fill-in-the-middle, and stream a suggestion back — all under a strict time budget, measured by whether developers accept it.

Step 1 · The skeleton

Keystroke to suggestion, on a deadline

The developer types; you have a few hundred milliseconds to show a useful completion before it’s stale. What’s the request path, and what makes it different from a normal API call?

EditorIDE pluginCompletion Svcorchestrator
New in this step: Editor, Completion Svc.

What defines the shape of an inline-completion request?

  1. Blindly blocking means slow suggestions arrive after the developer has typed past them. Completion is latency-bounded and cancellable — stale work must be abandoned.

  2. The editor debounces keystrokes, sends context to a completion service, and streams back a suggestion — cancelling in-flight requests the instant the developer keeps typing so only fresh work completes.

  3. Firing (and finishing) a request per keystroke floods the model and wastes compute on suggestions already invalidated. You debounce and cancel.

The Editor debounces keystrokes and sends the cursor context to the Completion Service, which streams a suggestion back under a tight latency budget and cancels in-flight requests as soon as the developer types more. Everything downstream is engineered around that deadline — a completion that’s late is worthless.

Why this piece earns its place

The cancel in this loop is easy to describe and easy to fake. Cancelling in the editor stops the ghost text appearing; it does not stop the tokens being generated. Unless the abort is propagated all the way into the inference request, you keep paying in full for suggestions nobody will ever see. It is worth being precise about what counts as invalidation, too, because a loop that only watches keystrokes is wrong in the ordinary cases: the developer switches file, presses undo, or moves the caret with the mouse, and the request in flight now describes a position that no longer exists. Then the measurement trap, which is what hides all of this. Latency measured inside the Completion Service excludes every request you cancelled, so the percentile on the dashboard looks excellent while the developer sits watching nothing appear. The number that matters is client-side: from the typing pause to painted ghost text, counting the debounce, the network and the render. Carry a per-document sequence number as well, so a response that wins the race against its own cancellation is still discarded on arrival — cancellation is best-effort, but validating at arrival is not.

What the new pieces do

Editorclient
The IDE extension. It watches the cursor, debounces keystrokes, requests completions, renders ghost-text suggestions, and reports whether the developer accepted them.
Completion Svcbackend
Orchestrates a completion under a tight latency budget: assemble context, call the model, cache, and stream a suggestion back — cancelling in-flight work the moment the developer keeps typing.

Step 2 · What to send

Context assembly and fill-in-the-middle

The model needs to know what you’re writing — but you can’t send the whole repo, and you have a small token budget. And unlike chat, code has text after the cursor too. What context do you assemble?

Completion SvcorchestratorContext Builderprefix · suffix
New in this step: Context Builder.

What’s the right context for completing at the cursor?

  1. Ignoring the code after the cursor (the suffix) makes completions that don’t fit what follows. Code models use fill-in-the-middle: prefix and suffix both matter.

  2. The whole repo blows the token budget and the latency budget. You send a small, curated slice — the most relevant signal, not everything.

  3. The Context Builder packs the highest-signal context: the prefix and suffix around the cursor (FIM), nearby open files, imports, and recent edits — fit into a small budget so the model completes into the actual surrounding code.

The Context Builder assembles a fill-in-the-middle prompt — the cursor’s prefix and suffix — plus open tabs, imports, and recent edits, packed into a small token budget by relevance. FIM lets the model complete a span that fits both what came before and after the cursor, which plain left-to-right prompting can’t do.

Why this piece earns its place

Two things about the packer surprise people who have only read about it. First, fill-in-the-middle is a wire format at serving time, not only a property of the model: each family has its own sentinel tokens and its own ordering of the two halves, and emitting them wrong raises no error anywhere. The model reads your markers as ordinary text and quietly reverts to predicting left to right. Completions get a little worse and nothing says why, which is part of the true cost of changing models — you rewrite and re-test the packer too. Second, the packer runs inside the budget it exists to protect. A ranking pass that opens and tokenizes files at request time has already spent the latency it was trying to save, so the inputs have to be kept warm and updated as the developer edits, leaving the request itself closer to a lookup than a gather. The dial worth arguing about is the split between the two halves: the lines immediately after the cursor constrain the completion far more than a page of them does, so the suffix takes the smaller share and gets cut at a structural boundary rather than a token count.

  • fill-in-the-middleprefix + suffix
  • open tabs + importslocal context
  • token budgetrank by relevance

What the new pieces do

Context Builderservice
Assembles the prompt from the cursor’s prefix and suffix (fill-in-the-middle), open tabs, imports and recent edits — packing the most relevant signal into a small token budget.

Step 3 · Know the whole repo

Repo-aware retrieval (RAG for code)

The function you’re calling is defined in another file the model has never seen. Local context alone makes the model guess its signature — and guess wrong. How does completion learn about the rest of the codebase?

Repo Retrievalrelevant snippetsCode Indexembeddings + symbols
New in this step: Repo Retrieval, Code Index.

How does the assistant know about code defined elsewhere in the repo?

  1. A background-built index (semantic embeddings plus a symbol/definition map) lets the assistant pull the definitions and usages relevant to the cursor into the prompt — RAG applied to your codebase.

  2. Per-repo fine-tuning is slow, expensive, and stale between runs. Retrieval gives fresh, precise repo knowledge at request time without training.

  3. Single-file completion is exactly what makes the model hallucinate cross-file APIs. Real assistants retrieve from the whole repo — the failure this page’s chaos button demonstrates.

Repo Retrieval pulls the snippets most relevant to the cursor from a Code Index — embedded chunks for semantic matches plus a symbol index for exact definitions/usages — and folds them into the context. It’s RAG for code: the model completes against the functions and types that actually exist in this repo, not plausible inventions.

Why this piece earns its place

The unit you index is the decision that matters, and it is not a token count. Code chunked every N tokens cuts functions in half, and half a function is close to unretrievable: the signature that would have matched the query is in the other chunk. Split on parsed structure — function, method, class — and keep the enclosing signature and the file path inside the chunk text, so a retrieved body still says what it is. Then decide what is in scope, because a working tree is not a codebase. It also holds vendored dependencies, build output and generated clients, and retrieval that surfaces machine-written code is worse than no retrieval, since the completion copying it looks entirely reasonable in review. Ignore rules are the cheap first filter and they are wrong about generated code that is committed, so expect to maintain an explicit exclusion list. Last, run the two halves of the index on different clocks: a symbol map can be rebuilt from the file that just saved and should be, because being offered the old name after a rename reads as the tool being broken, while re-embedding is batched work that trails the edit and nothing downstream notices the lag.

  • embeddingssemantic snippets
  • symbol indexexact definitions
  • RAG for codeground in the repo

What the new pieces do

Repo Retrievalservice
Finds code elsewhere in the repo relevant to the cursor — via embeddings and a symbol index — so completions know about functions and types defined in other files.
Code Indexstore
A background-built index of the repository: embedded chunks for semantic search plus a symbol/definition map, kept fresh as files change.

Step 4 · The model that fills the gap

A code-specialized, streaming model

Now the context is assembled. What model turns it into a suggestion, and why not just use a general chat model?

Completion SvcorchestratorContext Builderprefix · suffixRepo Retrievalrelevant snippetsCode Modelfill-in-the-middle
New in this step: Code Model.

What kind of model serves inline code completion best?

  1. The largest models are too slow for a sub-second inline budget and aren’t trained for fill-in-the-middle. Inline completion favors smaller, code-specialized, FIM-tuned models.

  2. A model trained on code with a FIM objective completes the span between prefix and suffix accurately, and a smaller/optimized variant meets the latency budget while streaming tokens as they generate.

  3. Templates can’t generalize to arbitrary code and context. The value is a learned model that adapts to your prefix, suffix, and retrieved repo context.

A code-specialized model trained for fill-in-the-middle generates the completion, tuned for low latency and token streaming. Inline completion often uses a smaller/faster model than chat, because meeting the deadline matters more than a marginally better suggestion that arrives too late. (Heavier reasoning lives in the chat/agent mode, step 7.)

Why this piece earns its place

The deadline is usually lost in the queue, not in the forward pass. Inline traffic is bursty — a whole company types during the same hours — and the continuous batching that makes the model affordable buys throughput by paying in queueing delay. A request that has already spent most of its budget waiting will produce a suggestion nobody sees, so admission control belongs in front of the model: put the deadline on the request and reject anything that cannot make it rather than enqueueing it. Shedding load here is unusually cheap, because the work being shed was worthless anyway. The other knob that belongs in the request rather than the config is where to stop. The model has no idea it was asked for one line; left alone it writes the rest of the function, and a block long enough that it has to be read before it can be accepted is a different product from a completion. A stop policy — end of statement, end of block, a hard cap on new tokens — is a per-language judgement, and it is a cost lever as well, because the tokens you never generate are the only ones that are free.

  • code-specializedtrained on code + FIM
  • low latencyfits the budget
  • streamingtokens as generated

What the new pieces do

Code Modelmodel
A code-specialized LLM trained for fill-in-the-middle: given prefix + suffix + context, predict the span between. Tuned for low latency and streaming.

Step 5 · Win the milliseconds

Debounce, cancel, cache

Even a fast model plus retrieval can blow the budget if you call it carelessly. The developer is typing quickly, invalidating requests constantly. How do you keep latency (and cost) down in a tight edit loop?

Completion Svccontext · model · cacheCode Modelfill-in-the-middleCompletion Cacheprefix keyed
New in this step: Completion Cache.

What keeps completion fast and cheap under rapid typing?

  1. More hardware helps throughput but doesn’t fix wasted work on stale requests or repeated identical contexts. The wins are algorithmic: debounce, cancel, cache.

  2. That maximizes wasted compute and shows suggestions the developer has already typed past. You must suppress and cancel invalidated work.

  3. Wait for a typing pause (debounce), cancel any in-flight request the moment new input arrives, and serve a Completion Cache hit when the context repeats — cutting both latency and model spend.

Cut latency and cost with three levers: debounce keystrokes to avoid firing mid-type, cancel superseded requests so only fresh work finishes, and a Completion Cache keyed on context so repeated cursor positions return instantly. In a fast edit loop these do more for perceived speed than any hardware upgrade.

Why this piece earns its place

The debounce interval is the number here that gets tuned by feel and then quietly forgotten. Too short and you pay for keystrokes in the middle of a word; too long and the suggestion lands after the developer has committed to typing the line themselves, which reads as the assistant being useless rather than merely slow. It is also not one number — a pause that means thinking in a verbose language means nothing in a terse one. The cache has a subtler problem: its key is whatever the packer emitted, which makes the packer and the cache one component in practice. Anything nondeterministic in the pack — iteration order over open tabs, an absolute path that differs per machine, a recent-edit list naming a file that was not actually sent — turns one cursor position into two keys, and the hit rate collapses with no failure to point at. So sort the inputs, hash the packed prompt, and put the packer and model versions inside that hash. Otherwise a rollout ships while the previous model’s completions keep being served from cache, and a change to either one becomes a flush you have to plan rather than a deploy you can roll back.

  • debouncefire on a pause
  • cancelabandon stale requests
  • cacheinstant on repeats

What the new pieces do

Completion Cachecache
Reuses completions for identical (or prefix-matching) context so common spots return instantly without a model call — a big latency win in tight edit loops.

Step 6 · Measure what matters

Acceptance rate, not benchmarks

Is the assistant actually good? Offline code benchmarks say one thing, but the only judgment that counts is whether developers keep your suggestions. How do you measure real quality?

Completion SvcCode ModelAcceptance Data
New in this step: Acceptance Data. · swipe to pan the diagram

What’s the north-star quality metric for a coding assistant?

  1. Benchmarks are a proxy that doesn’t capture real-world context, languages, or usefulness in a live editor. They can look great while acceptance is poor.

  2. Track shown-vs-accepted suggestions and whether accepted code survives — the direct measure of usefulness, and the signal you optimize context, model, and latency against.

  3. Volume says nothing about quality; a flood of ignored suggestions is worse than a few accepted ones. Acceptance is the metric.

Log acceptance data: suggestions shown vs accepted, and whether accepted code is retained (not immediately deleted). Acceptance rate is the north-star — it captures real usefulness that offline benchmarks miss, and it’s what you A/B test context strategies, models, and latency budgets against.

Why this piece earns its place

Retention is the expensive half of this, and it does not live on the server. Knowing whether accepted code survived means tracking the inserted range through every later edit — the developer types above it, reformats the file, renames a variable inside it — which is range-mapping work in the plugin, not a query you write afterwards. The window has to be chosen rather than assumed, too: still present at the next save is a different metric from still present at the next commit, and the two disagree most sharply on exactly the suggestions you care about. It also has to be content-free. You are measuring what happened to somebody’s private source, so the telemetry carries ranges, hashes and counts, never the code. And the denominator deserves as much argument as the numerator. A suggestion painted for an instant and wiped by the next keystroke was never really offered: count it as shown and every latency improvement reads as a quality regression, exclude it and a slow path has somewhere to hide. Write that boundary down and version it, because moving it later moves the headline number further than most model changes do, and nobody reading the chart six months on will know it moved.

  • accept-rateshown vs accepted
  • retentiondid it survive?
  • A/Boptimize against it

What the new pieces do

Acceptance Datastore
Logs shown-vs-accepted suggestions and retention. Acceptance rate is the north-star quality metric — offline benchmarks don’t capture real usefulness.

Step 7 · Beyond the next line

Agentic chat and multi-file edits

Inline completion suggests the next span, but developers also want "add tests for this," "refactor across these files," "why is this failing?" — tasks that need reasoning, multiple files, and running tools. How do you support that without breaking inline latency?

Completion SvcCode ModelCompletion CacheAcceptance DataAgent / Chat
New in this step: Agent / Chat. · swipe to pan the diagram

How do you add multi-file, tool-using help alongside inline completion?

  1. A distinct deliberate loop uses a more powerful model with a larger context, reads multiple files, runs tools (tests, search, build), and returns reviewable diffs — decoupled from the sub-second inline path.

  2. Heavy multi-step reasoning can’t fit the inline latency budget. Agentic tasks belong in a separate, slower mode, not the ghost-text path.

  3. Inline and agentic have opposite constraints (fast+small vs powerful+multi-step). Splitting modes is what lets each meet its own bar.

Add a separate Agent / Chat mode: a stronger model with a larger context that reads multiple files, runs tools (tests, search, build), and proposes reviewable diffs in a deliberate loop — decoupled from the latency-bound inline path. Same retrieval and privacy plumbing, different model and cadence. (This is the agent loop from the AI-agent design, specialized to code.)

Why this piece earns its place

Running tools quietly changes what you are operating. Running the tests means executing whatever code is in that repo, so the agent loop is arbitrary code execution by construction, and the only real question is whose machine pays for it. On the developer's laptop it inherits their shell, their credentials and their cloud logins, and a test the agent was talked into writing runs with all of them. On your infrastructure you own an isolated execution environment, a real cost per task, and the problem of getting the repository there in the first place. Neither is free, and treating the choice as a deployment detail is how it gets made badly. The other thing separating a demo from a tool is what happens when a run is abandoned, because it will be — the laptop closes, the developer changes their mind. Edits must accumulate somewhere stageable and land as one reviewable unit, never streaming into files as they are produced. A half-applied refactor is worse than none, and nobody can tell which half they got.

  • separate modedeliberate, slower
  • tools + multi-filetests, search, diffs
  • reviewable diffshuman approves

What the new pieces do

Agent / Chatmodel
The heavier mode: chat and agentic edits that read multiple files, run tools (tests, search), and propose diffs — a slower, deliberate loop distinct from inline completion.

Step 8 · Handle code responsibly

Privacy, secrets, and licensing

You’re shipping a company’s private source code to a model, and shipping model output back into their repo. That cuts both ways: secrets could leak out, and secrets or license-risky code could get suggested in. How do you keep it safe?

Completion Svccontext · model · cacheSecret Filterprivacy
New in this step: Secret Filter.

What must a coding assistant do to handle code responsibly?

  1. Training on customer code without explicit consent is a serious privacy and IP violation. Retention/training must be opt-in and clearly scoped.

  2. Code carries secrets (keys, tokens) and license obligations. Ignoring that leaks credentials and creates legal risk in both directions.

  3. A Secret Filter strips credentials/PII from outgoing context, respects code-retention and training-consent settings, and scans suggestions for leaked secrets or large verbatim copies that pose licensing risk.

A Secret Filter and privacy layer: redact secrets/PII from outgoing context, honor retention and training-consent settings (don’t train on code without opt-in), and scan suggestions for leaked secrets or license-risky verbatim copies. Handling proprietary code demands guarantees in both directions — nothing sensitive leaves, nothing unsafe comes back.

Why this piece earns its place

A redactor has two error directions and only one of them gets discussed. A missed credential is the loud failure. The quiet one is the false positive: the filter strips a constant that merely looked like a token, the model now sees a hole where a value was, and it completes against the hole — the same starvation this page's chaos button demonstrates, in miniature, on one line. So redact to a typed placeholder rather than deleting: something that still says string, still says assigned here, still parses. The model keeps the shape and loses only the secret. Then accept the bind that creates. You cannot log what you redacted in order to debug the redactor, so tuning runs on hashes, counters and a synthetic corpus of planted credentials instead of on production examples — budget for building that, because someone will ask you to explain a false positive and you will have nothing to look at. And resolve consent and retention server-side per request: the client is a plugin the developer can edit, so an organization's policy cannot be a value it sends you.

  • redactsecrets out of context
  • consentopt-in retention/training
  • filter outputno leaked/verbatim risk

What the new pieces do

Secret Filterservice
Redacts secrets/PII from outgoing context, honors code-retention/training consent, and filters suggestions for leaked secrets or license-risky verbatim copies.

The payoff

You built an AI coding assistant

From "suggest the next lines, instantly, correctly" to a real system: a debounced/cancellable completion service, fill-in-the-middle context assembly, repo-aware retrieval over a code index, a fast code-specialized model, latency wins from debounce/cancel/cache, acceptance-rate telemetry, a separate agentic chat mode, and a privacy/secret layer.

EditorCompletion SvcContext BuilderRepo RetrievalCode IndexCode ModelCompletion CacheAcceptance DataAgent / ChatSecret Filter
The finished design, end to end. · swipe to pan the diagram

Now starve the context — send only the current line — and watch the model confidently complete against APIs that don’t exist: invented functions, wrong imports, guessed signatures, all syntactically perfect and all wrong. That’s why context assembly and repo retrieval are the quality engine, and why grounding beats a bigger model when the goal is code that actually runs.

Everything you assembled, in order

  • Latency-bound — debounce, cancel, stream — a late completion is worthless
  • Context builder — fill-in-the-middle prefix/suffix + open tabs + imports, on a budget
  • Repo retrieval — RAG for code — embeddings + symbol index ground completions in the repo
  • Code model — code-specialized, FIM-tuned, fast and streaming
  • Latency tricks — debounce + cancel + cache beat brute-force hardware
  • Acceptance rate — the north-star metric — the developer is the eval
  • Agent mode — a separate powerful, multi-file, tool-using, diff-based loop
  • Privacy — redact secrets, opt-in training, filter risky output — trust is the moat
  • The failure — too little context → confident, hallucinated code, silently

Deep cut · 34:04

Twelve suggestions started. You saw one.

The interactive build above lays out the assistant: context assembly with fill-in-the-middle, repo retrieval, a small code model, debounce, cancel and cache, acceptance telemetry, an agent mode and a secret filter. This film is an autopsy. One minute of writing a checkout function starts twelve suggestions and shows one, and the film rebuilds the assistant one reasonable decision at a time, running the request loop and the context packer line by line as pseudo code, until the eleven in the bin read as the design working.

  • See the throwaway: every keystroke in a burst makes the pending suggestion stale, so the assistant waits for a pause, cancels on the next key and streams what survives. Eleven of twelve go in the bin on purpose.
  • Take it into the interview: pack the code before and after the cursor first, then the most useful slips per token inside a budget; retrieve by name and by embedding; judge by acceptance and retention, not benchmarks; cap the agent with a step budget; fail closed on secrets.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. Why is fill-in-the-middle a genuinely different training objective from standard left-to-right next-token prediction, not just a prompting trick?

    A model trained purely left-to-right learns to predict what comes next given only what came before — it has never learned to condition its generation on text that appears AFTER the span it's producing. Simply concatenating "prefix + suffix" into a left-to-right prompt doesn't teach the model to actually use the suffix as a constraint; it just adds more preceding text. FIM training specifically rearranges training documents (moving a suffix span to be predicted using a reordered prefix+suffix+marker format) so the model learns, during training itself, to generate a span that's consistent with text on both sides — which is why FIM-tuned models complete meaningfully better mid-file than a general chat model given the same prefix/suffix concatenated into one prompt.

  2. Why maintain a completion cache keyed by context, when code context changes constantly as the developer types?

    Not every keystroke meaningfully changes the RELEVANT context for a given completion — a developer pausing at the same function signature repeatedly while deciding what to type, or two developers independently writing similar boilerplate against the same well-known API, can hit an identical or near-identical context. The cache doesn't need every possible context to repeat to be valuable; it needs the COMMON ones to repeat often enough to matter, and in practice a meaningful fraction of completions are requested against contexts recently seen (the same cursor position revisited, a common pattern across a codebase) — enough for caching to be a real, cheap latency win even though most individual keystrokes do invalidate the exact previous context.

  3. Could acceptance rate alone be a misleading metric — is there a way for it to look good while the assistant is actually harmful?

    Yes — acceptance rate measures whether a developer accepted a suggestion in the moment, not whether that code was actually CORRECT or good practice; a plausible-looking but subtly buggy completion can absolutely get accepted (especially if the developer is moving fast and trusts the tool) and only surface as a problem later, well outside the telemetry's ability to attribute it back to that suggestion. This is why retention (did the accepted code survive subsequent edits, or get quickly deleted/rewritten) is tracked alongside raw acceptance — a completion accepted then immediately reverted is a much weaker positive signal than one accepted and kept, and a mature system watches both rather than optimizing acceptance rate in isolation.

  4. Why does agentic chat use a "stronger model with a larger context" while inline completion deliberately uses a smaller, faster one — couldn't one sufficiently good model serve both?

    The two modes optimize for opposite things on the same underlying trade-off: inline completion needs to return a usable suggestion in a few hundred milliseconds or it's worthless (the developer has already moved on), which favors a smaller, faster, FIM-specialized model even at some accuracy cost. Agentic chat has no such hard deadline — a developer asking "refactor across these three files" is willing to wait several seconds or longer for a genuinely well-reasoned, multi-file result, which favors a larger, more capable model even at higher latency and cost. Using one model for both would either make inline completion too slow to be useful or make agentic chat unnecessarily shallow — matching the model to each mode's actual constraint serves both far better than a one-size-fits-all compromise.

  5. How would this design need to adapt for a coding assistant used across many different customers' private repositories, rather than one team's codebase?

    The Code Index and retrieval logic already operate per-repository in spirit (retrieval is scoped to "this codebase"), but a genuinely multi-tenant version needs that scoping to be an ENFORCED boundary, not just an implicit assumption — one customer's embeddings and symbol index must never be retrievable by another customer's completion requests, which is exactly the kind of independently-enforced, multiple-layer isolation described in the Multi-Tenant AI SaaS design (checked at the retrieval/index layer itself, not just trusted from whichever request came in). The Secret Filter's importance also compounds at multi-tenant scale: a leak that exposed even one customer's secrets in a shared-infrastructure product is a breach with a much larger blast radius than a single-team internal tool.

Check yourself — the answers, and why

Nine steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. Inline completion is shaped above all by…

    Suggestions must arrive in a few hundred ms or the developer types past them; debounce, cancel, and streaming exist to hit that deadline.

  2. Fill-in-the-middle context includes the code after the cursor because…

    Code has text after the cursor; FIM lets the model complete a span consistent with what precedes and follows it.

  3. Repo retrieval (RAG for code) exists to…

    An embeddings + symbol index pulls relevant repo code into context so the model uses real APIs instead of hallucinating cross-file ones.

  4. The biggest latency wins under fast typing come from…

    Most in-flight completions are invalidated by the next keystroke; not computing throwaway work beats brute-forcing it.

  5. The north-star quality metric is…

    Whether developers accept and keep suggestions measures real usefulness that offline benchmarks miss; it’s the A/B target.

  6. Starving the model of context produces…

    Without file/repo context the model guesses functions and imports that don’t exist; grounding via context + retrieval is the fix.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Complete inline: suggest the next span as you type, in a few hundred milliseconds.
  • Assemble context: build a fill-in-the-middle prompt from cursor prefix/suffix, open tabs, imports.
  • Know the repo: retrieve relevant snippets from a code index so completions use real APIs.
  • Measure: log acceptance rate and retention of accepted code.
  • Agentic chat: a separate mode for multi-file, tool-using edits with reviewable diffs.

The qualities that shape everything

Each one names the mechanism that buys it.

A suggestion arrives before it’s stale
The Completion Service streams under a tight latency budget and cancels in-flight requests the moment the developer keeps typing.
Completions fit the surrounding code
The Context Builder assembles prefix + suffix (fill-in-the-middle) plus open tabs, imports and recent edits within a token budget.
No hallucinated cross-file APIs
Repo Retrieval pulls relevant snippets from a Code Index (embeddings + symbol map) into the context — RAG for code.
Fast and cheap under rapid typing
Debounce keystrokes, cancel superseded requests, and serve a Completion Cache hit on repeated context.
Optimize real usefulness, not benchmarks
Log acceptance rate (and retention) as the north-star, A/B-testing context, model and latency budgets against it.
Nothing sensitive leaks in or out
A Secret Filter redacts secrets/PII from outgoing context, honors training consent, and scans suggestions for leaked or license-risky code.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

Debounce + cancel + stream over a blocking call per keystroke

A blocking call returns a suggestion after the developer has typed past it. Stale work must be abandoned so only fresh work finishes — a late completion is worthless.

Fill-in-the-middle + curated context over the whole repo, or prefix only

The whole repo blows the token and latency budget; prefix-only ignores the code after the cursor. FIM plus a ranked slice fits both what precedes and follows.

Repo retrieval (RAG for code) over per-repo fine-tuning

Per-repo fine-tuning is slow, expensive, and stale between runs. Retrieval gives fresh, precise repo knowledge at request time without training.

A code-specialized FIM model over the biggest general chat model

The largest models are too slow for a sub-second inline budget and aren’t FIM-trained. Inline favors smaller, code-specialized, streaming models; heavy reasoning lives in chat mode.

Acceptance rate from telemetry over a public code benchmark

Benchmarks don’t capture real context, languages, or usefulness in a live editor, and can look great while acceptance is poor. Whether developers keep suggestions is the honest signal.

The answer, out loud

What a strong answer to “Design an AI Coding Assistant” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Pin down what we’re completing

    I’ll start by pinning scope: inline completion — the “ghost text” suggestion that appears at your cursor as you type — with chat added later as a separate mode. The number developers feel is latency: a suggestion has a few hundred milliseconds, or they’ve typed past it. The number that says we’re good is whether they accept it. More context and a bigger model both cost time, so the real question is: how do I make completion fast and grounded in their code?

  2. 3–8 min

    Keystroke to suggestion, on a deadline

    The skeleton is thin on purpose: an editor plugin that owns the cursor, the typing pause and the painting, and a completion service that does everything else — assemble the context, call the model, stream the span back. I’d keep it thin because it multiplies; every editor we support is another copy. And this loop is unusual: most of the requests it starts are meant to die. The plugin waits for a pause rather than firing on every key, and the service abandons work in flight the moment another key lands, so what reaches the model still describes where the cursor is.

    Built in step 1: Keystroke to suggestion, on a deadline
  3. 8–14 min

    What goes in the prompt

    What goes in the prompt is where I’d spend the most time, and it’s the cheapest thing to change later. The model gets the code on both sides of the cursor — prefix and suffix — and writes the span between, so a completion has to agree with the line it’s about to run into, not just the lines above. Around that I’d pack open tabs, the file’s imports and the edits just made, ranked into a small budget. Ranking is the job: recent edits say what this code is about to do, a tab open all week says nothing, and every token I keep is time the developer waits.

    Built in step 2: Context assembly and fill-in-the-middle
  4. 14–20 min

    Ground it in the rest of the repo

    Local context runs out the moment I call something defined in a file the model has never seen — it still answers, and it invents the signature. The tempting fix is to teach the model this repository directly, and I’d argue against it: that buys a model which knew the codebase when training stopped. So I’d retrieve instead, from a background-built index with two halves, because there are two questions: what is this function’s exact signature is a symbol lookup, and how does this codebase usually do this is an embedding search. One mechanism doesn’t answer both, and what I take on in exchange is an index I have to keep current.

    Built in step 3: Repo-aware retrieval (RAG for code)
  5. 20–26 min

    Right-size the model, then skip the waste

    For the model I’d filter on the shape of the task before the size: it’s a fill-in-the-middle job, so the first question about a candidate is whether it was trained to do that, and only then how big it can be inside the deadline. I’d take a real accuracy cost there and say so. Streaming changes which number I’d optimize — the developer reads the first line while the rest arrives, so time to the first token is what they feel. Then I’d stop paying for repeats: a cache keyed on the assembled context answers instantly when the same spot comes round.

    Built in step 4: A code-specialized, streaming model
  6. 26–31 min

    Let the developer be the judge

    To know whether it works I’d log suggestions shown against suggestions accepted, then whether the accepted code is still there later. I wouldn’t put a public benchmark beside it: that scores a model, and what the developer reacts to is the model plus the packer plus the index plus the deadline. And I wouldn’t read acceptance as one number — sliced by language and by repository it usually says we’re strong in one place and pointless in another, and an aggregate that moved because the user mix moved gets read as a quality change.

    Built in step 6: Acceptance rate, not benchmarks
  7. 31–37 min

    A second mode, and trust both ways

    Developers also ask for things that aren’t a next line — add tests here, change this across these files. I’d serve that as a second mode rather than stretching the first: the two have opposite budgets and share only plumbing — the same index, the same privacy layer, a different model on a different cadence. The part people skip is the output. Inline is accepted by not being deleted; an agent’s output is a diff somebody has to read and approve, and building that review surface is product work, not a model choice. Both modes carry private source out of the building, so I’d put a filter on that path and I’d rather it make the assistant worse than riskier.

    Built in step 7: Agentic chat and multi-file edits
  8. 37–42 min

    What I’d watch, and how it fails

    What I’d watch: client-side latency against the budget, acceptance and retention, cache hit rate, filter health. The failure I’d plan for is the one that pages nobody: every service stays up, the latency chart improves, and the suggestions get quietly worse, because what degraded was how much of the codebase reached the prompt, and nothing in the request path can tell a full prompt from a thin one. So I’d monitor context coverage directly: how often retrieval came back empty, how often the packer hit its ceiling. And I’d never read a latency win on its own: the cheapest way to get faster is to send less.

  9. 42–45 min

    Close on the trade-off

    So: a loop that throws most of its work away to hit the deadline, a prompt packed from both sides of the cursor on a budget, an index that grounds it in this repository, a small fast model with a cache in front, accept-and-keep behaviour as the scorecard, a separate deliberate mode for the bigger asks, and a filter on everything crossing the boundary. Every dial trades grounding against speed, and I’d rather name where on that line I’m sitting than pretend a setting wins both. With more time I’d take on the degraded case — the loop assumes a round trip inside a few hundred milliseconds, and a developer on a bad connection still types.

What this teaches

Learn AI system design by building an AI coding assistant (like Copilot or Cursor) step by step. An interactive guide to sub-second inline completion / tab prediction — context assembly with fill-in-the-middle, repo-aware retrieval, a specialized code model, latency tricks (debounce, cancel, cache), acceptance-rate telemetry, an agentic chat mode, and privacy/secret filtering — plus why starving the context produces confident, wrong code.

Key takeaways

  • Latency-bound — debounce, cancel, stream — a late completion is worthless
  • Context builder — fill-in-the-middle prefix/suffix + open tabs + imports, on a budget
  • Repo retrieval — RAG for code — embeddings + symbol index ground completions in the repo
  • Code model — code-specialized, FIM-tuned, fast and streaming
  • Latency tricks — debounce + cancel + cache beat brute-force hardware
  • Acceptance rate — the north-star metric — the developer is the eval
  • Agent mode — a separate powerful, multi-file, tool-using, diff-based loop
  • Privacy — redact secrets, opt-in training, filter risky output — trust is the moat
  • The failure — too little context → confident, hallucinated code, silently

Concepts covered

  • Autocomplete that understands your code
  • Keystroke to suggestion, on a deadline
  • Context assembly and fill-in-the-middle
  • Repo-aware retrieval (RAG for code)
  • A code-specialized, streaming model
  • Debounce, cancel, cache
  • Acceptance rate, not benchmarks
  • Agentic chat and multi-file edits
  • Privacy, secrets, and licensing
RUN IT YOURSELF

The budget ran out before the definition got in

One toy repository packed into the same context window three ways — most-recently-opened files, whole files by relevance, and a repo map of signatures plus the bodies that earn their tokens — running for real in your browser. It prints what each packing spends, how many of the repo’s 21 definitions reach the window, and whether the model gets charge_card’s real parameter list or has to invent it. Change BUDGET_TOKENS, hit Run, and watch the definition cross the line.

HOW TO READ THE CODE — 5 IDEAS
  1. The cursor’s prefix and suffix are already in the window — that is fill-in-the-middle. What the budget rations is everything else (step 2).
  2. Whole-file packing spends 86% of the budget and delivers 6 of 21 signatures, none of them the one being called: the tokens went on imports, tests and unrelated functions.
  3. Recency is not relevance. The file holding the definition was opened four files ago, so it needs 1211 tokens to arrive — the case step 3 exists to answer.
  4. A signature is the cheap unit: 94 tokens buys the exact parameter list, against 449 for the whole file carrying the same line — 4.8× (step 3’s symbol map, beside the embedding half).
  5. When that line is missing the model does not stall, it guesses the arguments and compiles in its own imagination — this page’s starve the context, one call wide.
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, starve the context, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs