Vibe Engines
YouTube
AI System Design

Design Multi-Agent Orchestration

Learn AI system design by building a multi-agent orchestration / workflow engine step by step.

The numbers to beatsplit ondistinct skills / parallelismisolated contexteach agent stays focusedcost of agentslatency + failure surface

The whole design, in writing

Learn AI system design by building a multi-agent orchestration / workflow engine step by step. An interactive guide to coordinating specialist agents — orchestrator-worker topology, shared state, a durable workflow engine with checkpoint/retry, typed handoffs, parallel fan-out/fan-in, budgets and termination conditions, and end-to-end tracing — plus why an uncapped multi-agent system burns unbounded cost.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

When one agent isn’t enough

A single agent works until the task gets broad: it needs many skills, a huge context, and dozens of steps — and its one prompt turns into an unfocused mess that loses the thread. The instinct is to split the work across specialist agents (a researcher, an analyst, a writer) coordinated toward one goal. But the moment you have many agents, you inherit distributed-systems problems: coordination, shared state, partial failure, and — worst — runaway loops. How do you orchestrate multiple agents reliably?

A multi-agent orchestration engine decomposes a goal, routes sub-tasks to specialist agents, coordinates them through shared state, and runs the whole thing as a durable workflow with hard budgets and termination. The discipline: gain the power of many focused agents without letting autonomy become an unbounded loop.

Step 1 · The skeleton

An orchestrator over worker agents

A broad goal comes in. Something has to break it down, decide who does what, and assemble the answer. What’s the basic coordination shape?

Clienta goalOrchestratorplan · route
New in this step: Client, Orchestrator.

What’s the minimal structure for coordinating multiple agents?

  1. A free-for-all mesh has no owner of the plan, no clear termination, and combinatorial coordination — it’s how multi-agent systems descend into chaos. Start with a coordinator.

  2. A coordinator plans the work, dispatches sub-tasks to specialist worker agents, tracks their results in shared state, and composes the final output — a clear owner of the workflow.

  3. That’s the single-agent design you’re outgrowing: too many skills and too much context in one prompt. The point of multi-agent is focused specialists under coordination.

Start with an Orchestrator over worker agents: it decomposes the goal, routes sub-tasks to specialists, and assembles their outputs into the answer. A clear coordinator owns the plan and the termination — the foundation every reliable multi-agent system is built on.

Why this piece earns its place

The client that submitted this goal will be gone before it finishes — tab closed, request retried, socket timed out — and a run living inside that connection dies with it. So the entry point records the goal, hands back a run id and returns; the answer is polled or streamed against that id. Everything added later hangs off it: the checkpoint, the spend ledger, the trace. Retrofitting that shape means changing every caller. The orchestrator’s other job is to keep its thinking and its doing in separate boxes. A model call produces a plan; ordinary deterministic code reads that plan and dispatches tasks. Re-planning is allowed and usually necessary — after a step returns something surprising, after a branch comes back empty — but it happens at defined points, as another recorded call emitting a new plan, rather than as one agent deciding and acting in the same breath. The difference shows up the day you want to replay a run exactly as it ran, show a person the plan before it executes, or simply answer what work is still outstanding. A fused loop has no plan to read: the outstanding work is inside a model’s head.

What the new pieces do

Clientclient
Submits a high-level goal ("research these vendors and draft a comparison") that’s too broad or multi-skill for a single agent to do well in one context.
Orchestratorbackend
The coordinator: decomposes the goal into sub-tasks, routes each to a specialist agent, manages shared state and the workflow, enforces budgets, and assembles the final result.

Step 2 · Don’t over-engineer

Specialists — and when NOT to go multi-agent

Multi-agent is powerful and fashionable — and often the wrong choice. Each agent adds latency, cost, and a coordination failure surface. When does splitting into agents actually help, and how do you scope them?

Orchestratorplan · routeResearcherspecialistAnalystspecialistWriterspecialist
New in this step: Researcher, Analyst, Writer.

When is multiple agents the right call over one good agent?

  1. Every agent multiplies cost, latency, and failure modes. Multi-agent is justified only when a task genuinely needs distinct skills or parallelism, not by default.

  2. Past a point, one agent’s context gets overloaded and unfocused; genuinely multi-skill or parallelizable tasks do benefit from specialists. It’s a judgment call, not a rule.

  3. Split when sub-tasks need different tools/expertise or can run independently, and one agent’s context would be a muddle. Each specialist gets a focused prompt, its own tools, and an isolated context.

Use specialists when a task splits into distinct skills or parallel sub-tasks that would overload one agent’s context — and not otherwise, because every agent adds cost, latency, and failure surface. Each worker gets a focused prompt, its own tools, and an isolated context, so it does one thing well instead of everything poorly.

Why this piece earns its place

Splitting the work invents a failure a single agent cannot have: mis-routing. A pricing question goes to the researcher, every downstream step succeeds, the handoff validates, and the answer is confidently about the wrong thing. Score routing separately from answer quality, or an aggregate number sags and never tells you which half moved. A specialist is also a versioned artifact rather than a prompt in a file — a prompt, a model choice and a tool allowlist, where editing any of the three changes how the whole system behaves. Pin all three into the run record, or yesterday’s run cannot be reproduced and a regression cannot be bisected to the change that caused it. That allowlist is the strongest reason to keep a specialist narrow. Context isolation is the quality argument; permission isolation is the safety one, and it buys something specific rather than everything. If only the writer holds the publish tool, an instruction smuggled into a page the researcher fetched cannot publish. It can still steer the researcher, and an outbound fetch is a channel, so what that agent holds can still leave. The split bounds the blast radius; it does not clean the input.

  • split ondistinct skills / parallelism
  • isolated contexteach agent stays focused
  • cost of agentslatency + failure surface

What the new pieces do

Researchermodel
A specialist agent with its own prompt, tools, and isolated context — e.g. gathers information. Focused scope means a smaller, cleaner context and better results than one do-everything agent.
Analystmodel
Another specialist — e.g. reasons over gathered data. Receives a typed handoff from upstream, not a raw dump, so quality doesn’t degrade through a game of telephone.
Writermodel
A specialist that produces the final artifact from the analysts’ structured output. Independent specialists can also run in parallel when their tasks don’t depend on each other.

Step 3 · How agents coordinate

Shared state vs. message passing

The researcher’s findings have to reach the analyst; the analyst’s conclusions have to reach the writer. If you just paste everything into every prompt, contexts explode and quality rots. How do agents share information?

Orchestratorplan · routeShared StateblackboardResearcherspecialistAnalystspecialist
New in this step: Shared State.

How should agents pass work between each other?

  1. A shared store holds task results and intermediate artifacts; agents read what they need and write structured outputs, so no single context has to carry everything and handoffs stay clean.

  2. Piling full transcripts forward blows up context length and drags irrelevant detail through the pipeline — cost and quality both suffer. Pass structured results, not raw history.

  3. Without shared state agents can’t build on each other, so the "team" produces disconnected fragments. Coordination requires a shared substrate.

Agents coordinate through Shared State (a blackboard): each reads what it needs and writes structured artifacts — results, not raw transcripts. The orchestrator can also message-pass explicit handoffs. Either way, keep contexts small by sharing distilled outputs, so adding an agent doesn’t multiply everyone’s prompt size.

Why this piece earns its place

Everything written here has to be plain data — no open handles, no live connections, no callback an agent means to invoke later. State is exactly what the next step checkpoints, and a run holding a socket on the blackboard cannot resume on another machine however good the workflow engine is. The rule is free to adopt now and miserable to retrofit once three agents are passing objects around. A read from the board is a selection, not a fetch. State grows with every step, so an agent that reads all of it has rebuilt the exact problem this step exists to avoid — the bloated context is merely assembled from a store now instead of concatenated by hand. Give each artifact a key and a type, have the orchestrator hand an agent only the keys its task names, and keep bulky material behind a reference so what lands in a prompt is an id and a type rather than a scraped page. And keep the board scoped to one run. Once a shared store exists, the pull is to let it double as memory across runs, which quietly makes one goal’s leftovers visible to the next goal that reads the same key.

  • blackboardshared read/write state
  • structured artifactsnot raw transcripts
  • small contextsshare distilled outputs

What the new pieces do

Shared Statestore
The shared context agents read from and write to — task results, intermediate artifacts, a scratchpad. How agents coordinate without cramming everything into one prompt.

Step 4 · Survive failure

The durable workflow engine

A run is twenty steps and five minutes long. Step 14 calls a flaky API and the process crashes. Do you really restart from step 1 — re-paying for all that work? Agentic runs are long and partial failure is constant. How do you make them reliable?

Orchestratordurable workflowWorkflow Enginedurable
New in this step: Workflow Engine.

How do you keep a long multi-agent run reliable across failures?

  1. Restarting a twenty-step run because step 14 failed re-does (and re-pays for) all prior work and may never finish if the flake recurs. You need to resume, not restart.

  2. A workflow engine persists state after each step; a crash or a failed tool call retries that step or resumes from the last checkpoint, instead of throwing away completed work.

  3. In-memory-only runs lose everything on any crash or deploy — unacceptable for long agentic workflows. Durability requires checkpointing.

Run the orchestration on a Workflow Engine with durable execution: checkpoint state after each step so a crash or a flaky tool retries that step or resumes from the last good state — never restarting a long run from scratch. Agentic workflows are long-lived and failure-prone; durability (à la Temporal/Step Functions) is what makes them production-grade.

Why this piece earns its place

Durable execution works by replaying a recorded history to rebuild state, and that puts one hard rule on agentic runs: a model call is non-deterministic, so it has to be recorded as a result and replayed from the record, never re-invoked. Get it wrong and a resumed run continues under a different plan than the half it already executed — the most confusing incident this design can produce, because every individual step still looks correct. Tools follow the same reasoning. The engine gives you at-least-once, not exactly-once, so a step that filed a ticket or sent a mail will eventually be retried; the idempotency key belongs on the tool call, and the completion record has to be written before anything acknowledges success. Surviving a deploy is the property that creates the next problem. A run that finishes in five minutes can simply be drained before a release; a run parked on a human approval waits days and comes back into code that is not the code it started in, and if the step sequence moved underneath it the replay diverges. So workflow definitions get versioned and pinned per run, and any change to the shape of a plan is a migration with in-flight runs to carry.

  • checkpointstate after each step
  • retry / resumefrom last good state
  • durablesurvives crashes + deploys

What the new pieces do

Workflow Engineservice
Runs the orchestration as a durable workflow: each step is checkpointed so a crash or a flaky tool can retry or resume from the last good state instead of restarting the whole run.

Step 5 · Clean handoffs

Typed contracts between agents

The researcher hands to the analyst hands to the writer. If each handoff is freeform prose, meaning blurs at every hop — a game of telephone that degrades over the chain. How do you keep quality across handoffs?

ClientOrchestratorShared StateWorkflow EngineResearcherAnalystWriter
The system as it stands at this step. · swipe to pan the diagram

How do you stop quality degrading as work passes between agents?

  1. Prose summaries lose and distort information at each hop — the telephone problem. Handoffs need structure the next agent can rely on, not a re-narration.

  2. Raw history bloats context and still buries the signal. The fix is a distilled, structured contract, not more data.

  3. Each agent emits a defined structured output (schema/fields) that the next agent consumes, so information passes precisely and predictably instead of blurring through freeform text.

Define typed handoff contracts: each agent produces a structured output (a schema — fields, not freeform prose) that the next agent consumes. Structured handoffs stop the telephone-game degradation, make each step testable in isolation, and let the orchestrator validate an agent’s work before passing it on.

Why this piece earns its place

A schema checks shape, not truth. A perfectly valid artifact can be entirely invented — the analyst returns every required field, with figures no source ever contained, validation passes, and the writer treats them as fact. So put provenance in the contract itself: which source, which tool call produced the claim, how sure the agent is. Then the next agent can weigh it, the orchestrator can refuse to forward an unsourced number, and the final artifact can point at something. Fields nobody reads on day one are what make that possible in month six. Versioning these contracts feels like ceremony between two components you own, right up to the deploy that kills a day-old run. They sit between agents whose prompts change on different days, and — because of the previous step — old-shape artifacts are sitting inside checkpoints belonging to runs that have not finished, so one new required field invalidates every one of them. Treat it like any wire format: new fields arrive optional, removals take two releases, and a consumer reads the version it is handed.

  • typed outputschema, not prose
  • no telephoneprecise handoffs
  • testablevalidate each step

Step 6 · Do things at once

Parallel fan-out / fan-in

Three of the sub-tasks don’t depend on each other — researching three vendors, say. Running them one after another triples the latency for no reason. How do you exploit independence?

ClientOrchestratorShared StateWorkflow EngineResearcherAnalystWriter
The system as it stands at this step. · swipe to pan the diagram

How do you speed up independent sub-tasks?

  1. Forcing sequence on independent tasks wastes time — they could run at once. Parallelism is a major latency win when sub-tasks don’t depend on each other.

  2. The orchestrator dispatches independent sub-tasks concurrently to multiple agents (fan-out), waits for them, and merges their structured outputs (fan-in) — cutting latency for the parallelizable part of the plan.

  3. Parallelizing dependent steps breaks correctness — the analyst can’t run before the researcher finishes. You parallelize only the independent parts; the DAG of dependencies dictates what can overlap.

Exploit independence with fan-out / fan-in: the orchestrator runs independent sub-tasks in parallel across agents, then aggregates their structured results. Dependencies form a DAG — parallelize the independent branches, sequence the dependent ones. It’s the map-reduce shape, applied to agents, and often the biggest latency lever in a multi-agent plan.

Why this piece earns its place

Fan-in finishes when the slowest branch finishes, so a parallel stage takes the maximum of its branches where the serial version took their sum. That helps at every percentile, but not evenly: each branch you add is another chance that one of them is having a bad minute, so the median improves far more than the tail does — and the tail is what a user feels. A per-branch deadline with a defined fallback — a cached result, a branch marked incomplete, a shorter prompt — matters more here than anywhere else in this design, because the run inherits the worst minute any branch had. Money doesn’t move at all. Three vendors researched at once bills what three researched in sequence bills; fan-out spends the same tokens sooner, and anyone arguing for it on efficiency grounds has confused latency with cost. What a test run won’t show you is contention. Fan-out turns one run into several concurrent model calls, and a provider’s rate limit is per account, not per run, so one wide plan starves every other run in the system and the symptom arrives as somebody else’s latency. That wants a concurrency bound per run and a fair share across runs — in practice a queue and a worker pool, not firing every independent branch the instant the plan says it may go.

  • fan-outindependent tasks in parallel
  • fan-inaggregate results
  • DAGparallel where independent

Step 7 · Put a ceiling on autonomy

Budgets, loop detection, termination

Agents that can spawn agents and retry tools have no natural stopping point. Two agents can correct each other forever; a planner can delegate infinitely; a retry can loop. Nothing errors — it just spins and spends. What stops a runaway?

OrchestratorAnalystBudget / Stop
New in this step: Budget / Stop. · swipe to pan the diagram

What keeps a multi-agent system from looping and burning unbounded cost?

  1. Hard caps on steps, recursion depth, and token/dollar spend, plus loop detection and approval gates for consequential actions, bound the run — so autonomy always has a ceiling and can’t spin forever.

  2. LLM agents frequently don’t recognize completion or detect their own loops — that’s exactly how runaways happen. Termination must be enforced by the system, not hoped for.

  3. A timeout alone still lets a run burn a fortune before it fires, and doesn’t catch tight loops early. You need step/depth/cost caps and loop detection, not just a wall clock.

Enforce termination at the system level: per-run budgets (max steps, max recursion depth, token/$ caps), loop detection (spot repeating states/ping-pong), and human-approval gates for consequential or expensive actions. Autonomy needs both a floor and a ceiling — this is the single most important safeguard, and removing it (the chaos button) is how multi-agent systems fail.

Why this piece earns its place

Where the cap is checked decides whether it holds. Spend is only known after a call returns, so a budget read at dispatch is always reading a stale number, and with branches running at once several of them can each see room that only one of them actually had. Put one counter on the run and write to it twice: debit an estimate atomically as you dispatch, then reconcile that debit against real usage when the call completes. Anywhere else and a wide plan overshoots its ceiling every time, by more the wider it is. Hitting a cap is not an error, which is the part that catches teams out. Nothing throws, no alert fires, the run simply ends with what it had — so in ordinary telemetry a capped run and a finished one look identical. Track the cap-hit rate as a first-class number, split by which cap fired and by the shape of the workflow that fired it. If a real share of runs end on the step ceiling rather than finishing, that ceiling has quietly become the product’s quality limit, and you would much rather read that on a dashboard than hear it from a user saying the answers got thinner.

  • budgetssteps · depth · token/$
  • loop detectioncatch ping-pong
  • approval gateshuman on big actions

What the new pieces do

Budget / Stopservice
Enforces the floor and ceiling on autonomy: max steps, max recursion depth, token/$ budget, loop detection, and human-approval gates — the difference between a bounded run and a runaway one.

Step 8 · See the whole run

End-to-end tracing and evaluation

A multi-agent run failed to produce a good answer. Which agent? Which handoff? Which tool call? With work spread across many agents and steps, a single log line tells you nothing. How do you debug and evaluate the whole thing?

OrchestratorShared StateResearcherAnalystBudget / StopRun Trace
New in this step: Run Trace. · swipe to pan the diagram

How do you make a multi-agent run debuggable and evaluable?

  1. Final outputs hide where a multi-step, multi-agent run went wrong — you can’t see the handoff or tool call that broke it. You need the full tree.

  2. One message can’t explain a failure distributed across agents and steps. Debugging needs the structured trace of the entire run.

  3. A structured trace of the whole run (each agent, prompt, tool call, handoff, token/cost) makes failures debuggable and lets you evaluate the end-to-end workflow, not just individual model calls.

Capture a Run Trace: the full tree of the workflow — every agent, prompt, tool call, handoff, and cost — so a failure is traceable to the exact step, and the whole workflow can be evaluated end to end (not just per model call). Multi-agent quality is a property of the system, so observability and eval have to span the entire run.

Why this piece earns its place

Sampling is where an agent trace stops behaving like a request trace. A low span sample rate tells you plenty about an HTTP service and nothing at all about a run, because half a tree is unreadable — you cannot see which handoff went wrong when you kept the handoff and dropped the step that produced it. Sample by run, never by span, and keep every failed or capped run whatever the rate says. The skeleton and the payloads then want different lifetimes. A run’s shape — its steps, timings, costs and which agent did what — is small and worth keeping for a long time. The prompts and tool responses hanging off it are bulky and far more sensitive, so they get a shorter clock of their own; keeping them forever is how an observability bill outgrows the model bill. This tree is also the only place unit economics can be computed. Cost per run, attributed to a tenant and to the plan shape that produced it, is what tells you which workflows deserve which caps — without it, the budgets in the previous step are numbers somebody guessed.

  • run treeagents · tools · handoffs · cost
  • debuggablefailure → exact step
  • eval end-to-endthe workflow, not one call

What the new pieces do

Run Tracestore
The full tree of the run — every agent, prompt, tool call, handoff and cost — so a multi-agent failure is debuggable and the whole workflow can be evaluated end to end.

The payoff

You built multi-agent orchestration

From "one agent can’t hold it all" to a coordinated system: an orchestrator over focused specialists (used only when the task warrants it), shared-state coordination via structured artifacts, a durable workflow engine with checkpoint/retry, typed handoffs, parallel fan-out/fan-in, hard budgets + termination, and end-to-end tracing.

ClientOrchestratorShared StateWorkflow EngineResearcherAnalystWriterBudget / StopRun Trace
The finished design, end to end. · swipe to pan the diagram

Now uncap the agents — remove budgets and termination — and watch the system loop: agents correcting each other, planners delegating endlessly, retries spinning, tokens burning, with no error and no answer. That’s why termination conditions are the most important safeguard in multi-agent design, and why autonomy must always sit between a floor and a ceiling.

Everything you assembled, in order

  • Orchestrator-worker — one owner of the plan, state, and stop condition
  • Use specialists wisely — split on distinct skills/parallelism — not by default
  • Shared state — coordinate via structured artifacts, not forwarded transcripts
  • Durable workflow — checkpoint each step; retry/resume, never restart
  • Typed handoffs — structured contracts stop telephone-game degradation
  • Fan-out / fan-in — the plan is a DAG — parallelize the independent branches
  • Budgets + termination — steps, depth, cost caps, loop detection, approval gates
  • Trace + eval — the run tree makes the workflow debuggable and evaluable
  • The failure — uncapped autonomy loops and burns unbounded cost, silently

Deep cut · 21:26

Every agent stopped. The team didn’t.

The interactive build above lays out multi-agent orchestration: an orchestrator that splits a goal across specialist workers, shared state holding structured results, typed handoffs, parallel fan-out and fan-in, a durable workflow engine that checkpoints and resumes, budgets and termination conditions, and a trace of the whole run. The film builds the same system one fix at a time in a made-up shop’s dungeon, then replays one weekend-long run step by step to find why a team whose every agent obeyed its limit never stopped.

  • See the wreck: every agent had a limit of ten steps and stopped at ten. But an agent out of steps handed its part to a fresh agent with a fresh ten, and a writer and a reviewer with conflicting rules kept undoing each other — forty thousand steps over a weekend, no error, no answer.
  • Take it into the interview: give the run one budget and pass it down as a shrinking slice, never a fresh allowance per agent; add a depth limit and loop detection that compares recent states; stop with the best answer so far, and ask a person before spending more.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. In a fan-out of three parallel workers, one fails while the other two succeed. What happens to the fan-in?

    The orchestrator has to decide per-workflow whether the failed branch is required or optional — a required failure blocks fan-in and triggers the retry/resume machinery from the durable workflow engine on just that branch, while an optional one lets fan-in proceed with a gap noted (the writer gets two vendor comparisons instead of three, flagged as incomplete rather than silently presented as complete). Treating "partial fan-in" as an explicit, typed state — not an exception to catch — is what keeps this from becoming an ad-hoc special case per workflow.

  2. An agent’s output fails schema validation against the next agent’s expected typed handoff. What actually happens?

    Not a crash and not a silent pass-through — the orchestrator typically retries that agent’s step with the validation error fed back into its prompt ("your output was missing field X"), since LLM agents are often able to self-correct against explicit structured feedback. Repeated validation failures on the same step should count against the step budget just like any other retry, so a persistently malformed agent triggers the same termination safeguards as a genuine infinite loop.

  3. How do you programmatically detect a "ping-pong" loop between two agents correcting each other, rather than just capping total steps?

    By tracking state similarity across recent steps — if the shared-state diff between step N and step N-2 is near-zero (agent A undoes what agent B just did, or vice versa), that is a much sharper loop signal than "we hit step 40." A pure step-count ceiling still lets a loop burn its whole budget before stopping; comparing recent state snapshots catches the ping-pong pattern early, often within a handful of cycles.

  4. Does the workflow engine checkpoint after every agent step, or after every tool call within a step?

    Tool-call granularity is usually worth the extra checkpointing overhead for expensive steps — an agent step can itself involve multiple tool calls (search, then read, then extract), and checkpointing only at the step boundary means a crash after two successful tool calls still re-does all of them. The tradeoff is checkpoint volume vs. redone work; cheap, fast steps can checkpoint coarsely, while a step with expensive external calls (search-engine hits, paid API calls) benefits from finer-grained checkpoints.

  5. What changes if an orchestrator’s "worker" is itself another orchestrator — recursive multi-agent systems?

    The budget has to propagate downward as a shrinking allowance, not reset at each level — a sub-orchestrator spawned with 50 remaining steps and a $2 remaining budget must enforce ITS OWN sub-workflow within that inherited ceiling, not start fresh with its own full budget, or nesting becomes a loophole that multiplies total spend by however many levels deep the recursion goes. Depth limits (from the existing "max recursion depth" cap) exist specifically to bound how many such levels can exist at all.

Check yourself — the answers, and why

Nine steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. The default multi-agent topology to reach for is…

    A coordinator owning the plan, state, and termination stays debuggable; a free-for-all mesh has no owner and descends into chaos.

  2. You should go multi-agent when…

    Each agent adds cost, latency, and failure surface; the real win is context isolation, justified only when the task needs it.

  3. Agents should coordinate by…

    A blackboard of distilled, structured outputs keeps contexts small and handoffs clean as the number of agents grows.

  4. A durable workflow engine matters because agentic runs…

    Long multi-step runs hit partial failures constantly; checkpointed durable execution resumes from the last good state instead of re-doing everything.

  5. Typed handoff contracts between agents prevent…

    Structured outputs each agent produces and the next consumes keep information precise and make each step testable in isolation.

  6. The defining failure mode of multi-agent systems is…

    Without budgets, loop detection, and approval gates, agents ping-pong or delegate endlessly — spinning and spending forever; termination is the key safeguard.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Decompose: take a broad goal and break it into sub-tasks, each routed to a specialist agent.
  • Coordinate: agents share work through structured artifacts in shared state, not forwarded transcripts.
  • Hand off: each agent emits a typed, structured output that the next agent consumes.
  • Parallelize: fan out independent sub-tasks to run at once, then fan in and aggregate their results.
  • Terminate: enforce per-run budgets, loop detection, and approval gates so a run always stops and stays affordable.

The qualities that shape everything

Each one names the mechanism that buys it.

One clear owner of the plan and stop condition
An orchestrator-worker topology — a coordinator decomposes the goal, routes to specialists, and owns the plan, state, and termination.
Each agent stays focused, not drowning in context
Specialists with isolated contexts, used only when a task splits into distinct skills or parallel sub-tasks — not by default.
Adding an agent doesn’t multiply everyone’s prompt
A shared-state blackboard where agents read and write distilled, structured artifacts rather than raw transcripts.
Survive crashes and flaky tools mid-run
A durable workflow engine checkpoints each step, so a failure retries that step or resumes from the last good state instead of restarting.
Quality doesn’t rot across handoffs
Typed handoff contracts — a schema each agent produces and the next consumes — stop telephone-game degradation and make each step testable.
Autonomy can never loop forever
Per-run budgets (max steps, depth, token/$), loop detection, and human-approval gates cap the run at the system level.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

Orchestrator-worker over a free-for-all agent mesh

A full mesh has no owner of the plan, no clear termination, and combinatorial coordination — how multi-agent descends into chaos. A coordinator owns the plan, state, and stop condition, and stays debuggable.

Specialists only when the task warrants it over more agents by default

Every agent multiplies cost, latency, and failure surface. The real win is context isolation — split on distinct skills or parallelism, not because multi-agent is fashionable.

Shared structured artifacts over forwarding full transcripts

Piling every agent’s transcript into the next prompt blows up context length and drags irrelevant detail through the pipeline. Sharing distilled outputs keeps token cost and quality from degrading as the team grows.

Checkpoint and resume over restart the run on any failure

Restarting a twenty-step run because step 14 flaked re-pays for all prior work and may never finish if the flake recurs. A durable engine retries the failed step or resumes from the last checkpoint.

Enforced budgets + loop detection over a single global timeout

A timeout alone lets a run burn a fortune before it fires and misses tight loops. Step/depth/cost caps plus loop detection bound autonomy early — the defining safeguard of multi-agent design.

The answer, out loud

What a strong answer to “Design Multi-Agent Orchestration” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Say what makes this multi-agent

    I’d start by checking whether this should be multi-agent at all. One agent holds up until the goal gets broad — several skills, a growing context, dozens of steps — and past that the one prompt stops holding the thread. A goal like research these vendors and draft a comparison sits right on that line. So I’d build a small team of focused specialists pointed at one goal, and saying that buys me a distributed system: coordination, shared state, partial failure, and runs that don’t know how to stop. That last one is where I’d spend most of my time, so I’ll flag it now and come back to it.

  2. 3–7 min

    One coordinator owns the plan

    The skeleton is a client submitting a goal, an orchestrator, and worker agents. The orchestrator breaks the goal into sub-tasks, sends each to a specialist, and puts the answer back together. I’d take that over a mesh where any agent can call any other because somebody has to be accountable — for the plan, for what the run may spend, for deciding it’s finished — and a mesh hands that to nobody. It’s also the shape I can still explain at two in the morning.

    Built in step 1: An orchestrator over worker agents
  3. 7–12 min

    Specialists only where they earn it

    Then I’d argue against my own diagram. Each box is another bill, another hop of latency and another way a run can break, so I’d split only where the work needs different skills or where sub-tasks can go at the same time. What splitting buys is a smaller, cleaner context per agent — not extra intelligence, which is how it gets sold. If one agent can carry the task without losing the plot, one agent wins, and I’d rather say that than defend the boxes I just drew.

    Built in step 2: Specialists — and when NOT to go multi-agent
  4. 12–17 min

    Share artifacts, not transcripts

    Next, how they coordinate. The tempting answer is to paste each agent’s transcript into the next one’s prompt, and that’s the thing that makes multi-agent expensive — every prompt ends up carrying everyone’s history, and the part that matters gets buried in it. I’d give them shared state instead, a blackboard: an agent reads what its task calls for and writes a structured result back. Add a fourth agent and nobody else’s prompt gets bigger. Results, not history.

    Built in step 3: Shared state vs. message passing
  5. 17–23 min

    Treat the run as a workflow

    Then durability, because this is a run rather than a request — twenty steps, five minutes. If step fourteen hits a flaky API and the process falls over, starting again means paying for everything that already worked, and if that API is still flaky the run may never reach the end. So I’d sit the orchestration on a durable workflow engine: state written down after every step, so a crash picks up where it left off instead of at the top. Once you call this a long-running workflow instead of a call, everything around it gets built differently.

    Built in step 4: The durable workflow engine
  6. 23–28 min

    Type the handoffs

    Handoffs I’d make typed. Researcher to analyst to writer — and if every hop is freeform prose, each agent re-narrates the last one and the meaning drifts a little further from whatever was actually found. So each agent produces a structured output with named fields, and the next one reads that. That gives the orchestrator something it can check before it forwards anything, and it lets me test one agent on its own rather than only judging the finished report.

    Built in step 5: Typed contracts between agents
  7. 28–33 min

    Parallelize what the graph allows

    With contracts in place, parallelism is easy to reason about. The plan is a dependency graph, not a line: three vendors can be researched at once and merged afterwards, while the analyst still has to wait on the researcher. I’d overlap only what the graph says is independent — parallelizing a dependent step trades correctness for a win I don’t need. And I’d be straight about what fan-out buys: it cuts the clock, not the bill.

    Built in step 6: Parallel fan-out / fan-in
  8. 33–39 min

    The part I flagged at the start

    Here’s the problem I flagged at the start. Nothing in this system has a natural reason to stop. Two agents can keep politely correcting each other, a planner can keep turning one task into three more, a retry can sit there retrying — and none of it raises an error, it just calls models and spends money. So termination gets enforced by the system rather than asked of the agents: a per-run budget on steps, on delegation depth and on token spend, loop detection for a run that has stopped making progress, and a human gate in front of anything expensive or irreversible. I wouldn’t lean on a wall clock alone — by the time it fires the money is spent, and a tight loop spins happily inside it. And I wouldn’t ask the agents to notice they’re done, because mostly they don’t.

    Built in step 7: Budgets, loop detection, termination
  9. 39–44 min

    Make the run visible

    For operating it, I’d want the whole run recorded as a tree — who was asked what, what came back, what each hop cost — because when a multi-agent answer comes out wrong, no single log line tells me whether it was the routing, a handoff, or a tool call. And I’d score the run end to end rather than grading model calls one at a time, since most of what goes wrong with a team of agents happens between them.

    Built in step 8: End-to-end tracing and evaluation
  10. 44–47 min

    The trade-off, and what’s left

    To close: the trade-off running through all of this is autonomy against control. Every specialist I add buys focus and costs coordination; every cap I add buys predictability and costs the run some room to actually solve the problem, and no setting is right for both. With more time I’d go at the orchestrator’s planning step, which I’ve been treating as given — how a goal gets decomposed is a quality problem of its own and deserves its own evaluation, separate from the workers. And I’d work through what each agent is permitted to do, since a team of agents holding tools is a much wider surface than one.

What this teaches

Learn AI system design by building a multi-agent orchestration / workflow engine step by step. An interactive guide to coordinating specialist agents — orchestrator-worker topology, shared state, a durable workflow engine with checkpoint/retry, typed handoffs, parallel fan-out/fan-in, budgets and termination conditions, and end-to-end tracing — plus why an uncapped multi-agent system burns unbounded cost.

Key takeaways

  • Orchestrator-worker — one owner of the plan, state, and stop condition
  • Use specialists wisely — split on distinct skills/parallelism — not by default
  • Shared state — coordinate via structured artifacts, not forwarded transcripts
  • Durable workflow — checkpoint each step; retry/resume, never restart
  • Typed handoffs — structured contracts stop telephone-game degradation
  • Fan-out / fan-in — the plan is a DAG — parallelize the independent branches
  • Budgets + termination — steps, depth, cost caps, loop detection, approval gates
  • Trace + eval — the run tree makes the workflow debuggable and evaluable
  • The failure — uncapped autonomy loops and burns unbounded cost, silently

Concepts covered

  • When one agent isn’t enough
  • An orchestrator over worker agents
  • Specialists — and when NOT to go multi-agent
  • Shared state vs. message passing
  • The durable workflow engine
  • Typed contracts between agents
  • Parallel fan-out / fan-in
  • Budgets, loop detection, termination
  • End-to-end tracing and evaluation
RUN IT YOURSELF

Every agent obeyed its limit. No agent ever reached it. The run still never stopped.

One plan — survey three vendors, draft a comparison, publish it — dispatched four ways, where a writer and a reviewer each politely hand the work back to the other. It prints messages sent, what stopped the run, how deep delegation nested, and how many stages actually finished. Change RUN_STEPS, hit Run, and watch a run that still says it completed do less of the job.

HOW TO READ THE CODE — 5 IDEAS
  1. The per-agent cap is not just too loose — it never binds at all. The busiest agent used 3 of its 10 steps and the out-of-steps hand-off fired 0 times, and the run still had to be killed by hand (step 7).
  2. What runs away is a cycle in the plan: draft → review → draft. Each agent on it is short and well behaved, so a cap scoped to one agent cannot see it. Delete the two back-edges and the same dispatcher finishes in 11 messages.
  3. Row 1’s message count and depth are the operator’s numbers, not the plan’s — drop the ceiling tenfold and the depth drops with it. What survives every ceiling is 6 of 12 stages, and publish is not one: it waits on compose, and the plan is a dependency graph (step 6), so a stalled stage starves every stage after it.
  4. The two fixes do different jobs: the depth bound is what makes the run terminate, the seen-set is what makes it affordable — 117 messages down to 13 for the identical 12 stages.
  5. Nothing here raises. At half the budget the run still reports terminated with 6/12 done, so a capped run and a finished one look the same unless you record which cap fired (step 7).
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, uncap the agents, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs