The whole design, in writing
Learn AI system design by building a multi-agent orchestration / workflow engine step by step. An interactive guide to coordinating specialist agents — orchestrator-worker topology, shared state, a durable workflow engine with checkpoint/retry, typed handoffs, parallel fan-out/fan-in, budgets and termination conditions, and end-to-end tracing — plus why an uncapped multi-agent system burns unbounded cost.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
When one agent isn’t enough
A single agent works until the task gets broad: it needs many skills, a huge context, and dozens of steps — and its one prompt turns into an unfocused mess that loses the thread. The instinct is to split the work across specialist agents (a researcher, an analyst, a writer) coordinated toward one goal. But the moment you have many agents, you inherit distributed-systems problems: coordination, shared state, partial failure, and — worst — runaway loops. How do you orchestrate multiple agents reliably?
A multi-agent orchestration engine decomposes a goal, routes sub-tasks to specialist agents, coordinates them through shared state, and runs the whole thing as a durable workflow with hard budgets and termination. The discipline: gain the power of many focused agents without letting autonomy become an unbounded loop.
Step 1 · The skeleton
An orchestrator over worker agents
A broad goal comes in. Something has to break it down, decide who does what, and assemble the answer. What’s the basic coordination shape?
What’s the minimal structure for coordinating multiple agents?
A free-for-all mesh has no owner of the plan, no clear termination, and combinatorial coordination — it’s how multi-agent systems descend into chaos. Start with a coordinator.
A coordinator plans the work, dispatches sub-tasks to specialist worker agents, tracks their results in shared state, and composes the final output — a clear owner of the workflow.
That’s the single-agent design you’re outgrowing: too many skills and too much context in one prompt. The point of multi-agent is focused specialists under coordination.
Start with an Orchestrator over worker agents: it decomposes the goal, routes sub-tasks to specialists, and assembles their outputs into the answer. A clear coordinator owns the plan and the termination — the foundation every reliable multi-agent system is built on.
Why this piece earns its place
The client that submitted this goal will be gone before it finishes — tab closed, request retried, socket timed out — and a run living inside that connection dies with it. So the entry point records the goal, hands back a run id and returns; the answer is polled or streamed against that id. Everything added later hangs off it: the checkpoint, the spend ledger, the trace. Retrofitting that shape means changing every caller. The orchestrator’s other job is to keep its thinking and its doing in separate boxes. A model call produces a plan; ordinary deterministic code reads that plan and dispatches tasks. Re-planning is allowed and usually necessary — after a step returns something surprising, after a branch comes back empty — but it happens at defined points, as another recorded call emitting a new plan, rather than as one agent deciding and acting in the same breath. The difference shows up the day you want to replay a run exactly as it ran, show a person the plan before it executes, or simply answer what work is still outstanding. A fused loop has no plan to read: the outstanding work is inside a model’s head.
What the new pieces do
- Clientclient
- Submits a high-level goal ("research these vendors and draft a comparison") that’s too broad or multi-skill for a single agent to do well in one context.
- Orchestratorbackend
- The coordinator: decomposes the goal into sub-tasks, routes each to a specialist agent, manages shared state and the workflow, enforces budgets, and assembles the final result.
Step 2 · Don’t over-engineer
Specialists — and when NOT to go multi-agent
Multi-agent is powerful and fashionable — and often the wrong choice. Each agent adds latency, cost, and a coordination failure surface. When does splitting into agents actually help, and how do you scope them?
When is multiple agents the right call over one good agent?
Every agent multiplies cost, latency, and failure modes. Multi-agent is justified only when a task genuinely needs distinct skills or parallelism, not by default.
Past a point, one agent’s context gets overloaded and unfocused; genuinely multi-skill or parallelizable tasks do benefit from specialists. It’s a judgment call, not a rule.
Split when sub-tasks need different tools/expertise or can run independently, and one agent’s context would be a muddle. Each specialist gets a focused prompt, its own tools, and an isolated context.
Use specialists when a task splits into distinct skills or parallel sub-tasks that would overload one agent’s context — and not otherwise, because every agent adds cost, latency, and failure surface. Each worker gets a focused prompt, its own tools, and an isolated context, so it does one thing well instead of everything poorly.
Why this piece earns its place
Splitting the work invents a failure a single agent cannot have: mis-routing. A pricing question goes to the researcher, every downstream step succeeds, the handoff validates, and the answer is confidently about the wrong thing. Score routing separately from answer quality, or an aggregate number sags and never tells you which half moved. A specialist is also a versioned artifact rather than a prompt in a file — a prompt, a model choice and a tool allowlist, where editing any of the three changes how the whole system behaves. Pin all three into the run record, or yesterday’s run cannot be reproduced and a regression cannot be bisected to the change that caused it. That allowlist is the strongest reason to keep a specialist narrow. Context isolation is the quality argument; permission isolation is the safety one, and it buys something specific rather than everything. If only the writer holds the publish tool, an instruction smuggled into a page the researcher fetched cannot publish. It can still steer the researcher, and an outbound fetch is a channel, so what that agent holds can still leave. The split bounds the blast radius; it does not clean the input.
- split ondistinct skills / parallelism
- isolated contexteach agent stays focused
- cost of agentslatency + failure surface
What the new pieces do
- Researchermodel
- A specialist agent with its own prompt, tools, and isolated context — e.g. gathers information. Focused scope means a smaller, cleaner context and better results than one do-everything agent.
- Analystmodel
- Another specialist — e.g. reasons over gathered data. Receives a typed handoff from upstream, not a raw dump, so quality doesn’t degrade through a game of telephone.
- Writermodel
- A specialist that produces the final artifact from the analysts’ structured output. Independent specialists can also run in parallel when their tasks don’t depend on each other.
Step 3 · How agents coordinate
Shared state vs. message passing
The researcher’s findings have to reach the analyst; the analyst’s conclusions have to reach the writer. If you just paste everything into every prompt, contexts explode and quality rots. How do agents share information?
How should agents pass work between each other?
A shared store holds task results and intermediate artifacts; agents read what they need and write structured outputs, so no single context has to carry everything and handoffs stay clean.
Piling full transcripts forward blows up context length and drags irrelevant detail through the pipeline — cost and quality both suffer. Pass structured results, not raw history.
Without shared state agents can’t build on each other, so the "team" produces disconnected fragments. Coordination requires a shared substrate.
Agents coordinate through Shared State (a blackboard): each reads what it needs and writes structured artifacts — results, not raw transcripts. The orchestrator can also message-pass explicit handoffs. Either way, keep contexts small by sharing distilled outputs, so adding an agent doesn’t multiply everyone’s prompt size.
Why this piece earns its place
Everything written here has to be plain data — no open handles, no live connections, no callback an agent means to invoke later. State is exactly what the next step checkpoints, and a run holding a socket on the blackboard cannot resume on another machine however good the workflow engine is. The rule is free to adopt now and miserable to retrofit once three agents are passing objects around. A read from the board is a selection, not a fetch. State grows with every step, so an agent that reads all of it has rebuilt the exact problem this step exists to avoid — the bloated context is merely assembled from a store now instead of concatenated by hand. Give each artifact a key and a type, have the orchestrator hand an agent only the keys its task names, and keep bulky material behind a reference so what lands in a prompt is an id and a type rather than a scraped page. And keep the board scoped to one run. Once a shared store exists, the pull is to let it double as memory across runs, which quietly makes one goal’s leftovers visible to the next goal that reads the same key.
- blackboardshared read/write state
- structured artifactsnot raw transcripts
- small contextsshare distilled outputs
What the new pieces do
- Shared Statestore
- The shared context agents read from and write to — task results, intermediate artifacts, a scratchpad. How agents coordinate without cramming everything into one prompt.
Step 4 · Survive failure
The durable workflow engine
A run is twenty steps and five minutes long. Step 14 calls a flaky API and the process crashes. Do you really restart from step 1 — re-paying for all that work? Agentic runs are long and partial failure is constant. How do you make them reliable?
How do you keep a long multi-agent run reliable across failures?
Restarting a twenty-step run because step 14 failed re-does (and re-pays for) all prior work and may never finish if the flake recurs. You need to resume, not restart.
A workflow engine persists state after each step; a crash or a failed tool call retries that step or resumes from the last checkpoint, instead of throwing away completed work.
In-memory-only runs lose everything on any crash or deploy — unacceptable for long agentic workflows. Durability requires checkpointing.
Run the orchestration on a Workflow Engine with durable execution: checkpoint state after each step so a crash or a flaky tool retries that step or resumes from the last good state — never restarting a long run from scratch. Agentic workflows are long-lived and failure-prone; durability (à la Temporal/Step Functions) is what makes them production-grade.
Why this piece earns its place
Durable execution works by replaying a recorded history to rebuild state, and that puts one hard rule on agentic runs: a model call is non-deterministic, so it has to be recorded as a result and replayed from the record, never re-invoked. Get it wrong and a resumed run continues under a different plan than the half it already executed — the most confusing incident this design can produce, because every individual step still looks correct. Tools follow the same reasoning. The engine gives you at-least-once, not exactly-once, so a step that filed a ticket or sent a mail will eventually be retried; the idempotency key belongs on the tool call, and the completion record has to be written before anything acknowledges success. Surviving a deploy is the property that creates the next problem. A run that finishes in five minutes can simply be drained before a release; a run parked on a human approval waits days and comes back into code that is not the code it started in, and if the step sequence moved underneath it the replay diverges. So workflow definitions get versioned and pinned per run, and any change to the shape of a plan is a migration with in-flight runs to carry.
- checkpointstate after each step
- retry / resumefrom last good state
- durablesurvives crashes + deploys
What the new pieces do
- Workflow Engineservice
- Runs the orchestration as a durable workflow: each step is checkpointed so a crash or a flaky tool can retry or resume from the last good state instead of restarting the whole run.
Step 5 · Clean handoffs
Typed contracts between agents
The researcher hands to the analyst hands to the writer. If each handoff is freeform prose, meaning blurs at every hop — a game of telephone that degrades over the chain. How do you keep quality across handoffs?
How do you stop quality degrading as work passes between agents?
Prose summaries lose and distort information at each hop — the telephone problem. Handoffs need structure the next agent can rely on, not a re-narration.
Raw history bloats context and still buries the signal. The fix is a distilled, structured contract, not more data.
Each agent emits a defined structured output (schema/fields) that the next agent consumes, so information passes precisely and predictably instead of blurring through freeform text.
Define typed handoff contracts: each agent produces a structured output (a schema — fields, not freeform prose) that the next agent consumes. Structured handoffs stop the telephone-game degradation, make each step testable in isolation, and let the orchestrator validate an agent’s work before passing it on.
Why this piece earns its place
A schema checks shape, not truth. A perfectly valid artifact can be entirely invented — the analyst returns every required field, with figures no source ever contained, validation passes, and the writer treats them as fact. So put provenance in the contract itself: which source, which tool call produced the claim, how sure the agent is. Then the next agent can weigh it, the orchestrator can refuse to forward an unsourced number, and the final artifact can point at something. Fields nobody reads on day one are what make that possible in month six. Versioning these contracts feels like ceremony between two components you own, right up to the deploy that kills a day-old run. They sit between agents whose prompts change on different days, and — because of the previous step — old-shape artifacts are sitting inside checkpoints belonging to runs that have not finished, so one new required field invalidates every one of them. Treat it like any wire format: new fields arrive optional, removals take two releases, and a consumer reads the version it is handed.
- typed outputschema, not prose
- no telephoneprecise handoffs
- testablevalidate each step
Step 6 · Do things at once
Parallel fan-out / fan-in
Three of the sub-tasks don’t depend on each other — researching three vendors, say. Running them one after another triples the latency for no reason. How do you exploit independence?
How do you speed up independent sub-tasks?
Forcing sequence on independent tasks wastes time — they could run at once. Parallelism is a major latency win when sub-tasks don’t depend on each other.
The orchestrator dispatches independent sub-tasks concurrently to multiple agents (fan-out), waits for them, and merges their structured outputs (fan-in) — cutting latency for the parallelizable part of the plan.
Parallelizing dependent steps breaks correctness — the analyst can’t run before the researcher finishes. You parallelize only the independent parts; the DAG of dependencies dictates what can overlap.
Exploit independence with fan-out / fan-in: the orchestrator runs independent sub-tasks in parallel across agents, then aggregates their structured results. Dependencies form a DAG — parallelize the independent branches, sequence the dependent ones. It’s the map-reduce shape, applied to agents, and often the biggest latency lever in a multi-agent plan.
Why this piece earns its place
Fan-in finishes when the slowest branch finishes, so a parallel stage takes the maximum of its branches where the serial version took their sum. That helps at every percentile, but not evenly: each branch you add is another chance that one of them is having a bad minute, so the median improves far more than the tail does — and the tail is what a user feels. A per-branch deadline with a defined fallback — a cached result, a branch marked incomplete, a shorter prompt — matters more here than anywhere else in this design, because the run inherits the worst minute any branch had. Money doesn’t move at all. Three vendors researched at once bills what three researched in sequence bills; fan-out spends the same tokens sooner, and anyone arguing for it on efficiency grounds has confused latency with cost. What a test run won’t show you is contention. Fan-out turns one run into several concurrent model calls, and a provider’s rate limit is per account, not per run, so one wide plan starves every other run in the system and the symptom arrives as somebody else’s latency. That wants a concurrency bound per run and a fair share across runs — in practice a queue and a worker pool, not firing every independent branch the instant the plan says it may go.
- fan-outindependent tasks in parallel
- fan-inaggregate results
- DAGparallel where independent
Step 7 · Put a ceiling on autonomy
Budgets, loop detection, termination
Agents that can spawn agents and retry tools have no natural stopping point. Two agents can correct each other forever; a planner can delegate infinitely; a retry can loop. Nothing errors — it just spins and spends. What stops a runaway?
What keeps a multi-agent system from looping and burning unbounded cost?
Hard caps on steps, recursion depth, and token/dollar spend, plus loop detection and approval gates for consequential actions, bound the run — so autonomy always has a ceiling and can’t spin forever.
LLM agents frequently don’t recognize completion or detect their own loops — that’s exactly how runaways happen. Termination must be enforced by the system, not hoped for.
A timeout alone still lets a run burn a fortune before it fires, and doesn’t catch tight loops early. You need step/depth/cost caps and loop detection, not just a wall clock.
Enforce termination at the system level: per-run budgets (max steps, max recursion depth, token/$ caps), loop detection (spot repeating states/ping-pong), and human-approval gates for consequential or expensive actions. Autonomy needs both a floor and a ceiling — this is the single most important safeguard, and removing it (the chaos button) is how multi-agent systems fail.
Why this piece earns its place
Where the cap is checked decides whether it holds. Spend is only known after a call returns, so a budget read at dispatch is always reading a stale number, and with branches running at once several of them can each see room that only one of them actually had. Put one counter on the run and write to it twice: debit an estimate atomically as you dispatch, then reconcile that debit against real usage when the call completes. Anywhere else and a wide plan overshoots its ceiling every time, by more the wider it is. Hitting a cap is not an error, which is the part that catches teams out. Nothing throws, no alert fires, the run simply ends with what it had — so in ordinary telemetry a capped run and a finished one look identical. Track the cap-hit rate as a first-class number, split by which cap fired and by the shape of the workflow that fired it. If a real share of runs end on the step ceiling rather than finishing, that ceiling has quietly become the product’s quality limit, and you would much rather read that on a dashboard than hear it from a user saying the answers got thinner.
- budgetssteps · depth · token/$
- loop detectioncatch ping-pong
- approval gateshuman on big actions
What the new pieces do
- Budget / Stopservice
- Enforces the floor and ceiling on autonomy: max steps, max recursion depth, token/$ budget, loop detection, and human-approval gates — the difference between a bounded run and a runaway one.
Step 8 · See the whole run
End-to-end tracing and evaluation
A multi-agent run failed to produce a good answer. Which agent? Which handoff? Which tool call? With work spread across many agents and steps, a single log line tells you nothing. How do you debug and evaluate the whole thing?
How do you make a multi-agent run debuggable and evaluable?
Final outputs hide where a multi-step, multi-agent run went wrong — you can’t see the handoff or tool call that broke it. You need the full tree.
One message can’t explain a failure distributed across agents and steps. Debugging needs the structured trace of the entire run.
A structured trace of the whole run (each agent, prompt, tool call, handoff, token/cost) makes failures debuggable and lets you evaluate the end-to-end workflow, not just individual model calls.
Capture a Run Trace: the full tree of the workflow — every agent, prompt, tool call, handoff, and cost — so a failure is traceable to the exact step, and the whole workflow can be evaluated end to end (not just per model call). Multi-agent quality is a property of the system, so observability and eval have to span the entire run.
Why this piece earns its place
Sampling is where an agent trace stops behaving like a request trace. A low span sample rate tells you plenty about an HTTP service and nothing at all about a run, because half a tree is unreadable — you cannot see which handoff went wrong when you kept the handoff and dropped the step that produced it. Sample by run, never by span, and keep every failed or capped run whatever the rate says. The skeleton and the payloads then want different lifetimes. A run’s shape — its steps, timings, costs and which agent did what — is small and worth keeping for a long time. The prompts and tool responses hanging off it are bulky and far more sensitive, so they get a shorter clock of their own; keeping them forever is how an observability bill outgrows the model bill. This tree is also the only place unit economics can be computed. Cost per run, attributed to a tenant and to the plan shape that produced it, is what tells you which workflows deserve which caps — without it, the budgets in the previous step are numbers somebody guessed.
- run treeagents · tools · handoffs · cost
- debuggablefailure → exact step
- eval end-to-endthe workflow, not one call
What the new pieces do
- Run Tracestore
- The full tree of the run — every agent, prompt, tool call, handoff and cost — so a multi-agent failure is debuggable and the whole workflow can be evaluated end to end.
The payoff
You built multi-agent orchestration
From "one agent can’t hold it all" to a coordinated system: an orchestrator over focused specialists (used only when the task warrants it), shared-state coordination via structured artifacts, a durable workflow engine with checkpoint/retry, typed handoffs, parallel fan-out/fan-in, hard budgets + termination, and end-to-end tracing.
Now uncap the agents — remove budgets and termination — and watch the system loop: agents correcting each other, planners delegating endlessly, retries spinning, tokens burning, with no error and no answer. That’s why termination conditions are the most important safeguard in multi-agent design, and why autonomy must always sit between a floor and a ceiling.
Everything you assembled, in order
- Orchestrator-worker — one owner of the plan, state, and stop condition
- Use specialists wisely — split on distinct skills/parallelism — not by default
- Shared state — coordinate via structured artifacts, not forwarded transcripts
- Durable workflow — checkpoint each step; retry/resume, never restart
- Typed handoffs — structured contracts stop telephone-game degradation
- Fan-out / fan-in — the plan is a DAG — parallelize the independent branches
- Budgets + termination — steps, depth, cost caps, loop detection, approval gates
- Trace + eval — the run tree makes the workflow debuggable and evaluable
- The failure — uncapped autonomy loops and burns unbounded cost, silently
