The whole design, in writing
Learn AI system design by building an autonomous AI agent step by step. An interactive guide covering the plan-act-observe loop, tool calling with typed schemas, sandboxed execution, short- and long-term memory, error recovery and retries, and the cost/latency/safety controls that keep a multi-step agent from running away.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
What makes it an "agent"?
A chatbot answers a question in one shot. An agent is handed a goal — "find the cheapest flight and book it" — and has to figure out the steps, take actions in the real world, react to what happens, and keep going until it’s done. How do you build something that acts, not just answers?
Wrap the model in a loop: ask it for the next action, run that action with real tools, feed the result back, and repeat. The model supplies the reasoning; the loop supplies the hands, the memory, and the brakes.
What the new pieces do
- Userclient
- A person who states a goal ("book me a flight under $400") rather than a single question.
Step 1 · The skeleton
A goal and a loop
The user gives a goal, not a question. A single model call can’t book a flight — it can only produce text. What structure turns "produce text" into "get something done"?
The user states a multi-step goal. What runs it?
The model can’t check prices, call APIs, or run code on its own — it only emits tokens. Real tasks need actions between thoughts.
The agent loop alternates thinking and doing: decide → act → observe → repeat, until the goal is met. This is the spine of every agent.
A hardcoded script can’t adapt when a flight is sold out or a tool errors. The point of an agent is deciding the next step dynamically.
An Agent Loop (orchestrator) drives everything: it asks the model what to do next, executes that action, observes the result, and loops — until the model says the goal is complete. The model decides; the loop does.
Why this piece earns its place
Write this loop as a function inside a request handler and it works until the first deploy lands mid-run. A run takes minutes, sometimes longer, so the process holding it will restart while runs are in flight, and that is not a crash you can replay from the top — the agent has already acted in the world, and starting over books the flight a second time. So the loop keeps its state outside itself: a persisted run record holding the goal, the step cursor and the observations so far, with every tool call carrying an idempotency key derived from run_id + step, so a resumed run re-sends a call it may already have made and gets the first answer back instead of a second booking. That record is also the only thing that can report progress while the run is still going, which is what stands between the user and a blank screen on a task that is working. None of this is expensive to build on day one. All of it is expensive to add later, against side effects already loose in production.
What the new pieces do
- Agent Loopbackend
- The control loop. Calls the model to decide the next action, runs it, feeds the result back, and repeats until done.
Step 2 · Let it decide
The model picks the next action
Inside the loop, something has to choose what to do next — search? query a DB? finish? That judgment is exactly what the LLM is good at. But how does free-form text become a concrete, runnable action?
How does the agent turn the model’s reasoning into an actual action?
Brittle and ambiguous — "I should probably search" isn’t a runnable call. You need structured output, not text mining.
Powerful but dangerous and unbounded. You want the model to request specific tools, validated before they run — not a blank shell.
The model outputs a function call — tool name and JSON arguments matching a schema — which the loop can validate and execute deterministically.
The LLM Planner is given the goal, the history so far, and the list of available tools (each with a typed schema). It responds with a structured tool call — a name and JSON arguments — or signals the task is done. Structured output is what makes the next step runnable.
Why this piece earns its place
Everything the loop puts in front of the model here is prompt, including the parts that look like configuration. The tool list is re-sent on every step of every run, so its size is a tax paid per step rather than per run, and it competes with the growing history for the same window — the right number of tools is as few as will do the job, and once several of them overlap, the model’s choice between them gets less reliable rather than more. Past that point the fix is to retrieve a relevant subset per step, or to put one façade over several tools, instead of listing everything. A tool’s description is prompt too: rewording it changes which tool gets picked, so it belongs under the same review and the same evals as the instructions. The schema is the part that is also an interface. Renaming an argument does not break a run in flight — the model is re-prompted with the current shape on its next step — but it does break any call that was recorded and is executed later, and every stored trace and eval fixture written against the old shape. Version the schema, stamp the version on the call, and keep the old version runnable until those have drained.
- N toolsexposed to the model
- 1 callper loop step
- typedargs validated
What the new pieces do
- LLM (Planner)service
- The reasoning model. Given the goal and history, it picks the next tool to call (or declares the task finished).
Step 3 · Give it hands
Execute tools — safely
The model asked to run a tool. But its arguments might be malformed, the tool might be dangerous (delete data, spend money), and it might hang. You can’t just blindly run whatever it emits. How do you execute safely?
The model requests a tool call. How do you run it?
The arguments may be invalid, the call destructive, or the tool may hang forever. Unchecked execution is how agents cause real damage.
A Tool Executor checks the arguments, runs the tool in a sandbox with limits, and returns a structured result OR a structured error — both readable by the model.
Approval matters for dangerous actions, but gating everything kills autonomy. Guard the risky calls (step 6); run the safe ones automatically.
A Tool Executor sits between the model and the world. It validates the requested arguments against the tool’s schema, runs the tool in a sandbox with a timeout, and returns a typed result — success or a structured error. The Tools/APIs are the agent’s hands: search, code, databases, third-party calls.
Why this piece earns its place
Both directions through this boundary need work and the inbound one usually gets none. Outbound, the model never holds a credential: it names a tool, and the executor attaches the key — so what a run is able to do is decided by one component rather than by the model’s judgment, and the authority worth attaching is the authority of whoever started the run, narrowed to the tools this run was granted. Inbound, whatever a tool returns is about to be read by the model as if it were part of the conversation, so a fetched web page or a support ticket can carry text addressed to the agent; the executor is the only place that can mark a result as data rather than instruction. It is also the only place that can bound a result’s size, because a single query can return more than the window holds: cap what goes back, keep the full payload in storage, return a handle the model can ask to expand. That is a different job from the trimming in the next step. The executor bounds one result at the moment it is produced and never sees the run’s history; working memory decides what survives across many steps and never sees the full payload.
What the new pieces do
- Tool Executorservice
- Validates the model’s requested call against a schema, runs it in a sandbox with timeouts, and returns a typed result.
- Tools / APIsservice
- The agent’s hands: web search, code runner, database queries, third-party APIs — anything that touches the world.
Step 4 · Close the loop
Observe, then think again
The tool ran and returned something — a result, or an error. The model doesn’t know what happened yet. How does the outcome get back into its reasoning so it can decide the next move?
A tool returned a result (or failed). What happens next?
If the agent ignores outcomes it can’t react — it’ll happily book a sold-out flight. The observation MUST feed the next decision.
The result (or typed error) is appended to the agent’s history and handed back to the model, which reasons about it and picks the next action. That’s the "observe" in plan-act-observe.
Blind infinite retries are exactly the death spiral guardrails exist to stop. Observe, adapt, and bound the attempts.
The observation — result or typed error — is appended to the run’s history and fed back to the model, which reasons over it and chooses the next action. Every loop iteration is also written to a Trace Log, so the whole chain of thought→action→result is replayable and debuggable.
Why this piece earns its place
The model re-reads everything before it on every step, so the tokens a run processes grow with the square of its length — step twenty pays for steps one through nineteen again. The bill does not have to follow that curve, because the re-read part is a stable prefix and that is what prompt caching exists for, which turns the cost question into a different one: caching only pays while the front of the prompt is byte-identical from step to step. A timestamp in the system message, tools listed in a new order, a history edited in the middle rather than appended to the end — each quietly cancels the discount, and nothing in the telemetry says so. Append-only is a cost decision as much as a correctness one. The trace is the other artifact here, and it should not share a code path with the history. The history is lossy and edited, because its job is to fit; the trace is complete and append-only, because its job is to explain the run afterwards. Keep it outside the run and give it a retention policy — it holds raw prompts and tool payloads, usually the most sensitive data the system touches.
What the new pieces do
- Trace / Event Logbus
- Records every thought, action and observation — for debugging, evals, cost tracking and replay.
Step 5 · Make it remember
Working memory vs long-term memory
Within one task the agent must remember what it has already tried. Across tasks, it’s wasteful to relearn the same facts every time. These are two different memory needs. How do you serve both?
The agent needs to remember within a task AND across tasks. How?
A long task overflows the context window, and nothing persists once the run ends. The prompt is working memory, not storage.
That loses the reasoning the agent needs mid-task and the reusable knowledge it could recall later. You need both scopes, not just the output.
A working-memory scratchpad holds the live task; a vector-backed long-term store lets the agent recall facts and past outcomes in future tasks.
Two memories. Working Memory is the scratchpad for the current run — goal, steps taken, observations — trimmed to fit the window. Long-term Memory is a vector store the agent can write durable facts and outcomes to and retrieve in later tasks (RAG, applied to the agent’s own experience).
Why this piece earns its place
An agent that writes its own conclusions can write a wrong one — an inference drawn from a misread tool result — and then recall it as settled fact in every later run, with nothing in those runs explaining where it came from. Two habits keep that survivable. Prefer recording what the world said (a confirmation code, a stated preference, an outcome a tool reported) over what the model concluded from it, and stamp every entry with the run that wrote it, when, and which observation produced it, so a bad memory can be found and deleted rather than argued with. Entries need a shelf life as well: a price or a policy is true for a while, not forever. Then there is when the write happens. Let the model save mid-run and a run that fails later leaves its conclusions behind; consolidating at the end of a run, as a separate job reading the trace, keeps abandoned reasoning out of the store. And label a retrieved memory as recall where it enters the prompt, or it arrives looking exactly like something the user said this run, and a stale one quietly outranks the person in front of it.
What the new pieces do
- Working Memorystore
- The scratchpad for the current task: the goal, the steps taken so far, and their observations.
- Long-term Memorystore
- Vector-backed memory of facts and past outcomes the agent can recall in future tasks.
Step 6 · Don’t let it run away
Budgets, limits and approvals
A looping model can spin forever, rack up a huge bill, or take a destructive action. Autonomy without limits is a liability. How do you keep an agent on a leash without making it useless?
What stops an agent from looping forever or doing something destructive?
Models miscount steps, get stuck in loops, and underestimate consequences. Safety can’t depend on the thing you’re trying to constrain.
Hard limits cap steps, tokens and cost; an approval gate pauses for human sign-off before destructive or expensive actions. Independent of the model’s judgment.
Then it’s not an agent. The goal is bounded autonomy — let it loop, but inside guardrails that make runaway impossible.
A Budget & Guard layer wraps the loop: hard caps on steps, tokens and spend, plus approval gates that pause for human sign-off before destructive or costly actions (sending money, deleting data, emailing customers). Bounded autonomy — free to loop, impossible to run away.
Why this piece earns its place
A gate is worth exactly as much as the decision it collects. Show a reviewer a tool name and a blob of arguments and approval becomes a reflex — the control keeps firing and stops meaning anything, which is the failure mode nobody notices, because the dashboard still shows every action reviewed. So the request has to carry the action in plain language, the arguments that will actually run, the reasoning that produced them, and what it costs to be wrong. Bind the decision to those arguments: an approval approves one payload, and a call that changes afterwards is a new call needing a new approval. Record who decided and when, since that trail is half of why the gate exists. Pending approvals need a default too — nobody answers by Monday, and the run should expire rather than sit open forever. Budgets have a scope problem rather than a value problem. The per-run cap is the one everyone implements, and it does nothing about one user, or one caller retrying a failed job, starting run after run; the caps that save you are per user and per tenant per day. Check the budget before a step that changes the world rather than after, and hold back enough allowance to write the ending.
What the new pieces do
- Budget & Guardservice
- Caps steps, tokens and spend; gates dangerous actions behind approval. The agent’s seatbelt.
Back of the envelope
- max steps / run
- hard stop on the loop — no infinite spinning
- token + $ budget
- cap spend per task; halt and report if exceeded
- approval gate
- human sign-off before irreversible/costly actions
- allow-list tools
- the agent can only call what you’ve granted
Step 7 · Finish cleanly
Done, or give up gracefully
Eventually the agent either achieves the goal or can’t. Both endings need to be deliberate — not "loop exhausted, returns garbage." How does a run actually end?
How should an agent run terminate?
Dumping a half-finished buffer at the user is the worst ending. Termination should be a decision, not an accident.
Some tasks are impossible (sold out, missing access). An agent that can’t admit defeat loops forever or lies. Graceful failure is a feature.
The model signals completion with the final answer; if it’s blocked or hits a limit, the loop stops and reports what it tried and why it couldn’t finish.
A run ends two ways: the model emits a "task complete" with the final answer, or the agent hits a wall (blocked, out of budget, repeated failures) and stops cleanly — reporting what it attempted and why it couldn’t finish. Both are returned to the user deliberately. No silent garbage.
Why this piece earns its place
The model is the wrong narrator for half of this. It can say what it was trying to do and where it got stuck, and that account is worth having — but what actually happened out in the world is something it believes rather than something it knows, and a run that has just failed is where that gap is widest. The executor recorded every call that succeeded, so the loop can state the side effects from the run record and let the model supply the intent around them: what was done, what was left half done, and what somebody now has to undo or finish by hand. The status itself belongs in a field rather than a sentence — done, blocked, budget exhausted, approval refused, tool unavailable — because that is what turns a pile of runs into a number worth acting on: what fraction finish, and which wall the rest hit. A closing paragraph aggregates into nothing, and adding the field later means mining it back out of old logs. Blocked is the one to design hardest, because it is a product surface: name the single missing thing, and a person can supply it and let the run carry on from where it stopped rather than restart against a goal that is now partly done.
The payoff
You built an AI agent
From a one-shot chatbot to an autonomous agent: a plan-act-observe loop, structured tool calls, a guarded executor, two scopes of memory, full tracing, hard budgets, approval gates, and graceful endings.
Now break a tool and watch the agent hit an error mid-loop — and see why the executor’s typed errors, bounded retries and step budget are the difference between recovery and a runaway death spiral.
Everything you assembled, in order
- Agent Loop — plan → act → observe → repeat until done
- Tool calling — model emits typed, schema-checked function calls
- Tool Executor — the airlock: validate, sandbox, timeout, typed result
- Observe + Trace — feed results back; log every step for debug/eval
- Two memories — working (this run) + long-term (across runs)
- Budget & Guard — step/token/$ caps + approval on risky actions
- Graceful end — "done", or stop cleanly and report when blocked
- Error handling — typed errors + bounded retries beat death spirals
