Vibe Engines
YouTube
AI System Design

Design an AI Agent

Learn AI system design by building an autonomous AI agent step by step.

The numbers to beatN toolsexposed to the model1 callper loop steptypedargs validated

The whole design, in writing

Learn AI system design by building an autonomous AI agent step by step. An interactive guide covering the plan-act-observe loop, tool calling with typed schemas, sandboxed execution, short- and long-term memory, error recovery and retries, and the cost/latency/safety controls that keep a multi-step agent from running away.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

What makes it an "agent"?

A chatbot answers a question in one shot. An agent is handed a goal — "find the cheapest flight and book it" — and has to figure out the steps, take actions in the real world, react to what happens, and keep going until it’s done. How do you build something that acts, not just answers?

Usersets a goal
New in this step: User.

Wrap the model in a loop: ask it for the next action, run that action with real tools, feed the result back, and repeat. The model supplies the reasoning; the loop supplies the hands, the memory, and the brakes.

What the new pieces do

Userclient
A person who states a goal ("book me a flight under $400") rather than a single question.

Step 1 · The skeleton

A goal and a loop

The user gives a goal, not a question. A single model call can’t book a flight — it can only produce text. What structure turns "produce text" into "get something done"?

Usersets a goalAgent Loopplan · act · obs
New in this step: Agent Loop.

The user states a multi-step goal. What runs it?

  1. The model can’t check prices, call APIs, or run code on its own — it only emits tokens. Real tasks need actions between thoughts.

  2. The agent loop alternates thinking and doing: decide → act → observe → repeat, until the goal is met. This is the spine of every agent.

  3. A hardcoded script can’t adapt when a flight is sold out or a tool errors. The point of an agent is deciding the next step dynamically.

An Agent Loop (orchestrator) drives everything: it asks the model what to do next, executes that action, observes the result, and loops — until the model says the goal is complete. The model decides; the loop does.

Why this piece earns its place

Write this loop as a function inside a request handler and it works until the first deploy lands mid-run. A run takes minutes, sometimes longer, so the process holding it will restart while runs are in flight, and that is not a crash you can replay from the top — the agent has already acted in the world, and starting over books the flight a second time. So the loop keeps its state outside itself: a persisted run record holding the goal, the step cursor and the observations so far, with every tool call carrying an idempotency key derived from run_id + step, so a resumed run re-sends a call it may already have made and gets the first answer back instead of a second booking. That record is also the only thing that can report progress while the run is still going, which is what stands between the user and a blank screen on a task that is working. None of this is expensive to build on day one. All of it is expensive to add later, against side effects already loose in production.

What the new pieces do

Agent Loopbackend
The control loop. Calls the model to decide the next action, runs it, feeds the result back, and repeats until done.

Step 2 · Let it decide

The model picks the next action

Inside the loop, something has to choose what to do next — search? query a DB? finish? That judgment is exactly what the LLM is good at. But how does free-form text become a concrete, runnable action?

Agent Loopplan · act · obsLLM (Planner)decides actions
New in this step: LLM (Planner).

How does the agent turn the model’s reasoning into an actual action?

  1. Brittle and ambiguous — "I should probably search" isn’t a runnable call. You need structured output, not text mining.

  2. Powerful but dangerous and unbounded. You want the model to request specific tools, validated before they run — not a blank shell.

  3. The model outputs a function call — tool name and JSON arguments matching a schema — which the loop can validate and execute deterministically.

The LLM Planner is given the goal, the history so far, and the list of available tools (each with a typed schema). It responds with a structured tool call — a name and JSON arguments — or signals the task is done. Structured output is what makes the next step runnable.

Why this piece earns its place

Everything the loop puts in front of the model here is prompt, including the parts that look like configuration. The tool list is re-sent on every step of every run, so its size is a tax paid per step rather than per run, and it competes with the growing history for the same window — the right number of tools is as few as will do the job, and once several of them overlap, the model’s choice between them gets less reliable rather than more. Past that point the fix is to retrieve a relevant subset per step, or to put one façade over several tools, instead of listing everything. A tool’s description is prompt too: rewording it changes which tool gets picked, so it belongs under the same review and the same evals as the instructions. The schema is the part that is also an interface. Renaming an argument does not break a run in flight — the model is re-prompted with the current shape on its next step — but it does break any call that was recorded and is executed later, and every stored trace and eval fixture written against the old shape. Version the schema, stamp the version on the call, and keep the old version runnable until those have drained.

  • N toolsexposed to the model
  • 1 callper loop step
  • typedargs validated

What the new pieces do

LLM (Planner)service
The reasoning model. Given the goal and history, it picks the next tool to call (or declares the task finished).

Step 3 · Give it hands

Execute tools — safely

The model asked to run a tool. But its arguments might be malformed, the tool might be dangerous (delete data, spend money), and it might hang. You can’t just blindly run whatever it emits. How do you execute safely?

Tool Executorsandbox + guardTools / APIssearch · code · db
New in this step: Tool Executor, Tools / APIs.

The model requests a tool call. How do you run it?

  1. The arguments may be invalid, the call destructive, or the tool may hang forever. Unchecked execution is how agents cause real damage.

  2. A Tool Executor checks the arguments, runs the tool in a sandbox with limits, and returns a structured result OR a structured error — both readable by the model.

  3. Approval matters for dangerous actions, but gating everything kills autonomy. Guard the risky calls (step 6); run the safe ones automatically.

A Tool Executor sits between the model and the world. It validates the requested arguments against the tool’s schema, runs the tool in a sandbox with a timeout, and returns a typed result — success or a structured error. The Tools/APIs are the agent’s hands: search, code, databases, third-party calls.

Why this piece earns its place

Both directions through this boundary need work and the inbound one usually gets none. Outbound, the model never holds a credential: it names a tool, and the executor attaches the key — so what a run is able to do is decided by one component rather than by the model’s judgment, and the authority worth attaching is the authority of whoever started the run, narrowed to the tools this run was granted. Inbound, whatever a tool returns is about to be read by the model as if it were part of the conversation, so a fetched web page or a support ticket can carry text addressed to the agent; the executor is the only place that can mark a result as data rather than instruction. It is also the only place that can bound a result’s size, because a single query can return more than the window holds: cap what goes back, keep the full payload in storage, return a handle the model can ask to expand. That is a different job from the trimming in the next step. The executor bounds one result at the moment it is produced and never sees the run’s history; working memory decides what survives across many steps and never sees the full payload.

What the new pieces do

Tool Executorservice
Validates the model’s requested call against a schema, runs it in a sandbox with timeouts, and returns a typed result.
Tools / APIsservice
The agent’s hands: web search, code runner, database queries, third-party APIs — anything that touches the world.

Step 4 · Close the loop

Observe, then think again

The tool ran and returned something — a result, or an error. The model doesn’t know what happened yet. How does the outcome get back into its reasoning so it can decide the next move?

Agent LoopTool ExecutorTools / APIsTrace / Event Log
New in this step: Trace / Event Log. · swipe to pan the diagram

A tool returned a result (or failed). What happens next?

  1. If the agent ignores outcomes it can’t react — it’ll happily book a sold-out flight. The observation MUST feed the next decision.

  2. The result (or typed error) is appended to the agent’s history and handed back to the model, which reasons about it and picks the next action. That’s the "observe" in plan-act-observe.

  3. Blind infinite retries are exactly the death spiral guardrails exist to stop. Observe, adapt, and bound the attempts.

The observation — result or typed error — is appended to the run’s history and fed back to the model, which reasons over it and chooses the next action. Every loop iteration is also written to a Trace Log, so the whole chain of thought→action→result is replayable and debuggable.

Why this piece earns its place

The model re-reads everything before it on every step, so the tokens a run processes grow with the square of its length — step twenty pays for steps one through nineteen again. The bill does not have to follow that curve, because the re-read part is a stable prefix and that is what prompt caching exists for, which turns the cost question into a different one: caching only pays while the front of the prompt is byte-identical from step to step. A timestamp in the system message, tools listed in a new order, a history edited in the middle rather than appended to the end — each quietly cancels the discount, and nothing in the telemetry says so. Append-only is a cost decision as much as a correctness one. The trace is the other artifact here, and it should not share a code path with the history. The history is lossy and edited, because its job is to fit; the trace is complete and append-only, because its job is to explain the run afterwards. Keep it outside the run and give it a retention policy — it holds raw prompts and tool payloads, usually the most sensitive data the system touches.

What the new pieces do

Trace / Event Logbus
Records every thought, action and observation — for debugging, evals, cost tracking and replay.

Step 5 · Make it remember

Working memory vs long-term memory

Within one task the agent must remember what it has already tried. Across tasks, it’s wasteful to relearn the same facts every time. These are two different memory needs. How do you serve both?

Agent LoopTool ExecutorTools / APIsWorking MemoryLong-term Memory
New in this step: Working Memory, Long-term Memory. · swipe to pan the diagram

The agent needs to remember within a task AND across tasks. How?

  1. A long task overflows the context window, and nothing persists once the run ends. The prompt is working memory, not storage.

  2. That loses the reasoning the agent needs mid-task and the reusable knowledge it could recall later. You need both scopes, not just the output.

  3. A working-memory scratchpad holds the live task; a vector-backed long-term store lets the agent recall facts and past outcomes in future tasks.

Two memories. Working Memory is the scratchpad for the current run — goal, steps taken, observations — trimmed to fit the window. Long-term Memory is a vector store the agent can write durable facts and outcomes to and retrieve in later tasks (RAG, applied to the agent’s own experience).

Why this piece earns its place

An agent that writes its own conclusions can write a wrong one — an inference drawn from a misread tool result — and then recall it as settled fact in every later run, with nothing in those runs explaining where it came from. Two habits keep that survivable. Prefer recording what the world said (a confirmation code, a stated preference, an outcome a tool reported) over what the model concluded from it, and stamp every entry with the run that wrote it, when, and which observation produced it, so a bad memory can be found and deleted rather than argued with. Entries need a shelf life as well: a price or a policy is true for a while, not forever. Then there is when the write happens. Let the model save mid-run and a run that fails later leaves its conclusions behind; consolidating at the end of a run, as a separate job reading the trace, keeps abandoned reasoning out of the store. And label a retrieved memory as recall where it enters the prompt, or it arrives looking exactly like something the user said this run, and a stale one quietly outranks the person in front of it.

What the new pieces do

Working Memorystore
The scratchpad for the current task: the goal, the steps taken so far, and their observations.
Long-term Memorystore
Vector-backed memory of facts and past outcomes the agent can recall in future tasks.

Step 6 · Don’t let it run away

Budgets, limits and approvals

A looping model can spin forever, rack up a huge bill, or take a destructive action. Autonomy without limits is a liability. How do you keep an agent on a leash without making it useless?

Agent LoopguardedBudget & Guardlimits · approval
New in this step: Budget & Guard.

What stops an agent from looping forever or doing something destructive?

  1. Models miscount steps, get stuck in loops, and underestimate consequences. Safety can’t depend on the thing you’re trying to constrain.

  2. Hard limits cap steps, tokens and cost; an approval gate pauses for human sign-off before destructive or expensive actions. Independent of the model’s judgment.

  3. Then it’s not an agent. The goal is bounded autonomy — let it loop, but inside guardrails that make runaway impossible.

A Budget & Guard layer wraps the loop: hard caps on steps, tokens and spend, plus approval gates that pause for human sign-off before destructive or costly actions (sending money, deleting data, emailing customers). Bounded autonomy — free to loop, impossible to run away.

Why this piece earns its place

A gate is worth exactly as much as the decision it collects. Show a reviewer a tool name and a blob of arguments and approval becomes a reflex — the control keeps firing and stops meaning anything, which is the failure mode nobody notices, because the dashboard still shows every action reviewed. So the request has to carry the action in plain language, the arguments that will actually run, the reasoning that produced them, and what it costs to be wrong. Bind the decision to those arguments: an approval approves one payload, and a call that changes afterwards is a new call needing a new approval. Record who decided and when, since that trail is half of why the gate exists. Pending approvals need a default too — nobody answers by Monday, and the run should expire rather than sit open forever. Budgets have a scope problem rather than a value problem. The per-run cap is the one everyone implements, and it does nothing about one user, or one caller retrying a failed job, starting run after run; the caps that save you are per user and per tenant per day. Check the budget before a step that changes the world rather than after, and hold back enough allowance to write the ending.

What the new pieces do

Budget & Guardservice
Caps steps, tokens and spend; gates dangerous actions behind approval. The agent’s seatbelt.

Back of the envelope

max steps / run
hard stop on the loop — no infinite spinning
token + $ budget
cap spend per task; halt and report if exceeded
approval gate
human sign-off before irreversible/costly actions
allow-list tools
the agent can only call what you’ve granted

Step 7 · Finish cleanly

Done, or give up gracefully

Eventually the agent either achieves the goal or can’t. Both endings need to be deliberate — not "loop exhausted, returns garbage." How does a run actually end?

Usersets a goalAgent Loopguarded
New in this step: Agent Loop → User.

How should an agent run terminate?

  1. Dumping a half-finished buffer at the user is the worst ending. Termination should be a decision, not an accident.

  2. Some tasks are impossible (sold out, missing access). An agent that can’t admit defeat loops forever or lies. Graceful failure is a feature.

  3. The model signals completion with the final answer; if it’s blocked or hits a limit, the loop stops and reports what it tried and why it couldn’t finish.

A run ends two ways: the model emits a "task complete" with the final answer, or the agent hits a wall (blocked, out of budget, repeated failures) and stops cleanly — reporting what it attempted and why it couldn’t finish. Both are returned to the user deliberately. No silent garbage.

Why this piece earns its place

The model is the wrong narrator for half of this. It can say what it was trying to do and where it got stuck, and that account is worth having — but what actually happened out in the world is something it believes rather than something it knows, and a run that has just failed is where that gap is widest. The executor recorded every call that succeeded, so the loop can state the side effects from the run record and let the model supply the intent around them: what was done, what was left half done, and what somebody now has to undo or finish by hand. The status itself belongs in a field rather than a sentence — done, blocked, budget exhausted, approval refused, tool unavailable — because that is what turns a pile of runs into a number worth acting on: what fraction finish, and which wall the rest hit. A closing paragraph aggregates into nothing, and adding the field later means mining it back out of old logs. Blocked is the one to design hardest, because it is a product surface: name the single missing thing, and a person can supply it and let the run carry on from where it stopped rather than restart against a goal that is now partly done.

The payoff

You built an AI agent

From a one-shot chatbot to an autonomous agent: a plan-act-observe loop, structured tool calls, a guarded executor, two scopes of memory, full tracing, hard budgets, approval gates, and graceful endings.

UserAgent LoopLLM (Planner)Tool ExecutorTools / APIsWorking MemoryLong-term MemoryBudget & GuardTrace / Event Log
The finished design, end to end. · swipe to pan the diagram

Now break a tool and watch the agent hit an error mid-loop — and see why the executor’s typed errors, bounded retries and step budget are the difference between recovery and a runaway death spiral.

Everything you assembled, in order

  • Agent Loop — plan → act → observe → repeat until done
  • Tool calling — model emits typed, schema-checked function calls
  • Tool Executor — the airlock: validate, sandbox, timeout, typed result
  • Observe + Trace — feed results back; log every step for debug/eval
  • Two memories — working (this run) + long-term (across runs)
  • Budget & Guard — step/token/$ caps + approval on risky actions
  • Graceful end — "done", or stop cleanly and report when blocked
  • Error handling — typed errors + bounded retries beat death spirals

Deep cut · 25:41

It said booked. Nobody booked it.

The interactive build above lays out an AI agent: a plan-act-observe loop around the model, structured tool calls checked against a schema, one tool executor that validates, times out and returns a typed result or a typed error, every observation fed back, a trace of every step, working and long-term memory, and a budget, a step limit and an approval gate around it all. The film builds the same agent one fix at a time in a made-up travel app’s courtyard, then replays one Friday night lap by lap to find why an agent that read every result announced a booking nobody made.

  • See the wreck: the booking call timed out on a swamped airline, and the executor sent back whatever the airline had said — nothing — in the same envelope as a success. The model read a success after “book”, filled the gap with the flight, seat and price it already knew, and said done; the loop delivered it, because the model decides when it is done. Every brake held.
  • Take it into the interview: a timeout is an error, not an empty result, and for a write it means unknown — look before you retry; “done” is checked by the loop against ground truth from the world (a confirmation code), not trusted from the model.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. Why does the tool schema validation happen in the Tool Executor rather than just trusting the model to always emit well-formed arguments?

    Models are probabilistic text generators — even a well-trained one occasionally emits an argument of the wrong type, a missing required field, or a value outside a valid range, especially deep into a long agent run where context has accumulated and drifted. Validating against the schema BEFORE execution catches these malformed calls deterministically and returns a clear, typed error the model can actually reason about and correct — rather than either crashing on bad input or, worse, silently executing a tool with subtly wrong arguments that produces a plausible-looking but incorrect result.

  2. Why maintain two separate memory scopes (working memory and long-term memory) instead of just always writing everything to the long-term vector store?

    Working memory's value is being complete and immediately available for THIS run — every step and observation, in order, without needing a retrieval step, because reasoning about what just happened benefits from having it directly in context. Long-term memory's value is being selective and durable ACROSS runs — only facts and outcomes worth remembering later, retrieved by relevance rather than recency. Writing everything indiscriminately to long-term memory would flood it with run-specific noise (intermediate tool outputs, abandoned approaches) that pollutes future retrieval — the two scopes serve genuinely different purposes and mixing them degrades both.

  3. Could the step/token/cost budget alone (without a separate approval gate) sufficiently protect against a destructive action?

    No — a budget caps HOW MUCH the agent can do, but says nothing about WHETHER a specific action, even well within budget, should be allowed to happen without oversight. An agent could delete a customer's account or send a large payment in a single, cheap, fast tool call that consumes almost none of its budget — the action is destructive because of what it DOES, not because of how expensive it was to attempt. This is why approval gates are a separate mechanism scoped to action CONSEQUENCE, layered on top of budgets that are scoped to resource CONSUMPTION — the same two-axis distinction (how much vs. how risky) that shows up in several other agent designs in this series.

  4. Why does the agent write a trace of "thoughts" (the model's reasoning), not just the actions it took and their results?

    When an agent produces a wrong or surprising final outcome, the actions and results alone often don't explain WHY it chose that path — two runs could take the exact same sequence of tool calls for completely different (one sound, one flawed) reasons. Capturing the model's stated reasoning at each step lets a developer debugging a bad run see not just what happened but what the model THOUGHT was happening, which is usually where the actual bug or misunderstanding lived — a wrong action is often the symptom, and the reasoning trace is what reveals the actual cause.

  5. How would this design need to change to support multiple agents collaborating on one goal, rather than a single agent looping alone?

    The core loop (plan → act → observe) stays the same for each individual agent, but a multi-agent version needs an additional coordination layer above the single-agent orchestrator — deciding how sub-goals are divided, how one agent's output becomes another's input, and how the shared budget and approval gate apply across ALL the agents combined rather than just one. This is a genuinely different system design problem (see Multi-Agent Orchestration for the coordination-specific concerns) — this design is intentionally scoped to what one agent, alone, needs to act safely and effectively, which is itself a prerequisite for getting multi-agent coordination right.

Check yourself — the answers, and why

Eight steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. What turns a chatbot into an agent?

    Agents alternate thinking and doing — the plan-act-observe loop, with real tools — rather than answering in one shot.

  2. Why does the model emit a structured tool call instead of prose?

    A typed tool name + JSON arguments can be schema-checked and run; free-form prose can’t be executed reliably.

  3. The Tool Executor exists mainly to…

    Centralizing execution gives you one place to validate arguments, sandbox, time out, and return typed results/errors.

  4. Budgets and approval gates provide…

    Step/token/$ caps stop runaway loops; approval gates put a human in front of irreversible actions.

  5. The agent's trace log records the model's reasoning, not just its actions and results, because…

    Debugging a bad run needs to see WHY the model chose a path, not just what it did — the reasoning is usually where the actual cause lived.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Take a goal: accept a multi-step goal, not a single question, and figure out the steps.
  • Decide + act: the model emits a structured tool call; the loop runs it and feeds back the result.
  • Execute safely: validate arguments against a schema, run tools sandboxed with a timeout.
  • Remember: working memory for this run, long-term memory across runs.
  • End deliberately: finish with a "done", or stop cleanly and report when blocked.

The qualities that shape everything

Each one names the mechanism that buys it.

Handle tasks no single call could
An Agent Loop (orchestrator) plans, acts, observes and repeats until the model says the goal is met.
Turn reasoning into runnable actions
The model emits a structured tool call (name + typed args matching a schema) the loop can validate and execute deterministically.
Run tools without causing real damage
A Tool Executor validates arguments against a schema, sandboxes with a timeout, and returns a typed result or a typed error.
Debug a non-deterministic, multi-step run
Every thought/action/observation is written to a Trace Log, so the whole chain is replayable and auditable.
Remember within a task and across tasks
Working memory for the current run plus a vector long-term store the agent recalls in future tasks.
Free to loop, impossible to run away
A Budget & Guard layer caps steps/tokens/spend and gates destructive actions behind human approval.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

A loop that acts over one big model call

The model only emits tokens — it can’t check prices, call APIs, or run code. Real tasks need actions between thoughts, decided dynamically.

A structured tool call over parsing the model’s prose

Free-form text isn’t runnable and regex intent-mining is brittle. A typed name + JSON args can be schema-checked and executed deterministically.

Validate, sandbox, timeout over executing exactly what the model emitted

The arguments may be malformed, the call destructive, or the tool may hang. One guarded airlock is where limits, logging and safety live.

Two memory scopes over keeping everything in the prompt

A long task overflows the window and nothing persists once the run ends. Working memory holds this run; long-term memory carries knowledge across runs.

Budgets + approval gates over trusting the model to stop

Models miscount steps, get stuck in loops, and underestimate consequences. Safety can’t depend on the thing you’re trying to constrain.

The answer, out loud

What a strong answer to “Design an AI Agent System” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Pin down what makes it an agent

    Before I draw anything, I want to agree on what we’re building. A chatbot answers one question in one shot. An agent is handed a goal — “book me a flight under $400” — and has to work out the steps, act in the world, react to what happens and keep going until it’s done. What the user feels is whether the task got done, and if it couldn’t be, an honest account of why. So the real question is: how do we let a model act on its own, over many steps, without it doing damage or running forever?

  2. 3–8 min

    The skeleton is a loop

    The skeleton is the user and an agent loop, or orchestrator. It asks the model for the next action, runs it, feeds the result back and repeats until the model says the goal is met: plan, act, observe. Not one big model call — a model only emits tokens; it can’t check a price or call an API. Not a fixed script either — a script can’t adapt when the flight is sold out or a tool errors. The cost: one task becomes many model calls, so I’ll need brakes later.

    Built in step 1: A goal and a loop
  3. 8–14 min

    Turn reasoning into a runnable action

    Inside the loop, the model plans. It gets the goal, the history so far and the tools it may use, each described as a typed function with a schema — a declared shape for its arguments. It answers with a structured tool call: a tool name plus JSON arguments, or a signal that it’s finished. Parsing its prose with regexes is brittle — “I should probably search” isn’t a runnable call — and letting it run arbitrary code is unbounded. A typed call is a checkable decision: which function, with what arguments.

    Built in step 2: The model picks the next action
  4. 14–20 min

    One guarded airlock for every action

    Between the model and the world — search, code, databases, APIs — sits a tool executor. What the model produced is a request, not a command, so I treat it as untrusted input: check the arguments against the schema before anything runs, run the tool in a sandbox — an isolated environment — with a timeout, and hand back a typed result or a typed error the model can read either way. I’m not putting a person in front of every call; that kills the autonomy. One component in front of every call is how the credentials, the limits and the logging end up with somewhere to live.

    Built in step 3: Execute tools — safely
  5. 20–25 min

    Close the loop, and trace everything

    The observation — result or typed error — goes into the run’s history and back to the model, which picks the next move. Assume each call worked and the agent books a sold-out flight; retry the same call forever and it’s a death spiral. Every iteration also goes to a trace log: the thought, the action and what came back. I can’t reproduce a bad run by running it again, because it won’t make the same choices twice — so reading that log back is the whole of my debugging, my evals and my cost reporting.

    Built in step 4: Observe, then think again
  6. 25–31 min

    Two kinds of memory

    Memory has two scopes. Working memory is this run’s scratchpad — goal, steps, results — trimmed to fit the context window, the text the model can read at once. Long-term memory is a vector store, a database that looks things up by similarity of meaning, holding durable facts and past outcomes the agent retrieves in later tasks. Everything in the prompt overflows on a long task and vanishes the moment the run ends. Push everything into the long-term store instead and later runs start recalling this run’s dead ends as though they were knowledge. So the price of two scopes is deciding what is worth keeping, and I’d rather that were an explicit rule than a default.

    Built in step 5: Working memory vs long-term memory
  7. 31–38 min

    Brakes, and a deliberate ending

    I can’t trust the model to stop — models miscount steps, get stuck in loops and underestimate consequences. So a budget-and-guard layer wraps the loop: hard caps on steps, tokens and spend, an allow-list of tools, and approval gates that pause for a human before irreversible actions like sending money or deleting data. Those are two axes: a budget limits how much the agent does, but one cheap call can still delete an account. And a run ends on purpose — “done” with the answer, or a clean stop when blocked or out of budget, reporting what it tried and why it couldn’t finish.

    Built in step 6: Budgets, limits and approvals
  8. 38–42 min

    What I’d watch, and how it fails

    On the dashboard, from the trace: steps, tokens and spend per run, tool errors, runs that end blocked rather than done, and actions waiting on approval. Two failures I’d plan for. A tool starts failing: the executor returns a typed error, retries are bounded and spaced further apart each time, the step budget backstops them, and the agent tries another path or stops cleanly and tells the user. Or the model concludes that deleting something is the shortest route to the goal — that is the case the gate is there for, and the reasoning behind it carries no weight at that boundary. The thing I’d watch there is how often the approval queue sits empty: a gate nothing ever reaches is usually pointed at the wrong actions.

  9. 42–45 min

    Close on the trade-off

    To close, in one breath: a plan-act-observe loop, typed tool calls, one guarded executor, a trace of every step, two scopes of memory, budgets with approval gates, and a deliberate ending. Every knob trades freedom against control — too loose and it runs away, too tight and it’s useless, and where I set it depends on how reversible the actions underneath are. With more time I’d spend it on evaluation: I already keep a trace of every run, so I’d turn a set of those into a regression suite, and then changing a tool, a description or the model is something I measure rather than something a user discovers for me.

What this teaches

Learn AI system design by building an autonomous AI agent step by step. An interactive guide covering the plan-act-observe loop, tool calling with typed schemas, sandboxed execution, short- and long-term memory, error recovery and retries, and the cost/latency/safety controls that keep a multi-step agent from running away.

Key takeaways

  • Agent Loop — plan → act → observe → repeat until done
  • Tool calling — model emits typed, schema-checked function calls
  • Tool Executor — the airlock: validate, sandbox, timeout, typed result
  • Observe + Trace — feed results back; log every step for debug/eval
  • Two memories — working (this run) + long-term (across runs)
  • Budget & Guard — step/token/$ caps + approval on risky actions
  • Graceful end — "done", or stop cleanly and report when blocked
  • Error handling — typed errors + bounded retries beat death spirals

Concepts covered

  • What makes it an "agent"?
  • A goal and a loop
  • The model picks the next action
  • Execute tools — safely
  • Observe, then think again
  • Working memory vs long-term memory
  • Budgets, limits and approvals
  • Done, or give up gracefully
RUN IT YOURSELF

The timeout that comes back as a success

The whole loop on this page in one runnable file — plan, call a tool, observe, repeat, under a step budget and a per-tool timeout — running for real in your browser. It runs the same booking task twice: once with the airline's timeout surfaced as a typed error, once with it swallowed into an empty-but-successful envelope. Flip swallow, or set BUDGET_MS to 20 so even the read-back times out, hit Run, and read which ending is true.

HOW TO READ THE CODE — 4 IDEAS
  1. The loop is plan → execute → observe under a budget: it ends on a finish or on the step cap, never by dumping a buffer (steps 1 and 7).
  2. Every call goes through one airlock: execute checks the arguments against the tool's schema, checks the cost against the timeout, and hands back anything a tool raises as a typed error (step 3).
  3. envelope() is (True, None) for a real success and for a swallowed timeout — nothing that branches on it, rule table or frontier model, can tell those apart (step 4).
  4. The ending is built clause by clause out of the history, then checked claim by claim against the world — so a read-back that itself times out ends unresolved, not "nothing to undo" (steps 6 and 7).
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, break a tool, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs