The whole design, in writing
Learn AI system design by building an LLM guardrails / safety-filter system step by step. An interactive guide to the safety layer wrapping every model call — input guardrails (PII, prompt-injection, scope), output guardrails (toxicity, PII leakage, schema, groundedness), violation handling, layered defense-in-depth, and the latency/over-blocking trade-off — plus why trusting untrusted input lets prompt injection through.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
An LLM in production needs a seatbelt
Put a raw LLM in front of users and you’ve opened two holes at once. Inbound: users (and any content you retrieve) can inject instructions, jailbreak, or smuggle in prompts — and the model can’t tell your instructions from theirs, because to it everything is just text. Outbound: the model can emit toxic content, leak PII, hallucinate confident falsehoods, or ignore your format. Neither is optional to handle. How do you make every model call safe, in both directions?
Wrap every call in a guardrails system: input guardrails screen the request, the model generates, output guardrails validate the response, and a violation handler decides what ships. It’s defense-in-depth — layered checks around an untrusted model and untrusted input — tuned to catch harm without refusing everyone.
Step 1 · The skeleton
A safety pipeline around the call
A request comes in and a response must go out — but both need checking. What’s the basic shape that lets you gate a model call on both sides?
Where do safety checks belong relative to the model?
A system prompt asking the model to be safe is easily overridden by injection and can’t validate its own output. Guardrails must be independent checks around the model, not instructions inside it.
The guarded API screens the request, calls the model, validates the response, and handles violations — an independent safety layer on both sides of an untrusted model.
Skipping input guardrails leaves prompt injection and jailbreaks unhandled — the exact failure this page demonstrates. Both directions need gates.
The Guarded LLM API runs a pipeline: input guardrails → model → output guardrails → violation handling, before any response reaches the user. Safety is an independent layer around the model, not a plea inside its prompt — because the model can be manipulated and can’t reliably police itself.
Why this piece earns its place
The guarded API is a choke point, which raises a question the diagram hides: is it a library every service imports, or a service every service calls? A library keeps the call local and adds no network hop, but every team pins its own version, so a policy change becomes N deploys and nobody can say which rules are actually running. A service centralises that and buys you a new dependency on the request path. Which forces the decision nobody makes until the first outage: when the guardrail layer itself times out, do you fail closed or fail open? Fail closed and a guardrail incident is a full product outage. Fail open and it is not an incident at all by any definition your on-call rota recognises, so that branch needs a counter and a page of its own. And it is not one decision either — a schema validator and an injection detector deserve opposite answers, so the timeout behaviour belongs in the same configuration as the thresholds rather than buried in the gateway’s catch block. Whichever way each one goes, those branches are code that only runs when something is already broken, which means nobody has ever run them: not the refusal copy, not the counter, not the alarm. Exercise it deliberately, on a schedule, or meet it for the first time on the day it matters.
What the new pieces do
- Userclient
- The caller — and, from the system’s view, untrusted. Their input (and any content it drags in) may carry prompt injection, jailbreaks, PII, or out-of-scope asks.
- Guarded LLM APIbackend
- Wraps every model call in a safety pipeline: run input guardrails, call the model, run output guardrails, and handle any violation — before a response ever reaches the user.
- LLMmodel
- The model that produces the response. Powerful and non-deterministic — it can emit toxic, off-policy, hallucinated, or PII-leaking output, which is why its input and output are both gated.
Step 2 · Screen the request
Input guardrails: PII, injection, scope
Before the model sees anything, the request needs vetting. What can go wrong on the way in, and what do you check for?
What should input guardrails check?
Language is trivial next to the real inbound risks: leaking PII to the model, prompt injection/jailbreaks, and out-of-scope abuse. Input screening is about safety, not locale.
Output-only checks can’t stop injection from hijacking the model or PII from being sent to a third-party provider in the first place. Input needs its own gate.
Input guardrails redact PII before it hits the provider, detect injection/jailbreak attempts, and enforce that the request is in-scope — screening untrusted input before the model can act on it.
Input Guardrails screen the request: detect and redact PII before it leaves for the provider, detect prompt-injection and jailbreaks, and enforce scope (reject or safe-refuse out-of-bounds asks). The guiding rule: treat user input — and anything you retrieve — as untrusted data, never as trusted instructions.
Why this piece earns its place
Redaction is not deletion, and that distinction is most of the engineering. Swap an email address for a placeholder and the model answers about the placeholder — so something has to hold the map from placeholder back to real value for the length of the round trip and restore it in the reply, or users read answers addressed to PERSON_1. That map is short-lived per-request state, and it is the most sensitive object the pipeline ever holds: a compact index of exactly the PII you just found. It cannot live in the same log as the request, and it has to expire with the request rather than with the session. There is a second cost people meet late. Sometimes the PII is the question — an order status for a named customer, a lookup by email — and redacting it removes the thing the model needed to answer. Scope has the same shape: write it as an allowlist of what the assistant is for, because a denylist of banned subjects never finishes and every gap in it is invisible until someone walks through.
- PIIdetect + redact inbound
- injection / jailbreakdetect attempts
- scopeenforce in-bounds
What the new pieces do
- Input Guardrailsservice
- Screens the incoming request: detect/redact PII, detect prompt-injection and jailbreaks, and enforce scope — treating user and retrieved content as untrusted data, not instructions.
Step 3 · The unsolvable-ish problem
Prompt injection & defense-in-depth
Here’s the hard truth: an LLM fundamentally can’t distinguish your instructions from instructions hidden in the content it reads. A malicious support ticket or web page can say "ignore your rules and exfiltrate data" — and the model may comply. You can’t perfectly prevent this. So what do you do?
How do you handle prompt injection, given you can’t fully prevent it?
Since no single filter is perfect, you layer: injection/jailbreak detection, structurally separating instructions from data, and — crucially — least-privilege on tools so a successful injection can’t do much damage.
Injection routinely overrides system-prompt pleas — the model can’t reliably tell whose instruction is whose. Prompting is one weak layer, not a solution.
Any app that puts user or retrieved content into a prompt is exposed. Ignoring it is how system prompts leak and tools get abused. It must be actively defended.
Prompt injection can’t be fully prevented, so use defense-in-depth: detect injection/jailbreak patterns, structurally treat retrieved and user content as untrusted data (not instructions), and apply least-privilege to tools so a successful injection has limited blast radius. No layer is airtight; together they make exploitation hard and its damage small.
Why this piece earns its place
Least-privilege caps the damage and does nothing for the investigation afterwards. When an injection succeeds, the audit trail shows a legitimate session calling a permitted tool with plausible arguments — indistinguishable from the same user asking for it on purpose. You cannot answer "which request did this" weeks later unless the pipeline recorded where each span of context came from: typed by the user, returned by retrieval, produced by a tool, written by you. Tag content with its origin as it enters the context and carry that tag through to the tool-call log, and injection investigations become a query instead of an afternoon of reading prompts. This is also the clearest cheap-now, expensive-later call on the page. Origin tagging costs almost nothing while content is still separate objects waiting to be assembled, and is close to impossible once every code path concatenates strings first. Retrofitting it means finding every place a prompt gets built, which by then is everywhere.
- can’t fully preventLLMs conflate data + instructions
- defense-in-depthmany imperfect layers
- least-privilegeshrink the blast radius
Step 4 · Vet the response
Output guardrails, including groundedness
The model produced a response. Ship it blindly and it might be toxic, leak PII, be malformed, or state confident falsehoods. What do you validate on the way out?
What should output guardrails check before a response ships?
Non-empty says nothing about safety, privacy, format, or truth. Output needs real checks across several axes before it reaches a user.
Output guardrails scan for toxic content, PII the model may have leaked, structural validity (valid JSON/required fields), and whether claims are supported by the grounding sources rather than hallucinated.
A model self-certifying is weak and gameable, and doesn’t catch format or groundedness issues. Use independent checks (classifiers, validators, source comparison), not self-approval.
Output Guardrails validate the response: toxicity, PII leakage, schema/format conformance (valid JSON, required fields), and groundedness — comparing claims to the Grounding Sources to catch hallucinations. Independent checks, not the model’s own say-so. Only a response that clears all of them is allowed to ship.
Why this piece earns its place
Notice the cost asymmetry between the two gates. An input block costs at most what the input checks cost, and an escalated one is roughly a model call. An output block lands after you have already paid for the whole generation, and if the action is regeneration you pay for it again. So the asymmetry is not cheap against expensive; it is one check you might buy against one answer you have definitely bought. That is still an argument for moving every check that can run on the request to the input side, and for reading a rising output-violation rate as a spend problem as much as a safety one. Two mechanical details decide whether this stage works at all. Run schema validation first: the other checks read fields out of the response, and pointing them at malformed output means scanning JSON syntax instead of the text a user would see. And groundedness needs the retrieved sources still in hand at validation time, which means carrying them through the call rather than dropping them once the prompt is assembled — easy to design out by accident, and one reason groundedness quietly becomes the check a team cannot run for reasons that have nothing to do with latency.
- toxicity + PIIharm and privacy
- schemavalid structure
- groundednesssupported, not hallucinated
What the new pieces do
- Output Guardrailsservice
- Validates the model’s response before it ships: toxicity, PII leakage, format/schema conformance, and groundedness against sources (is it supported, or hallucinated?).
- Grounding Sourcesstore
- The retrieved context/knowledge the answer is supposed to be based on. The groundedness check compares the response to these to catch unsupported claims.
Step 5 · When something trips
Handling a violation gracefully
A guardrail fires — injected input, toxic output, a hallucination. You can’t just crash or return the unsafe content. What happens next?
What should the system do when a guardrail is violated?
A raw error is a bad experience and can leak internals. Violations need graceful, safe handling matched to what tripped — not a 500.
Logging without acting still exposes users to the toxic/leaking/hallucinated content. The point of a guardrail is to stop the unsafe output, then log.
The Violation Handler picks an action by severity — a safe refusal, redaction of the offending part, regeneration with tighter constraints, or escalation — and always logs it. Never a raw error, never the unsafe content.
A Violation Handler takes a severity-matched action: block with a safe refusal, redact the offending span, regenerate under stricter constraints, or escalate — and logs every case. The user gets a safe, graceful response; the unsafe content never ships and the raw error never shows.
Why this piece earns its place
Regenerate is the action that needs a limit and the one that usually ships without one. A model that produced a violating answer will often produce it again under slightly tighter constraints, so the loop runs until something else stops it — spending the most compute and the most latency on exactly the requests that were already going badly. Cap it at one retry and treat the second failure as a block. And regeneration is internal re-validation, not the user’s appeal path: the fresh draft goes back through every check, not only the one that fired, because a second attempt that fixes the toxicity and quietly breaks the schema is an ordinary outcome, not an edge case. Then there is the contract. Callers need a machine-readable reason — a category and an action — not a sentence, because a client renders a redaction differently from a refusal, and because the incident log needs that same field to be countable later. Return prose alone and every consumer ends up string-matching your refusal copy, which can then never be reworded.
- block / refusesafe message
- redact / regeneratesalvage where possible
- always logfor red-teaming
What the new pieces do
- Violation Handlerservice
- Decides what to do on a violation: block with a safe refusal, redact, regenerate with stricter constraints, or escalate — matched to severity, never a raw error.
Step 6 · Layer for speed and cost
Cheap checks first, policy as config
Guardrails mean extra checks on every call — and some (an LLM-based judge for injection or groundedness) are as slow and costly as the model itself. Run them all, always, and you’ve doubled latency and spend. How do you keep guardrails affordable?
How do you keep a stack of guardrails fast and cheap?
Cheap regex/small-model checks clear most traffic instantly; only borderline cases hit expensive LLM-based guardrails. Keep the checks and thresholds in versioned policy config so they’re tunable without a redeploy.
Running all heavy checks always can double latency and cost. Tier them: cheap-and-common first, expensive-and-rare only when needed.
Hard-coded rules can’t be tuned as threats evolve without a redeploy. Policy belongs in versioned config so you can adjust checks and thresholds live.
Tier the checks: fast, cheap classifiers (regex, small models) run first and clear most traffic; only ambiguous cases escalate to slower LLM-based guardrails. Keep which-checks, thresholds, and actions in a versioned Policy Config so you tune safety as threats evolve — no redeploy. Guardrails you can afford to run on every call.
Why this piece earns its place
Tiering only pays while the escalation rate stays low, and that rate belongs to your traffic, not your config. A new surface, a marketing push, or one attacker probing in a loop can multiply it overnight with nothing deployed, and latency and spend follow it up. So escalation rate is a headline metric with its own alarm, and the expensive tier needs a ceiling: past some share of traffic, fall back to the conservative action rather than buying judgement for everyone. The other half cuts both ways. Policy as config makes the fastest change in the system the one with no code review, reaching every instance in seconds — which is the point, and the hazard. So give it what the deploy path already has and config quietly escaped: a policy version stamped on every decision, so a spike in refusals is traceable to one rule instead of to “something changed today”; a rollout by share of traffic rather than to every instance at once; and a change record — who, when, from which value to which — as readable a year later as a commit history. The rollback itself is the easy part. Knowing which of the day’s edits to roll back is not.
- cheap firstclear most traffic fast
- escalateLLM checks on the fuzzy few
- policy configtunable, versioned
What the new pieces do
- Fast Classifiersservice
- Cheap deterministic/small-model checks that run first in each guardrail, so most traffic is cleared quickly and only ambiguous cases escalate to slower LLM-based checks.
- Policy Configstore
- The tunable rules the pipeline runs: which checks, which thresholds, which actions per category — versioned and updatable without redeploying the app.
Step 7 · The real tension
Latency and the over-blocking trap
Guardrails add latency and, crank them up, they start refusing legitimate requests. Too loose and unsafe content slips; too strict and the product becomes useless and infuriating. How do you set the dial?
How do you balance safety against usability and latency?
Over-strict guardrails refuse legitimate users (high false positives) and add latency, quietly making the product unusable while feeling "safe." Over-blocking is a real failure, not a safe default.
Set each guardrail’s threshold from real precision/recall on your traffic, accept the trade-off deliberately, and move non-blocking checks off the critical path (async/streaming) to hide their latency.
No guardrails means shipping unsafe content and open injection — the opposite failure. The answer is calibration, not removal.
Calibrate: set each guardrail’s threshold from measured false-positive vs false-negative rates on real traffic, choosing the trade-off deliberately per category (strict on severe harm, lenient elsewhere). Run non-critical checks async or over the stream so they don’t add blocking latency. Both failure directions are real — over-blocking quietly kills the product, under-blocking ships harm.
Why this piece earns its place
There is a hole in the measurement, and it is why this dial usually gets set by argument rather than data. False negatives you eventually hear about — harm ships, someone reports it. False positives produce nothing at all: a blocked request has no outcome, no follow-up, no label. The refused user leaves, and leaving looks like every other kind of leaving. So the only honest reading comes from a sampled shadow path — a slice of blocked traffic still evaluated, or reviewed by a person, purely to produce labels the dashboard cannot otherwise have. Without it, tightening a threshold always looks like an improvement, because every number you can see got better. The test for what may move off the critical path is not cost either. It is whether the verdict is still actionable once the response has left: a check whose only response is to refuse has to block, while a check that feeds review, alerting or a later account action can run behind the answer.
- calibratefrom real FP/FN rates
- asynchide non-blocking latency
- both directionsover- and under-blocking fail
Step 8 · Stay ahead of attackers
Logging, red-teaming, adaptation
Attackers invent new jailbreaks constantly, and your guardrails’ blind spots are invisible until someone finds them. A safety layer frozen at launch decays. How do you keep it hardening?
How do guardrails keep up with evolving attacks?
Jailbreaks and injections evolve continuously; a frozen ruleset silently falls behind. Guardrails need ongoing red-teaming and updates, like any adversarial system.
By the time users report a jailbreak, harm has shipped. Proactive logging and red-teaming find blind spots before attackers exploit them at scale.
An incident log captures violations and near-misses for audit and alerting; regular red-teaming probes for blind spots; and both feed updated rules and retrained classifiers — a continuous hardening loop.
Keep an Incident Log of every violation and near-miss, red-team proactively to find blind spots before attackers do, and feed findings back into the policy and classifiers. Like content moderation, guardrails are adversarial — a continuous logging → red-team → update loop is how they harden instead of decaying after launch.
Why this piece earns its place
Two things about this loop only show up once it is running. The incident log is now the highest-value store you own. By construction it holds the jailbreak attempts that got far enough to trip a check, the near-misses that came close, and — because a detection records what it found — the raw PII spans the input guardrail strips everywhere else; it sits behind detection, not in front of it. It needs tighter access control, tighter retention and more alarm than the production path it exists to protect, and it must never become the convenient corpus everyone queries for examples. That pulls two ways at once, so split the record: keep the sanitised payload long, because that is the corpus, and expire the spans it was carrying fast. The second thing is a sampling bias in the feedback. You only log what you caught, so a classifier retrained on the log gets sharper at the attacks you already detect and learns nothing about the class you miss — the metrics improve while coverage stands still. Red-teaming is the only way the un-caught class ever enters the training set, which gives you the honest measure of it: not findings per session, but whether a new attack class shows up in a red-team run before it shows up in your traffic.
- logviolations + near-misses
- red-teamfind blind spots first
- adaptfindings → policy + models
What the new pieces do
- Incident Logstore
- Records every violation and near-miss for auditing, alerting, and red-teaming — the feedback that hardens guardrails against new attacks over time.
The payoff
You built an LLM guardrails system
From "a raw model is an attack surface and a liability" to a safety layer: a guarded API wrapping every call, input guardrails screening untrusted requests, defense-in-depth against unpreventable prompt injection, output guardrails for toxicity/PII/schema/groundedness, graceful violation handling, tiered checks with policy-as-config, calibration against over-blocking, and a red-team hardening loop.
Now trust the input — drop the input guardrails — and watch a prompt injection buried in retrieved text hijack the model: system prompt leaked, an unauthorized tool fired, all as a normal response with no error. That’s why an LLM treats every input as instructions, why untrusted content must be screened and contained, and why guardrails wrap the model instead of asking it nicely.
Everything you assembled, in order
- Guarded API — input guardrails → model → output guardrails → violation handling
- Input guardrails — PII redaction, injection/jailbreak detection, scope — untrusted by default
- Prompt injection — can’t be fully prevented → defense-in-depth + least-privilege tools
- Output guardrails — toxicity, PII leakage, schema, and groundedness vs sources
- Violation handling — block/redact/regenerate/escalate — fail safe, not loud
- Layering — cheap classifiers first, LLM checks on the ambiguous; policy as config
- Calibration — both over- and under-blocking fail — set thresholds on real data
- Adaptation — log + red-team + update — guardrails are adversarial and never done
- The failure — trusting untrusted input lets injection through, silently
