Vibe Engines
YouTube
AI System Design

Design an LLM Guardrails System

Learn AI system design by building an LLM guardrails / safety-filter system step by step.

The numbers to beatPIIdetect + redact inboundinjection / jailbreakdetect attemptsscopeenforce in-bounds

The whole design, in writing

Learn AI system design by building an LLM guardrails / safety-filter system step by step. An interactive guide to the safety layer wrapping every model call — input guardrails (PII, prompt-injection, scope), output guardrails (toxicity, PII leakage, schema, groundedness), violation handling, layered defense-in-depth, and the latency/over-blocking trade-off — plus why trusting untrusted input lets prompt injection through.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

An LLM in production needs a seatbelt

Put a raw LLM in front of users and you’ve opened two holes at once. Inbound: users (and any content you retrieve) can inject instructions, jailbreak, or smuggle in prompts — and the model can’t tell your instructions from theirs, because to it everything is just text. Outbound: the model can emit toxic content, leak PII, hallucinate confident falsehoods, or ignore your format. Neither is optional to handle. How do you make every model call safe, in both directions?

Wrap every call in a guardrails system: input guardrails screen the request, the model generates, output guardrails validate the response, and a violation handler decides what ships. It’s defense-in-depth — layered checks around an untrusted model and untrusted input — tuned to catch harm without refusing everyone.

Step 1 · The skeleton

A safety pipeline around the call

A request comes in and a response must go out — but both need checking. What’s the basic shape that lets you gate a model call on both sides?

UserGuarded LLM APILLM
New in this step: User, Guarded LLM API, LLM. · swipe to pan the diagram

Where do safety checks belong relative to the model?

  1. A system prompt asking the model to be safe is easily overridden by injection and can’t validate its own output. Guardrails must be independent checks around the model, not instructions inside it.

  2. The guarded API screens the request, calls the model, validates the response, and handles violations — an independent safety layer on both sides of an untrusted model.

  3. Skipping input guardrails leaves prompt injection and jailbreaks unhandled — the exact failure this page demonstrates. Both directions need gates.

The Guarded LLM API runs a pipeline: input guardrails → model → output guardrails → violation handling, before any response reaches the user. Safety is an independent layer around the model, not a plea inside its prompt — because the model can be manipulated and can’t reliably police itself.

Why this piece earns its place

The guarded API is a choke point, which raises a question the diagram hides: is it a library every service imports, or a service every service calls? A library keeps the call local and adds no network hop, but every team pins its own version, so a policy change becomes N deploys and nobody can say which rules are actually running. A service centralises that and buys you a new dependency on the request path. Which forces the decision nobody makes until the first outage: when the guardrail layer itself times out, do you fail closed or fail open? Fail closed and a guardrail incident is a full product outage. Fail open and it is not an incident at all by any definition your on-call rota recognises, so that branch needs a counter and a page of its own. And it is not one decision either — a schema validator and an injection detector deserve opposite answers, so the timeout behaviour belongs in the same configuration as the thresholds rather than buried in the gateway’s catch block. Whichever way each one goes, those branches are code that only runs when something is already broken, which means nobody has ever run them: not the refusal copy, not the counter, not the alarm. Exercise it deliberately, on a schedule, or meet it for the first time on the day it matters.

What the new pieces do

Userclient
The caller — and, from the system’s view, untrusted. Their input (and any content it drags in) may carry prompt injection, jailbreaks, PII, or out-of-scope asks.
Guarded LLM APIbackend
Wraps every model call in a safety pipeline: run input guardrails, call the model, run output guardrails, and handle any violation — before a response ever reaches the user.
LLMmodel
The model that produces the response. Powerful and non-deterministic — it can emit toxic, off-policy, hallucinated, or PII-leaking output, which is why its input and output are both gated.

Step 2 · Screen the request

Input guardrails: PII, injection, scope

Before the model sees anything, the request needs vetting. What can go wrong on the way in, and what do you check for?

Guarded LLM APIsafety pipelineInput Guardrailsscreen request
New in this step: Input Guardrails.

What should input guardrails check?

  1. Language is trivial next to the real inbound risks: leaking PII to the model, prompt injection/jailbreaks, and out-of-scope abuse. Input screening is about safety, not locale.

  2. Output-only checks can’t stop injection from hijacking the model or PII from being sent to a third-party provider in the first place. Input needs its own gate.

  3. Input guardrails redact PII before it hits the provider, detect injection/jailbreak attempts, and enforce that the request is in-scope — screening untrusted input before the model can act on it.

Input Guardrails screen the request: detect and redact PII before it leaves for the provider, detect prompt-injection and jailbreaks, and enforce scope (reject or safe-refuse out-of-bounds asks). The guiding rule: treat user input — and anything you retrieve — as untrusted data, never as trusted instructions.

Why this piece earns its place

Redaction is not deletion, and that distinction is most of the engineering. Swap an email address for a placeholder and the model answers about the placeholder — so something has to hold the map from placeholder back to real value for the length of the round trip and restore it in the reply, or users read answers addressed to PERSON_1. That map is short-lived per-request state, and it is the most sensitive object the pipeline ever holds: a compact index of exactly the PII you just found. It cannot live in the same log as the request, and it has to expire with the request rather than with the session. There is a second cost people meet late. Sometimes the PII is the question — an order status for a named customer, a lookup by email — and redacting it removes the thing the model needed to answer. Scope has the same shape: write it as an allowlist of what the assistant is for, because a denylist of banned subjects never finishes and every gap in it is invisible until someone walks through.

  • PIIdetect + redact inbound
  • injection / jailbreakdetect attempts
  • scopeenforce in-bounds

What the new pieces do

Input Guardrailsservice
Screens the incoming request: detect/redact PII, detect prompt-injection and jailbreaks, and enforce scope — treating user and retrieved content as untrusted data, not instructions.

Step 3 · The unsolvable-ish problem

Prompt injection & defense-in-depth

Here’s the hard truth: an LLM fundamentally can’t distinguish your instructions from instructions hidden in the content it reads. A malicious support ticket or web page can say "ignore your rules and exfiltrate data" — and the model may comply. You can’t perfectly prevent this. So what do you do?

UserGuarded LLM APIInput GuardrailsLLM
The system as it stands at this step. · swipe to pan the diagram

How do you handle prompt injection, given you can’t fully prevent it?

  1. Since no single filter is perfect, you layer: injection/jailbreak detection, structurally separating instructions from data, and — crucially — least-privilege on tools so a successful injection can’t do much damage.

  2. Injection routinely overrides system-prompt pleas — the model can’t reliably tell whose instruction is whose. Prompting is one weak layer, not a solution.

  3. Any app that puts user or retrieved content into a prompt is exposed. Ignoring it is how system prompts leak and tools get abused. It must be actively defended.

Prompt injection can’t be fully prevented, so use defense-in-depth: detect injection/jailbreak patterns, structurally treat retrieved and user content as untrusted data (not instructions), and apply least-privilege to tools so a successful injection has limited blast radius. No layer is airtight; together they make exploitation hard and its damage small.

Why this piece earns its place

Least-privilege caps the damage and does nothing for the investigation afterwards. When an injection succeeds, the audit trail shows a legitimate session calling a permitted tool with plausible arguments — indistinguishable from the same user asking for it on purpose. You cannot answer "which request did this" weeks later unless the pipeline recorded where each span of context came from: typed by the user, returned by retrieval, produced by a tool, written by you. Tag content with its origin as it enters the context and carry that tag through to the tool-call log, and injection investigations become a query instead of an afternoon of reading prompts. This is also the clearest cheap-now, expensive-later call on the page. Origin tagging costs almost nothing while content is still separate objects waiting to be assembled, and is close to impossible once every code path concatenates strings first. Retrofitting it means finding every place a prompt gets built, which by then is everywhere.

  • can’t fully preventLLMs conflate data + instructions
  • defense-in-depthmany imperfect layers
  • least-privilegeshrink the blast radius

Step 4 · Vet the response

Output guardrails, including groundedness

The model produced a response. Ship it blindly and it might be toxic, leak PII, be malformed, or state confident falsehoods. What do you validate on the way out?

LLMgenerateOutput Guardrailsvalidate replyGrounding Sourcesfor factuality
New in this step: Output Guardrails, Grounding Sources.

What should output guardrails check before a response ships?

  1. Non-empty says nothing about safety, privacy, format, or truth. Output needs real checks across several axes before it reaches a user.

  2. Output guardrails scan for toxic content, PII the model may have leaked, structural validity (valid JSON/required fields), and whether claims are supported by the grounding sources rather than hallucinated.

  3. A model self-certifying is weak and gameable, and doesn’t catch format or groundedness issues. Use independent checks (classifiers, validators, source comparison), not self-approval.

Output Guardrails validate the response: toxicity, PII leakage, schema/format conformance (valid JSON, required fields), and groundedness — comparing claims to the Grounding Sources to catch hallucinations. Independent checks, not the model’s own say-so. Only a response that clears all of them is allowed to ship.

Why this piece earns its place

Notice the cost asymmetry between the two gates. An input block costs at most what the input checks cost, and an escalated one is roughly a model call. An output block lands after you have already paid for the whole generation, and if the action is regeneration you pay for it again. So the asymmetry is not cheap against expensive; it is one check you might buy against one answer you have definitely bought. That is still an argument for moving every check that can run on the request to the input side, and for reading a rising output-violation rate as a spend problem as much as a safety one. Two mechanical details decide whether this stage works at all. Run schema validation first: the other checks read fields out of the response, and pointing them at malformed output means scanning JSON syntax instead of the text a user would see. And groundedness needs the retrieved sources still in hand at validation time, which means carrying them through the call rather than dropping them once the prompt is assembled — easy to design out by accident, and one reason groundedness quietly becomes the check a team cannot run for reasons that have nothing to do with latency.

  • toxicity + PIIharm and privacy
  • schemavalid structure
  • groundednesssupported, not hallucinated

What the new pieces do

Output Guardrailsservice
Validates the model’s response before it ships: toxicity, PII leakage, format/schema conformance, and groundedness against sources (is it supported, or hallucinated?).
Grounding Sourcesstore
The retrieved context/knowledge the answer is supposed to be based on. The groundedness check compares the response to these to catch unsupported claims.

Step 5 · When something trips

Handling a violation gracefully

A guardrail fires — injected input, toxic output, a hallucination. You can’t just crash or return the unsafe content. What happens next?

LLMgenerateOutput Guardrailsvalidate replyViolation Handlerblock · redact
New in this step: Violation Handler.

What should the system do when a guardrail is violated?

  1. A raw error is a bad experience and can leak internals. Violations need graceful, safe handling matched to what tripped — not a 500.

  2. Logging without acting still exposes users to the toxic/leaking/hallucinated content. The point of a guardrail is to stop the unsafe output, then log.

  3. The Violation Handler picks an action by severity — a safe refusal, redaction of the offending part, regeneration with tighter constraints, or escalation — and always logs it. Never a raw error, never the unsafe content.

A Violation Handler takes a severity-matched action: block with a safe refusal, redact the offending span, regenerate under stricter constraints, or escalate — and logs every case. The user gets a safe, graceful response; the unsafe content never ships and the raw error never shows.

Why this piece earns its place

Regenerate is the action that needs a limit and the one that usually ships without one. A model that produced a violating answer will often produce it again under slightly tighter constraints, so the loop runs until something else stops it — spending the most compute and the most latency on exactly the requests that were already going badly. Cap it at one retry and treat the second failure as a block. And regeneration is internal re-validation, not the user’s appeal path: the fresh draft goes back through every check, not only the one that fired, because a second attempt that fixes the toxicity and quietly breaks the schema is an ordinary outcome, not an edge case. Then there is the contract. Callers need a machine-readable reason — a category and an action — not a sentence, because a client renders a redaction differently from a refusal, and because the incident log needs that same field to be countable later. Return prose alone and every consumer ends up string-matching your refusal copy, which can then never be reworded.

  • block / refusesafe message
  • redact / regeneratesalvage where possible
  • always logfor red-teaming

What the new pieces do

Violation Handlerservice
Decides what to do on a violation: block with a safe refusal, redact, regenerate with stricter constraints, or escalate — matched to severity, never a raw error.

Step 6 · Layer for speed and cost

Cheap checks first, policy as config

Guardrails mean extra checks on every call — and some (an LLM-based judge for injection or groundedness) are as slow and costly as the model itself. Run them all, always, and you’ve doubled latency and spend. How do you keep guardrails affordable?

Guarded LLM APIInput GuardrailsLLMOutput GuardrailsGrounding SourcesFast ClassifiersViolation HandlerPolicy Config
New in this step: Fast Classifiers, Policy Config. · swipe to pan the diagram

How do you keep a stack of guardrails fast and cheap?

  1. Cheap regex/small-model checks clear most traffic instantly; only borderline cases hit expensive LLM-based guardrails. Keep the checks and thresholds in versioned policy config so they’re tunable without a redeploy.

  2. Running all heavy checks always can double latency and cost. Tier them: cheap-and-common first, expensive-and-rare only when needed.

  3. Hard-coded rules can’t be tuned as threats evolve without a redeploy. Policy belongs in versioned config so you can adjust checks and thresholds live.

Tier the checks: fast, cheap classifiers (regex, small models) run first and clear most traffic; only ambiguous cases escalate to slower LLM-based guardrails. Keep which-checks, thresholds, and actions in a versioned Policy Config so you tune safety as threats evolve — no redeploy. Guardrails you can afford to run on every call.

Why this piece earns its place

Tiering only pays while the escalation rate stays low, and that rate belongs to your traffic, not your config. A new surface, a marketing push, or one attacker probing in a loop can multiply it overnight with nothing deployed, and latency and spend follow it up. So escalation rate is a headline metric with its own alarm, and the expensive tier needs a ceiling: past some share of traffic, fall back to the conservative action rather than buying judgement for everyone. The other half cuts both ways. Policy as config makes the fastest change in the system the one with no code review, reaching every instance in seconds — which is the point, and the hazard. So give it what the deploy path already has and config quietly escaped: a policy version stamped on every decision, so a spike in refusals is traceable to one rule instead of to “something changed today”; a rollout by share of traffic rather than to every instance at once; and a change record — who, when, from which value to which — as readable a year later as a commit history. The rollback itself is the easy part. Knowing which of the day’s edits to roll back is not.

  • cheap firstclear most traffic fast
  • escalateLLM checks on the fuzzy few
  • policy configtunable, versioned

What the new pieces do

Fast Classifiersservice
Cheap deterministic/small-model checks that run first in each guardrail, so most traffic is cleared quickly and only ambiguous cases escalate to slower LLM-based checks.
Policy Configstore
The tunable rules the pipeline runs: which checks, which thresholds, which actions per category — versioned and updatable without redeploying the app.

Step 7 · The real tension

Latency and the over-blocking trap

Guardrails add latency and, crank them up, they start refusing legitimate requests. Too loose and unsafe content slips; too strict and the product becomes useless and infuriating. How do you set the dial?

UserGuarded LLM APIInput GuardrailsLLMOutput GuardrailsGrounding SourcesFast ClassifiersViolation HandlerPolicy Config
The system as it stands at this step. · swipe to pan the diagram

How do you balance safety against usability and latency?

  1. Over-strict guardrails refuse legitimate users (high false positives) and add latency, quietly making the product unusable while feeling "safe." Over-blocking is a real failure, not a safe default.

  2. Set each guardrail’s threshold from real precision/recall on your traffic, accept the trade-off deliberately, and move non-blocking checks off the critical path (async/streaming) to hide their latency.

  3. No guardrails means shipping unsafe content and open injection — the opposite failure. The answer is calibration, not removal.

Calibrate: set each guardrail’s threshold from measured false-positive vs false-negative rates on real traffic, choosing the trade-off deliberately per category (strict on severe harm, lenient elsewhere). Run non-critical checks async or over the stream so they don’t add blocking latency. Both failure directions are real — over-blocking quietly kills the product, under-blocking ships harm.

Why this piece earns its place

There is a hole in the measurement, and it is why this dial usually gets set by argument rather than data. False negatives you eventually hear about — harm ships, someone reports it. False positives produce nothing at all: a blocked request has no outcome, no follow-up, no label. The refused user leaves, and leaving looks like every other kind of leaving. So the only honest reading comes from a sampled shadow path — a slice of blocked traffic still evaluated, or reviewed by a person, purely to produce labels the dashboard cannot otherwise have. Without it, tightening a threshold always looks like an improvement, because every number you can see got better. The test for what may move off the critical path is not cost either. It is whether the verdict is still actionable once the response has left: a check whose only response is to refuse has to block, while a check that feeds review, alerting or a later account action can run behind the answer.

  • calibratefrom real FP/FN rates
  • asynchide non-blocking latency
  • both directionsover- and under-blocking fail

Step 8 · Stay ahead of attackers

Logging, red-teaming, adaptation

Attackers invent new jailbreaks constantly, and your guardrails’ blind spots are invisible until someone finds them. A safety layer frozen at launch decays. How do you keep it hardening?

Violation Handlerblock · redactIncident Logred-team
New in this step: Incident Log.

How do guardrails keep up with evolving attacks?

  1. Jailbreaks and injections evolve continuously; a frozen ruleset silently falls behind. Guardrails need ongoing red-teaming and updates, like any adversarial system.

  2. By the time users report a jailbreak, harm has shipped. Proactive logging and red-teaming find blind spots before attackers exploit them at scale.

  3. An incident log captures violations and near-misses for audit and alerting; regular red-teaming probes for blind spots; and both feed updated rules and retrained classifiers — a continuous hardening loop.

Keep an Incident Log of every violation and near-miss, red-team proactively to find blind spots before attackers do, and feed findings back into the policy and classifiers. Like content moderation, guardrails are adversarial — a continuous logging → red-team → update loop is how they harden instead of decaying after launch.

Why this piece earns its place

Two things about this loop only show up once it is running. The incident log is now the highest-value store you own. By construction it holds the jailbreak attempts that got far enough to trip a check, the near-misses that came close, and — because a detection records what it found — the raw PII spans the input guardrail strips everywhere else; it sits behind detection, not in front of it. It needs tighter access control, tighter retention and more alarm than the production path it exists to protect, and it must never become the convenient corpus everyone queries for examples. That pulls two ways at once, so split the record: keep the sanitised payload long, because that is the corpus, and expire the spans it was carrying fast. The second thing is a sampling bias in the feedback. You only log what you caught, so a classifier retrained on the log gets sharper at the attacks you already detect and learns nothing about the class you miss — the metrics improve while coverage stands still. Red-teaming is the only way the un-caught class ever enters the training set, which gives you the honest measure of it: not findings per session, but whether a new attack class shows up in a red-team run before it shows up in your traffic.

  • logviolations + near-misses
  • red-teamfind blind spots first
  • adaptfindings → policy + models

What the new pieces do

Incident Logstore
Records every violation and near-miss for auditing, alerting, and red-teaming — the feedback that hardens guardrails against new attacks over time.

The payoff

You built an LLM guardrails system

From "a raw model is an attack surface and a liability" to a safety layer: a guarded API wrapping every call, input guardrails screening untrusted requests, defense-in-depth against unpreventable prompt injection, output guardrails for toxicity/PII/schema/groundedness, graceful violation handling, tiered checks with policy-as-config, calibration against over-blocking, and a red-team hardening loop.

UserGuarded LLM APIInput GuardrailsLLMOutput GuardrailsGrounding SourcesFast ClassifiersViolation HandlerPolicy ConfigIncident Log
The finished design, end to end. · swipe to pan the diagram

Now trust the input — drop the input guardrails — and watch a prompt injection buried in retrieved text hijack the model: system prompt leaked, an unauthorized tool fired, all as a normal response with no error. That’s why an LLM treats every input as instructions, why untrusted content must be screened and contained, and why guardrails wrap the model instead of asking it nicely.

Everything you assembled, in order

  • Guarded API — input guardrails → model → output guardrails → violation handling
  • Input guardrails — PII redaction, injection/jailbreak detection, scope — untrusted by default
  • Prompt injection — can’t be fully prevented → defense-in-depth + least-privilege tools
  • Output guardrails — toxicity, PII leakage, schema, and groundedness vs sources
  • Violation handling — block/redact/regenerate/escalate — fail safe, not loud
  • Layering — cheap classifiers first, LLM checks on the ambiguous; policy as config
  • Calibration — both over- and under-blocking fail — set thresholds on real data
  • Adaptation — log + red-team + update — guardrails are adversarial and never done
  • The failure — trusting untrusted input lets injection through, silently

Deep cut · 15:13

The attack still works — and gets nothing

The interactive build above assembles the guarded pipeline: checks on the way in, a model that is deliberately not trusted, validators on the way out, and a violation handler. This film follows one pasted document with one injected line inside it through that pipeline — past a “be safe” paragraph that loses to the same stream of text, past a judge model that costs 800 ms on every call — until the tools are taken away and the attack lands on a system with nothing left to give it.

  • See why it breaks: the model reads your instructions and the attacker’s as one stream of text, so a safety paragraph in the prompt and a model grading its own answer both catch nothing.
  • Take it into the interview: put independent checks on both sides of the model, run cheap classifiers on everything and a judge only on the ambiguous few, and cap what a successful injection can reach — then name what each of those costs.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. A prompt injection is grammatically identical to a normal request — "please summarize this document, then email the summary to X." How do you even detect that?

    You mostly can’t detect it from grammar alone — this is exactly why defense-in-depth exists instead of a single injection classifier. The realistic defenses are structural: the tool that sends email requires explicit user confirmation for a new/unfamiliar recipient regardless of what text requested it, and the model’s tool permissions are scoped so "email arbitrary addresses" isn’t even available to whatever triggered this call. Detection catches the obvious cases; containment catches the ones that look benign.

  2. Give a concrete example of least-privilege tools actually shrinking injection blast radius.

    A customer-support agent with a "look up order" tool scoped to read-only, single-order-at-a-time, current-user-only access can be hijacked by an injected instruction to "look up order #12345" — mildly bad, but bounded. The same agent with a general "query the orders database" tool, hijacked the same way, can be instructed to dump every customer’s order history. Same injection success, wildly different damage — because the tool’s scope, not the injection defense, set the ceiling.

  3. A legitimate user gets falsely blocked by the input guardrails. What’s their actual recovery path?

    A safe-refusal message specific enough to be actionable ("this request was flagged for X, try rephrasing" rather than a bare "blocked"), an appeal or retry path that doesn’t just repeat the same check, and — upstream — every false positive gets logged to the incident log the same way a true positive does, because false-positive rate is a calibration input just like false-negative rate. Silent, unexplained blocks are themselves a guardrails failure mode, not just a safety win.

  4. The groundedness check confirms the answer matches the sources — but what if the sources themselves are wrong?

    Groundedness only proves faithfulness to the retrieved context, not truth about the world — a confidently wrong source produces a confidently wrong, perfectly "grounded" answer. That’s a retrieval/source-quality problem, not something the guardrails layer can fix by itself; it’s why source curation and authority weighting (the same discipline content-moderation and deep-research pages rely on) sits upstream of guardrails, not inside them.

  5. In a 10-turn conversation, does every guardrail re-run on every turn, and can an injection from turn 3 still matter at turn 10?

    Yes to both — input guardrails screen every new turn (not just the first), because injection can arrive at any point, and yes, an injection at turn 3 can still be "live" at turn 10 if the conversation history (including the injected text) is still in context, which is exactly why guardrails need to treat conversation history as untrusted data too, not just the newest message.

Check yourself — the answers, and why

Nine steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. A guardrails system is fundamentally…

    The model can be manipulated and can’t police itself; guardrails are independent layers around it, not instructions inside its prompt.

  2. Input guardrails treat user and retrieved content as…

    Anything entering the prompt can be adversarial; input guardrails redact PII, detect injection/jailbreaks, and enforce scope before the model acts.

  3. The right stance on prompt injection is…

    LLMs can’t reliably separate instructions from injected content, so you layer imperfect defenses and shrink the blast radius rather than seeking a silver bullet.

  4. The groundedness output guardrail catches…

    It compares the response to the retrieved sources so unsupported (hallucinated) claims are blocked before a user believes them.

  5. To keep guardrails affordable you…

    Tiering high-precision cheap filters ahead of expensive LLM judges (with policy in config) keeps latency and cost manageable on every call.

  6. Uniquely, guardrails can fail by…

    Over-strict guardrails quietly refuse real users while feeling safe; calibrate thresholds per category on real false-positive/negative data.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Screen input: vet every request for PII, injection, jailbreaks, and scope before the model sees it.
  • Validate output: check the reply for toxicity, PII leakage, schema, and groundedness before it ships.
  • Handle violations: block, redact, regenerate, or escalate — never a raw error or unsafe content.
  • Stay affordable: run cheap checks first, escalate only the ambiguous to LLM judges.
  • Adapt: log every violation and near-miss, and feed red-team findings back into policy.

The qualities that shape everything

Each one names the mechanism that buys it.

Safety the model can’t override
A guarded API runs input guardrails → model → output guardrails → violation handling as an independent layer around the model, not a plea inside its prompt.
Untrusted input never reaches the model unchecked
Input guardrails redact PII before it leaves for the provider, detect injection/jailbreaks, and enforce scope — treating user and retrieved content as untrusted data.
Limit the damage of an unpreventable injection
Defense-in-depth: detect injection patterns, structurally treat content as data, and least-privilege the tools so a successful injection has a small blast radius.
No unsafe or fabricated response ships
Output guardrails validate toxicity, PII leakage, schema, and groundedness against the sources — independent checks, only a clean response is allowed out.
Guardrails you can afford on every call
Tier the checks — cheap classifiers first clear most traffic, only ambiguous cases escalate to slow LLM judges — with thresholds and actions in versioned policy config.
Safe without refusing real users
Calibrate each threshold from measured false-positive/negative rates per category and run non-critical checks async, since both over- and under-blocking are real failures.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

Independent checks around the model over a "be safe" system prompt

A prompt asking the model to behave is easily overridden by injection and can’t validate its own output. Guardrails must be independent checks on both sides of an untrusted model, not instructions inside it.

Defense-in-depth + least-privilege over a "never follow injected instructions" prompt

Injection routinely overrides system-prompt pleas — the model can’t tell whose instruction is whose. Layer detection, treat content as data, and scope tools so a successful injection simply can’t do anything catastrophic.

Independent output validators over the model self-certifying "is this safe?"

A model approving its own output is weak and gameable, and catches nothing about format or groundedness. Classifiers, schema validators and source comparison judge the response independently.

Cheap classifiers first, escalate over every LLM judge on every request

Running heavy LLM-based checks always can double latency and spend. High-precision cheap filters clear most traffic instantly and only borderline cases hit the expensive judges — with rules in versioned config, tunable without a redeploy.

Calibrate thresholds per category over maxing out every guardrail

Over-strict guardrails refuse legitimate users and add latency, quietly killing the product while feeling "safe." Set each threshold from real precision/recall and run non-critical checks async — both directions of failure are real.

The answer, out loud

What a strong answer to “Design an LLM Guardrails System” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Pin down what we’re protecting

    Before I draw anything, I want to agree on what we’re protecting. An LLM in front of users opens two holes. Inbound, users — and anything we retrieve — can smuggle in instructions, and the model can’t tell ours from theirs. Outbound, it can emit toxic content, leak personal data, break our format, or state confident falsehoods. And my three numbers pull against each other: harm that slips through, real users wrongly refused, and latency. So: safe both ways, without refusing the people we built it for.

  2. 3–7 min

    Put the safety outside the model

    The shape I’d draw is a guarded API around every call — input guardrails, the model, output guardrails, violation handling, and only then does anything reach the user. I want to be explicit about why safety sits outside the model instead of inside its prompt. A system prompt has no standing the attacker’s text doesn’t also have — same window, same priority. And a paragraph of good intentions isn’t something I can version, measure or switch off. A check that runs outside the model is all three.

    Built in step 1: A safety pipeline around the call
  3. 7–12 min

    Screen the request before anyone else sees it

    On the way in, three things. Personal data — PII — detected and redacted, and I’d stress that this gate is the last moment that’s still our decision; once the request is at the provider, it isn’t. Second, prompt injection and jailbreaks: text trying to talk the model out of its rules. Third, scope — is this even a question we’re here to answer? None of those detectors is reliable alone, so I’m not selling input screening as the defense. It’s what makes the next layer’s job small.

    Built in step 2: Input guardrails: PII, injection, scope
  4. 12–18 min

    Contain what I can’t prevent

    Then the hard part. Prompt injection is instructions hidden in content the model reads — a support ticket, a fetched page — and it can’t separate those from mine. I can’t prevent that. What I can do is make succeeding worthless. Keep user and retrieved text marked as content rather than folded in with my instructions. Then list everything the model can reach — tools, data, actions — and cut it until the worst case of a fully successful injection is something I’d be willing to read in a postmortem. That question has a definite answer. Detection never does.

    Built in step 3: Prompt injection & defense-in-depth
  5. 18–23 min

    The model’s output is untrusted too

    Coming back out, four checks: toxicity, PII the model leaked on its own, schema — valid JSON with the fields the caller expects — and groundedness: does each claim hold up against the sources the answer was meant to rest on. I wouldn’t hand the response back to the same model and ask if it’s safe; if it was willing to write the sentence, it’ll defend it. And the four aren’t equally solid. Schema is deterministic; toxicity and PII I’d trust to act alone. Groundedness is a judgement call — the one I’d tune hardest and trust least.

    Built in step 4: Output guardrails, including groundedness
  6. 23–28 min

    When a check trips

    When something fires, the two things I won’t do are return a raw error and let the content through with a log line. A violation handler picks by severity: block with a safe refusal, redact the offending span, regenerate under tighter constraints, or escalate — and records every case. The one I’d argue about is escalate. The others are automatic; escalate means a person, and a person has a queue and a working day. I’d only route a category there if I can name who is on the other end — otherwise it’s a slower block.

    Built in step 5: Handling a violation gracefully
  7. 28–33 min

    Cheap checks first, policy as config

    Now cost, because the strongest checks here are themselves model calls. So I’d tier them: cheap classifiers — pattern rules and small models — clear most traffic, and only ambiguous cases reach an LLM judge. Which checks run, at what threshold, with what action, lives in versioned policy config, not app code, so retuning as attacks change isn’t a deploy. What I’d flag is that tiering reshapes the latency curve rather than lowering it — most requests get faster, the ambiguous ones get a lot slower. I’d quote a p99 here, never an average.

    Built in step 6: Cheap checks first, policy as config
  8. 33–38 min

    Set the dial per category

    I’d refuse to set one number here. Severe harm and mild profanity are different decisions with different acceptable error rates; collapse them into one safety level and you are too strict for one and too loose for the other. So it’s per category, from measured false-positive and false-negative rates on our own traffic, and checks that needn’t block move off the critical path. I’d record the category on every decision — “the guardrails are too strict” is an argument nobody wins; “PII screening refuses tickets that name a customer” is something you can go and fix.

    Built in step 7: Latency and the over-blocking trap
  9. 38–43 min

    What I’d watch, and how it fails

    On a dashboard: violations and near-misses per category, false-positive rate beside false-negative rate, and latency per tier. Every case lands in an incident log, and we red-team on purpose and feed findings back into policy and classifiers. The failure I’d plan for isn’t technical: the guardrails keep working and the people stop. The log fills, nobody reads it, red-teaming slips a quarter, and the layer still passes the tests it passed at launch. The other is a ratchet — legal wants a phrase blocked, support wants a topic refused, and the refusal rate climbs one uncontroversial rule at a time, none big enough to argue about.

    Built in step 8: Logging, red-teaming, adaptation
  10. 43–45 min

    Close on the trade-off

    To close: a guarded API, input screening that trusts nothing, containment not prevention, independent output checks including groundedness, severity-matched handling, cheap checks first with policy in config, thresholds per category, and a red-team loop. Every dial in there is the same trade — catch more harm, refuse more real people, add more latency. With more time I’d open up streaming: everything I’ve described assumes I hold the whole response and then judge it, and the moment we stream tokens to make it feel fast, a block becomes a retraction.

What this teaches

Learn AI system design by building an LLM guardrails / safety-filter system step by step. An interactive guide to the safety layer wrapping every model call — input guardrails (PII, prompt-injection, scope), output guardrails (toxicity, PII leakage, schema, groundedness), violation handling, layered defense-in-depth, and the latency/over-blocking trade-off — plus why trusting untrusted input lets prompt injection through.

Key takeaways

  • Guarded API — input guardrails → model → output guardrails → violation handling
  • Input guardrails — PII redaction, injection/jailbreak detection, scope — untrusted by default
  • Prompt injection — can’t be fully prevented → defense-in-depth + least-privilege tools
  • Output guardrails — toxicity, PII leakage, schema, and groundedness vs sources
  • Violation handling — block/redact/regenerate/escalate — fail safe, not loud
  • Layering — cheap classifiers first, LLM checks on the ambiguous; policy as config
  • Calibration — both over- and under-blocking fail — set thresholds on real data
  • Adaptation — log + red-team + update — guardrails are adversarial and never done
  • The failure — trusting untrusted input lets injection through, silently

Concepts covered

  • An LLM in production needs a seatbelt
  • A safety pipeline around the call
  • Input guardrails: PII, injection, scope
  • Prompt injection & defense-in-depth
  • Output guardrails, including groundedness
  • Handling a violation gracefully
  • Cheap checks first, policy as config
  • Latency and the over-blocking trap
  • Logging, red-teaming, adaptation
RUN IT YOURSELF

The harmful request that scores exactly what a legitimate one scores

A guardrail scores text against a rule set and blocks above a threshold — and that one number is the entire product decision. This file sweeps every threshold over a small labelled set and prints both columns at each: legitimate users refused, harmful replies shipped. Change what a shipped harm costs in pick(ROWS, 1, 10), or add your own request to EXAMPLES, hit Run, and watch the dial move.

HOW TO READ THE CODE — 4 IDEAS
  1. Each row is one threshold. Read both columns — a false positive refuses a real user, a false negative ships harm (step 7).
  2. Fewest mistakes lands at 0.593 with perfect precision, and still ships 3 of 7 harmful examples. Price harm at 10x and it moves to 0.393: zero misses, 3 real users refused.
  3. Neither is a floor you can tune away: 3 harmful requests score exactly what a legitimate one scores, so no threshold anywhere splits them. That is arithmetic, not a weak classifier.
  4. So band() keeps the cheap rules on the ends and escalates the overlap to a paid judge (step 6) — 0 mistakes on the 8 requests it decides, and a 43% escalate rate as the bill.
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, trust the input, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs