Vibe Engines
YouTube
AI System Design

Design Content Moderation

Learn AI system design by building a content-moderation pipeline (text + image) step by step.

The numbers to beathash matchknown-bad, fastexact + perceptualcatch altered re-uploadsfunnelcheap+certain → costly+fuzzy

The whole design, in writing

Learn AI system design by building a content-moderation pipeline (text + image) step by step. An interactive guide to the staged funnel — hash-matching known-bad content, text and multimodal classifiers, per-category thresholds, a human-in-the-loop review queue, a policy/enforcement engine, and an appeals + retraining loop against adversarial evasion — plus why fully trusting the classifier fails.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

Moderating content at platform scale

A platform accepts millions of posts, comments, and images a day. Some tiny fraction is harmful — abuse, hate, violence, spam, CSAM. You can’t have humans read all of it (impossible at scale), and you can’t fully trust a model (it’s wrong in both directions: censoring the innocent and passing the harmful). Users actively try to evade you. How do you catch harm at scale without either drowning in review or silently over-censoring?

Build a staged moderation funnel: cheap high-precision filters first, ML classifiers next, and humans on the uncertain middle. A policy engine turns scores into actions; an appeals + retraining loop keeps up with adversaries. The core discipline: automate the confident extremes, send the ambiguous to people.

Step 1 · The skeleton

Content in, action out

A piece of content arrives and needs a decision — allow, limit, block, or escalate — often before it’s shown. What sits between the upload and the verdict?

Content Inpost · uploadModeration APIorchestrator
New in this step: Content In, Moderation API.

What’s the minimal shape of a moderation request?

  1. Purely reactive moderation leaves harmful content visible until someone reports it — unacceptable for severe categories. Serious platforms screen proactively, before or at publish.

  2. An orchestrator takes the content, runs the funnel (fast filters → classifiers → policy), records the decision for audit/appeal, and returns allow/limit/block/escalate.

  3. One model’s binary verdict has no notion of category, severity, confidence, or appeal — and it’s wrong in both directions. Moderation needs a staged funnel and a policy layer, not a coin flip.

A Moderation API orchestrates the decision: run content through the funnel, record a case (scores, action, reviewer) for audit and appeals, and return an action. Screening is largely proactive — decided at or before publish for anything severe — with latency budgets tighter for pre-publish checks.

Why this piece earns its place

Recording the case is the easy half of the paper trail. The half that matters six weeks later is which versions decided it: the checkpoint that produced the scores, and the policy revision that turned them into an action. Without both, an overturned appeal is unreadable — you cannot tell whether the model was wrong, the threshold has since moved, or the rule itself changed, and those three have different owners and different repairs. Store them beside the scores; they cost bytes. The other field to make first-class is the surface. One API serves posts, comments, private messages, profile names and uploads, and everything behind it wants to key off which one this is: the thresholds that apply, the severity a category carries there, and which actions even exist — you cannot age-gate a direct message. Let it arrive as an optional header nobody sets and you also lose the ability to answer the question that always comes, which is where we are wrong most. And treat the case write as part of the decision rather than as telemetry: make it fire-and-forget and you will eventually enforce an action that has no record, which is an action nobody can appeal, audit or learn from.

What the new pieces do

Content Inclient
User-generated content entering the platform — a comment, a post, an image. The moderation pipeline decides whether it can be shown, and how fast that decision must be.
Moderation APIbackend
Receives content, runs it through the staged funnel (hash filter → classifiers → policy), records a case, and returns an action. Balances latency (proactive, pre-publish) against thoroughness.

Step 2 · Cheapest check first

The funnel: hash-match known-bad

Running heavy ML on every item is expensive, and a lot of harmful content is re-uploaded, previously-seen material. Why pay a classifier to rediscover what you already know is banned?

Moderation APICase Store
New in this step: Case Store. · swipe to pan the diagram

What should run before the ML classifiers?

  1. That wastes compute re-judging known-bad content and misses the highest-precision signal you have: exact matches to previously-actioned material.

  2. Humans can’t pre-screen platform-scale volume; they’re the scarce resource you reserve for genuinely ambiguous cases, not a first-pass filter.

  3. Hash-matching against a database of known violating material (exact and perceptual/PhotoDNA-style) is fast and near-perfect precision — it catches re-uploads instantly before any model runs.

Front the funnel with a Hash Filter: match content against hashes of known-violating material — exact hashes for identical files and perceptual hashes for slightly-altered images. It’s cheap, near-perfect precision, and catches re-uploads (including legally-mandated categories like CSAM) before spending a cent on ML. The funnel’s principle: cheap and certain first, expensive and fuzzy later.

Why this piece earns its place

A hash hit is enforcement without a score, so every entry needs provenance: where it came from — a shared industry list, or one of your own takedowns — and when it was written. Matching a shared list means enforcing another platform's judgment at your precision; that is the right call for the legally mandated categories, and it is still someone else's call, so when a block is disputed the entry's record is the only thing that can explain it. Two consequences people miss. Adding a hash is not retroactive: it stops the next upload and does nothing about the copies already published, so a new entry needs a sweep over stored media if you want those gone. And perceptual matching has a distance threshold, which means it has false positives — an unrelated image can land inside the radius of a banned one. That block lands before any classifier or reviewer has looked at the content, so the appeal is the first time a person sees it, and the radius deserves the same tuning discipline as a classifier cutoff. A wrong entry is quiet and permanent: nothing in the funnel ever re-examines a match.

  • hash matchknown-bad, fast
  • exact + perceptualcatch altered re-uploads
  • funnelcheap+certain → costly+fuzzy

What the new pieces do

Case Storestore
Every decision with its scores, action, and reviewer — the audit trail behind appeals, transparency reporting, and retraining datasets.

Step 3 · Judge novel text

Text classifiers, multi-label, per-category

Most content is new, so hashing won’t catch it. You need to judge unseen text across many kinds of harm — and "is it bad?" is the wrong question. What does a text classifier actually output?

Moderation APIclassifyText Classifiermulti-label
New in this step: Text Classifier.

What should a text moderation classifier produce?

  1. One label collapses distinct policies (spam vs self-harm vs hate) that need different thresholds and actions. Moderation is multi-label with per-category severity, not binary.

  2. The model scores each policy category independently, so a post can trigger several, each with its own probability — feeding per-category thresholds and severity in the policy engine.

  3. Moderation classifies and acts; it doesn’t edit users’ speech. The output is scores that drive an allow/limit/block/escalate decision.

The Text Classifier emits per-category probabilities (harassment, hate, violence, spam, self-harm, …) — multi-label, since content can violate several policies at once. Crucially it outputs scores, not verdicts: the policy engine (step 7) turns those probabilities into actions using category-specific thresholds. Keep the model’s judgment and the policy separate.

Why this piece earns its place

Two things decide whether a per-category score is worth thresholding at all. The first is how many labeled positives that category actually has. Harm categories differ in base rate by orders of magnitude — spam is everywhere, the severe categories are rare — so the categories where a wrong call costs most are the ones trained and calibrated on the fewest examples, and their scores are noisiest exactly where you need to cut. Positives per category, not overall accuracy, is the number to ask for. The second is what the model is allowed to see. Here it gets a string, and a string is context-free: the same words are abuse, a quotation condemning abuse, or an in-group reclamation, depending on who wrote it, to whom, and under what. So decide deliberately whether context — the parent post, the account age, the audience size — is a feature of the model or an input to the policy layer. In the model it is stronger, and it couples your classifier to systems that change under you. In policy it is blunter, and you can change it tonight.

  • multi-labelmany policies at once
  • per-categoryindependent scores
  • scores → policynot a raw verdict

What the new pieces do

Text Classifiermodel
Scores text across policy categories (harassment, hate, violence, spam, self-harm) with per-category probabilities. Multi-label, not one verdict — content can violate several policies at once.

Step 4 · See the other modalities

Image + multimodal, and text-in-image

Plenty of abuse hides where a text model can’t look: graphic images, and — cleverly — harmful text baked into an image to dodge text classifiers. How do you cover non-text content?

Moderation APIclassifyImage Modelmultimodal
New in this step: Image Model.

How do you catch harm in images (including text hidden in them)?

  1. Captions miss the image itself and the classic evasion of putting the slur or threat inside the picture. You need to analyze the pixels and any embedded text.

  2. Blocking a whole modality destroys the product. You classify images, not ban them.

  3. Vision models score images/video frames for policy violations, and OCR extracts text-in-image so it’s judged by the text classifier too — closing the "hide the words in a picture" loophole.

Add an Image / Multimodal Model that scores images and video frames (nudity, violence, weapons) and OCRs embedded text back into the text pipeline — so abuse hidden inside an image is still judged. Multimodal coverage closes the biggest evasion gap; each modality feeds the same policy engine.

Why this piece earns its place

OCR does not hand the text classifier the kind of text it was trained on. Extracted strings arrive without punctuation or case, with characters misread and words fused, and a model tuned on typed prose scores that distribution differently — usually lower, which is the direction that hurts, since closing an evasion was the whole point. So OCR output wants either its own thresholds downstream or training data that contains OCR text, and pretending the two channels are one is how the loophole reopens quietly. The cost side has a fix worth naming precisely. Vision inference is the most expensive call in the funnel, and a popular image is posted by many accounts, so cache the score against the exact hash the filter already computes — that one is a map key, and byte-identical re-posts are free from then on. What the cache cannot cover is the interesting traffic: a crop or a resave produces a different exact hash, and matching it is a radius search on the perceptual side, which is a nearest-neighbor lookup rather than a key. Keying a cache on that would serve one image's score for another. So scoring scales with distinct files, and the gap up to distinct media is the transformation volume this design assumes everywhere else.

  • image + videovision classifiers
  • OCRtext hidden in images
  • one policy engineall modalities feed it

What the new pieces do

Image Modelmodel
Classifies images/video frames (nudity, violence, weapons) and OCRs embedded text so abuse hidden inside images is caught. Handles the modality text classifiers can’t see.

Step 5 · Turn scores into decisions

The policy engine: thresholds + severity

You have hash hits and per-category probabilities. But a 0.7 "hate" score isn’t self-executing — is that block, limit, or review? And a 0.7 on self-harm should behave very differently from 0.7 on spam. Where does that logic live?

Moderation APIclassifyHash Filterknown-bad
New in this step: Hash Filter.

What converts classifier scores into an actual action?

  1. The policy engine maps scores + hash hits to actions using category-specific thresholds and severity: high-confidence severe → block, medium → age-gate/limit, uncertain → human review. Policy lives here, separate from the models.

  2. One threshold ignores that categories differ wildly in harm and base rate — self-harm and spam can’t share a cutoff. Thresholds are per-category.

  3. The top label discards confidence and severity, and hard-codes policy into the model. Decisions belong in a tunable policy layer, not the classifier.

A Policy Engine maps scores + hash hits to an action via per-category thresholds and severity tiers: high-confidence + severe → block; medium → limit/age-gate; low-confidence or high-stakes → escalate to human review; else allow. Keeping policy separate from the models lets you retune enforcement without retraining — and defines the uncertain band that goes to people.

Why this piece earns its place

Once the table holds a threshold per category, a severity tier, a surface and a jurisdiction, it is not configuration any more — it is a rules engine edited by people who do not ship code, and a typo in it reaches all traffic in seconds with no compiler in the way. Give it what code gets: a version, a named reviewer, and a staged rollout by share of traffic. The shape of that table is what surprises people. Cells multiply — categories times surfaces times regions — so nobody holds it in their head, and most cells were never written by anyone. Decide early whether an unwritten cell inherits a global default or fails closed, because the alternative is a rule that fires only for one small market and that no one has ever read. The engine should also emit which rule fired, not just the action. Every conversation downstream is about a rule — the notice a user receives, the context a reviewer needs, the argument about whether a decision was right — and an action with a score attached cannot name one.

  • per-categorythresholds
  • severity tiersallow/limit/block/escalate
  • tunableno retrain to adjust

What the new pieces do

Hash Filtercache
Matches content against hashes of known violating material (exact + perceptual/PhotoDNA-style). Extremely fast and high-precision — catches re-uploads of previously-actioned content before any ML runs.

Step 6 · People on the hard cases

The human-in-the-loop review queue

Classifiers are confident at the extremes and unsure in the middle — sarcasm, reclaimed slurs, context-dependent threats, novel evasions. Auto-actioning that middle at scale means confident wrong decisions on real people. Who decides the ambiguous cases?

Human Reviewqueue
New in this step: Human Review.

What happens to low-confidence and high-severity cases?

  1. Auto-approving the uncertain band lets borderline harm through and provides no correction signal. Ambiguity is exactly what needs a human.

  2. Low-confidence and high-severity cases go to human moderators via a prioritized queue; their rulings resolve the case and feed back as labeled data to improve the classifiers.

  3. Blocking all ambiguity mass-censors legitimate content (the false-positive failure). The point of review is to get the hard calls right, not to default to removal.

Route low-confidence and high-severity cases to a prioritized Human Review queue. Humans resolve the ambiguity the model can’t, and their decisions become labeled training data. Automate only the confident extremes; the uncertain middle is where people add irreplaceable judgment — and where removing them (the chaos button) breaks everything.

Why this piece earns its place

The queue produces the labels, so the queue's selection rule quietly decides what the next model is able to learn. Everything a person sees was chosen because it was uncertain or severe, which is a narrow band around the current thresholds. Train only on that and you sharpen a boundary you already had while learning nothing about the regions you never sample — including the confident auto-allows, which is precisely where a new evasion lands first and where the model believes it is right. The fix is unglamorous and has to be budgeted for: route a small random sample of auto-actioned and auto-allowed content into the queue on purpose, and pay reviewers to look at obvious content, because that sample is the only unbiased data the system ever collects. Capacity is the other constraint that bites. Reviewer headcount is fixed for months; arrival rate is not, and a news event or a coordinated campaign moves it in an afternoon. The band is the only lever that responds in minutes, so decide in advance how it narrows under load.

  • route uncertainlow-confidence → human
  • prioritizeby severity
  • labels backreviews retrain the model

What the new pieces do

Human Reviewservice
A prioritized queue of low-confidence and high-severity cases for human moderators. Their decisions resolve the case and become labeled data to retrain the classifiers.

Step 7 · Enforce, and let users appeal

Actions, appeals, and accountability

A decision is made — now act on it. But moderation systems make mistakes in both directions, and users deserve recourse. How do you enforce while staying accountable?

Text Classifiermulti-labelImage ModelmultimodalPolicy Enginedecide actionHuman ReviewqueueEnforcementallow · blockCase Storeaudit trail
New in this step: Policy Engine, Enforcement.

What does responsible enforcement require beyond taking the action?

  1. Silent removal with no notice or recourse erodes trust and hides false positives. Enforcement needs transparency and an appeal path.

  2. Warnings alone can’t address severe, illegal, or high-harm content. Enforcement spans a range of actions matched to severity.

  3. Enforcement applies a severity-matched action (limit, block, account action), notifies the user, and provides an appeal that sends the case (back) to human review — catching false positives and feeding the loop.

Enforcement applies a severity-matched action (limit, block, account-level action), notifies the user, and offers an appeal that routes the case to human review. Appeals are a feature, not an annoyance: they surface false positives, correct them, and generate the highest-value labeled data — cases the system got wrong.

Why this piece earns its place

For content caught at or before publish, blocking really is close to one write, and that is the path this design optimizes. The hard path is everything decided late: an escalation a reviewer resolves an hour after the post went up, an appeal that reverses last week's call, a sweep after a new hash lands. By then the post exists in fan-out copies — home feeds, notifications already delivered, a search index, a CDN edge holding the image, a cached embed on another site — and enforcement has to reach all of them. The failure is asymmetric: your dashboard records the action as successful because the source of truth changed, while the content keeps being served from a copy nobody revoked. Nothing errors and no alert fires. That is what a reconciliation sweep is for, re-checking actioned content against where it is genuinely still visible, with takedown latency measured at the edge rather than at the API. Account-level actions carry the other hidden cost. Strikes compose — enough of them become a suspension — so reversing one action on appeal is not one update, it is a recomputation of everything decided using that strike as an input. A pipeline that cannot unwind the chain leaves a user carrying a penalty derived from a decision the platform has already admitted was wrong.

  • graduated actionmatched to severity
  • notifyno silent removals
  • appealre-routes to human review

What the new pieces do

Policy Engineservice
Maps classifier scores + hash hits to an action using per-category thresholds and severity tiers: allow, limit/age-gate, block, or escalate to human review. Policy, separated from the models.
Enforcementbackend
Executes the decided action on the content and, where relevant, the account — with a record so it can be appealed and audited.

Step 8 · Keep up with adversaries

The retraining loop vs. evasion

Moderation is adversarial: the moment you block a pattern, users mutate it — leet-speak, new slang, coded language, adversarial images. A model frozen at launch decays fast. How do you stay current?

Policy Enginedecide actionHuman ReviewqueueFeedback Loopretrain
New in this step: Feedback Loop.

How does the system keep up as evasion tactics evolve?

  1. A frozen model rots against adversaries who probe it daily — recall silently drops as new evasions spread. Continuous retraining is mandatory.

  2. The feedback loop turns human rulings and appeal reversals into fresh labeled data, retraining classifiers to catch new evasions and correct past errors — a continuous arms race, not a one-time launch.

  3. Reports help but lag the harm and miss what users don’t flag. The durable defense is a retraining loop fed by expert review, with reports as one input.

A Feedback Loop turns reviewer decisions and appeal outcomes into fresh labeled data, continuously retraining the classifiers to catch new evasions and fix past mistakes. Moderation is a permanent arms race — the loop (plus updated hashes and policies) is how the system tracks a moving adversary instead of decaying after launch.

Why this piece earns its place

What the loop ships is not a model. Thresholds were tuned against the score distribution of one checkpoint, and a new checkpoint moves that distribution — often for the better, which still means every cutoff downstream now means something different. So a retrain is a joint rollout: re-derive the cutoffs against the candidate's scores on real traffic, then move model and policy together. Shipping the model alone is how a genuine accuracy improvement arrives as a week of over-blocking. Notice what that costs next to a threshold edit. The stored scores came from the old model, so they cannot answer the question, and the candidate has to score traffic again — which is why a retrain is scheduled and budgeted while a policy change is not. Measurement has the same shape of problem. You cannot hold out a control group here the way you would for a ranking model, because the control group is people left exposed on purpose. And any evaluation set was collected against the model being replaced, so the adversary has had that long to adapt to it — recall measured on an aged set reads high, and reads highest exactly when evasion is moving fastest. Re-sample it on a schedule, and watch the reviewer overturn rate as the earlier signal.

  • adversarialevasion evolves daily
  • reviews → labelsretrain continuously
  • appeals → labelslearn from mistakes

What the new pieces do

Feedback Loopbus
Turns reviewer decisions and appeal outcomes into fresh labeled data, retraining classifiers to keep up with adversarial evasion and shifting norms.

The payoff

You built content moderation

From "millions of posts, some harmful, users evading" to a staged system: a Moderation API with an audit trail, a hash filter for known-bad, text and multimodal classifiers emitting per-category scores, a policy engine mapping scores to graduated actions, a human review queue on the uncertain middle, appeals for recourse, and a retraining loop against adversarial drift.

Content InModeration APIHash FilterText ClassifierImage ModelPolicy EngineHuman ReviewEnforcementCase StoreFeedback Loop
The finished design, end to end. · swipe to pan the diagram

Now drop human review — auto-action every classifier call — and watch it fail in both directions at once: legitimate users banned on sarcasm and reclaimed speech, obfuscated abuse passing as clean, no appeal ever seen. That’s why a classifier is a probability, not a verdict; why you automate only the confident extremes; and why the ambiguous middle belongs to people.

Everything you assembled, in order

  • Moderation API — orchestrate the funnel; record every case for audit and appeal
  • Hash filter — known-bad by exact + perceptual hash — cheap, high-precision, first
  • Text classifier — per-category, multi-label scores — not a single verdict
  • Multimodal — image/video models + OCR close the text-in-image evasion
  • Policy engine — per-category thresholds + severity → allow/limit/block/escalate, tunable
  • Human review — low-confidence + high-severity cases go to people; their calls retrain
  • Appeals — graduated, notified, reversible — recourse surfaces false positives
  • Retraining loop — reviewer + appeal labels track adversarial drift
  • The failure — automating the uncertain middle = confident wrong actions, silently

Deep cut · 24:38

The model caught it. The rulebook let it spread.

The interactive build above lays out the moderation pipeline: hash matching of known-bad content, a classifier with one score per harm, OCR for text inside images, a policy engine with per-category thresholds and severity, a prioritized human review queue, appeals, and a retraining loop. This film is an autopsy. A scam post is caught by the model in one second, shared ten thousand times anyway, and 312 people send money. The film builds the system one reasonable fix at a time, with the checker written out line by line as pseudo code, then replays the night station by station to find why the post spread for nine hours when the model had it right.

  • See the wreck: the scam score of 0.83 sat between the lines, so the post correctly went to a person. But scam was filed at low severity, in the same drawer as spam, so it waited at the back of the review tray behind forty thousand spam posts and stayed on the feed, collecting shares, while it waited.
  • Take it into the interview: match known-bad content by perceptual hash before any model runs; score each harm separately and read text out of images; keep thresholds and severity in a policy engine outside the model, so a rule changes tonight without retraining; decide what happens to a post while it waits (limit, not allow); sort the review queue by severity times reach; turn reviewer decisions and appeals into fresh labels. Time to a person went from nine hours to four minutes, with no change to the model.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. Why does the hash filter need perceptual hashes, not just exact hashes?

    Exact hashes only match byte-identical files — resave an image at a different quality, crop one pixel, or add a watermark, and an exact hash misses it entirely. Perceptual hashes (PhotoDNA-style) are robust to those small transformations, matching content that LOOKS the same to a human even though the file bytes changed — which is exactly how known-bad material gets re-uploaded to evade detection.

  2. Content scores 0.8 on "spam" and 0.3 on "hate," both below their respective action thresholds. Does the spam score affect the hate decision?

    No — multi-label means each category is scored and thresholded independently. A high spam probability doesn’t inflate or discount the hate score; the policy engine evaluates each category against its own threshold and takes the most severe resulting action across all categories that clear their bar. Independence is what lets you tune one category’s sensitivity without side effects on the others.

  3. Who decides where the "uncertain middle" boundary actually sits, and how?

    It’s a precision/recall tradeoff tuned per category against labeled data — set the confidence band too narrow and you overload human review with easy cases; set it too wide and confident-but-wrong classifier calls get auto-actioned. Teams tune this using reviewer agreement rates: categories where the model and humans consistently agree can narrow the human-review band, categories with more disagreement need a wider one.

  4. Why are appeal-corrected cases more valuable training data than routine review-queue cases?

    A routine review case is the system correctly routing something uncertain to a human — expected behavior. An appeal-corrected case is the system being CONFIDENTLY WRONG: high enough classifier confidence to auto-action, but wrong enough that a human overturned it on appeal. That’s the exact failure mode retraining needs to fix, so those cases get weighted more heavily in the feedback loop than routine queue resolutions.

  5. How would this pipeline change for a livestream, where content can’t be held before publish?

    The proactive/pre-publish model breaks down — you can’t gate a live video frame-by-frame before broadcast without destroying the product. The funnel shifts to near-real-time: hash-match and classify frames as they’re encoded with a small buffer (seconds, not the pre-publish model’s unlimited hold), auto-cut only on the highest-confidence severe hits, and route everything else to review AFTER broadcast with faster takedown/ban tooling — proactive becomes "fastest possible reactive."

Check yourself — the answers, and why

Nine steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. The core discipline of a moderation funnel is…

    Classifiers are confident at the extremes and unsure in the middle; auto-actioning the middle produces confident wrong decisions at scale.

  2. Hash matching runs first because it…

    Exact + perceptual hashes catch previously-seen violating content instantly, so classifiers and humans focus on novel, ambiguous items.

  3. A text moderation classifier should output…

    Content can violate several policies at once; per-category scores plus a separate policy layer keep judgment and enforcement decoupled and tunable.

  4. OCR in the image pipeline exists to…

    Putting slurs/threats inside an image is a classic evasion; OCR extracts that text back into the text pipeline so it’s judged too.

  5. The policy engine is kept separate from the classifiers so you can…

    Enforcement (per-category thresholds, severity tiers, the escalate-to-human band) changes often; separating it from the model makes it tunable and auditable.

  6. Content moderation needs a continuous retraining loop because…

    Evasion tactics evolve daily; reviewer decisions and appeal outcomes retrain classifiers to track the moving adversary instead of decaying.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Screen content: take a post, comment, or image and return an action — allow, limit, block, or escalate — often before it’s shown.
  • Catch re-uploads: match content against known-violating material so previously-actioned posts are stopped before any ML runs.
  • Classify novel content: score unseen text and images across policy categories (harassment, hate, violence, spam, self-harm).
  • Escalate the ambiguous: route low-confidence and high-severity cases to a prioritized human review queue.
  • Appeal & audit: record every decision, notify the user, and offer an appeal that re-routes to human review.

The qualities that shape everything

Each one names the mechanism that buys it.

Don’t pay a classifier to rediscover known-bad
A hash filter matches known-violating material (exact + perceptual/PhotoDNA-style) before any model runs — cheap, near-perfect precision on re-uploads.
Don’t collapse distinct harms into one verdict
The text classifier emits per-category probabilities (multi-label), so each policy gets its own threshold instead of a single safe/unsafe label.
Close the text-in-image evasion
Image/multimodal models plus OCR pull text baked into pictures back into the text pipeline, so the modality text classifiers can’t see is still judged.
Retune enforcement without retraining
A policy engine maps scores + hash hits to actions via per-category thresholds and severity tiers, kept separate from the models.
Get the ambiguous cases right
Low-confidence and high-severity cases go to a prioritized human queue; automating that uncertain middle produces confident wrong actions at scale.
Track a moving adversary
Reviewer decisions and appeal outcomes become fresh labels, continuously retraining classifiers so recall doesn’t silently erode.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

Hash-match first over send everything to the classifiers

Much harmful content is re-uploaded, previously-seen material. Exact + perceptual hashing catches it at near-perfect precision before spending a cent on ML; classifiers then focus on the novel and ambiguous.

Per-category multi-label scores over a single safe/unsafe label

One label collapses spam, self-harm and hate — policies that need different thresholds and actions. Multi-label scores feed a tunable policy layer; content can violate several at once.

Per-category thresholds + severity over a single global threshold

Categories differ wildly in harm and base rate — self-harm and spam can’t share a cutoff. Per-category thresholds encode risk tolerance and law, and carve out the band that goes to humans.

Human review on the uncertain band over auto-actioning the middle for throughput

A classifier is a probability, not a verdict. Auto-actioning its uncertain middle produces confident wrong decisions on real people; humans resolve the ambiguity and their calls become training data.

A continuous retraining loop over shipping the classifier once

Moderation is adversarial — users mutate patterns daily (leet-speak, coded language). A frozen model silently loses recall; reviewer and appeal labels keep it tracking the moving target.

The answer, out loud

What a strong answer to “Design Content Moderation” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–3 min

    Name both ways it fails

    I'd start by naming the two ways this system fails, because they pull in opposite directions. It can leave harm up, and it can take legitimate speech down, and any single knob that fixes one makes the other worse. So I'm not going to design a model that's right — I'm going to design a funnel that is cheap and certain at the top, fuzzy in the middle, and has people where being wrong is expensive. I'd also ask which surfaces we screen, and how much of the decision has to happen before the content is ever visible.

  2. 3–8 min

    Content in, action out, with a case

    The skeleton is content arriving at a Moderation API that runs the funnel and returns allow, limit, block or escalate. I'd screen proactively for anything severe — at or before publish — and record a case for every decision, with the scores and the action, not just the outcome. That case store isn't bookkeeping. It's what makes appeals possible, it's what feeds retraining later, and it's what lets me test a policy change without shipping it. I'd also pin down, per surface, what the API returns when the funnel doesn't answer in time.

    Built in step 1: Content in, action out
  3. 8–12 min

    The cheapest check goes first

    Before any model runs, I'd hash-match against known-violating material — exact hashes for identical files, perceptual hashes for the crop-and-resave re-uploads. It's fast, precision is close to perfect, and for the legally mandated categories it's the serious path. The principle I'd say out loud is cheap and certain first, expensive and fuzzy later: the filter takes the fraction we already know about, so classifiers and people spend their time on what's genuinely novel. And our own reviewers' takedowns should be writing new hashes back into it.

    Built in step 2: The funnel: hash-match known-bad
  4. 12–18 min

    Scores, not a verdict

    For novel text I want per-category probabilities — harassment, hate, violence, spam, self-harm — multi-label, because one post can violate several at once. What I'd resist is a single safe-or-unsafe label. It collapses categories that deserve different thresholds and different actions, and it bakes policy into the weights, where changing it means a training run. The classifier's job is to say how likely each harm is. Deciding what to do about that number is a different job, one layer down.

    Built in step 3: Text classifiers, multi-label, per-category
  5. 18–23 min

    Cover the modality text cannot see

    Then images and video, because adversaries move to whatever you aren't checking. A vision model scores the pixels for nudity, violence, weapons, and OCR pulls text baked into the picture back into the text pipeline — that's the classic evasion, put the slur inside the image. Both modalities feed the same policy engine. I specifically don't want a second, divergent set of rules growing inside the image path.

    Built in step 4: Image + multimodal, and text-in-image
  6. 23–29 min

    Policy is a dial, not a weight

    The policy engine is where scores become actions, with per-category thresholds and severity tiers. High confidence and severe blocks; medium limits or age-gates; the uncertain band escalates to people. The reason I'd keep this out of the model is that enforcement moves on a different clock than training does — a new law lands, a category turns out to have been wrong since Tuesday — and I want that to be an edit someone makes tonight, not a training run. And because every case kept its raw scores, a proposed threshold change can be replayed over recorded decisions first: how many actions flip, in which categories, in which direction, and how much extra work lands in the review queue. That turns a threshold argument into a diff.

    Built in step 5: The policy engine: thresholds + severity
  7. 29–35 min

    People on the ambiguous middle

    Then the part I'd insist on: people on the hard cases, through a prioritized queue. Two admission rules, not one. Low confidence, because that's where the model is genuinely unsure. And high severity even when the model is confident, because that's where being wrong costs the most and I want a person's name on the decision. I'd also be specific about what a reviewer hands back — not a yes or a no, but which category and which action. Those rulings are the labels the next classifier trains on, and a binary verdict can't train a multi-label model, so a queue that records only a bare removal quietly degrades the thing it exists to improve.

    Built in step 6: The human-in-the-loop review queue
  8. 35–41 min

    Enforce reversibly, then keep learning

    Enforcement is graduated — limit, age-gate, block, account-level action — and I'd let reversibility set the bar for each one. A limit I can undo in a second, so I'm comfortable applying it on a middling score while a case waits for a human. An account-level action follows a person rather than a post, and it's the hardest thing here to walk back, so I'd want either high confidence or a reviewer behind it. Everything is notified and appealable, and the appeal has to reach a person: routing it back through the model that made the call just produces the same answer twice. What comes out of that path is what the retraining loop runs on, and it has to keep running, because the adversary doesn't stop when we ship.

    Built in step 7: Actions, appeals, and accountability
  9. 41–45 min

    Close on the trade-off

    So: a funnel with a hash filter in front, multi-label classifiers across modalities, a separate policy engine, people on the uncertain middle, appeals, and a loop that turns both into labels. The trade-off I'd name at the end is that certainty costs time, and the pre-publish latency budget is the ceiling on how much of it we can buy before anyone sees the content. Every accurate thing in this design — a vision model, OCR, a reviewer — is slower than the answer we owe the poster, so the real argument is how much of the judgment happens before publish and how much lands after. With more time I'd go after coordinated behavior, since judging one post at a time can't see a campaign spread across many accounts, and per-region policy, because the same score has to produce different actions in different jurisdictions.

    Built in step 8: The retraining loop vs. evasion

What this teaches

Learn AI system design by building a content-moderation pipeline (text + image) step by step. An interactive guide to the staged funnel — hash-matching known-bad content, text and multimodal classifiers, per-category thresholds, a human-in-the-loop review queue, a policy/enforcement engine, and an appeals + retraining loop against adversarial evasion — plus why fully trusting the classifier fails.

Key takeaways

  • Moderation API — orchestrate the funnel; record every case for audit and appeal
  • Hash filter — known-bad by exact + perceptual hash — cheap, high-precision, first
  • Text classifier — per-category, multi-label scores — not a single verdict
  • Multimodal — image/video models + OCR close the text-in-image evasion
  • Policy engine — per-category thresholds + severity → allow/limit/block/escalate, tunable
  • Human review — low-confidence + high-severity cases go to people; their calls retrain
  • Appeals — graduated, notified, reversible — recourse surfaces false positives
  • Retraining loop — reviewer + appeal labels track adversarial drift
  • The failure — automating the uncertain middle = confident wrong actions, silently

Concepts covered

  • Moderating content at platform scale
  • Content in, action out
  • The funnel: hash-match known-bad
  • Text classifiers, multi-label, per-category
  • Image + multimodal, and text-in-image
  • The policy engine: thresholds + severity
  • The human-in-the-loop review queue
  • Actions, appeals, and accountability
  • The retraining loop vs. evasion
RUN IT YOURSELF

The one dial that pays for the rare category out of everyone else’s posts

A day of traffic scored by one fixed multi-label classifier, with the policy engine tuned over it twice — one global cutoff, then one cutoff per category — running for real in your browser. It prints both failure columns for each, and they move in opposite directions: the per-category table removes 212 fewer real posts and leaves up 20 more violations, while missing the severe category less than half as often. Change self_harm’s severity of 40 in CATS, hit Run, and watch the signed deltas flip.

HOW TO READ THE CODE — 5 IDEAS
  1. The classifier is identical in A and B — same scores, same day. The only edit is the policy table, which is exactly why step 5 keeps thresholds and severity outside the model: it ships tonight, not next training run.
  2. One dial means only a post’s loudest score counts. The median self_harm violation scores 0.36, under the best global cutoff of 0.41 — so the single number gets dragged down chasing it and removes 455 clean posts on the way.
  3. Per-category cutoffs (.56/.49/.25) do not win every column. They win the two that carry cost — 212 fewer real posts removed, self_harm from 11/23 left up to 4/23 — and pay for it in +13 spam and +14 harassment left up. Severity-weighted, that left-up column still falls 480 → 325: step 5’s severity tier choosing which failure to take.
  4. 300 reviewer-shifts spent on the misplaced cutoff buy back only 96, because the band is defined relative to the cutoff (step 6). Review cannot rescue a dial sitting in the wrong part of the distribution; it compounds a dial that is already right.
  5. Run, not argued: put the three heads on one scale and only 78 of the 367 survives. Most of the gap was 0.50 meaning three different things — and that fix is a retrain, while the table is an edit.
CPython · WebAssembly
built to be reasoned about, not memorized — make the calls, drop human review, run the quiz.
Finished this one? 0 / 61 AI System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More AI System Designs