The whole design, in writing
Learn AI system design by building a content-moderation pipeline (text + image) step by step. An interactive guide to the staged funnel — hash-matching known-bad content, text and multimodal classifiers, per-category thresholds, a human-in-the-loop review queue, a policy/enforcement engine, and an appeals + retraining loop against adversarial evasion — plus why fully trusting the classifier fails.
Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.
The big idea
Moderating content at platform scale
A platform accepts millions of posts, comments, and images a day. Some tiny fraction is harmful — abuse, hate, violence, spam, CSAM. You can’t have humans read all of it (impossible at scale), and you can’t fully trust a model (it’s wrong in both directions: censoring the innocent and passing the harmful). Users actively try to evade you. How do you catch harm at scale without either drowning in review or silently over-censoring?
Build a staged moderation funnel: cheap high-precision filters first, ML classifiers next, and humans on the uncertain middle. A policy engine turns scores into actions; an appeals + retraining loop keeps up with adversaries. The core discipline: automate the confident extremes, send the ambiguous to people.
Step 1 · The skeleton
Content in, action out
A piece of content arrives and needs a decision — allow, limit, block, or escalate — often before it’s shown. What sits between the upload and the verdict?
What’s the minimal shape of a moderation request?
Purely reactive moderation leaves harmful content visible until someone reports it — unacceptable for severe categories. Serious platforms screen proactively, before or at publish.
An orchestrator takes the content, runs the funnel (fast filters → classifiers → policy), records the decision for audit/appeal, and returns allow/limit/block/escalate.
One model’s binary verdict has no notion of category, severity, confidence, or appeal — and it’s wrong in both directions. Moderation needs a staged funnel and a policy layer, not a coin flip.
A Moderation API orchestrates the decision: run content through the funnel, record a case (scores, action, reviewer) for audit and appeals, and return an action. Screening is largely proactive — decided at or before publish for anything severe — with latency budgets tighter for pre-publish checks.
Why this piece earns its place
Recording the case is the easy half of the paper trail. The half that matters six weeks later is which versions decided it: the checkpoint that produced the scores, and the policy revision that turned them into an action. Without both, an overturned appeal is unreadable — you cannot tell whether the model was wrong, the threshold has since moved, or the rule itself changed, and those three have different owners and different repairs. Store them beside the scores; they cost bytes. The other field to make first-class is the surface. One API serves posts, comments, private messages, profile names and uploads, and everything behind it wants to key off which one this is: the thresholds that apply, the severity a category carries there, and which actions even exist — you cannot age-gate a direct message. Let it arrive as an optional header nobody sets and you also lose the ability to answer the question that always comes, which is where we are wrong most. And treat the case write as part of the decision rather than as telemetry: make it fire-and-forget and you will eventually enforce an action that has no record, which is an action nobody can appeal, audit or learn from.
What the new pieces do
- Content Inclient
- User-generated content entering the platform — a comment, a post, an image. The moderation pipeline decides whether it can be shown, and how fast that decision must be.
- Moderation APIbackend
- Receives content, runs it through the staged funnel (hash filter → classifiers → policy), records a case, and returns an action. Balances latency (proactive, pre-publish) against thoroughness.
Step 2 · Cheapest check first
The funnel: hash-match known-bad
Running heavy ML on every item is expensive, and a lot of harmful content is re-uploaded, previously-seen material. Why pay a classifier to rediscover what you already know is banned?
What should run before the ML classifiers?
That wastes compute re-judging known-bad content and misses the highest-precision signal you have: exact matches to previously-actioned material.
Humans can’t pre-screen platform-scale volume; they’re the scarce resource you reserve for genuinely ambiguous cases, not a first-pass filter.
Hash-matching against a database of known violating material (exact and perceptual/PhotoDNA-style) is fast and near-perfect precision — it catches re-uploads instantly before any model runs.
Front the funnel with a Hash Filter: match content against hashes of known-violating material — exact hashes for identical files and perceptual hashes for slightly-altered images. It’s cheap, near-perfect precision, and catches re-uploads (including legally-mandated categories like CSAM) before spending a cent on ML. The funnel’s principle: cheap and certain first, expensive and fuzzy later.
Why this piece earns its place
A hash hit is enforcement without a score, so every entry needs provenance: where it came from — a shared industry list, or one of your own takedowns — and when it was written. Matching a shared list means enforcing another platform's judgment at your precision; that is the right call for the legally mandated categories, and it is still someone else's call, so when a block is disputed the entry's record is the only thing that can explain it. Two consequences people miss. Adding a hash is not retroactive: it stops the next upload and does nothing about the copies already published, so a new entry needs a sweep over stored media if you want those gone. And perceptual matching has a distance threshold, which means it has false positives — an unrelated image can land inside the radius of a banned one. That block lands before any classifier or reviewer has looked at the content, so the appeal is the first time a person sees it, and the radius deserves the same tuning discipline as a classifier cutoff. A wrong entry is quiet and permanent: nothing in the funnel ever re-examines a match.
- hash matchknown-bad, fast
- exact + perceptualcatch altered re-uploads
- funnelcheap+certain → costly+fuzzy
What the new pieces do
- Case Storestore
- Every decision with its scores, action, and reviewer — the audit trail behind appeals, transparency reporting, and retraining datasets.
Step 3 · Judge novel text
Text classifiers, multi-label, per-category
Most content is new, so hashing won’t catch it. You need to judge unseen text across many kinds of harm — and "is it bad?" is the wrong question. What does a text classifier actually output?
What should a text moderation classifier produce?
One label collapses distinct policies (spam vs self-harm vs hate) that need different thresholds and actions. Moderation is multi-label with per-category severity, not binary.
The model scores each policy category independently, so a post can trigger several, each with its own probability — feeding per-category thresholds and severity in the policy engine.
Moderation classifies and acts; it doesn’t edit users’ speech. The output is scores that drive an allow/limit/block/escalate decision.
The Text Classifier emits per-category probabilities (harassment, hate, violence, spam, self-harm, …) — multi-label, since content can violate several policies at once. Crucially it outputs scores, not verdicts: the policy engine (step 7) turns those probabilities into actions using category-specific thresholds. Keep the model’s judgment and the policy separate.
Why this piece earns its place
Two things decide whether a per-category score is worth thresholding at all. The first is how many labeled positives that category actually has. Harm categories differ in base rate by orders of magnitude — spam is everywhere, the severe categories are rare — so the categories where a wrong call costs most are the ones trained and calibrated on the fewest examples, and their scores are noisiest exactly where you need to cut. Positives per category, not overall accuracy, is the number to ask for. The second is what the model is allowed to see. Here it gets a string, and a string is context-free: the same words are abuse, a quotation condemning abuse, or an in-group reclamation, depending on who wrote it, to whom, and under what. So decide deliberately whether context — the parent post, the account age, the audience size — is a feature of the model or an input to the policy layer. In the model it is stronger, and it couples your classifier to systems that change under you. In policy it is blunter, and you can change it tonight.
- multi-labelmany policies at once
- per-categoryindependent scores
- scores → policynot a raw verdict
What the new pieces do
- Text Classifiermodel
- Scores text across policy categories (harassment, hate, violence, spam, self-harm) with per-category probabilities. Multi-label, not one verdict — content can violate several policies at once.
Step 4 · See the other modalities
Image + multimodal, and text-in-image
Plenty of abuse hides where a text model can’t look: graphic images, and — cleverly — harmful text baked into an image to dodge text classifiers. How do you cover non-text content?
How do you catch harm in images (including text hidden in them)?
Captions miss the image itself and the classic evasion of putting the slur or threat inside the picture. You need to analyze the pixels and any embedded text.
Blocking a whole modality destroys the product. You classify images, not ban them.
Vision models score images/video frames for policy violations, and OCR extracts text-in-image so it’s judged by the text classifier too — closing the "hide the words in a picture" loophole.
Add an Image / Multimodal Model that scores images and video frames (nudity, violence, weapons) and OCRs embedded text back into the text pipeline — so abuse hidden inside an image is still judged. Multimodal coverage closes the biggest evasion gap; each modality feeds the same policy engine.
Why this piece earns its place
OCR does not hand the text classifier the kind of text it was trained on. Extracted strings arrive without punctuation or case, with characters misread and words fused, and a model tuned on typed prose scores that distribution differently — usually lower, which is the direction that hurts, since closing an evasion was the whole point. So OCR output wants either its own thresholds downstream or training data that contains OCR text, and pretending the two channels are one is how the loophole reopens quietly. The cost side has a fix worth naming precisely. Vision inference is the most expensive call in the funnel, and a popular image is posted by many accounts, so cache the score against the exact hash the filter already computes — that one is a map key, and byte-identical re-posts are free from then on. What the cache cannot cover is the interesting traffic: a crop or a resave produces a different exact hash, and matching it is a radius search on the perceptual side, which is a nearest-neighbor lookup rather than a key. Keying a cache on that would serve one image's score for another. So scoring scales with distinct files, and the gap up to distinct media is the transformation volume this design assumes everywhere else.
- image + videovision classifiers
- OCRtext hidden in images
- one policy engineall modalities feed it
What the new pieces do
- Image Modelmodel
- Classifies images/video frames (nudity, violence, weapons) and OCRs embedded text so abuse hidden inside images is caught. Handles the modality text classifiers can’t see.
Step 5 · Turn scores into decisions
The policy engine: thresholds + severity
You have hash hits and per-category probabilities. But a 0.7 "hate" score isn’t self-executing — is that block, limit, or review? And a 0.7 on self-harm should behave very differently from 0.7 on spam. Where does that logic live?
What converts classifier scores into an actual action?
The policy engine maps scores + hash hits to actions using category-specific thresholds and severity: high-confidence severe → block, medium → age-gate/limit, uncertain → human review. Policy lives here, separate from the models.
One threshold ignores that categories differ wildly in harm and base rate — self-harm and spam can’t share a cutoff. Thresholds are per-category.
The top label discards confidence and severity, and hard-codes policy into the model. Decisions belong in a tunable policy layer, not the classifier.
A Policy Engine maps scores + hash hits to an action via per-category thresholds and severity tiers: high-confidence + severe → block; medium → limit/age-gate; low-confidence or high-stakes → escalate to human review; else allow. Keeping policy separate from the models lets you retune enforcement without retraining — and defines the uncertain band that goes to people.
Why this piece earns its place
Once the table holds a threshold per category, a severity tier, a surface and a jurisdiction, it is not configuration any more — it is a rules engine edited by people who do not ship code, and a typo in it reaches all traffic in seconds with no compiler in the way. Give it what code gets: a version, a named reviewer, and a staged rollout by share of traffic. The shape of that table is what surprises people. Cells multiply — categories times surfaces times regions — so nobody holds it in their head, and most cells were never written by anyone. Decide early whether an unwritten cell inherits a global default or fails closed, because the alternative is a rule that fires only for one small market and that no one has ever read. The engine should also emit which rule fired, not just the action. Every conversation downstream is about a rule — the notice a user receives, the context a reviewer needs, the argument about whether a decision was right — and an action with a score attached cannot name one.
- per-categorythresholds
- severity tiersallow/limit/block/escalate
- tunableno retrain to adjust
What the new pieces do
- Hash Filtercache
- Matches content against hashes of known violating material (exact + perceptual/PhotoDNA-style). Extremely fast and high-precision — catches re-uploads of previously-actioned content before any ML runs.
Step 6 · People on the hard cases
The human-in-the-loop review queue
Classifiers are confident at the extremes and unsure in the middle — sarcasm, reclaimed slurs, context-dependent threats, novel evasions. Auto-actioning that middle at scale means confident wrong decisions on real people. Who decides the ambiguous cases?
What happens to low-confidence and high-severity cases?
Auto-approving the uncertain band lets borderline harm through and provides no correction signal. Ambiguity is exactly what needs a human.
Low-confidence and high-severity cases go to human moderators via a prioritized queue; their rulings resolve the case and feed back as labeled data to improve the classifiers.
Blocking all ambiguity mass-censors legitimate content (the false-positive failure). The point of review is to get the hard calls right, not to default to removal.
Route low-confidence and high-severity cases to a prioritized Human Review queue. Humans resolve the ambiguity the model can’t, and their decisions become labeled training data. Automate only the confident extremes; the uncertain middle is where people add irreplaceable judgment — and where removing them (the chaos button) breaks everything.
Why this piece earns its place
The queue produces the labels, so the queue's selection rule quietly decides what the next model is able to learn. Everything a person sees was chosen because it was uncertain or severe, which is a narrow band around the current thresholds. Train only on that and you sharpen a boundary you already had while learning nothing about the regions you never sample — including the confident auto-allows, which is precisely where a new evasion lands first and where the model believes it is right. The fix is unglamorous and has to be budgeted for: route a small random sample of auto-actioned and auto-allowed content into the queue on purpose, and pay reviewers to look at obvious content, because that sample is the only unbiased data the system ever collects. Capacity is the other constraint that bites. Reviewer headcount is fixed for months; arrival rate is not, and a news event or a coordinated campaign moves it in an afternoon. The band is the only lever that responds in minutes, so decide in advance how it narrows under load.
- route uncertainlow-confidence → human
- prioritizeby severity
- labels backreviews retrain the model
What the new pieces do
- Human Reviewservice
- A prioritized queue of low-confidence and high-severity cases for human moderators. Their decisions resolve the case and become labeled data to retrain the classifiers.
Step 7 · Enforce, and let users appeal
Actions, appeals, and accountability
A decision is made — now act on it. But moderation systems make mistakes in both directions, and users deserve recourse. How do you enforce while staying accountable?
What does responsible enforcement require beyond taking the action?
Silent removal with no notice or recourse erodes trust and hides false positives. Enforcement needs transparency and an appeal path.
Warnings alone can’t address severe, illegal, or high-harm content. Enforcement spans a range of actions matched to severity.
Enforcement applies a severity-matched action (limit, block, account action), notifies the user, and provides an appeal that sends the case (back) to human review — catching false positives and feeding the loop.
Enforcement applies a severity-matched action (limit, block, account-level action), notifies the user, and offers an appeal that routes the case to human review. Appeals are a feature, not an annoyance: they surface false positives, correct them, and generate the highest-value labeled data — cases the system got wrong.
Why this piece earns its place
For content caught at or before publish, blocking really is close to one write, and that is the path this design optimizes. The hard path is everything decided late: an escalation a reviewer resolves an hour after the post went up, an appeal that reverses last week's call, a sweep after a new hash lands. By then the post exists in fan-out copies — home feeds, notifications already delivered, a search index, a CDN edge holding the image, a cached embed on another site — and enforcement has to reach all of them. The failure is asymmetric: your dashboard records the action as successful because the source of truth changed, while the content keeps being served from a copy nobody revoked. Nothing errors and no alert fires. That is what a reconciliation sweep is for, re-checking actioned content against where it is genuinely still visible, with takedown latency measured at the edge rather than at the API. Account-level actions carry the other hidden cost. Strikes compose — enough of them become a suspension — so reversing one action on appeal is not one update, it is a recomputation of everything decided using that strike as an input. A pipeline that cannot unwind the chain leaves a user carrying a penalty derived from a decision the platform has already admitted was wrong.
- graduated actionmatched to severity
- notifyno silent removals
- appealre-routes to human review
What the new pieces do
- Policy Engineservice
- Maps classifier scores + hash hits to an action using per-category thresholds and severity tiers: allow, limit/age-gate, block, or escalate to human review. Policy, separated from the models.
- Enforcementbackend
- Executes the decided action on the content and, where relevant, the account — with a record so it can be appealed and audited.
Step 8 · Keep up with adversaries
The retraining loop vs. evasion
Moderation is adversarial: the moment you block a pattern, users mutate it — leet-speak, new slang, coded language, adversarial images. A model frozen at launch decays fast. How do you stay current?
How does the system keep up as evasion tactics evolve?
A frozen model rots against adversaries who probe it daily — recall silently drops as new evasions spread. Continuous retraining is mandatory.
The feedback loop turns human rulings and appeal reversals into fresh labeled data, retraining classifiers to catch new evasions and correct past errors — a continuous arms race, not a one-time launch.
Reports help but lag the harm and miss what users don’t flag. The durable defense is a retraining loop fed by expert review, with reports as one input.
A Feedback Loop turns reviewer decisions and appeal outcomes into fresh labeled data, continuously retraining the classifiers to catch new evasions and fix past mistakes. Moderation is a permanent arms race — the loop (plus updated hashes and policies) is how the system tracks a moving adversary instead of decaying after launch.
Why this piece earns its place
What the loop ships is not a model. Thresholds were tuned against the score distribution of one checkpoint, and a new checkpoint moves that distribution — often for the better, which still means every cutoff downstream now means something different. So a retrain is a joint rollout: re-derive the cutoffs against the candidate's scores on real traffic, then move model and policy together. Shipping the model alone is how a genuine accuracy improvement arrives as a week of over-blocking. Notice what that costs next to a threshold edit. The stored scores came from the old model, so they cannot answer the question, and the candidate has to score traffic again — which is why a retrain is scheduled and budgeted while a policy change is not. Measurement has the same shape of problem. You cannot hold out a control group here the way you would for a ranking model, because the control group is people left exposed on purpose. And any evaluation set was collected against the model being replaced, so the adversary has had that long to adapt to it — recall measured on an aged set reads high, and reads highest exactly when evasion is moving fastest. Re-sample it on a schedule, and watch the reviewer overturn rate as the earlier signal.
- adversarialevasion evolves daily
- reviews → labelsretrain continuously
- appeals → labelslearn from mistakes
What the new pieces do
- Feedback Loopbus
- Turns reviewer decisions and appeal outcomes into fresh labeled data, retraining classifiers to keep up with adversarial evasion and shifting norms.
The payoff
You built content moderation
From "millions of posts, some harmful, users evading" to a staged system: a Moderation API with an audit trail, a hash filter for known-bad, text and multimodal classifiers emitting per-category scores, a policy engine mapping scores to graduated actions, a human review queue on the uncertain middle, appeals for recourse, and a retraining loop against adversarial drift.
Now drop human review — auto-action every classifier call — and watch it fail in both directions at once: legitimate users banned on sarcasm and reclaimed speech, obfuscated abuse passing as clean, no appeal ever seen. That’s why a classifier is a probability, not a verdict; why you automate only the confident extremes; and why the ambiguous middle belongs to people.
Everything you assembled, in order
- Moderation API — orchestrate the funnel; record every case for audit and appeal
- Hash filter — known-bad by exact + perceptual hash — cheap, high-precision, first
- Text classifier — per-category, multi-label scores — not a single verdict
- Multimodal — image/video models + OCR close the text-in-image evasion
- Policy engine — per-category thresholds + severity → allow/limit/block/escalate, tunable
- Human review — low-confidence + high-severity cases go to people; their calls retrain
- Appeals — graduated, notified, reversible — recourse surfaces false positives
- Retraining loop — reviewer + appeal labels track adversarial drift
- The failure — automating the uncertain middle = confident wrong actions, silently
