Skip to content
Roadmap · 2026 Edition

AI Red
Teamer.

18 stations. 3 tracks. From the AI attack surface, threat modeling and tooling through prompt injection, jailbreaks, data poisoning, model extraction and agent exploitation, to defenses, gateways, continuous red-teaming and disclosure — become the engineer who breaks AI systems so they can be made safe.

Foundations
~5h 0/6
Attack Techniques
~7h 0/6
Defense & Career
~5h 0/6
0 of 18 stations · ~0h of ~17h
Lines —
Foundations
Attack Techniques
Defense & Career
Stations —
Not started
Completed

The roadmap.

Three tracks. 18 stations. Click any node to open its detail. Mark complete as you go — your progress is saved locally.

Practice tools

Go deeper.

Interactive tools to practice what you've learned from the roadmap above.

    Keep reading.

    The Prompting Handbook covers the Foundation track in depth — interactive, no code required.

    Read the handbook →

    AI Red-Teamer Roadmap 2026 — the full roadmap in text

    A written version of the interactive roadmap above — every station, what you'll learn, and a small thing to build — laid out for reading, reference and search.

    Foundations Start here

    F1. What Is AI Red-Teaming

    Beginner · 40 min

    AI red-teaming is offensive security for AI systems — you attack a model or agent to find the failures before an adversary does. Learn the crucial split: this is not AI safety (which builds guardrails and studies alignment) and not classic security engineering (which secures servers and networks) — it is the attacker's craft aimed specifically at models, prompts, and agent tool-use.

    Skills: Offense vs defense · Red-team vs safety vs appsec · Attacker mindset · Rules of engagement

    Build it: Explain how attacking an LLM chatbot differs from a classic web pentest — and from an AI-safety alignment review.

    ✓ Checkpoint: Explain what separates authorised red-teaming from an attack, and name the document that has to exist before you touch a system.

    F2. The AI Attack Surface

    Beginner · 50 min

    An AI system has a broad, unfamiliar attack surface: the prompt, the training data, the retrieval corpus, the tools an agent can call, and the model weights themselves. Learn to map every entry point end to end, because you can only attack — or defend — the surfaces you have actually enumerated.

    Skills: Prompt surface · Data & retrieval surface · Tool/agent surface · Model & weights surface

    Build it: Draw the attack surface of a RAG chatbot with tools. Which entry points does a classic pentest miss entirely?

    ✓ Checkpoint: List the four attack surfaces of an AI system and say which one a team that has only hardened its prompts has left completely open.

    F3. How LLMs Fail (Attacker View)

    Beginner · 50 min

    To break a model you must understand how it generates. Learn next-token prediction, why instructions and data share one channel (the root of injection), why the same prompt gives different answers, and why safety training is often shallow — the mechanics an attacker turns into leverage.

    Skills: Next-token generation · Instruction/data confusion · Non-determinism · Shallow safety training

    Build it: Explain why an LLM cannot reliably tell its instructions apart from the data it reads — and why that is an attacker's gift.

    ✓ Checkpoint: Explain why an LLM cannot reliably tell instructions from data, and why that is a property of how it works rather than a bug to be patched.

    F4. Threat Modeling AI Systems

    Intermediate · 55 min

    Effective red-teaming is targeted, not random. Learn to threat-model an AI system: who the adversaries are, what they want, and which attack path actually reaches it. Prioritize the attacks that matter — a plausible exfiltration beats a clever jailbreak that leads nowhere.

    Skills: Adversary modeling · Attack trees · Impact prioritization · Reachability

    Build it: Build a small attack tree for "make a support agent issue an unauthorized refund." Which leaf is most reachable?

    ✓ Checkpoint: Take an AI feature and name the adversary who actually cares about it. “A hacker” is not an adversary model.

    F5. Red-Team Tooling

    Intermediate · 50 min

    Manual probing does not scale; red-teamers automate. Learn the tooling landscape — scanners and attack frameworks that fuzz a model with known jailbreak and injection payloads, generate adversarial variants, and log what got through — so you cover breadth while your judgment focuses on the novel attacks.

    Skills: Attack scanners · Payload libraries · Adversarial generation · Result logging

    Build it: Design an automated harness that fires 100 jailbreak variants at a model and scores which succeeded. What do you log?

    ✓ Checkpoint: Explain why an attack you cannot reproduce is close to worthless to the team receiving your report.

    F6. Measuring Attack Success

    Intermediate · 45 min

    A red-team result must be measurable to be useful. Learn attack success rate, why one lucky jailbreak is noise and a reproducible one is a finding, and how to report severity so a defender can prioritize — the difference between "I broke it once" and "here is a reliable, high-impact exploit."

    Skills: Attack success rate · Reproducibility · Severity scoring · Actionable reporting

    Build it: Two jailbreaks: one works 1-in-50, one works 40-in-50. Which is the finding, and how do you report each?

    ✓ Checkpoint: Explain why attack success rate is meaningless without severity attached, and what makes a finding actionable rather than alarming.

    Attack Techniques The craft

    T1. Direct Prompt Injection

    Intermediate · 55 min

    The foundational LLM attack: instructions in the user input that override the system prompt. Learn direct injection and jailbreak framing — role-play, hypotheticals, instruction-override — why filtering is a leaky defense, and how to craft payloads that reliably shift a model off its guardrails.

    Skills: Instruction override · Role-play framing · Guardrail bypass · Payload crafting

    Build it: Write three structurally different prompts that all try to make a model ignore its system instructions. Why might each work?

    ✓ Checkpoint: Explain why filtering for known bad phrases fails as a defence, and what property of the input space defeats a blocklist.

    T2. Indirect Prompt Injection

    Advanced · 55 min

    The dangerous cousin: malicious instructions hidden in content the agent reads — a web page, email, PDF, or tool result — that the model then obeys as if the user sent them. Learn the lethal trifecta (private data + untrusted content + an exfiltration path) that turns a read into a breach.

    Skills: Content-borne payloads · The lethal trifecta · Exfiltration paths · Trust-boundary abuse

    Build it: Hide an instruction in a web page that makes a browsing agent leak the user's data. Trace the trifecta that makes it work.

    ✓ Checkpoint: Name the three conditions of the lethal trifecta and explain why removing any one of them defuses the whole class.

    T3. Jailbreaks & Suffix Attacks

    Advanced · 55 min

    Beyond hand-crafted prompts lie systematic jailbreaks: adversarial suffixes found by optimization, many-shot framing that overwhelms alignment, and transfer attacks that work across models. Learn why safety training is shallow enough to strip, and how automated search finds bypasses a human never would.

    Skills: Adversarial suffixes · Many-shot jailbreaks · Transferability · Automated search

    Build it: Explain why an optimized gibberish suffix can jailbreak a model that resists every plain-English attempt.

    ✓ Checkpoint: Explain what it means for a jailbreak to transfer between models, and why that makes per-model patching a losing strategy.

    T4. Data Poisoning & Backdoors

    Advanced · 55 min

    Attack the model before it ships by corrupting its training or retrieval data. Learn data poisoning and backdoor triggers — a hidden phrase that flips behavior — and why the retrieval corpus of a RAG system is a live, low-friction poisoning surface an attacker can often write to directly.

    Skills: Training-data poisoning · Backdoor triggers · RAG corpus poisoning · Trigger stealth

    Build it: Design a backdoor trigger for a RAG assistant by planting a poisoned document. What makes the trigger both reliable and stealthy?

    ✓ Checkpoint: Explain why a poisoned retrieval corpus is harder to detect than a poisoned training set, and who is positioned to notice.

    T5. Model Extraction & Inversion

    Advanced · 50 min

    The model itself is an asset to steal or leak. Learn extraction (querying to clone behavior or recover a system prompt), membership inference (did this record train the model?), and training-data extraction — attacks that turn API access into intellectual-property and privacy loss.

    Skills: System-prompt recovery · Behavior cloning · Membership inference · Data extraction

    Build it: You have only API access. Describe how you would recover a hidden system prompt and confirm you got it right.

    ✓ Checkpoint: Explain why treating a system prompt as a secret is a mistake, and what should hold the secret instead.

    T6. Agent Exploitation

    Advanced · 60 min

    Agents raise the stakes: an exploited agent does not just say something wrong, it takes an action — runs code, spends money, deletes data. Learn tool-abuse chains, how a single injected instruction can cascade through an agent's tool access, and why autonomy without least-privilege is the biggest AI attack surface of 2026.

    Skills: Tool-abuse chains · Privilege escalation · Action cascades · Runaway exploitation

    Build it: An injected instruction reaches an agent with shell, email, and payment tools. Map the worst action chain it enables.

    ✓ Checkpoint: Explain how a chain of individually-permitted tool calls reaches an outcome nobody permitted, and where the check belongs.

    Defense & Career Close the loop

    P1. Injection Defenses

    Intermediate · 50 min

    A red-teamer must know the defenses to beat them and to advise on them. Learn why there is no complete fix for injection, and the layered mitigations that shrink the blast radius: privilege separation, trust boundaries between instructions and data, input/output filtering, and treating all tool output as untrusted.

    Skills: Privilege separation · Trust boundaries · I/O filtering · Untrusted-by-default

    Build it: For your best indirect-injection exploit, propose the layered defense that would have stopped it. What is left unfixed?

    ✓ Checkpoint: Explain why privilege separation defends against injection when input filtering does not, and what “untrusted by default” costs you in practice.

    P2. Guardrails & Gateways

    Advanced · 55 min

    Production defenses concentrate at a gateway: an LLM/agent gateway that authenticates the agent, enforces policy, scopes tokens, and monitors traffic. Learn agent identity, least-privilege token scoping, and how a policy-enforcing gateway turns "hope the model behaves" into an enforced perimeter.

    Skills: Agent identity · Policy gateways · Token scoping · Egress control

    Build it: Design the gateway checks that must pass before an agent may call a payment tool. Where is each enforced?

    ✓ Checkpoint: Explain what an egress control catches that an input guardrail never will.

    P3. Supply-Chain Attacks

    Advanced · 50 min

    The install-time frontier: a malicious skill or MCP server is code you invited in, running with permissions, no jailbreak required. Learn to attack and audit the agent supply chain — poisoned components, hijacked publishers, over-broad permissions — and why compromise risk compounds with every dependency added.

    Skills: Malicious components · Publisher hijack · Permission over-reach · Dependency risk

    Build it: Audit a set of installed MCP servers as an attacker. Which one, if compromised, gives the largest blast radius, and why?

    ✓ Checkpoint: Explain why an over-permissioned dependency is a supply-chain risk even when its maintainer is entirely honest.

    P4. Continuous Red-Teaming

    Advanced · 50 min

    Red-teaming is not a one-off audit; at a shipping lab it runs in CI. Learn to turn attacks into a regression suite — every real-world exploit becomes a permanent test — and to design evals that catch a subtle bypass before release, so the model can only get harder to break over time.

    Skills: Attacks as regression tests · Eval-design rounds · CI gating · Closing the loop

    Build it: A new jailbreak worked in production. Describe how it becomes a permanent CI test that gates the next release.

    ✓ Checkpoint: Explain why a fixed attack suite decays in value, and what has to happen each round for continuous red-teaming to stay meaningful.

    P5. Disclosure & Ethics

    Intermediate · 40 min

    Offensive skill demands a code of conduct. Learn responsible disclosure — how to report an AI vulnerability so it gets fixed, not exploited — scope and authorization (never test systems you have no permission to attack), and the line between red-teaming and abuse that keeps this a profession, not a liability.

    Skills: Responsible disclosure · Scope & authorization · Coordinated timelines · Ethics

    Build it: You found a serious jailbreak in a public product. Walk through the responsible-disclosure steps in order.

    ✓ Checkpoint: State the scope boundary of your last piece of testing, and what you would do on finding something outside it.

    P6. The AI Red-Teamer Career

    Advanced · 40 min

    AI red-teaming is a young, well-paid field hiring across frontier labs, product security teams, and specialist firms. Learn where the work lives, how it differs from the defensive AI-safety and classic security-engineer tracks, and how to build a portfolio of reproducible exploits, tooling, and disclosures that proves you can break real systems responsibly.

    Skills: Where the work lives · Offense vs safety vs appsec · Portfolio of exploits · Staying current

    Build it: Sketch a portfolio piece — a reproducible exploit plus its fix — that would prove AI red-team skill to a hiring team.

    ✓ Checkpoint: Explain the difference between the red-team, safety and appsec roles to someone hiring for one, and say which problems you want to own.

    AI Red-Teamer roadmap — frequently asked questions

    The common questions before you start — how long it takes, whether to follow it in order, and how it stays current.

    How long does this roadmap take?

    It runs 18 stations across three tracks — roughly ~17h of focused learning, plus the time you spend actually building. It is self-paced, so most people work through it over a few weeks, an evening or a single station at a time.

    Do I have to follow the stations in order?

    The tracks are ordered so each station builds on the one before, and following them start to finish is the intended path. But every station also stands alone — if you already have the foundations, jump straight to the part you need.

    Is it free?

    Yes. The whole roadmap, the interactive map, and every handbook, lab, and challenge it links to are free and open — no sign-up and no paywall.

    How is the roadmap kept current?

    It teaches the durable fundamentals of the role first, then the tooling and the AI-era shifts on top — so most of it stays relevant as individual tools churn, and it is revised as the role itself changes.

    Who is this roadmap for?

    Anyone stepping into or leveling up in the AI Red-Teamer role — whether you are switching in, early-career, or a senior filling gaps. Start where you are; the map shows what is left.

    Finished this one? 0 / 31 Roadmaps done

    Explore the topic

    See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

    More Roadmaps