Skip to content
Roadmap · 2026 Edition

Generative
AI.

18 stations. 3 tracks. From how text and images are generated to shipping a real generative product — at your own pace.

Foundations
~5h 0/6
Techniques
~6h 0/6
Production
~6h 0/6
0 of 18 stations · ~0h of ~17h
Lines —
Foundations
Techniques
Production
Stations —
Not started
Completed

The roadmap.

Three tracks. 18 stations. Click any node to open its detail. Mark complete as you go — your progress is saved locally.

Practice tools

Go deeper.

Interactive tools to practice what you've learned from the roadmap above.

Interactive tool

Build a CRISP prompt.

Fill in each component. Your prompt assembles live on the right.

Your prompt
Start typing in any field to see your prompt take shape…
Interactive lab

The Technique Lab.

Pick an experiment. See how each technique transforms a prompt — before and after.

Adding CRISP components to a vague prompt produces dramatically better output.

Before Vague prompt
Write a blog post about AI.
After CRISP prompt
Context: Our audience is non-technical founders who are curious about AI but intimidated by jargon.

Role: Act as a tech journalist who explains complex topics in plain English.

Instruction: Write a 600-word blog post introduction about how AI is changing content creation.

Specifics: Start with a hook anecdote. Use short paragraphs. Avoid technical terms — if you must use one, define it.

Proof: Tone example: "AI doesn't replace writers. It gives them a co-pilot that never sleeps."
AI response
Click "Simulate response" to see the difference.

    Keep reading.

    The Prompting Handbook covers the Foundation track in depth — interactive, no code required.

    Read the handbook →

    Generative AI Roadmap 2026 — the full roadmap in text

    A written version of the interactive roadmap above — every station, what you'll learn, and a small thing to build — laid out for reading, reference and search.

    Foundations Start here

    F1. What is Generative AI

    Beginner · 30 min

    A discriminative model tells cat from dog; a generative model paints a cat that never existed. Generative AI learns the underlying distribution of data — text, images, audio — and samples new examples from it. That one shift is what powers ChatGPT, Midjourney, and every model that creates.

    Skills: Generative vs discriminative · Learning a distribution · Sampling new data · Modalities: text / image / audio

    Build it: List five products you use that generate content. For each, name the modality and whether it feels more like autocomplete or like painting from noise.

    ✓ Checkpoint: Explain what it means for a model to learn a distribution rather than a boundary, and why that is what lets it generate.

    F2. How LLMs Generate Text

    Beginner · 60 min

    Text generation is next-token prediction on a loop: the model reads everything so far and picks the most likely next token, then repeats. Under it is the Transformer and self-attention. Understanding this makes prompting, streaming, and context limits all click.

    Skills: Next-token prediction · Self-attention · Autoregressive generation · Streaming output

    Build it: Generate the same prompt at temperature 0 and 1.0 several times. Explain, in terms of next-token probabilities, why one is repetitive and one is wild.

    ✓ Checkpoint: Explain why next-token prediction produces coherent paragraphs at all — the objective looks far too local for the result.

    F3. How Diffusion Generates Images

    Beginner · 60 min

    Image models work backwards from noise. In training they add noise to real images until they are static; at generation they start from pure noise and a trained network removes it step by step, sculpting a picture. That is the idea behind Stable Diffusion, DALL·E, and Midjourney.

    Skills: Forward noising · Reverse denoising · Sampling steps · Text conditioning

    Build it: Explain to a friend why "add noise, then learn to remove it" is easier to train than generating a full image in one shot.

    ✓ Checkpoint: Explain what the reverse denoising process is actually undoing, and what the number of sampling steps buys you.

    F4. Embeddings & Multimodal

    Intermediate · 45 min

    To connect words and images you put both in one shared space, where a photo and its caption land close together. That is what CLIP does — and it is the bridge that lets a text prompt steer an image model. Multimodal generation lives on this idea.

    Skills: Shared embedding space · Image–text alignment (CLIP) · Cosine similarity · Text-to-image guidance

    Build it: Explain how "a photo of a cat" can steer an image generator even though the model was never told what a cat looks like in labels.

    ✓ Checkpoint: Explain how an image and a caption end up close together in one embedding space, and what that enables.

    F5. Prompting for Generation

    Beginner · 45 min

    A prompt is the steering wheel of a generative model. For text: roles, structure, examples. For images: subject, style, composition, and negative prompts. Learning to describe what you want — and what you do not — is the core creative skill of GenAI.

    Skills: Descriptive prompts · Negative prompts · Style & structure control · Iteration

    Build it: Take one image idea and write three prompts: bare, detailed, and detailed-with-negatives. Compare how much control each gives.

    ✓ Checkpoint: Explain what a negative prompt is doing, and why it is not the same as omitting the thing.

    F6. Tokens, Context & Cost

    Intermediate · 30 min

    Generation is metered. Text is billed per token and bounded by a context window; images cost per generation and per resolution/steps. Knowing what drives cost and latency is what separates a fun demo from a product you can actually afford to run.

    Skills: Tokenization · Context windows · Cost per generation · Latency drivers

    Build it: Estimate the cost of generating 1,000 product descriptions vs 1,000 images. Which is pricier, and what knob would you turn to cut it?

    ✓ Checkpoint: Compute the cost of one generation from tokens and price, then say which part of the prompt you would cut first.

    Techniques Level up

    T1. Text Generation & Sampling

    Intermediate · 45 min

    How a model picks the next token shapes everything. Temperature scales randomness, top-p and top-k prune the candidates, and beam search hunts for the most likely sequence. Tuning these trades creativity against reliability — the dial every text feature needs set right.

    Skills: Temperature · Top-p / top-k · Beam search · Creativity vs reliability

    Build it: For a legal-summary tool and a brainstorming tool, pick temperature and top-p for each and justify why they differ.

    ✓ Checkpoint: Explain what temperature and top-p each change about sampling, and why raising both at once is usually a mistake.

    T2. Image Generation

    Intermediate · 60 min

    Modern image models run diffusion in a compressed latent space for speed, guided by your prompt. Guidance scale controls how hard it follows the text; ControlNet-style conditioning adds pose, depth, or edges so you can direct composition, not just describe it.

    Skills: Latent diffusion · Guidance scale · ControlNet conditioning · Img2img & inpainting

    Build it: Explain what raising the guidance scale does to an image, and when a lower value actually produces a better result.

    ✓ Checkpoint: Explain what guidance scale trades off, and what an extreme value does to the image.

    T3. Multimodal Models

    Advanced · 60 min

    Vision-language models take an image and text together and reason about both — describing a photo, answering questions about a chart, reading a screenshot. They fuse a vision encoder with a language model, and they are the fastest-moving frontier of GenAI.

    Skills: Vision-language models · Image + text in · Visual question answering · Grounding to pixels

    Build it: List three product features that only become possible once a model can see an image and read text at the same time.

    ✓ Checkpoint: Explain what a vision-language model must align before it can answer a question about an image.

    T4. Audio & Video Generation

    Advanced · 45 min

    The same generative ideas extend to sound and motion: text-to-speech that sounds human, music from a prompt, and the young, fast-improving field of generated video. These push the hardest on compute, consistency over time, and evaluation.

    Skills: Text-to-speech · Music generation · Text-to-video · Temporal consistency

    Build it: Why is generating a 10-second video far harder than generating 10 separate images? Name the constraint that ties the frames together.

    ✓ Checkpoint: Explain why temporal coherence is the hard part of video generation rather than image quality.

    T5. Fine-Tuning Generative Models

    Advanced · 60 min

    To teach a model your style, brand, or a specific subject, fine-tune it. LoRA trains tiny adapters cheaply; DreamBooth teaches an image model a new subject from a handful of photos. The skill is knowing when a few example prompts would do instead.

    Skills: LoRA adapters · DreamBooth · Style vs subject tuning · When not to fine-tune

    Build it: You want a model to always draw your mascot. Argue whether prompting, LoRA, or DreamBooth fits — and what data each needs.

    ✓ Checkpoint: Say whether your problem is style or subject, and defend the tuning approach that follows.

    T6. RAG & Grounded Generation

    Intermediate · 60 min

    Left alone, generative models confidently make things up. Grounding them in retrieved sources — RAG — keeps generated text tied to real documents you control, so answers are checkable and citable. It is the main defence against hallucination in generative products.

    Skills: Retrieval-augmented generation · Grounding · Citations · Hallucination control

    Build it: Design a "chat with our help docs" feature. How do you make sure every answer is grounded in a real doc, with a link?

    ✓ Checkpoint: Explain why grounding reduces hallucination without eliminating it, and what a citation is really asserting.

    Production Ship it

    P1. Evaluating Generative Output

    Advanced · 90 min

    How do you score something open-ended — a paragraph, an image? There is no single right answer. You combine human review, automated metrics, and LLM-as-judge into a repeatable eval so you can tell whether a change actually improved quality.

    Skills: Human evaluation · Automated metrics · LLM-as-judge · Golden sets & regression

    Build it: Design an eval for an image generator: what mix of human ratings and automated checks tells you a new model version is better?

    ✓ Checkpoint: Explain why an LLM judge that agrees with you 95% of the time may still be useless, and what you would check first.

    P2. Guardrails & Content Safety

    Intermediate · 60 min

    Generative models can be prompted into unsafe or off-brand output. Safety is layered: filter incoming prompts, moderate generated text and images, block disallowed content, and handle refusals gracefully — especially critical for anything public-facing.

    Skills: Prompt & output filtering · Media moderation · Disallowed content · Refusal handling

    Build it: For a public image generator, list the checks you run before showing a result to a user, and what happens when one trips.

    ✓ Checkpoint: Explain why output filtering is needed even when input filtering is in place.

    P3. Serving & Latency

    Advanced · 60 min

    Generation is slow and bursty. Streaming tokens as they are produced hides latency; batching improves throughput; and long generations or high-resolution images dominate cost. Serving GenAI well is about making the wait feel short and the bill stay small.

    Skills: Token streaming · Batching · Time-to-first-token · Long-generation cost

    Build it: Why does streaming make a 20-second answer feel fast? And what does batching buy you that streaming does not?

    ✓ Checkpoint: Explain why time-to-first-token matters more than total generation time for a streamed UI.

    P4. Cost & Caching

    Intermediate · 45 min

    Repeated or similar generations are everywhere — cache them. Route simple requests to smaller, cheaper models, cache identical prompts, and cap spend per user. In generative products, cost control is a feature users never see but the business always feels.

    Skills: Prompt / result caching · Model routing · Spend caps · Cheap-model fallbacks

    Build it: Design a cache for a text generator: when is a cached result safe to reuse, and when must you regenerate?

    ✓ Checkpoint: Explain why a semantic cache can serve a confidently wrong answer, and what threshold decision that risk lives behind.

    P5. Watermarking & Provenance

    Intermediate · 45 min

    As generated media floods the world, proving what is AI-made — and where real media came from — matters. Watermarking embeds invisible signals in output; provenance standards (like content credentials) attach a tamper-evident history to a file.

    Skills: Invisible watermarking · Content credentials · Provenance metadata · Detection limits

    Build it: Why is watermarking generated text much harder than watermarking generated images? What can an attacker do to each?

    ✓ Checkpoint: Explain what a watermark can and cannot prove about a piece of generated media.

    P6. Ship a Generative Product

    Advanced · 60 min

    The finish line: fold everything — generation, grounding, evals, guardrails, serving, and cost control — into a real product people can use safely and you can afford to run. A conversational assistant is the canonical example that ties the whole roadmap together.

    Skills: End-to-end architecture · Safety + cost + quality · User experience · Launch checklist

    Build it: Draw the architecture of a generative assistant: where do grounding, guardrails, caching, and evals each sit in the request path?

    ✓ Checkpoint: Name the trade you made between safety, cost and quality in something you shipped — all three cannot be maximised.

    Generative AI roadmap — frequently asked questions

    The common questions before you start — how long it takes, whether to follow it in order, and how it stays current.

    How long does this roadmap take?

    It runs 18 stations across three tracks, and it is self-paced — so most people work through it over a few weeks, an evening or a single station at a time. There is no clock; the map shows what is left.

    Do I have to follow the stations in order?

    The tracks are ordered so each station builds on the one before, and following them start to finish is the intended path. But every station also stands alone — if you already have the foundations, jump straight to the part you need.

    Is it free?

    Yes. The whole roadmap, the interactive map, and every handbook, lab, and challenge it links to are free and open — no sign-up and no paywall.

    How is this roadmap kept current?

    It teaches the durable fundamentals first, then the tooling and the AI-era shifts on top — so most of it stays relevant as individual tools churn, and it is revised as the field itself changes.

    Who is this roadmap for?

    Anyone stepping into or leveling up in this area — whether you are switching in, early-career, or a senior filling gaps. Start where you are; the map shows what is left.

    Finished this one? 0 / 31 Roadmaps done

    Explore the topic

    See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

    More Roadmaps