AI research papers, in order.
110 written breakdowns, from 2011 to 2026, arranged as 19 reading ladders rather than a wall of cards. Each rung assumes the one below it — which is the whole point, because reading order changes what a paper teaches: Chinchilla measures exactly what Scaling Laws measured, in the same units, and gets a different answer.
Foundations
The machinery every paper after 2017 quietly assumes you already know: how a deep net is made to converge at all, and what it learns when it does.
Making deep nets train
Seven papers that turned "deeper is better in theory" into a network that actually converges. Every later architecture paper assumes all of them.
- Dropout2014 Start here Randomly switch off half your neurons on every training step — and overfitting collapses.
- Batch Normalization2015The layer that made deep networks train fast and forgivingly.
- Adam2015The optimizer that trains almost every modern network.
- ResNet2016The skip connection that made deep learning deep.
- LayerNorm & RMSNorm2016The normalization inside every transformer.
- mixup2018The two-line data augmentation that trains on CONVEX COMBINATIONS of example pairs — blending both inputs and labels: x̃ = λ·x_i + (1−λ…
- The Lottery Ticket Hypothesis2019The finding that a dense, randomly-initialized network hides a small "winning ticket" — a sparse subnetwork that, trained in isolation…
Learning to see
From the network that started the deep-learning era to one you can segment anything with — and the shift from labels to self-supervision in between.
- AlexNet2012 Start here The paper that started the deep learning revolution.
- Vision Transformer2021When an image became 16×16 words.
- SimCLR2020Contrastive learning that rivaled supervised vision with no labels.
- Masked Autoencoders2022BERT for images, finally working.
- DINOv22023Label-free visual features by self-distillation: a student learns to match a teacher that is an EMA of itself, with centering and sharp…
- Segment Anything2023A foundation model for cutting out objects.
Learning by trial and error
Reward, policy, search. The line from Atari to a Go engine that was never shown a human game — and the algorithm (PPO) that RLHF later borrowed wholesale.
- DQN2015 Start here The Deep Q-Network that learned Atari from raw pixels and launched deep reinforcement learning.
- PPO2017Proximal Policy Optimization — the RL algorithm behind RLHF.
- AlphaGo2016The system that beat a Go world champion by fusing Monte Carlo Tree Search with deep networks: a policy network to propose moves and a…
- AlphaZero2018The single algorithm that mastered Go, chess, and shogi from ZERO human data — pure self-play, same code for all three.
- MuZero2020Planning without knowing the rules.
Structure, not sequence
What to do when the data is a graph rather than a line of tokens — message passing, and the protein-folding system built on it.
Language models
One architecture, then fifteen years of arguing about how to scale it, sparsify it, align it and make it reason.
The Transformer
The architecture itself: what it replaced, how position is encoded once recurrence is gone, and what it turns out to have learned inside.
- Seq2Seq2014 Start here Teaching neural nets to translate.
- Attention Is All You Need2017Key paperThe paper that introduced the Transformer — the architecture underneath GPT, BERT, Claude, and Gemini.
- RoPE2021The position scheme inside LLaMA, GPT-NeoX, and most modern LLMs.
- Concept Circuits2026The foundation of mechanistic interpretability — how to read a model.
Making the context long
Attention costs quadratic in sequence length. Six attempts to escape that, ending with two 2026 papers that give the model a memory instead.
- Transformer-XL2019 Start here The transformer that broke the fixed-context limit.
- Longformer2020The attention pattern that let transformers read long documents.
- Ring Attention2023Context that scales with the number of devices.
- Mamba2023The linear-time challenger to attention.
- Titans2025Models that memorize at test time.
- Engram2026The memory half of the long-context problem.
Pre-training, and the scaling laws
The ladder where rung 5 overturns rung 4. Read `scaling-laws` before `chinchilla` — the second measures the same axis in the same units and gets a different answer.
- BERT2019 Start here The paper that made "pre-train once, fine-tune anywhere" the default for NLP.
- T52020The model that turned every NLP task into text-to-text.
- GPT-32020The paper that made models learn from the prompt.
- Scaling Laws2020The physics of making models bigger.
- Chinchilla2022The compute-optimal scaling law that said models were too big.
- Textbooks Are All You Need2023A 1.3B model trained on curated, textbook-quality data rivaled models 10× larger.
- Muon2024The optimizer that orthogonalizes weight-matrix updates instead of scaling them element-wise like Adam.
One model, many experts
Sparsity: route each token to a few experts and buy parameter count without paying for it at inference. 2017 idea, 2026 frontier default.
- Mixture of Experts2017 Start here How models got huge without getting slow.
- Switch Transformer2022The mixture-of-experts model that scaled to 1.6 trillion parameters by routing each token to a single expert.
- DeepSeek-V32024Key paperFrontier performance at a tenth of the cost.
- DeepSeek-V42026The efficiency sequel to V3, aimed at the real cost of long context.
- Qwen4-Coder2026An open, Apache-2.0 coding model that reportedly hits ~82% on SWE-Verified while running on a MacBook — the first open, Mac-runnable mo…
Teaching a base model to follow
A pre-trained model completes text; it does not take instructions. Nine papers on getting from one to the other, from RLHF to a loss with no reward model in it.
- InstructGPT2022 Start here How RLHF taught models to follow instructions — the paper behind ChatGPT.
- Constitutional AI2022Key paperAlignment from principles, not labels.
- Self-Instruct2023How to bootstrap a large instruction-tuning dataset from a language model's OWN generations, starting from just 175 human-written seed…
- LoRA2022The paper that made fine-tuning affordable.
- QLoRA2023Fine-tune a 65B model on one GPU.
- Direct Preference Optimization2023RLHF without the RL.
- GRPO2024The RL algorithm behind DeepSeek-R1.
- STaR2022A model teaching itself to reason.
- Let's Verify Step by Step2023Reward every reasoning step, not just the final answer.
Reasoning you can prompt, then train
It began as a prompting trick and ended as a training objective. Read in order, the progression from "let's think step by step" to a model trained to.
- Chain-of-Thought Prompting2022 Start here Show a model a few worked examples with their reasoning and it starts solving problems it used to fail.
- Self-Consistency2022Sample many diverse reasoning paths, marginalize out the reasoning, majority-vote the answer.
- Tree of Thoughts2023When LLMs learn to search.
- OpenAI o12024The reasoning model trained with reinforcement learning to think before it answers — and the discovery of a second scaling axis.
- DeepSeek-R12025How a model learned to reason from reward alone.
The open-weights line
Five releases you can actually download, in the order each one answered the last. This is where "what does a frontier recipe look like" is documented in public.
- Llama 22023 Start here The paper that opened up the LLM.
- Llama 32024A herd of open dense transformers trained far past the Chinchilla point.
- Qwen32025An open model family that unifies a thinking mode (chain-of-thought) and a non-thinking mode (direct answer) in one model, with a contr…
- Kimi K22025A trillion parameters built for agents.
- Kimi K32026The sequel to Kimi K2, and a different bet: attack decode SPEED, not KV memory.
What the labs report
Four papers that are reports, not methods: they state capability and withhold recipe. Read for the evaluation design and the omissions, not the architecture.
- PaLM2022 Start here A 540-billion-parameter dense model trained with Pathways that made emergent abilities concrete.
- GPT-4 Technical Report2023The report is famous for what it withholds, but its key methodological claim is predictable scaling: OpenAI predicted GPT-4's final los…
- Gemini 1.52024A million-token context window with near-perfect needle-in-a-haystack recall.
- GPT-5.6 System Card2026A modern system card is really a COMPUTE-ALLOCATION POLICY, not one model.
Systems & applications
What gets built on top: retrieval, agents, cheap serving, image and speech generation — and what the industry numbers actually say.
From a word vector to RAG
The longest ladder here, and the most linear: nine papers that carry one idea — text as geometry — from a 2013 word embedding to a graph-structured retrieval system.
- Word2Vec2013 Start here How words became math.
- GloVe2014Word vectors from GLOBAL co-occurrence statistics — the count-based cousin of word2vec.
- Sentence-BERT2019The model that turned BERT into fast, comparable sentence embeddings — the ancestor of every embedding model behind vector search and RAG.
- Product Quantization2011The compression that makes billion-scale vector search fit in memory.
- HNSW2016The graph index behind fast approximate nearest-neighbor search — the default in nearly every vector database.
- Dense Passage Retrieval2020The dual-encoder method that made dense retrieval beat BM25 for open-domain QA.
- ColBERT2020Keep one vector per token, not one per document, and score by late interaction: each query token takes its best match anywhere in the d…
- Retrieval-Augmented Generation2020Key paperThe paper behind every "chat with your docs" product.
- GraphRAG2024RAG for the global questions vector retrieval can't answer.
Agents that act
Give the model a tool, then a loop, then a memory, then a world. Ends on the 2026 paper measuring how badly the whole stack degrades over many turns.
- ReAct2023 Start here The blueprint for AI agents.
- Toolformer2023How a language model taught itself to use APIs.
- Reflexion2023Agents that learn from their own mistakes — reinforcement through words instead of weights.
- Voyager2023The Minecraft agent that never stops learning.
- Generative Agents2023The Smallville paper: 25 LLM characters living in a simulated town.
- MemGPT2023An operating system for LLM memory.
- LLMs Get Lost in Multi-Turn Conversation2026The ICLR 2026 Outstanding Paper that measured a failure everyone had felt: give a top model a fully-specified task in ONE prompt (conca…
Serving it cheaply
Five papers about the bill. Each one buys throughput back from a different part of the stack: the weights, the memory hierarchy, the KV cache, the decode loop.
- Knowledge Distillation2014 Start here How to compress a big, accurate "teacher" into a small, fast "student" — by training on the teacher's soft probability distribution, no…
- FlashAttention2022Attention was slow for the wrong reason.
- PagedAttention2023The idea that made LLM serving several times cheaper: store the KV cache in fixed-size blocks like an OS pages memory, instead of reser…
- Speculative Decoding2023Make a big model generate 2–3× faster with zero change to its output.
- Medusa2024Speculative decoding without a separate draft model.
Generating pixels
Ten papers from the first latent-variable generator to video and real-time 3D. The diffusion half is strictly cumulative — each rung removes a cost the last one paid.
- VAE2014 Start here The Variational Autoencoder — how to turn a compressor into a generator.
- GANs2014The counterfeiter and the detective.
- VQ-VAE2017The autoencoder that made the latent DISCRETE — and quietly became the tokenizer behind image generation.
- DDPM2020Key paperThe original denoising diffusion math — the paper that made diffusion work.
- DDIM2021Denoising Diffusion Implicit Models — the trick that made diffusion sampling deterministic and 10–50× faster, using the SAME trained DD…
- Classifier-Free Guidance2021The "guidance scale" slider behind every diffusion image tool.
- Latent Diffusion2022The paper behind Stable Diffusion.
- Sora2024Text-to-video by treating video as a diffusion transformer over spacetime patches.
- NeRF2020The paper that launched the radiance-field era.
- 3D Gaussian Splatting2023Real-time photorealistic radiance fields.
One model, many senses
Joining text to images, audio, video and finally to a robot arm. Starts with the contrastive objective the whole field standardised on.
- CLIP2021 Start here How AI learned to connect images and words.
- Flamingo2022The visual language model that bridges a frozen vision encoder and a frozen LLM with a small trained connector.
- LLaVA2023Visual instruction tuning the simple way.
- wav2vec 2.02020Self-supervised speech: learn from raw, untranscribed audio, then fine-tune on as little as TEN MINUTES of labels.
- Whisper2022Robust speech recognition from the open web.
- V-JEPA2024A self-supervised video world model that predicts in representation space, not pixels.
- RT-22023A Vision-Language-Action model that controls a robot by emitting actions as text tokens.
- Genie2024A generative interactive environment that turns an image into a frame-by-frame playable world — trained on unlabeled video with no acti…
Models that write code
Completion, then competition, then a benchmark made of real GitHub issues — the point at which "can it code?" got an answer you could audit.
The industry numbers
Not research papers: a consultancy census, a VC essay and a job-posting analysis. Included because their figures get quoted as if they were, and read here for their method.
- The 95% Number, Examined2026 Start here The most-quoted statistic in enterprise AI, read carefully: roughly 95% of pilots reportedly produced no MEASURABLE profit-and-loss imp…
- Trading Margin for Moat2026Why a software company would deliberately hire expensive engineers to do customer work.
- 1,000 FDE Jobs, Analysed2026Reading a thousand job postings beats reading a thousand opinions.
Paper breakdowns — frequently asked questions
What is a paper breakdown?
A written walk-through of one research paper: the problem it was published against, the mechanism it introduced, the numbers it reported, what its authors admitted it could not do, and a runnable Python exercise that re-derives its central result. It is not a summary of the abstract, and it is not a replacement for the paper — every breakdown links the original on arXiv.
Are the breakdowns free?
Yes — every breakdown is free to read, with no paywall and no sign-up. There is nothing to watch: these are written deep-dives with an interactive exercise, not videos.
Which paper should I read first?
"Attention Is All You Need" — rung 2 of the Transformer ladder, and the paper that introduced the architecture underneath GPT, Claude and Gemini. If you want the ground under it first, rung 1 of that ladder (Seq2Seq) is what the Transformer replaced.
Why are the papers in ladders instead of categories?
Because reading order changes what a paper teaches. The clearest case is Scaling Laws and Chinchilla: the second measures the same quantity as the first, in the same units, and reaches a different answer. Read the wrong way round, you learn the conclusion that was overturned. A category list cannot express that; an ordered ladder can, so every track here is a reading order and each rung is written assuming the one below it.
Do I need a research background to follow along?
No. Each breakdown is written for engineers and builders rather than academics — every mechanism is explained in plain terms before the formal version, and where a paper depends on an earlier one, that earlier paper is the rung below it on the same ladder.
A paper missing?
If there's a paper you'd like broken down — or a rung you think sits in the wrong place on one of the ladders above — send it through. Genuine requests shape what gets covered next.