Vibe Engines
YouTube
PAPER BREAKDOWNS · READ THEM IN ORDER

AI research papers, in order.

110 written breakdowns, from 2011 to 2026, arranged as 19 reading ladders rather than a wall of cards. Each rung assumes the one below it — which is the whole point, because reading order changes what a paper teaches: Chinchilla measures exactly what Scaling Laws measured, in the same units, and gets a different answer.

110Papers
19Ladders
15Years covered
100%Free · no sign-up
Start hereAttention Is All You NeedNeurIPS 2017 — the architecture underneath GPT, Claude and Gemini, and rung 2 of the Transformer ladder. Everything in the two arcs below it either leads here or starts from it.Read it →
110 papers

Foundations

The machinery every paper after 2017 quietly assumes you already know: how a deep net is made to converge at all, and what it learns when it does.

Making deep nets train

Seven papers that turned "deeper is better in theory" into a network that actually converges. Every later architecture paper assumes all of them.

0 / 7 read · 2014–2019
  1. Dropout2014 Start here Randomly switch off half your neurons on every training step — and overfitting collapses.Srivastava, Hinton, Krizhevsky, Sutskever, Salakhutdinov~11 min
  2. Batch Normalization2015The layer that made deep networks train fast and forgivingly.Ioffe, Szegedy~11 min
  3. Adam2015The optimizer that trains almost every modern network.Kingma, Ba~12 min
  4. ResNet2016The skip connection that made deep learning deep.He, Zhang, Ren, Sun~7 min
  5. LayerNorm & RMSNorm2016The normalization inside every transformer.Ba, Kiros, Hinton · Zhang, Sennrich~11 min
  6. mixup2018The two-line data augmentation that trains on CONVEX COMBINATIONS of example pairs — blending both inputs and labels: x̃ = λ·x_i + (1−λ…Zhang, Cisse, Dauphin & Lopez-Paz~10 min
  7. The Lottery Ticket Hypothesis2019The finding that a dense, randomly-initialized network hides a small "winning ticket" — a sparse subnetwork that, trained in isolation…Frankle & Carbin~11 min

Learning to see

From the network that started the deep-learning era to one you can segment anything with — and the shift from labels to self-supervision in between.

0 / 6 read · 2012–2023
  1. AlexNet2012 Start here The paper that started the deep learning revolution.Krizhevsky, Sutskever, Hinton~8 min
  2. Vision Transformer2021When an image became 16×16 words.Dosovitskiy, Beyer, Kolesnikov, et al. (Google)~8 min
  3. SimCLR2020Contrastive learning that rivaled supervised vision with no labels.Chen, Kornblith, Norouzi, Hinton~11 min
  4. Masked Autoencoders2022BERT for images, finally working.He, Chen, Xie, Li, Dollár, Girshick~11 min
  5. DINOv22023Label-free visual features by self-distillation: a student learns to match a teacher that is an EMA of itself, with centering and sharp…Oquab, Darcet, Moutakanni, et al. (Meta)~11 min
  6. Segment Anything2023A foundation model for cutting out objects.Kirillov, Mintun, Ravi, Mao, et al. (Meta)~8 min

Learning by trial and error

Reward, policy, search. The line from Atari to a Go engine that was never shown a human game — and the algorithm (PPO) that RLHF later borrowed wholesale.

0 / 5 read · 2015–2020
  1. DQN2015 Start here The Deep Q-Network that learned Atari from raw pixels and launched deep reinforcement learning.Mnih et al.~11 min
  2. PPO2017Proximal Policy Optimization — the RL algorithm behind RLHF.Schulman et al.~11 min
  3. AlphaGo2016The system that beat a Go world champion by fusing Monte Carlo Tree Search with deep networks: a policy network to propose moves and a…Silver, Huang, Maddison, et al. (DeepMind)~11 min
  4. AlphaZero2018The single algorithm that mastered Go, chess, and shogi from ZERO human data — pure self-play, same code for all three.Silver et al.~11 min
  5. MuZero2020Planning without knowing the rules.Schrittwieser, Antonoglou, Hubert, et al. (DeepMind)~11 min

Structure, not sequence

What to do when the data is a graph rather than a line of tokens — message passing, and the protein-folding system built on it.

0 / 2 read · 2017–2021
  1. Graph Neural Networks2017 Start here How a neural network learns on graphs — molecules, social networks, citation webs — where a convolution has nothing to slide over.Kipf & Welling~11 min
  2. AlphaFold 22021The system that solved the 50-year protein-folding problem — by reading evolution.Jumper, Evans, Pritzel, et al. (DeepMind)~11 min

Language models

One architecture, then fifteen years of arguing about how to scale it, sparsify it, align it and make it reason.

The Transformer

The architecture itself: what it replaced, how position is encoded once recurrence is gone, and what it turns out to have learned inside.

0 / 4 read · 2014–2026
  1. Seq2Seq2014 Start here Teaching neural nets to translate.Sutskever, Vinyals, Le~8 min
  2. Attention Is All You Need2017Key paperThe paper that introduced the Transformer — the architecture underneath GPT, BERT, Claude, and Gemini.Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin~10 min
  3. RoPE2021The position scheme inside LLaMA, GPT-NeoX, and most modern LLMs.Su, Lu, Pan, Murtadha, Wen, Liu~12 min
  4. Concept Circuits2026The foundation of mechanistic interpretability — how to read a model.Mechanistic Interpretability~9 min

Making the context long

Attention costs quadratic in sequence length. Six attempts to escape that, ending with two 2026 papers that give the model a memory instead.

0 / 6 read · 2019–2026
  1. Transformer-XL2019 Start here The transformer that broke the fixed-context limit.Dai, Yang, Yang, Carbonell, Le, Salakhutdinov~11 min
  2. Longformer2020The attention pattern that let transformers read long documents.Beltagy, Peters, Cohan~11 min
  3. Ring Attention2023Context that scales with the number of devices.Liu, Zaharia, Abbeel~11 min
  4. Mamba2023The linear-time challenger to attention.Gu, Dao~9 min
  5. Titans2025Models that memorize at test time.Behrouz, Zhong, Mirrokni (Google)~8 min
  6. Engram2026The memory half of the long-context problem.DeepSeek-AI~8 min

Pre-training, and the scaling laws

The ladder where rung 5 overturns rung 4. Read `scaling-laws` before `chinchilla` — the second measures the same axis in the same units and gets a different answer.

0 / 7 read · 2019–2024
  1. BERT2019 Start here The paper that made "pre-train once, fine-tune anywhere" the default for NLP.Devlin, Chang, Lee, Toutanova~8 min
  2. T52020The model that turned every NLP task into text-to-text.Raffel, Shazeer, Roberts, Lee, et al. (Google)~8 min
  3. GPT-32020The paper that made models learn from the prompt.Brown, Mann, Ryder, Subbiah, et al. (OpenAI)~9 min
  4. Scaling Laws2020The physics of making models bigger.Kaplan, McCandlish, Henighan, Brown, et al. (OpenAI)~8 min
  5. Chinchilla2022The compute-optimal scaling law that said models were too big.Hoffmann et al. (DeepMind)~7 min
  6. Textbooks Are All You Need2023A 1.3B model trained on curated, textbook-quality data rivaled models 10× larger.Gunasekar, Zhang, Aneja, et al. (Microsoft)~11 min
  7. Muon2024The optimizer that orthogonalizes weight-matrix updates instead of scaling them element-wise like Adam.Jordan, Jin, Boza, You, Cesista, Newhouse, Bernstein~11 min

One model, many experts

Sparsity: route each token to a few experts and buy parameter count without paying for it at inference. 2017 idea, 2026 frontier default.

0 / 5 read · 2017–2026
  1. Mixture of Experts2017 Start here How models got huge without getting slow.Shazeer, Mirhoseini, Maziarz, Davis, Le, Hinton, Dean~7 min
  2. Switch Transformer2022The mixture-of-experts model that scaled to 1.6 trillion parameters by routing each token to a single expert.Fedus, Zoph, Shazeer~11 min
  3. DeepSeek-V32024Key paperFrontier performance at a tenth of the cost.DeepSeek-AI~8 min
  4. DeepSeek-V42026The efficiency sequel to V3, aimed at the real cost of long context.DeepSeek-AI~9 min
  5. Qwen4-Coder2026An open, Apache-2.0 coding model that reportedly hits ~82% on SWE-Verified while running on a MacBook — the first open, Mac-runnable mo…Qwen Team~8 min

Teaching a base model to follow

A pre-trained model completes text; it does not take instructions. Nine papers on getting from one to the other, from RLHF to a loss with no reward model in it.

0 / 9 read · 2022–2024
  1. InstructGPT2022 Start here How RLHF taught models to follow instructions — the paper behind ChatGPT.Ouyang, Wu, Jiang, et al. (OpenAI)~9 min
  2. Constitutional AI2022Key paperAlignment from principles, not labels.Bai, Kadavath, Kundu, et al. (Anthropic)~8 min
  3. Self-Instruct2023How to bootstrap a large instruction-tuning dataset from a language model's OWN generations, starting from just 175 human-written seed…Wang et al.~10 min
  4. LoRA2022The paper that made fine-tuning affordable.Hu, Shen, Wallis, Allen-Zhu, Li, Wang, Wang, Chen~8 min
  5. QLoRA2023Fine-tune a 65B model on one GPU.Dettmers, Pagnoni, Holtzman, Zettlemoyer (UW)~8 min
  6. Direct Preference Optimization2023RLHF without the RL.Rafailov, Sharma, Mitchell, Ermon, Manning, Finn~7 min
  7. GRPO2024The RL algorithm behind DeepSeek-R1.Shao, Wang, Zhu, et al. (DeepSeek)~8 min
  8. STaR2022A model teaching itself to reason.Zelikman, Wu, Mu, Goodman~11 min
  9. Let's Verify Step by Step2023Reward every reasoning step, not just the final answer.Lightman, Kosaraju, Cobbe, et al.~11 min

Reasoning you can prompt, then train

It began as a prompting trick and ended as a training objective. Read in order, the progression from "let's think step by step" to a model trained to.

0 / 5 read · 2022–2025
  1. Chain-of-Thought Prompting2022 Start here Show a model a few worked examples with their reasoning and it starts solving problems it used to fail.Wei, Wang, Schuurmans, Bosma, Ichter, Xia, Chi, Le, Zhou~7 min
  2. Self-Consistency2022Sample many diverse reasoning paths, marginalize out the reasoning, majority-vote the answer.Wang, Wei, Schuurmans, et al. (Google)~7 min
  3. Tree of Thoughts2023When LLMs learn to search.Yao, Yu, Zhao, et al. (Princeton / DeepMind)~8 min
  4. OpenAI o12024The reasoning model trained with reinforcement learning to think before it answers — and the discovery of a second scaling axis.OpenAI~10 min
  5. DeepSeek-R12025How a model learned to reason from reward alone.DeepSeek-AI~9 min

The open-weights line

Five releases you can actually download, in the order each one answered the last. This is where "what does a frontier recipe look like" is documented in public.

0 / 5 read · 2023–2026
  1. Llama 22023 Start here The paper that opened up the LLM.Touvron, Martin, Stone, et al. (Meta)~9 min
  2. Llama 32024A herd of open dense transformers trained far past the Chinchilla point.Dubey, Jauhri, Pandey, et al. (Meta AI)~10 min
  3. Qwen32025An open model family that unifies a thinking mode (chain-of-thought) and a non-thinking mode (direct answer) in one model, with a contr…Qwen Team (Alibaba)~10 min
  4. Kimi K22025A trillion parameters built for agents.Moonshot AI~8 min
  5. Kimi K32026The sequel to Kimi K2, and a different bet: attack decode SPEED, not KV memory.Moonshot AI~9 min

What the labs report

Four papers that are reports, not methods: they state capability and withhold recipe. Read for the evaluation design and the omissions, not the architecture.

0 / 4 read · 2022–2026
  1. PaLM2022 Start here A 540-billion-parameter dense model trained with Pathways that made emergent abilities concrete.Chowdhery, Narang, Devlin, et al. (Google)~10 min
  2. GPT-4 Technical Report2023The report is famous for what it withholds, but its key methodological claim is predictable scaling: OpenAI predicted GPT-4's final los…OpenAI~10 min
  3. Gemini 1.52024A million-token context window with near-perfect needle-in-a-haystack recall.Gemini Team (Google DeepMind)~10 min
  4. GPT-5.6 System Card2026A modern system card is really a COMPUTE-ALLOCATION POLICY, not one model.OpenAI~8 min

Systems & applications

What gets built on top: retrieval, agents, cheap serving, image and speech generation — and what the industry numbers actually say.

From a word vector to RAG

The longest ladder here, and the most linear: nine papers that carry one idea — text as geometry — from a 2013 word embedding to a graph-structured retrieval system.

0 / 9 read · 2011–2024
  1. Word2Vec2013 Start here How words became math.Mikolov, Chen, Corrado, Dean~6 min
  2. GloVe2014Word vectors from GLOBAL co-occurrence statistics — the count-based cousin of word2vec.Pennington, Socher & Manning~10 min
  3. Sentence-BERT2019The model that turned BERT into fast, comparable sentence embeddings — the ancestor of every embedding model behind vector search and RAG.Reimers, Gurevych~11 min
  4. Product Quantization2011The compression that makes billion-scale vector search fit in memory.Jégou, Douze, Schmid~11 min
  5. HNSW2016The graph index behind fast approximate nearest-neighbor search — the default in nearly every vector database.Malkov & Yashunin~10 min
  6. Dense Passage Retrieval2020The dual-encoder method that made dense retrieval beat BM25 for open-domain QA.Karpukhin, Oğuz, Min, Lewis, Wu, Edunov, Chen, Yih~10 min
  7. ColBERT2020Keep one vector per token, not one per document, and score by late interaction: each query token takes its best match anywhere in the d…Khattab, Zaharia~11 min
  8. Retrieval-Augmented Generation2020Key paperThe paper behind every "chat with your docs" product.Lewis, Perez, Piktus, Petroni, Karpukhin, Goyal, et al.~8 min
  9. GraphRAG2024RAG for the global questions vector retrieval can't answer.Edge, Trinh, Cheng, et al. (Microsoft)~11 min

Agents that act

Give the model a tool, then a loop, then a memory, then a world. Ends on the 2026 paper measuring how badly the whole stack degrades over many turns.

0 / 7 read · 2023–2026
  1. ReAct2023 Start here The blueprint for AI agents.Yao, Zhao, Yu, Du, Shafran, Narasimhan, Cao~7 min
  2. Toolformer2023How a language model taught itself to use APIs.Schick, Dwivedi-Yu, Dessì, et al. (Meta)~8 min
  3. Reflexion2023Agents that learn from their own mistakes — reinforcement through words instead of weights.Shinn, Cassano, Gopinath, et al.~7 min
  4. Voyager2023The Minecraft agent that never stops learning.Wang, Xie, Jiang, et al. (NVIDIA)~8 min
  5. Generative Agents2023The Smallville paper: 25 LLM characters living in a simulated town.Park, O’Brien, Cai, et al. (Stanford)~8 min
  6. MemGPT2023An operating system for LLM memory.Packer, Wooders, Lin, et al. (Berkeley)~8 min
  7. LLMs Get Lost in Multi-Turn Conversation2026The ICLR 2026 Outstanding Paper that measured a failure everyone had felt: give a top model a fully-specified task in ONE prompt (conca…Laban, Hayashi, Zhou & Neville~8 min

Serving it cheaply

Five papers about the bill. Each one buys throughput back from a different part of the stack: the weights, the memory hierarchy, the KV cache, the decode loop.

0 / 5 read · 2014–2024
  1. Knowledge Distillation2014 Start here How to compress a big, accurate "teacher" into a small, fast "student" — by training on the teacher's soft probability distribution, no…Hinton, Vinyals & Dean~11 min
  2. FlashAttention2022Attention was slow for the wrong reason.Dao, Fu, Ermon, Rudra, Ré~7 min
  3. PagedAttention2023The idea that made LLM serving several times cheaper: store the KV cache in fixed-size blocks like an OS pages memory, instead of reser…Kwon, Li, Zhuang, Sheng, Zheng, et al.~11 min
  4. Speculative Decoding2023Make a big model generate 2–3× faster with zero change to its output.Leviathan, Kalman, Matias~12 min
  5. Medusa2024Speculative decoding without a separate draft model.Cai, Li, Geng, Peng, Lee, Chen, Dao~11 min

Generating pixels

Ten papers from the first latent-variable generator to video and real-time 3D. The diffusion half is strictly cumulative — each rung removes a cost the last one paid.

0 / 10 read · 2014–2024
  1. VAE2014 Start here The Variational Autoencoder — how to turn a compressor into a generator.Kingma & Welling~11 min
  2. GANs2014The counterfeiter and the detective.Goodfellow, Pouget-Abadie, Mirza, Xu, et al.~8 min
  3. VQ-VAE2017The autoencoder that made the latent DISCRETE — and quietly became the tokenizer behind image generation.van den Oord, Vinyals & Kavukcuoglu~11 min
  4. DDPM2020Key paperThe original denoising diffusion math — the paper that made diffusion work.Ho, Jain, Abbeel~9 min
  5. DDIM2021Denoising Diffusion Implicit Models — the trick that made diffusion sampling deterministic and 10–50× faster, using the SAME trained DD…Song, Meng & Ermon~11 min
  6. Classifier-Free Guidance2021The "guidance scale" slider behind every diffusion image tool.Ho & Salimans~10 min
  7. Latent Diffusion2022The paper behind Stable Diffusion.Rombach, Blattmann, Lorenz, Esser, Ommer~9 min
  8. Sora2024Text-to-video by treating video as a diffusion transformer over spacetime patches.Brooks, Peebles, et al. (OpenAI)~10 min
  9. NeRF2020The paper that launched the radiance-field era.Mildenhall, Srinivasan, Tancik, Barron, Ramamoorthi, Ng~11 min
  10. 3D Gaussian Splatting2023Real-time photorealistic radiance fields.Kerbl, Kopanas, Leimkühler, Drettakis~11 min

One model, many senses

Joining text to images, audio, video and finally to a robot arm. Starts with the contrastive objective the whole field standardised on.

0 / 8 read · 2020–2024
  1. CLIP2021 Start here How AI learned to connect images and words.Radford et al. (OpenAI)~7 min
  2. Flamingo2022The visual language model that bridges a frozen vision encoder and a frozen LLM with a small trained connector.Alayrac, Donahue, Luc, Miech, et al. (DeepMind)~11 min
  3. LLaVA2023Visual instruction tuning the simple way.Liu, Li, Wu, Lee~10 min
  4. wav2vec 2.02020Self-supervised speech: learn from raw, untranscribed audio, then fine-tune on as little as TEN MINUTES of labels.Baevski, Zhou, Mohamed & Auli~11 min
  5. Whisper2022Robust speech recognition from the open web.Radford, Kim, Xu, Brockman, McLeavey, Sutskever~8 min
  6. V-JEPA2024A self-supervised video world model that predicts in representation space, not pixels.Bardes, Garrido, Ponce, Assran, Ballas, et al. (Meta AI)~10 min
  7. RT-22023A Vision-Language-Action model that controls a robot by emitting actions as text tokens.Brohan, Brown, Carbajal, et al. (Google DeepMind)~10 min
  8. Genie2024A generative interactive environment that turns an image into a frame-by-frame playable world — trained on unlabeled video with no acti…Bruce, Dennis, Edwards, Parker-Holder, et al. (Google DeepMind)~10 min

Models that write code

Completion, then competition, then a benchmark made of real GitHub issues — the point at which "can it code?" got an answer you could audit.

0 / 3 read · 2021–2023
  1. Codex2021 Start here The GPT-on-code model behind GitHub Copilot — and the evaluation it standardized.Chen, Tworek, Jun, Yuan, et al. (OpenAI)~10 min
  2. AlphaCode2022Median-human competitive programming by scale plus selection.Li, Choi, Chung, Kushman, et al. (DeepMind)~10 min
  3. SWE-bench & SWE-agent2023Can a model do real software engineering?Jimenez et al. · Yang et al.~11 min

The industry numbers

Not research papers: a consultancy census, a VC essay and a job-posting analysis. Included because their figures get quoted as if they were, and read here for their method.

0 / 3 read · 2026–2026
  1. The 95% Number, Examined2026 Start here The most-quoted statistic in enterprise AI, read carefully: roughly 95% of pilots reportedly produced no MEASURABLE profit-and-loss imp…MIT NANDA~11 min
  2. Trading Margin for Moat2026Why a software company would deliberately hire expensive engineers to do customer work.a16z~11 min
  3. 1,000 FDE Jobs, Analysed2026Reading a thousand job postings beats reading a thousand opinions.Bloomberry~10 min

Paper breakdowns — frequently asked questions

What is a paper breakdown?

A written walk-through of one research paper: the problem it was published against, the mechanism it introduced, the numbers it reported, what its authors admitted it could not do, and a runnable Python exercise that re-derives its central result. It is not a summary of the abstract, and it is not a replacement for the paper — every breakdown links the original on arXiv.

Are the breakdowns free?

Yes — every breakdown is free to read, with no paywall and no sign-up. There is nothing to watch: these are written deep-dives with an interactive exercise, not videos.

Which paper should I read first?

"Attention Is All You Need" — rung 2 of the Transformer ladder, and the paper that introduced the architecture underneath GPT, Claude and Gemini. If you want the ground under it first, rung 1 of that ladder (Seq2Seq) is what the Transformer replaced.

Why are the papers in ladders instead of categories?

Because reading order changes what a paper teaches. The clearest case is Scaling Laws and Chinchilla: the second measures the same quantity as the first, in the same units, and reaches a different answer. Read the wrong way round, you learn the conclusion that was overturned. A category list cannot express that; an ordered ladder can, so every track here is a reading order and each rung is written assuming the one below it.

Do I need a research background to follow along?

No. Each breakdown is written for engineers and builders rather than academics — every mechanism is explained in plain terms before the formal version, and where a paper depends on an earlier one, that earlier paper is the rung below it on the same ladder.

Request · Open Channel

A paper missing?

If there's a paper you'd like broken down — or a rung you think sits in the wrong place on one of the ladders above — send it through. Genuine requests shape what gets covered next.