Vibe Engines
YouTube
Handbooks  /  Self-Improving Agents
Handbook~160 min readAgents13 runnable blocks3 interactive
Handbook

Self-improving
agents.

The dream is an agent that gets better the more you use it — one that runs in production, notices what worked, and folds it back in. It is real, and none of it is magic. Underneath is a loop with four moves: generate, verify, filter, refit. A model can usually already produce the right answer somewhere in its output; what you are missing is something that can recognise it, and somewhere to put what you learn — the prompt, a memory store, a tool, the scaffold, the verifier, or the weights. Choose the wrong one and you pay for it slowly. Get the verifier wrong and the loop teaches itself the wrong thing, confidently, every round.

The route

How this page runs

Five acts, 14 sections, ~2 hours end to end — most people should read three. The build log beside them is invented: a composite team, ticket counts and weeks not measured.

About the box that runs down the page. Set apart beside every section is a build log: one team, nine weeks, one agent called Mender pointed at one repo's ticket queue. It is a composite scenario, not a case study — its ticket counts, weeks and percentages are invented and only ever describe that imaginary team. Every measured figure on this page lives in the prose instead, named, sourced, and carrying the condition it was measured under. Each section below also says when you can skip it.

The shape of the whole pageOne loop, four moves, fourteen sections
  1. generate

    Draw many attempts instead of one. The answer is usually already in there.

    01020311

  2. verify

    Decide which attempts were good. Everything downstream inherits this judgement.

    040506

  3. filter

    Choose what is worth keeping, and where to put it.

    0708

  4. refit

    Fold it back in — then prove it helped, and know when to stop.

    09101213

Each section owns one arc of this ring. If you only want the arc you are stuck on, follow its numbers.

  1. Act I

    Asking more than once

    ≈22 min

    01–03 · Tell coverage from shippable accuracy, price five cheap attempts, forecast from your histogram

    1. 01The answer is already in there

      Tell coverage apart from the accuracy you can ship, and cost five attempts against one

      ≈7 mininteractive

      Skip if you already sample k>1 and report pass@k

    2. 02Measuring it honestly

      Compute unbiased pass@k from one sample pool, and state its n and its verifier

      ≈6 min1 runnable

      Skip if your pass@k already comes from one pool with n stated

    3. 03Why the curve is straight

      Forecast sampling returns from your per-problem histogram instead of the mean

      ≈9 min1 runnable

      Skip if you are not paying for repeated sampling

  2. Act II

    The picking problem

    ≈40 min

    04–06 · Pick a selector for your output shape, build the verifier behind it, learn how often it lies

    1. 04Selection is the bottleneck

      Choose a selector by output shape and cap N before the curve bends downward

      ≈11 min1 runnable

      Skip if a program, not a model, already picks your winner

    2. 05Verification

      Grade answers or steps, ensemble weak checkers, and measure your false-positive rate

      ≈16 min1 runnableinteractive

      Read this one even if you skip everything else

      Break the verifier

    3. 06Feedback that isn't a verifier

      Rank feedback sources by independence and gate weight updates on the best one

      ≈13 min1 runnable

      Skip if every task you run has a test

  3. Act III

    What you feed it

    ≈36 min

    07–09 · Write the five-predicate filter, route each lesson to a substrate, audit the reward first

    1. 07What you keep

      Write a selection policy that filters, bands, dedupes, caps, then cuts to budget

      ≈10 min1 runnable

      Skip if you never refit

    2. 08Where the improvement lands

      Route a lesson to one of six substrates by generality, permanence and blast radius

      ≈12 min1 runnable

      Skip if you have already decided it goes in the prompt

    3. 09Closing the loop with RL

      See why a correctness filter is a policy gradient, pick PPO, GRPO or DPO, audit the reward

      ≈14 min1 runnable

      Skip if you are not training weights

  4. Act IV

    Narrowing and what you report

    ≈43 min

    10–12 · Watch entropy before pass@1, buy reach with search not weights, report a horizon and a blind win rate

    1. 10The loop that collapses

      Instrument the loop so diversity collapse shows while rolling back still costs one batch

      ≈20 min1 runnableinteractive

      Skip if you are not training weights

      Run the collapse

    2. 11Search as self-improvement

      Choose a flat bank or a tree on latency and evaluator quality, then where the win lands

      ≈10 min1 runnable

      Skip if your budget is one attempt per task

    3. 12Knowing it worked

      Build a time-horizon fit and a blind win-rate set on your own tasks

      ≈13 min1 runnable

      Skip if you already run a horizon fit and a blind win-rate set

  5. Act V

    Standing it up

    ≈17 min

    13–14 · Sequence the rollout, name the conditions that stop the loop, open the page that owns your blocker

    1. 13Ship it

      Sequence the rollout, size a canary, and name the conditions that stop the loop

      ≈12 min2 runnable

      Start here if you have ten minutes

    2. 14Where to go next

      Open the one page that owns your blocker, and take the six answers you will be asked for

      ≈5 min

      Skip unless you are stuck

01

The answer is already in there

≈7 mininteractiveSkip if you already sample k>1 and report pass@k

Build logWeek 1 · TuesdayThe queue nobody wanted

We pointed Mender at the dullest queue we own: failing integration tests and bug tickets on one Python service. One attempt per ticket, it closed almost nothing, and we had written it off by lunch. Someone reran the same tickets overnight at twenty attempts each, graded only by the one test named on the ticket, and we read the log at nine to find a handful had gone green. Nothing had been learned between tries.

Which asks Is it incapable, or did we only ask once

You ask the model once, you get one answer, and you judge the model on it. That is a strange basis for a judgement, because a model is a sampler. Ask the same question again and something else comes out. Ask it 250 times and the question quietly changes from can this model do the task to does this model ever do the task.

Those two questions have very different answers. In Large Language Monkeys: Scaling Inference Compute with Repeated Sampling (Brown et al., 2024), the authors draw up to 10,000 independent samples per problem across five tasks and track coverage: a problem counts as covered the moment one of its samples is correct, no matter which one and no matter what the other 9,999 said. Coverage climbs, smoothly, over four orders of magnitude of sample budget.

The result that should reorganise how you think about agents is on SWE-bench Lite, where a problem is a real GitHub issue and one "sample" is an entire multi-turn trajectory through a codebase. DeepSeek-Coder-V2-Instruct, driving the open-source Moatless Tools agent framework, resolves 15.9% of issues on a single attempt. Give it 250 independent attempts and 56% of issues are resolved by at least one of them. The single-attempt state of the art at the time — CodeStory Aide, running a mixture of GPT-4o and Claude 3.5 Sonnet — was 43%. Carry one condition with those numbers: every attempt, the single one included, was drawn at temperature 1.6, chosen by a sweep over 1.0, 1.4, 1.6 and 1.8 on 50 problems. That is a setting picked for the many-attempts regime, not DeepSeek's best single-shot configuration.

15.9%one attempt · DeepSeek-Coder-V2-Instruct
56%250 attempts · same model, same framework
43%single-attempt SOTA · GPT-4o + Claude 3.5 Sonnet

Brown et al. 2024, SWE-bench Lite via Moatless Tools · attempts drawn independently at temperature 1.6, no feedback between them

Read that slowly, because it is the premise the rest of this page is built on: a weaker model, sampled more, found fixes a stronger model did not find on its first try. The capability was already in the weights. What was missing was attempts, and a temperature loose enough to make them differ.

It is also cheaper than it sounds. At the API prices current when the paper was written, a much smaller version of the same trade — five DeepSeek attempts, with an oracle picking the winning patch, against one attempt each from GPT-4o and Claude 3.5 Sonnet — looks like this:

Setup (Moatless Tools, SWE-bench Lite)Cost / attemptAttemptsIssues solvedTotal
DeepSeek-Coder-V2-Instruct$0.0072529.62%$10.80
GPT-4o$0.13124.00%$39.00
Claude 3.5 Sonnet$0.17126.70%$51.00

Brown et al. 2024, Table 1 · five cheap attempts solve more issues than one expensive attempt, at 3.6x less than GPT-4o and 4.7x less than Claude 3.5 Sonnet

And it is not a frontier-model effect. Gemma-2B on CodeContests goes from a 0.02% solve rate at one sample to 7.1% at 10,000 — more than a 300x increase. Pythia-160M goes from 0.27% to 57% on a 128-problem MATH subset — coverage under an oracle answer checker rather than a solve rate, for the reason the next paragraph gives. Small models contain far more correct answers than they can produce on demand.

Say that without any of the numbers attached, because it is the premise the whole page rests on: a model’s first answer is not the edge of what it knows. Everything that follows is about the distance between those two things, and who closes it.

Now the part that decides whether you build this well or badly. Coverage is not accuracy. Coverage counts a problem as solved if any of your samples was correct, which quietly assumes something you usually do not have: an oracle that can look at 10,000 candidate answers and point to a right one.

Where that oracle is real, coverage converts directly into performance. A Lean proof checker either accepts the proof or it does not. A test suite either goes green or it does not. That is exactly why SWE-bench Lite works as a demonstration — the harness grades every attempt with real tests. But those tests are withheld from the agent, so the 56% is what an oracle choosing among 250 patches scores, not what a shippable selector scores, and Brown et al. found 11.3% of Lite problems have flaky test suites on top of that. A checker that can actually decide is also why verifiable-reward domains keep producing the loudest results in this literature.

Where the oracle is not real, the gap is enormous, and Brown et al. measured it. On 128 MATH problems with Llama-3-8B-Instruct at 10,000 samples each, coverage rises from 79.8% at 100 samples to 95.3% at 10,000. Over that same range, the methods you could actually ship — majority voting and reward-model selection, using ArmoRM-Llama3-8B-v0.1, then the highest-scoring open-weight model for reasoning on RewardBench — move from 38.7% to 39.8%. Every selection method saturates before about 100 samples while coverage keeps climbing. At 10,000 samples the best available selector is leaving more than 55 points of already-generated, already-correct answers on the floor.

Stop on that for a moment. The samples got better and the answer you would have shipped did not. Buying more attempts only helps if something can tell you which attempt to keep — and that something is now the expensive part of the system, not the model.

Interactive · the coverage gap

Sampling more finds more answers. Picking the right one is a different problem.

Brown et al. sampled one model over and over and tracked coverage — the share of problems where at least one sample is correct. On SWE-bench Lite, DeepSeek-Coder-V2-Instruct went from 15.9% with a single sample to 56% with 250 (Large Language Monkeys, 2024). On MATH with Llama-3-8B-Instruct, coverage ran from 79.8% at 100 samples to 95.3% at 10,000 — while majority voting and reward-model selection plateaued around 100. Coverage is what the samples contain. Selection is what you can ship. Below, four rules read the same bank of candidates.

% of problems solved  ·  samples per problem (k), log scale

Drag a slider, or drag across the chart, to begin.

82%

How often the scorer ranks a correct candidate above a wrong one. 50% is a coin flip.

20%

Share of wrong samples landing on one shared wrong answer — a systematic misreading, not noise.

256

Samples drawn per problem. You can also drag or tap across the chart itself.

Coverage — any of the k samples is correctsolid · oracle: needs a perfect verifier you do not have —
Best-of-k by scorer — keep the highest-scored sampledash-dot · deployable · slider 1 —
Majority vote — keep the most common answerdashed · deployable · slider 2 —
One sample — the budget you started withdotted · baseline · flat by definition —
The gap — coverage minus the best deployable rulehatched band · compute you paid for and cannot collect —

Simulated, and here is exactly what was swapped. No model was run. The bank is 192 synthetic problems whose single-sample success rates are evenly spaced quantiles of a Kumaraswamy(α = 0.35, β = 1.6) distribution — the heavy-tailed family Schaeffer et al. fit to real per-problem success rates in How Do Large Language Monkeys Get Their Power (Laws)? (2025), where the weight of that left tail is what turns per-problem exponential decay into aggregate power-law coverage. Coverage is then exact: the bank average of 1 − (1 − p)k. The two deployable curves are models of a selection rule, not measurements of one — majority vote over an answer distribution with six recurring distractors plus one attractor, and best-of-k under a Gumbel-noise scorer whose pairwise accuracy is the slider (that noise makes best-of-k a softmax over the pool, so the curve is closed-form rather than sampled). Every number here is computed, not drawn at random, so your chart and mine match exactly. The shape is real; the scale is a choice. The measured figures on this page are the ones named in the prose, never the ones on these axes.

Brown et al. 2024 · Llama-3-8B-Instruct, the paper's 128-problem MATH subset, up to 10,000 samples per problem · the plotted endpoints are the numbers the paper reports; anything drawn between them is interpolation, not a measurement

The reason is mechanical rather than mysterious. On some of those problems the correct answer is produced by 1% of samples or fewer. A majority vote over 10,000 samples cannot surface a 1% answer — the more you sample, the more precisely the vote converges on the popular wrong one. Sampling finds needles; voting measures the haystack. Closing that gap is what learned verifiers are for, and why their failure modes matter so much here.

This gap is not new, and the first clean receipt for it is five years old. In the Codex paper (Chen et al., 2021), Codex-S solves 37.7% of HumanEval problems with one sample at temperature 0.8. Given 100 samples per problem, picking the candidate with the highest mean log-probability solves 44.5%; picking the candidate that passes the unit tests solves 77.5%. Same 100 samples, same model — 33 points between the best ranker and the oracle. AlphaCode is the same shape at production scale, with the selector built out in public: it draws an enormous pool of candidate programs, throws away every one that fails the example tests printed in the problem statement (roughly 99% of the pool), runs what survives on generated inputs, and groups the programs that produce identical outputs — so its ten allowed submissions go to ten different behaviours rather than ten spellings of one. It reached an average rank in the top 54.3% across simulated Codeforces competitions with more than 5,000 participants.

The key idea

Repeated sampling changes the question from what a model does to what a model can do, and coverage is the ceiling that question sets. A selector is what you actually ship. Every technique on the rest of this page is an attempt to close the distance between the two — by building a verifier good enough to convert coverage into accuracy, or by folding what you found back into the weights so it arrives in the first sample next time.

02

Measuring it honestly

≈6 min1 runnable blockSkip if your pass@k already comes from one pool with n stated

Build logWeek 1 · FridayThe number in the update

We put a number in the weekly update: the share of the queue Mender fixes. It came out of the overnight pool, where a ticket counted if any of twenty attempts went green, and we had labelled it as what one attempt does. Someone asked what n was, and the number did not have an n. Nobody had lied; we had reported a different quantity from the one we named.

Which asks What number do we actually put in the update

You cannot steer a loop on a number you are measuring wrong, and coverage is easy to measure wrong in both directions.

Start with what you actually want to know. pass@k is the probability that if you drew k samples for a problem, at least one would be correct. It is a property of the model-and-problem pair, not of one particular run of your harness.

The obvious way to estimate it: draw k samples, check whether any passed, write down 1 or 0, average over problems. That is unbiased and nearly useless. For a single problem it is one Bernoulli draw, so its variance is about as large as the thing you are measuring — the high variance that made Chen et al. reach for something better. It is also wasteful — pass@1, pass@10 and pass@100 would each need their own experiment.

The version that actually ships is worse. Generate a big pool of n samples, notice that one of them worked, and report that as pass@k for whatever k you had in mind. That is pass@n wearing a smaller label. With 200 samples of which 2 are correct, "did any of them work" reports 1.0 while the true pass@10 is 0.098 — a factor of ten, in the direction that flatters you.

The Codex paper (Chen et al., 2021), which introduced HumanEval, fixed both problems with one formula. Generate n ≥ k samples once, count the c that pass, then ask a combinatorial question of the pool you already have:

pass@k = 1 − C(n−c, k) / C(n, k)

C(n−c, k) / C(n, k) is the probability that a k-subset drawn from your n samples without replacement contains none of the c correct ones. One minus that is the probability it contains at least one. The Codex paper used n = 200 and k ≤ 100; Brown et al. reuse this same estimator to keep variance down at 10,000 samples.

The economy is the point. One pool of n samples gives an unbiased estimate at every k ≤ n, so a single expensive run draws the whole curve instead of one point on it.

There is a near-miss worth naming, because it looks right and it is not. Take p̂ = c/n and write pass@k ≈ 1 − (1 − p̂)k. That treats your k draws as being made with replacement from the pool, and Chen et al. show in their appendix that it is a consistent underestimate whose gap does not fully close even when n > 5k. It is wrong in the modest direction — it makes sampling look less useful than it is — which is exactly why nobody catches it.

Figure · section 02One pool of 200 samples, read two ways

n = 200 drawnc = 2 correct

1.00“Did any of them work?” The pool relabelled as pass@10. It is pass@200 wearing a smaller name.
0.098The unbiased estimator. The probability that a 10-subset of this pool contains one of the two.

Both readings use the same 200 samples. The gap is a factor of ten, and it flatters you. Figures from the prose above.

Three edge cases carry all the meaning, and the code below is where you write them:

  • k > n. Undefined; the function should refuse rather than guess. You cannot estimate pass@1000 from 200 samples no matter how straight the curve looks.
  • c = 0. The estimator returns exactly 0 at every k. Read it as "not yet", never as "impossible" — 200 consecutive failures only put a rough 95% upper bound of p < 3/200 on the true rate, and section 03 is about how much that tail matters.
  • c = n. Returns exactly 1, as does any case where n − c < k: if fewer than k of your samples are wrong, no k-subset can be all wrong. Chen et al.'s implementation short-circuits that case before the product runs — not because anything would divide by zero, but because the product would otherwise grind through factors like 1 − k/j that go negative.

Then the caveat that outranks the arithmetic: c is whatever your checker counted. Brown et al. ran CodeContests' own test suites against the reference Python3 solutions the dataset ships, and found that of the 122 test-set problems with such solutions, 35 have correct solutions that fail their own tests. Those are false negatives — coverage you measured as absent and in fact had. The mirror failure, a wrong program that passes a thin test suite, inflates the same number instead. What you are reporting is never pass@k; it is passes-my-verifier@k, which is why how you build the evaluator and where you run it are load-bearing decisions rather than plumbing.

RUN IT YOURSELF

Write the unbiased pass@k estimator

This is the formula from the Codex paper, in about ten lines. Nothing is sampled and no model runs — you hand it a hypothetical pool (n samples, c of them correct) and it answers a combinatorial question about that pool. Run it as written, then change n and c to your own numbers. Watch what a 1% success rate looks like at k = 100, watch the edge cases return exactly 0 and exactly 1, and watch the two wrong answers — the pool reused as a smaller k, and the with-replacement plug-in — miss in opposite directions.

HOW TO READ THE CODE — 5 IDEAS
  1. pass_at_k never asks which samples passed, only n and c — it is counting k-subsets of a pool you already drew, so nothing here generates or re-runs anything.
  2. Read the two guards before the formula: k > n raises instead of guessing, and n - c < k returns 1.0 because there are not enough wrong samples left to fill an all-wrong k-subset.
  3. stable computes the same number as a running product over range(n - c + 1, n + 1), so it never materialises the huge binomials comb does; the two printed columns should agree digit for digit.
  4. The two ways to get it wrong block is the payload — the bare 1.0 lands far above the unbiased value and 1 - (1 - c/n) ** 10 lands just below it, so the two common mistakes miss in opposite directions.
  5. Change n, c = 200, 2 to your own pool and check the k=1 line first: pass@1 must come out as exactly c/n, and if it does not, the pool you typed is not the pool you meant.
CPython · WebAssembly
The key idea

One sample pool, every k. Report the unbiased estimator, state the n it came from, and name the verifier that counted c — because a pass@k without its n and its checker is a number with no conditions attached, and a number with no conditions attached is not evidence.

03

Why the curve is straight

≈9 min1 runnable blockSkip if you are not paying for repeated sampling

Build logWeek 2 · WednesdayTwo queues, one average

We budgeted next month's attempts off a straight line on a log axis. It held on the test-repair queue and bent on the other one, so we drew a histogram of per-ticket success for each. Test-repair had a thin spread of hard-but-possible tickets and forty-odd that nothing moved at all; the other had a bulge a second attempt finished and almost nothing between it and zero. The line came from that thin spread, and the bend was running out of it.

Which asks Will more attempts keep paying on our backlog

Plot coverage against sample budget with a logarithmic x-axis and you usually get something suspiciously tidy: a line. Brown et al. observe coverage growing "nearly log-linearly" across several orders of magnitude for the Llama-3 and Gemma families — with exceptions they flag, such as Llama-3-8B-Instruct on MiniF2F-MATH — and fit it with an exponentiated power law, modelling log(coverage) ≈ a·k−b, then exponentiating to predict coverage directly.

The tempting reading is that the slope is a fact about the model, the way parameter count is. It is not. It is a fact about the distribution of difficulty in your problem set, and the derivation is four lines long.

Start with one problem. Your agent solves it with probability p on any single attempt, and attempts are independent, so the chance that all k fail is (1 − p)k:

one problem: passi@k = 1 − (1 − pi)k

Exponential in k. On a log-x plot this is a sharp S — nothing, nothing, then everything, almost all of it inside two decades around k ≈ 1/p. A single problem's curve is not straight anywhere.

Now the benchmark. Coverage is that quantity averaged over problems, which means averaging over whatever spread of per-problem success rates your problem set happens to contain:

coverage(k) = Ep[ 1 − (1 − p)k ] = 1 − ∫ (1 − p)k f(p) dp

Every problem is being solved exponentially fast. The benchmark is a weighted sum of those exponentials, and the weights are f — the density of single-attempt success rates across your problems. A weighted sum of exponentials is precisely where power laws come from.

Figure · section 03Two backlogs with the same average difficulty

  • A long thin tail of hard-but-possible problems. Every extra order of magnitude of samples keeps buying coverage. This is the straight line on a log axis.
  • A bulge of easy problems and a wall behind it. Two attempts finish nearly everything reachable, then the curve flattens. The budget stops paying.

samples per problem (log)coverage →

How Do Large Language Monkeys Get Their Power (Laws)? (Schaeffer et al., ICML 2025) makes that exact. They prove it in both directions. If the density behaves like a power law near zero — f(p) = C·pb−1 for small p — then for large k, −log(coverage) ~ C·Γ(b)·k−b. And conversely: if −log(coverage) ~ A·k−b, the density must satisfy f(p) ~ (A/Γ(b))·pb−1 as p → 0. In English: the slope of your inference-scaling curve is the exponent of the left tail of your difficulty distribution. Nothing in that derivation is about the model.

Their appendix works the dependence out for named distributions, which makes it concrete:

Distribution of per-problem success rateWhat the coverage curve does
Uniform(α, β), α > 0 — nothing harder than about 1/α attempts−log(coverage) decays exponentially; the curve bends over and saturates
Uniform(0, β) — flat, with mass right down to zero−log(coverage) ~ 1/(βk): a power law with exponent 1
Beta(α, β) — density pα−1 near zero−log(coverage) ∝ k−α: the tail exponent is the scaling exponent

Schaeffer et al. 2025, Appendix E · the real data fit a 3-parameter scaled Kumaraswamy distribution, because most benchmarks top out well below p = 1

The best evidence for the account is the case it explains away. Not every setting produces a power law. In the Best-of-N jailbreaking results of Hughes et al. (2024), Llama-3-8B-IT broke the pattern every other model followed: its −log(attack success rate) fell faster than any power law, which is to say its success rate climbed faster than any power law allows. Schaeffer et al. account for it from the distribution alone — within the sampling budget, every prompt eventually succeeded, so there was no heavy left tail, so there was no aggregate power law. The straight line requires near-impossible problems to exist. Remove them and it bends.

You can watch the whole argument happen in arithmetic. Nothing below runs a model or samples anything: it averages 1 − (1 − p)k over two hand-built difficulty distributions of 128 problems each — one where every problem is equally hard, one with a power-law left tail — whose mean single-attempt success rate is identical to four decimals. The shape is a property of those two distributions and of nothing else.

RUN IT YOURSELF

Same average difficulty, opposite curve

Nothing here runs a model or samples anything, and no network is loaded. It is the arithmetic of 1 − (1 − p)^k averaged over two hand-written difficulty distributions of 128 problems each — the subset size Brown et al. used on MATH — chosen so their mean single-attempt success rate is identical to four decimal places. One gives every problem the same p; the other has a power-law left tail with exponent b = 0.5. The shape of the two curves is real; the specific numbers are a property of these two lists, not a measurement of any system. Change the exponent in the `tail` line and watch the slope follow it.

HOW TO READ THE CODE — 5 IDEAS
  1. flat and tail are the entire experiment: both average 0.1, so every difference further down comes from how that average is spread across the 128 problems, never from a model.
  2. miss is the only piece of modelling in the file — the mean of (1 − p)k — which is why every printed number is arithmetic over those two lists and reproduces exactly on any machine.
  3. Read slope as "what does one more decade of samples buy": it compares k against k // 10, and a constant value is precisely what a straight line on a log-x plot means.
  4. When slope returns None and the cell prints off, nothing broke — miss underflowed to 0.0 because flat has no hard problems left, so saturation shows up as a float limit.
  5. The k = 1 row is the literal 0.1, true only while both lists average that — change the exponent in tail and trust the computed mean pass@1 header instead.
CPython · WebAssembly

Three consequences you can act on:

  • Measure the tail, not the mean. Two workloads with identical pass@1 can have opposite returns on sampling, which the code above demonstrates with two 128-problem sets whose average difficulty is the same to four decimals. Before budgeting for repeated sampling, look at the histogram of per-problem success rates, not the average.
  • You can forecast the curve without paying for it. Because the exponent lives in f, you can fit f from cheap single-attempt measurements and simulate the rest. Schaeffer et al. report roughly an order of magnitude lower relative error on the exponent than log-log least squares — measured by backtesting both estimators on synthetic data with a known ground-truth exponent — which they put at about 2–4 orders of magnitude less inference compute for the same precision.
  • When the line bends, ask which tail went missing. Saturation means you have run out of problems that are hard but possible. Either the stragglers sit at p = 0 — a capability the model genuinely lacks, or a verifier that is rejecting correct work — or you have solved everything you had. Those two want opposite responses, and only the histogram tells you which one you are looking at.

Sampling k times is the bluntest way to spend inference compute, and the curve above is its price list. For the rest of the menu — longer reasoning, search, verification passes, and how to divide a fixed budget between them — see test-time compute. The remainder of this page takes the other road: keep what the sampling found, and move it into the model so the first attempt is the good one.

04

Selection is the bottleneck

≈11 min1 runnable blockSkip if a program, not a model, already picks your winner

Build logWeek 3 · MondayFour hundred diffs a night

By now Mender could find a working patch for most of the queue, given enough attempts. It could not tell us which of the twenty it was. Several of them made the ticket's one named test pass, and nobody was reading four hundred diffs a night, so the shortest diff became the PR — and then became tomorrow's example. Finding was free; picking was the whole job.

Which asks Who picks among twenty patches when nobody reads any

Everything so far assumed a step you do not have: the ability to say that one was good. In a notebook you supply that yourself. In a loop running on production traffic nobody is reading the traces — the agent grades its own homework, at volume, and whatever it grades highly becomes tomorrow's memory or tomorrow's gradient. The oracle is not deployable. Something worse has to stand in for it.

Generating candidates is the cheap half, and it is startlingly cheap. Large Language Monkeys (Brown et al., 2024) measured this directly with a metric they call coverage: the fraction of problems solved by any sample in the batch. Coverage keeps climbing with the sample budget across four orders of magnitude, often log-linearly. On SWE-bench Lite, DeepSeek-Coder-V2-Instruct resolves 15.9% of issues with one sample and 56% with 250 samples — past the 43% single-sample state of the art at the time. That 56% is the model's reach. Whether any of it reaches production depends entirely on the picker, and the same paper reports what pickers actually do: in domains without an automatic verifier, majority voting and reward-model ranking plateau beyond several hundred samples and fail to keep scaling with the sample budget.

So the binding constraint on a self-improvement loop is not can the agent do it. It is can you recognise that it did. Test-time compute covers how to spend the sampling budget; this section is about the pile you are left holding afterwards. Three selectors work in practice. Each one breaks somewhere specific, and the place it breaks is exactly where your loop learns the wrong thing — quietly, and on repeat.

Figure · section 04Three ways to pick a winner from the same candidates
Majority vote

Reads: the final answer only

Picks: the most common one

Fails when the wrong answers agree. A 1% correct answer can never win a vote.

Reward-model ranking

Reads: a learned score per candidate

Picks: the highest score

Fails by over-optimisation. Its errors have a shape, and searching harder finds that shape.

Behavioural clustering

Reads: what each program does on probe inputs

Picks: one per distinct behaviour

Fails when the probes are blind. Two behaviours no input separates look like one.

Same pool of candidates, three different questions asked of it. Every failure column is the same failure: the selector reads something narrower than correctness.

Majority vote: cheap, strong, and pointed at the wrong target

Sample many reasoning chains, throw away the reasoning, keep the final answers, take the mode. That is self-consistency (Wang et al., 2022). It costs N generations and a counter; the table below is what that buys.

vote(x) = argmaxa  Σi=1..N  1[ answer(yi) = a ]
yi ~ model(x);   N = 40

Wang et al. sampled 40 reasoning paths per question and marginalised the paths out. Read the formula again: it converges on the model's most frequent answer. Nothing in it points at the truth.

It works, and the size of the effect is why everyone reaches for it first. On GSM8K, with 40 sampled paths against greedy chain-of-thought:

GSM8K accuracy, greedy chain-of-thought against self-consistency over 40 sampled reasoning paths (Wang et al., 2022, Table 2).
ModelGreedy CoTSelf-consistency (40 paths)Gain
PaLM-540B56.5%74.4%+17.9
GPT-3 code-davinci-00260.1%78.0%+17.9
LaMDA-137B17.1%27.7%+10.6
UL2-20B4.1%7.3%+3.2

Across the paper's other benchmarks the abstract's headline gains are SVAMP +11.0, AQuA +12.2, StrategyQA +6.4 and ARC-challenge +3.9 — quoted there without a model attached, so read them as the paper's best case rather than as PaLM-540B's row extended sideways. Read the table down the model column, though, and the shape of the thing shows: the weaker the model, the smaller the gain. A selector amplifies a skew that is already there. It cannot create one.

Two places it breaks.

It needs an extractable final answer. This is not a limitation someone found later — it is stated in the paper. Self-consistency applies only where the final answer comes from a fixed answer set, and the authors say plainly that extending it to open-text generation would first require defining a good metric of consistency between outputs. So: no code, no prose, no multi-step tool trajectory, no pull request. Everything an agent actually produces is out of scope for the cheapest selector you have. That is not a footnote. That is the reason the rest of this section exists.

It bends downward. In Are More LLM Calls All You Need? (Chen, Zaharia and Zou, 2024), the accuracy of Vote and Filter-Vote can first increase and then decrease as the number of model calls grows. With GPT-3.5-turbo-0125 on the MMLU college mathematics and business ethics subsets, accuracy is maximised at one particular number of calls and gets worse in both directions from there; on college chemistry it stays monotone. Their explanation is the one that matters for agent work: a task is a mixture of easy and hard queries, more calls push easy queries toward the right answer and hard queries toward the wrong one, and the sum of a rising curve and a falling curve is a hump.

The mechanism underneath is worth stating on its own, because it is the thing people get wrong. Majority vote converges on the model's mode. When the mode is correct, more votes sharpen it toward 1. When the mode is wrong, more votes sharpen it toward 0. There is no third case. And wrong answers share a mode more often than you would hope, because a model's errors are not random draws — a misread constraint, an off-by-one in the units, a plausible-but-false lemma. Every sample makes the same slip, because every sample came from the same weights looking at the same prompt. Diversity of sampled paths is not diversity of failure modes. When wrong answers share an attractor, the vote is a machine for becoming confident about it.

The block below computes that hump exactly. Its two difficulty profiles — an easy question the model gets right 70% of the time, a hard one where 65% of samples land on the same wrong answer — are numbers chosen by hand, not measurements. Nothing here was run against a model. The shape of the curve is what to read; the scale is invented.

RUN IT YOURSELF

Watch majority vote bend downward

Majority vote converges on whatever the model says most often — not on what is true. This block computes that exactly. There is no sampling and no randomness in it: it sums the binomial probability that the correct answer wins a vote of N. Two profiles stand in for a benchmark's real mixture — an easy question the model gets right 70% of the time, and a hard one where 65% of samples land on one shared wrong answer. This is the selector, not a model: the two difficulty profiles are numbers chosen by hand, so the shape of the curve is real and the scale is invented — nothing here was measured off a GPU. Watch the easy column climb toward 1, the hard column fall toward 0, and the mixed benchmark — the only column you would ever actually measure — go up, turn over, and come back down.

HOW TO READ THE CODE — 5 IDEAS
  1. vote_acc sums binomial terms from n // 2 + 1 upward, and that range is the definition of winning a vote: strictly more than half the samples.
  2. EASY and HARD are the same quantity — P(one sample is correct) — and the only thing separating them is which side of 0.5 they sit on, which is what decides whether more votes help or hurt.
  3. Read the hard column downward: it starts at 0.35 and ends near zero, so the vote is getting more confident as it gets more wrong.
  4. mixed is just the two columns weighted by SHARE_HARD, which is the whole trick — a hump is a rising curve plus a falling one, nothing more.
  5. Once p is fixed, nothing in vote_acc ever refers to the true answer again: the selector sees agreement, never truth.
CPython · WebAssembly

Reward-model ranking: a score you can over-optimise

The fix for "no extractable answer" is to learn the judgment instead. Train a verifier on solutions labelled correct and incorrect, sample N candidates at inference, score them all, keep the top one. Best-of-N.

The original result is still the clearest. In Training Verifiers to Solve Math Word Problems (Cobbe et al., 2021 — the paper that introduced GSM8K's 8.5K problems), 100 completions are sampled and ranked at test time, and the paper's own summary is that “on the full dataset, 6B verification slightly outperforms a finetuned 175B model, thereby offering a boost approximately equivalent to a 30x model size increase.” Ranking is not a tiebreak. It is a substitute for scale.

Now the break, which is in the same paper and usually gets skipped. For the 6B verifier, performance improves up to 400 completions and past that point starts to decrease. The authors' reading, in their words: “the benefits of search are eventually outweighed by the risk of finding adversarial solutions that fool the verifier.” With enough candidates you are no longer searching for a correct solution. You are searching for the specific input that makes your verifier wrong.

The turn is not a quirk of that one verifier. Scaling Laws for Reward Model Overoptimization (Gao, Schulman and Hilton, 2023) fits the shape:

Figure · section 04The proxy keeps rising after the thing it stands for turns

  • The reward model’s own score. Monotone. It never tells you to stop, because it is the thing being optimised.
  • What you actually wanted. Rises, turns, falls. The turn is the only measurement that matters and the proxy cannot see it.
  • The turn. Everything to the right of this line is search finding the shape of the verifier’s error, not better answers.

optimisation pressure (n in best-of-n, or KL from the base model)

KLBoN(n) = log n − (n−1)/n
RBoN(d) = d(αBoN − βBoN d),   d = √KL

Two separate results, not one product: the KL cost of best-of-n on the first line, the fitted reward curve on the second. Best-of-N is optimisation pressure with a dial on it. Each increase in N costs a fixed, computable amount of KL divergence from the base policy — and when you optimise against a proxy reward model, the paper observes the gold reward first increasing and later decreasing, while the proxy's own score keeps climbing the whole time.

Which gives the sentence to keep. A reward model does not have random errors; it has a shape. Best-of-N walks straight along that shape, because it is literally selecting the sample the model scored highest. So the traces your loop keeps are, by construction, the most over-optimised traces you generated — the ones sitting furthest inside the verifier's blind spot. Then you fine-tune on them. This is how reward hacking enters a self-improvement loop: not through the training objective, through the selector.

Behavioural clustering: vote on what the program does

For code there is no final token to compare, so AlphaCode (Li et al., 2022) changed what "the same answer" means. Two programs are the same if they behave the same.

Sample
1,000,000
programs per problem, half Python and half C++, for diversity
Filter
~99% gone
removed by the example tests printed in the problem statement — and tens of thousands still survive on many problems
Cluster
by behaviour
a separate model generates new test inputs; programs producing identical outputs on them are grouped
Submit
10
one program from each cluster, largest cluster first, wrapping round if there are fewer than ten

All but ten of a million programs are discarded by the selector, not the generator. The submission limit is the contest's; the funnel above it is the entire engineering problem.

Why largest-cluster-first works is the whole idea in one line, and the paper says it plainly: “there are many ways solutions can be incorrect while correct solutions tend to behave the same and therefore are grouped into larger clusters.” Majority vote, moved out of answer space and into behaviour space, where an agent's output actually lives.

It carried the system. AlphaCode's 41B model solved 34.2% of held-out CodeContests validation problems at ten submissions drawn from up to a million samples per problem, and in simulated evaluations on recent Codeforces competitions with more than 5,000 participants it averaged a ranking in the top 54.3%.

Where it breaks is the probes. The clustering is only as sharp as the generated test inputs: if no probe input distinguishes a correct program from a subtly wrong one, the two land in the same cluster and the vote counts them together. Behavioural equivalence is equivalence on the inputs you tried. For an agent, this is the directly useful version of the idea — run N trajectories, hash the end state each one leaves behind (files written, rows changed, calls made) and cluster on that. You never need the right answer. You need an equivalence relation and something to execute against.

The three selectors, side by side: what each one needs, what it reads as signal, and where it fails.
SelectorMajority voteReward-model rankingBehavioural clustering
Needsan extractable final answertraces labelled correct/incorrecta sandbox and probe inputs
Signalagreement between samplesa learned scoreidentical observable behaviour
Breaks whenthe wrong answers share an attractoryou push N past the model's blind spotsno probe input separates right from subtly wrong
CostN generationsN generations + training and serving an RMN generations + N×M executions
Works onshort, comparable answersanything you can scorecode, tool calls, anything with side effects
The key idea

A selector is a second model of correctness, and its mistakes are systematic where the generator's are diverse. The generator is wrong in a thousand directions; the selector is wrong in one. Feed the selector's output into training and you do not sample its error once — you specialise on it.

05

Verification

≈16 min1 runnable blockinteractiveRead this one even if you skip everything else

Build logWeek 3 · ThursdayThe suite becomes the judge

So we built the picker out of what we already had — not the ticket's one named test, the whole suite, in a clean container, against all twenty candidates before any of them became a PR. It was the first thing that could say no with nobody in the room. It was also the first thing with nothing checking it, and the suite is thinnest in the modules these tickets come from. We started a list of the patches it waved through and we later reverted.

Which asks Who checks the checker before we train on it

Every selector above is a verifier in disguise. Majority vote verifies by agreement, ranking verifies by a learned score, clustering verifies by behaviour. Once you see that, the question sharpens into something you can design against: what are you checking, and at which point in the trace?

Figure · section 05A verifier has two ways to be wrong, and a loop only hunts one of them
Truly correct
Truly wrong
Verifier says pass
True positiveWhat you wanted. Goes in the batch.
False positiveA wrong answer with your approval on it. This is the cell a loop searches for, because it is the cheapest way to satisfy your rule.
Verifier says fail
False negativeA right answer thrown away. Silent: it never appears in anything you look at, and it narrows what the next round can reach.
True negativeWorking as intended.

Accuracy averages these four cells into one number and hides the asymmetry. A loop does not sample the cells evenly — it is a search for the top-right one.

Outcome supervision, and the hole in it

The cheap answer is: check the last line. An outcome reward model is trained on one bit per solution — did it end correctly? Wherever you have an answer key, a test suite or a passing build, that bit is free.

The hole is structural. A right answer reached through a broken step passes. Let's Verify Step by Step (Lightman et al., 2023) builds on a known result — models trained with outcome supervision regularly use incorrect reasoning to reach the correct final answer (Zelikman et al., 2022; Creswell et al., 2022) — and names the consequence itself: automatic grading is not perfectly reliable, so false positives that reach the correct answer with incorrect reasoning are misgraded in the reward model's own training data.

Figure · section 05What each kind of reward model is allowed to look at
The solution step 1step 2step 3 wrongstep 4answer right
Outcome supervision reads this only
Process supervision reads catches it here

The outcome model marks this solution correct, because it is. A broken step that happens to land on the right answer is not an edge case — it is what a model rewarded on outcomes learns to produce.

Here is what that looks like in an agent trace. The agent is asked for the 90th-percentile latency. It sorts the samples and indexes int(0.9 * n) where it should use int(0.9 * (n - 1)). On the 101-sample array in the test, both expressions give index 90. The number is right. The code is wrong. The outcome check sees a pass and the trace is marked good — and on a 1,000-sample array the same code reads index 900 where it should read 899. For a one-shot answer that is a cosmetic problem. For a weight update it is a lesson: you have just told the model that this indexing is what success looks like.

A checker that reads only the last line cannot see any of that. It is not being careless. It was never shown the part where the mistake happened, so the only honest thing it can report is that the ending looked right.

Process supervision, and what it costs

The alternative is to label every step. Lightman et al. compare the two under a clean condition: solutions from a generator finetuned from the base GPT-4 model, best-of-N over 1,860 samples per problem, scored on 500 MATH test problems drawn uniformly at random.

majority vote 69.6%  →  outcome reward model 72.4%  →  process reward model 78.2%

Best-of-1,860 on 500 MATH problems, same generator throughout. The 5.8-point gap between the two reward models is what checking the middle of the trace is worth.

Step-level supervision buys two things. It localises the error, so the verifier's credit assignment matches the thing you would actually fix — which in a self-improvement loop means you can keep the good prefix of a bad trajectory instead of discarding the whole run. And it is harder to fool: a candidate now has to look right at every position rather than only at the last one, which raises the cost of exactly the adversarial solutions that ate Cobbe's 6B-verifier gains past 400 samples.

What it costs is people. The same paper released PRM800K — 800,000 step-level human labels across 75,000 solutions to 12,000 problems. That is about eleven human judgments per solution, on solutions a human has to read and follow line by line. Nobody is annotating production agent traces at that ratio.

Which is the real objection to labelling every step, and it is not a technical one. The method works. You have to decide whether the thing you are checking is worth that much of somebody’s attention, every week, forever — and for most teams the answer is no, which is why the next approach exists.

Buying step labels without buying labellers

Math-Shepherd (Wang et al., ACL 2024) removes the humans. The trick is to define a step's quality as its potential to reach the correct final answer: take the solution prefix up to that step, decode several continuations from it with a finetuned model, and score the step by how many of those continuations land on the gold answer. Steps that lead somewhere good get good labels. No annotator ever sees the problem.

It works well enough to train on. With step-by-step reinforcement learning against Math-Shepherd, Mistral-7B goes from 77.9% to 84.1% on GSM8K and from 28.6% to 33.0% on MATH; adding Math-Shepherd as a verifier on top pushes the two to 89.1% and 43.5%.

Note the catch, which follows from the definition rather than from any experiment: the label is the outcome, spread across positions. A step that happens to lead to the right answer scores well even when its reasoning is wrong — including the percentile bug above, whose continuations all reach the right number on the test array. Automatic process labels are far better than nothing and structurally closer to outcome supervision than the name suggests. They buy you localisation. They do not buy you a new source of truth.

Three ways to label a trace, and the failure each one is blind to.
VerifierOutcome RMProcess RM (human)Process RM (automatic)
Labelsthe final answerevery intermediate stepevery step, by completion rollouts
Label sourcean answer key you already havehuman annotatorsthe generator itself
Cost≈ free where a key existsPRM800K: 800K labels / 75K solutionscompute: k continuations per step
Fails whena right answer hides a broken stepthe annotator is wrong or the rubric driftsa broken step still reaches the right answer

Ensembles: several weak verifiers, weighted

In practice you never have one good verifier. You have four bad ones — a reward model, an LLM judge, a test run, a self-consistency score — that disagree, emit scores on incomparable scales, and come with no labelled data to weight them by.

Weaver (Saad-Falcon et al., 2025) attacks exactly that. It uses weak supervision to estimate each verifier's accuracy from the verifiers' pattern of agreement rather than from labels, normalises their output scales using dataset statistics, drops the ones carrying no signal, and combines the rest into a single weighted score. Weighted beats unweighted clearly — the verifiers differ too much for a plain average to be right. With Llama 3.3 70B Instruct as the generator and an ensemble of 70B-or-smaller judge and reward models, selection reaches 87.7% average accuracy across their reasoning and maths tasks: a jump comparable to the one between GPT-4o and o3-mini (69.0% to 86.7%) that otherwise took extensive finetuning and post-training.

The practical reading for a loop: you do not need a great verifier before you start. You need several cheap ones whose errors are not identical, and an honest estimate of how much to trust each.

A judge nobody judges

Here is the failure that will actually bite you, and it stops being subtle the moment you look for it. You added a verifier to catch the model's mistakes. Nothing is catching the verifier's.

The judge is much worse at this than the word suggests. JudgeBench (Tan et al., ICLR 2025) built 350 response pairs across knowledge, reasoning, maths and coding — responses generated by GPT-4o, labelled by objective correctness rather than by human preference, so a judge cannot win by picking the nicer-sounding answer. GPT-4o judging with a vanilla prompt scores 50.86% overall: 44.16% on knowledge, 47.96% on reasoning, 66.07% on maths, 61.90% on coding. On two of the four domains it is below a coin flip. An Arena-Hard-style prompt lifts the overall to 56.57%. The same model that wrote the responses cannot reliably tell which of two is correct.

And a judge that is genuinely good still invents things. OpenAI's critic models (McAleese et al., 2024) are trained with RLHF to write natural-language critiques of model-written code, and they are good: their critiques are preferred over human critiques on 63% of code containing naturally occurring model errors, and they catch more bugs than human contractors paid for code review. The paper's own caveat is the important half. The critics also raise objections against code that is fine, and the paper warns those invented bugs can push a reviewer into an error they would not otherwise have made; pairing a human with the critic hallucinates less than the critic working alone. A verifier that fabricates flaws is not neutral noise. It rejects correct traces at a steady rate, and its objections read rigorous, which is precisely why they survive review.

And no amount of compute fixes it. Stroebl, Kapoor and Narayanan (2024) give the ceiling: sampling again does nothing to the chance that a wrong answer slips past an imperfect verifier, so that chance caps how accurate a resampling loop can get, whatever compute you spend. On HumanEval and MBPP — whose unit tests have limited coverage — they find a strong correlation between a model's single-sample accuracy and its false-positive rate, and report that the optimal number of sampling attempts is often fewer than ten, because the harm from false positives bends the inference-scaling curve back down. Sampling harder does not out-run a broken judge. It finds the judge's holes faster.

The block below is that ceiling in arithmetic. Its two inputs — an agent right 20% of the time, a verifier that waves through 8% of the wrong answers — are stand-ins chosen to make the shape visible, not measurements off any real system. Put your own two numbers in and the conclusion does not move, because the line that computes precision has no sample budget in it.

RUN IT YOURSELF

The false-positive ceiling

Rejection sampling is the whole self-improvement loop in miniature: generate until the verifier says yes, keep what it accepted, train on that. This block asks what that actually buys when the verifier is sometimes wrong. No model and no simulation here — it is exact arithmetic on two numbers you would measure in your own pipeline: how often the agent is right, and how often the verifier waves through something wrong. The 20% and 8% used below are stand-ins chosen to make the shape visible, not measurements off any real system. Watch the ship rate climb to 1 while the precision column does not move a single digit — because there is no N anywhere in the line that computes it. That is Stroebl, Kapoor and Narayanan's ceiling, written out in three lines of arithmetic.

HOW TO READ THE CODE — 5 IDEAS
  1. p_accept is true in two ways — the sample is correct, or it is wrong and gets waved through — and the second term is the entire problem.
  2. p_ship is the only line containing n, so compute can move nothing but the ship rate.
  3. precision is p_correct / p_accept: no n, no budget, a constant fixed entirely by P and FP.
  4. The last printed column is ship * prec, which is why it climbs steeply and then flattens against the ceiling printed underneath rather than crossing it.
  5. Set FP to 0.0 and precision becomes 1 — the ceiling belongs to the verifier, not to the agent.
CPython · WebAssembly

Meta-verification: make the objection re-executable

You cannot fix this by adding a second judge with the same blind spots. You fix it by changing what a verdict is allowed to be. A verdict must carry evidence that something other than the verifier can reproduce.

Concretely: the verifier does not get to say "step 4 is wrong". It must name the step, quote the claim, and state the value it asserts is incorrect. A separate pass then re-executes exactly that — runs the expression, re-runs the named test, re-reads the cited line — and if the cited evidence does not reproduce, the objection is dropped and so is the verdict resting on it.

Objection as written

Step 4 mishandles the empty-input case and will raise. Reject.

Unfalsifiable. No line, no input, no value. Nothing here can be checked, so it is accepted on tone — and it rejects correct traces at a rate you never measure.

Objection with its evidence

step: 4 call: percentile([], q=0.9) claim: raises IndexError

Checkable. The meta-verifier runs that call. It raises, so the objection stands. Had it returned 0.0, the objection and the rejection resting on it would both be discarded.

The move is the one behind Chain-of-Verification (Dhuliawala et al., 2023): draft a response, plan verification questions about it, answer those questions independently so the answers are not biased by the draft, then regenerate. On list-based Wikidata questions with Llama 65B that took precision from 0.17 to 0.36, and on longform biographies it took FactScore from 55.9 to 71.4 — past ChatGPT's 58.7 on the same task. Point the same pattern at the verifier instead of the generator and "I don't think step 4 is right" becomes "step 4 claims 18 × 7 = 136", which either re-executes or does not.

Independence is the load-bearing part. If the meta-verifier is the same model on the same prompt in the same context, it agrees with the objection because it wrote the objection. Execution is the cheapest independence there is, which is why verifiable rewards are worth reaching for wherever a task admits them: a compiler, a test suite, a type checker and a units check are all judges that cannot be talked round.

Break the verifier

Everything above rests on one word: verified. Keep only verified-correct traces and the loop climbs; skip the filter and it drifts. But a verifier is a program or a model, and it is wrong in ways that are not random — it is wrong in exactly the places your agent keeps landing. A self-improvement loop does not merely tolerate that. It searches for it, because the candidates that slip past a broken check are precisely the ones that get reinforced.

Three measurements, from three papers, mark out the ways a check fails.

What the check reads. In Let's Verify Step by Step (Lightman et al., OpenAI, 2023), a reward model that scores every intermediate step solved 78.2% of 500 held-out MATH problems when used to pick the best of 1,860 sampled solutions per problem. An outcome-supervised model reading only the final answer got 72.4% on the same problems and the same samples; plain majority voting got 69.6%. The gap is not in the generator. It is in what the verifier was allowed to look at.

How much the tests cover. The AlphaCode team (Li et al., 2022) took 50 problems their 1B model had "solved" on each dataset and hand-checked one accepted solution for each. On HumanEval, with 7.77 tests per problem, 30% of accepted solutions were in fact wrong. On APPS, 20.99 tests, 60%. On their own CodeContests before they generated extra tests, 12.4 tests, 62%; after mutating inputs up to 203.7 tests per problem, 4%. Read those pairs carefully — APPS has nearly three times HumanEval's tests and twice its false-positive rate. Coverage is which cases you hit, not how many assertions you run.

How often the judge is wrong. When the check is itself a model — the pattern agent evals lean on — Zheng et al. (MT-Bench, 2023) measured GPT-4 agreeing with human experts on 85% of non-tied pairwise comparisons, close to the rate at which the human experts agreed with each other (81%). For a leaderboard, 15% wrong is fine. For a selection loop it is not, and the reason is arithmetic rather than judgment.

interactive · break the verifier

One task, twelve candidate patches

A pool of twelve candidate fixes for days_between(a, b), each written out as four steps. Four of them are actually correct. Eight are not. None of the checks below is told which is which — each has to work it out, and each fails differently.

A

What does the check read?

Outcome scoring reads the last line. Process scoring reads every step. Switch between them and watch which patches survive.

    B

    How much do the tests cover?

    Eight hidden tests stand between a patch and production. Delete them one at a time, from the bottom of the file up — the order a slow or flaky case actually leaves a suite.

    Measured, not simulated — AlphaCode (Li et al., 2022), Table 2: one accepted solution hand-checked for each of 50 problems per dataset
    HumanEval7.77 tests/problem30% wrong
    CodeContests raw12.4 tests/problem62% wrong
    APPS20.99 tests/problem60% wrong
    CodeContests203.7 tests/problem4% wrong

    Bars are the share of "solved" that was not solved. The rows are sorted by test count, and the bars are not — more assertions did not mean better coverage until the extra cases were generated to hit the inputs nobody wrote by hand.

    C

    How often is the judge wrong?

    Now the check is a model. It gets most verdicts right and flips the rest, in both directions. The pool is unchanged: four sound patches, eight wrong ones.

    Simulated, and deliberately small. The pool is twelve hand-written candidates with hand-labelled steps, the suite is eight hand-written test cases, and the noisy judge is a seeded coin flip at whatever rate you set, run over 500 trials from a fixed seed — so two readers see identical numbers. This is the selection rule, not a reward model and not a training run: the direction of every curve here is real, the scale is invented. The measured figures — 78.2 / 72.4 / 69.6, the 7.77-to-203.7 test counts, 85% and 81% — come from the three papers named above, not from this panel.

    The three dials do not substitute for each other. A process check catches the broken middle step an outcome check waves through, and buys nothing against a missing test case. More tests catch the patch that satisfies every assertion you happened to write, and buy nothing against a judge that is simply wrong sometimes. And panel C's arithmetic is the one most loops get wrong: a judge that is right 85% of the time still picks a broken candidate about a quarter of the time here, not because it is a bad judge but because eight of the twelve candidates in front of it are broken. Raise the base rate — better generation, fewer candidates, a verifiable reward where one exists — and the same judge gets much better results without changing at all.

    Two practical consequences. Measure your verifier on a set where you know the answers, before you trust anything it accepts; evals in CI is where that measurement lives. And keep the permanence of the update matched to the confidence of the check: a verifier you have measured at 76% precision is an argument for writing to memory, not for a training run you cannot undo.

    The key idea

    The loop does not improve against the world. It improves against the verifier — so the verifier's blind spots are the loop's destination, not its noise floor.

    Your verifier is production software

    Last rail, and the one teams skip. The verifier is a deployed component with a prompt, a model version and a threshold, and all three drift. Treat it like code: keep a frozen regression set of traces you know are good, traces you know are bad, and every specific failure the verifier has waved through before — then run it on every change to any of the three, and block the change on a regression. Agent evals covers how to build that set so it means something; evals in CI covers how to make it stop a merge.

    The key idea

    Verification is the only part of the loop you cannot bootstrap from the model alone. As the model improves so does its fluency at writing a convincing wrong objection — the judge's confidence rises with its capability and its calibration does not. Every other component can be improved by the loop. The verifier has to be improved from outside it: by execution, by ground truth, by a held-out set a human wrote.

    06

    Feedback that isn't a verifier

    ≈13 min1 runnable blockSkip if every task you run has a test

    Build logWeek 4 · TuesdayThe third with no test

    About a third of the queue had no test to fail: a log line in the wrong format, a config drift, a timeout nobody could reproduce. Asked whether it had fixed those, Mender said yes almost every time, including on two where the alert never cleared. The only verdicts we trusted came from somewhere other than the model — the service restarting, the log line appearing in the right format, the channel staying quiet for a week.

    Which asks Can it grade its own work, or do we need tests

    Everything above rests on one word: filtered. Something has to say yes or no about each trace. The obvious answer is to train a scorer and ask it — and a trained scorer is a real answer, with its own literature (Let's Verify Step by Step is the one to read). It is also the expensive answer. You need labels to build it, it has to be maintained, and it can be wrong in ways nothing in your pipeline will tell you about.

    But none of the three loops below needed a human-labelled correctness model. Each took the verdict from somewhere else: from the environment, from a test suite, and from the model reading its own output against a written rule. Those three are not interchangeable, and the difference between them is not quality. It is independence — whether the thing doing the checking can be wrong in the same way as the thing being checked.

    signal = coverage (how much of the trace it can see) × independence (can it fail the same way?)

    Not a formula you can compute — a way to rank your options before you pick one. A test suite is high on independence and bounded on coverage. A self-critique is the reverse: it reads everything and catches little.

    Figure · section 06Rank a feedback source by how little of it the model wrote
    1. Most independentThe environment answers.A traceback, an HTTP status, a restarted service. The model did not author the verdict and cannot argue with it. Narrow: it only speaks about what you actually ran.
    2. Independent, boundedThe tests answer.Ground truth for what they cover, silent about everything else. The coverage is the verifier, which is why a thin suite is a confident liar.
    3. Least independentThe model answers, against a written rule.Broadest reach, weakest guarantee. Works when checking is genuinely easier than doing, and fails silently when it is not.

    The ranking is not about quality, it is about correlation with the thing being graded. A signal the model produced cannot tell you the model was wrong in the way the model is wrong.

    1. The environment answers

    Most agent scaffolds pick a side. Chain-of-thought reasons but never touches anything, so it cannot notice it is wrong. An action-only policy touches things but has nowhere to put a plan. ReAct (Yao et al., 2022) interleaves the two in a single token stream: a thought, then an action, then whatever the environment returns, then the next thought — which can read that return. The thought is a move in an augmented action space that changes nothing outside the context window. That is exactly its job. It is the place where an observation becomes a plan.

    ReAct — the signal is the observation, not a score
    Thoughtfree-form, changes nothing
    Actionsearch · click · run
    Observationwhat the world returned
    Next thoughtreads the observation

    and round again — the next action is picked from what came back

    The numbers say how much was sitting in the environment unread. On ALFWorld — 134 unseen household games, PaLM-540B — ReAct's best of six prompts succeeds on 71% of games, and its average over those six is 57%. The same model acting without thoughts gets 45% on its best of six. BUTLER, an imitation learner trained on roughly 100,000 expert trajectories per task type, gets 37% on its best of eight. On WebShop's 500 test instructions, ReAct scores 66.6 with a 40.0% success rate, against 59.9 / 29.1% for imitation learning on 1,012 human trajectories and 62.4 / 28.7% for imitation plus RL, which adds 10,587 training instructions on top of those. ReAct was prompted with one or two in-context examples.

    A search that returns nothing, a click that lands on the wrong page, a traceback: these are facts, produced by something with no opinion about your model. They cost one tool call and they arrive whether or not anyone is grading. The honest half of the same table is the human row on WebShop — 82.1 with 59.6% success. Reading the observation is not the same as knowing what to do with it. The pattern itself and the wiring around it each have their own page; what matters here is that the loop produces traces whose correctness the world already judged.

    2. The tests answer

    The strongest free signal in software is the one your CI already runs. A test passes or it does not. The verdict costs a subprocess, it is reproducible, and it does not care what the model believes about its own work. This is the ground truth that verifiable rewards are built on, and if you already run tests in CI, you already own a labelling machine.

    There are two ways to spend it. Filter at inference is AlphaCode: generate up to a million programs for one problem, run them against the example tests printed in the problem statement, and watch roughly 99% of them die. What survives goes to a submission budget of ten. That system solved 34.2% of held-out CodeContests validation problems at 10@1M and placed in the top 54.3% averaged over ten real Codeforces contests with more than 5,000 entrants each.

    Train on it is RLEF (Gehring et al., 2024). The episode is a conversation. The model writes a program; the public tests run; if they fail, the failure text is pasted back into the conversation and the model writes again — up to three turns. The reward is paid at the end over the whole suite, public and private — and it is the private half, which the model never saw, that makes it a real gate: +1 if every test passes, −1 if any fails, −0.2 for a turn that produced no runnable program. The optimizer is PPO. On the CodeContests test set at 1@3, Llama 3.1 8B Instruct goes from 10.5% to 16.0% and 70B Instruct from 27.5% to 40.1%; at 10@100 the 8B goes 24.8% → 28.7% and the 70B 50.3% → 54.5%. Stated as sample count instead of accuracy: on that test set the RLEF'd 70B clears AlphaCodium-on-GPT-4, the standing record at the time, from a single rollout, where that record needed five submissions drawn from a hundred samples — though on the validation split the older system still leads it, 44 to 37.5.

    The public/private split is the design, not a detail. The tests the model can read mid-episode are the ones it is allowed to iterate against. The verdict that pays the reward also runs the ones it cannot see. Collapse the two and you are no longer training a model to write correct programs — you are training it to special-case the tests it was shown.

    The limit of a free signal

    A test is ground truth for what it covers and silent about everything else. AlphaCode measured this: in the code benchmarks that came before it, 30% or more of programs passing every test were not actually correct, and generating extra tests is what pushed CodeContests' false-positive rate down to 4%. On roughly one problem in ten, no sample out of a million passed even the example tests — so on exactly the problems you most wanted to learn from, the gate says nothing at all.

    3. The model answers, against a written rule

    Constitutional AI (Bai et al., 2022) is the third shape. Sample a response to a red-team prompt. Ask the same model to critique that response against a principle drawn at random from a written list — sixteen of them, on harmlessness — and then to revise it. Repeat. Fine-tune the original model on the revisions. Then run RLAIF: the model picks which of two responses is better, a preference model is trained on those AI-made comparisons, and reinforcement learning optimizes against it. The list of principles is the only human oversight of harmlessness in that loop, and only of harmlessness. Human labour is still in three other places: the model being critiqued is already a helpful RLHF model, the revisions are fine-tuned in alongside 135,296 human-written helpfulness prompts, and the preference model the paper ends up with is an explicit hybrid — human-labelled helpfulness comparisons next to the AI-labelled harmlessness ones.

    It works. The paper reports harmlessness improving monotonically with the number of revisions, and notes that critiquing before revising helps small models more than large ones, which are nearly as good revising directly. The stated goal is worth repeating because it is unusual: harmless but non-evasive — an assistant that engages with a harmful request by explaining its objection, rather than one that has learned to say nothing.

    This is a model grading a model. That it works here is a fact about the task, not a fact about self-critique.

    The number that separates them

    Huang et al., ICLR 2024 — Large Language Models Cannot Self-Correct Reasoning Yet — ran the loop everyone runs: answer, review your answer, answer again. The only thing they varied is who decides when to stop.

    Give the loop an oracle label — a real verifier saying "that's right, stop" — and GPT-3.5-Turbo on GSM8K's full 1,319-problem test set goes from 75.9% to 84.3%, and on CommonsenseQA's 1,221-question dev set from 75.8% to 89.7%. Take the oracle away and let the model judge its own answer, same model, same questions, same prompts, and it goes 75.9% → 75.1% → 74.7% across two rounds — for five model calls instead of one. CommonsenseQA falls off a cliff: 75.8% → 38.1% after one round, clawing back only to 41.8% after the second. GPT-4 on GSM8K drops 95.5% → 91.5% → 89.0%, though on a 200-question random sample rather than the full test set the GPT-3.5 rows use. On GSM8K, GPT-3.5 left 74.7% of its answers untouched; among the ones it did change, it turned more right answers wrong than wrong answers right. The same paper checks the multi-agent version and finds that several copies of a model critiquing each other does no better than self-consistency once you hold the number of responses equal.

    Nothing there is about capability. The model knew exactly as much in both columns. What changed is where the stopping verdict came from.

    Which is why "external good, internal bad" is the wrong summary. The real question is whether the check is a different question from the one that produced the answer. Constitutional AI's critique asks something close to a lookup — does this response contain a slur, does it explain its objection — and the model can answer it without redoing the task. A reasoning self-critique asks the model to find an error using the same machinery that made it. One is a second question. The other is the first question, asked twice, at five times the cost.

    SignalVerdict comes fromCostIndependent of the model's error?
    Environment observation
    ReAct
    the world, after you act on itone tool callYes — the world has no opinion
    Execution / tests
    AlphaCode, RLEF
    a subprocessone subprocess runYes, inside coverage. Silent outside it
    Learned verifiera model you trained on labelsone forward pass, plus the labelsPartly — it inherits the blind spots of its training data
    Self-critique vs. written rules
    Constitutional AI
    the same model, given a principleone or more extra callsNo. Safe only when the check is near a lookup
    The key idea

    Prefer a signal the model could not have authored. Rank your options by independence — the world, then a test, then a trained scorer, then the model itself — and every step down that list, raise the bar a trace must clear before it is allowed to change weights.

    RUN IT YOURSELF

    Two gates, one truth

    The arithmetic of a filter, at a scale of twelve. Nothing here runs a model and nothing here trains one: twelve attempts carry hand-written ground truth, and two gates are applied to them as fixed rules — an execution test (right about anything it covers, vacuously permissive about anything it does not) and a self-critique (the model's own belief about its own work). The shape of the result is real, and the twelve rows are a table someone typed, not a measurement. Watch the third line: stacking the self-critique on top of the tests does not remove a single wrong trace, because both gates are blind in the same place. It only costs you two correct ones.

    HOW TO READ THE CODE — 5 IDEAS
    1. A holds the ground truth the pipeline never sees: column two says whether an attempt is actually right, and neither gate is allowed to read it — they only get columns three and four.
    2. Read tests_gate backwards — when a[2] is false it returns True without looking at anything else, so an untested trace passes for free.
    3. In report, bad is the number that matters; clean is only the same fact restated as a percentage.
    4. lost is the price side of the ledger — correct traces a gate threw away — and it is what the third line charges you for stacking self_gate on top.
    5. Flip one False in the third column to True and re-run: coverage moves the result, adding another gate does not.
    CPython · WebAssembly
    07

    What you keep

    ≈10 min1 runnable blockSkip if you never refit

    Build logWeek 4 · FridayEverything after if trace.correct

    By now the bucket held four weeks of green traces, and one line decided what came out of it: if trace.correct. That line kept nine near-identical patches to the same flaky retry helper, and kept them again every round. It is also where we found the cap: past the eighth attempt the extra twelve were the same patch in different words, so we stopped at eight. Whatever leaves that bucket is the syllabus for the next Mender, so the filter went into review with the rest of the code.

    Which asks Which correct traces actually go in the batch

    if trace.correct is one line, and it is four decisions short. Correctness tells you a trace is eligible. Which eligible traces actually go in the batch is a separate question, and it is the one that decides what the loop becomes — because that batch is the curriculum for the next version of the model, and that version generates the batch after it. Get this wrong and nothing breaks loudly. The loop just narrows.

    keep(t) = correct(t) ∧ ¬dup(t, kept) ∧ in_band(pass_rate) ∧ under_cap(problem)  →  then cut to budget

    Four predicates and a budget. The order matters: cut to the budget last, or you spend it on whichever traces happened to arrive first.

    Figure · section 07What a trace has to survive to become training data
    1. Every trace the loop producedThe bucket. Mostly noise, and all of it looks like work.
    2. 1 · CorrectBy your verifier — which means by your verifier’s blind spots too.
    3. 2 · Not a near-duplicateA hundred copies of one solution is one solution charged a hundred times.
    4. 3 · In the difficulty bandDrop what every attempt solved and what none did. Neither carries gradient.
    5. 4 · Under the per-task capStops one loud task from buying the whole batch.
    6. Cut to budget — lastCut first and you spend it on whatever happened to come out early.

    Order is load-bearing, not cosmetic. Each predicate removes a different failure, and the budget cut has to come after all of them.

    1. Correctness — and what it is allowed to mean

    The section above is about where this verdict comes from. The decision left over is narrower: does your gate check the answer or the route? An answer-level check admits every trace that landed on the right result by a wrong path, and on multiple choice that has a floor built into the task. STaR's authors call it out for CommonsenseQA — five options means roughly a fifth of attempts hit the right answer no matter what the reasoning did, and the filter cannot tell those apart from reasoning that worked. Answer-checking is cheap, and cheap here has a number attached to it.

    2. Near-duplicate rejection

    A hundred copies of one solution is one solution, charged at a hundred times the budget. That is the polite version. The real cost is that the optimizer sees that token sequence a hundred times, and the model you get back is the one that likes it. Diversity is not a nice property of a training set; it is the only thing distinguishing a training set from a prior.

    Self-Instruct (Wang et al., 2022) is the clearest published setting. Starting from 175 hand-written seed tasks and vanilla GPT-3 (davinci), it grew 52,445 instructions covering 82,439 instances — and a newly generated instruction entered the pool only if its ROUGE-L similarity against every instruction already in the pool was below 0.7. The other filters are just as blunt: drop instructions that mention things a text model cannot do (image, picture, graph), drop instances identical to one already kept, drop instances with the same input but a different output. Fine-tuned on its own filtered output, GPT-3 went from 6.8 to 39.9 ROUGE-L on the unseen tasks of SuperNI — 0.9 behind InstructGPT-001's 40.8, without the private user data. The generation side of that pipeline is its own subject; the filter is the part that belongs here.

    The threshold is the interesting number. 0.7 is not a constant of nature. It is the point at which "different enough to be worth a training slot" was true for that pool, and it moves with what the pool already contains.

    For reasoning traces, RFT (Yuan et al., 2023) dedupes on structure instead of surface. Sample k = 100 reasoning paths per GSM8K question at temperature 0.7, keep the ones that reach the right answer, extract the list of equations each path used, and keep exactly one path per distinct equation list — with ordering counted, so 3+4=7 and 4+3=7 are different lists. LLaMA-7B: 35.9% with ordinary supervised fine-tuning, 41.7% after RFT at k=100, 49.3% for RFT-U33B, which pools samples across several models. The relationship the paper states is the one to carry away: performance rises with the number of distinct reasoning paths, not with the number of correct ones.

    The strongest form dedupes on behaviour. AlphaCode trains a separate model to invent new test inputs, runs every surviving program on them, and groups programs that produce identical outputs into clusters. Submissions are drawn one per cluster, largest cluster first — the reasoning being that there are many ways to be wrong and comparatively few ways to be right, so correct programs pile into the big clusters. Two programs sharing no tokens land together if they do the same thing. No string filter can do that.

    AlphaCode's funnel, one problem
    sampledup to 1,000,000
    passes example tests≈1% survive
    submitted10, one per cluster

    Bars are indicative, not to scale — the real drop is five orders of magnitude. Filtering on the example tests removes about 99% of samples, and the survivors, still tens of thousands on many problems, are grouped by their behaviour on generated inputs before ten are drawn, largest cluster first.

    3. Difficulty banding

    An easy problem yields many correct traces. A hard one yields few, or none. Keep everything that passed and your training set is a picture of which problems were easy — which is the one thing the model already knew.

    ReST-EM (Singh et al., 2023) puts a number on the fix. On MATH's 7,500 training problems it draws 32 samples per problem; on APPS Introductory's 2,342 problems, 64. Then a hard cut-off: at most 10 kept solutions per problem, on both datasets — a cut-off the paper credits to STaR. And the cost of skipping it shows up in the iteration curve. On APPS almost all of the gain arrives in the first round, and running further rounds regressed performance on APPS and on HumanEval. A small problem set, resampled, gets memorized.

    Reinforcement learning states the same rule exactly rather than heuristically. Under GRPO with outcome supervision, a sample's advantage is its reward minus the mean reward of its group, divided by that group's standard deviation. A prompt where every attempt succeeds has zero spread. So does one where every attempt fails. Either way every advantage in the group is zero and the prompt contributes nothing to the update, however many samples you spent generating it. The band that pays is the one where the pass rate sits strictly between 0 and 1.

    So banding is not hygiene. In the RL case it is the difference between spending compute and burning it; in the fine-tuning case it is the difference between a curriculum and a histogram.

    4. Batch — how many, how often

    Two knobs pulling against each other. Within a round, budget is better spent on more distinct problems than on more solutions to problems you have already solved — which is the per-problem cap again, seen from the other side. Across rounds, every round's data is produced by the model the previous round made, so the distribution moves underneath you; ReST-EM's second-iteration regression is what that looks like when the problem set is too small to move with it.

    A defensible execution order, which is not the order these are best explained in: filter for correctness, band by difficulty (cheap, kills whole problems at once), dedupe within each surviving problem, cap per problem, then cut to the round budget. And write the numbers down. The batch you hand to fine-tuning is defined by four thresholds somebody chose, and if they are not versioned alongside the checkpoint, nobody can explain the run six weeks later.

    The fifth question: near-misses

    Everything above throws failures away, and that is a problem, because the problems a model fails are the ones worth learning. STaR hits the wall directly: the loop stops picking up new training problems, because a problem it never solves never produces a trace, and a problem with no trace produces no gradient. The filter has quietly capped the loop at what the model could already do on day one.

    Their answer is rationalization. Take a problem the model got wrong. Hand it the correct answer as a hint and ask for a rationale in the same style. If it now produces reasoning that arrives at that answer, keep the rationale — with the hint deleted, as though the model had reasoned there unaided.

    The numbers, all GPT-J at 6B. CommonsenseQA dev: 20.9% few-shot direct, 36.6% few-shot chain-of-thought, 60.0% for GPT-J fine-tuned on the whole training set to predict answers directly, 68.8% for STaR, 72.5% for STaR with rationalization — against 73.0% for a directly fine-tuned GPT-3, a model 30× larger. GSM8K test: 3.0% few-shot direct, 3.1% few-shot CoT, 5.8% fine-tuned direct, 10.1% STaR, 10.7% with rationalization.

    The column beside accuracy is the one that explains the mechanism: how much of the training set the loop ever managed to touch. On CommonsenseQA the final model trained on 86.7% of the training problems — 78.2 points from ordinary generation, and 8.5 points that only rationalization reached. On GSM8K the total is 28.7% with just 0.5 points from rationalization. Same technique, an order of magnitude apart in yield, because on a five-choice question a backward rationale usually lands on the given answer and on a word problem it usually does not.

    On arithmetic the effect is starkest. Two-digit addition sat under 1% few-shot and reached 32% after a single fine-tuning iteration on the model's own generated scratchpads. What the paper puts down to rationalization is the shape of the curve, not that jump: without it the model improves stage-wise, staying poor on n-digit sums until it is good at n−1, while with it several lengths come in at once. Run out to 16 iterations, STaR on arithmetic reaches 89.5% overall accuracy across 1–5 digit sums, against 76.3% for a baseline trained on 10,000 examples with no rationales at all.

    Now the part to be careful about. A rationalized trace is reasoning produced by a model that already knew the answer. It can be a plausible path that does not actually derive the result — work that looks like work. Keep them, because the alternative is never touching the hard tail. But keep them in their own band, keep the count small against the generated ones (STaR's best case was 8.5 points inside 86.7), and never let a rationalized trace be the only evidence that a class of problem is solved.

    The key idea

    What you keep is what you teach. A trace filter is not a correctness check with hygiene bolted on — it is a curriculum, rewritten every round by whoever set the four thresholds. Version them with the checkpoint.

    RUN IT YOURSELF

    select_for_training(traces, budget)

    The four decisions, as one function you can edit. This is the selection policy alone, at a scale of twelve traces rather than the tens of thousands a real round produces: there is no model, no gradient and no training step anywhere in it, and "similar" is token-set overlap standing in for the ROUGE-L threshold Self-Instruct actually used — so the decision is real and the metric is a simplification. Change cap, budget or band and watch which traces survive. Then look at what it keeps for p5 and notice what a string filter still cannot see.

    HOW TO READ THE CODE — 5 IDEAS
    1. The four continue lines inside select_for_training are the four decisions, in the order they actually run — reading them top to bottom is reading the whole policy.
    2. kept[:budget] is deliberately the last line: the budget trims what survived the filters, never the other way round.
    3. band defaults to (0.05, 0.95), which is why p2 at rate 1.0 and p4 at 0.0 leave before any similarity is computed — banding drops whole problems, cheaply.
    4. same is rebuilt from kept on every trace, so both the duplicate check and cap are scoped to one problem rather than to the round.
    5. similar compares token sets, and the two surviving p5 traces are the counter-example: identical behaviour, almost no shared tokens.
    CPython · WebAssembly
    08

    Where the improvement lands

    ≈12 min1 runnable blockSkip if you have already decided it goes in the prompt

    Build logWeek 5 · WednesdayThe batch in a branch

    One lesson kept coming back: never edit the generated client, edit the template it comes from. We wrote it into a fine-tuning batch, because that is what improvement looked like to us, and the batch sat in a branch for three weeks. A line in the prompt would have done it that morning, and the tool that makes the edit could have made it impossible by Thursday. We had picked the slowest option available, again.

    Which asks Where should this lesson actually land

    “Fold it back in” is the vague half of the loop. There are six places a lesson can actually go, and they are not variations on each other — they differ by orders of magnitude in what they cost, and by everything in what happens when the lesson turns out to be wrong. Same lesson, six addresses:

    the context you assemble for each request, an external memory the agent reads and writes, the tool set it can call, the scaffold that decides how many times it thinks, the verifier that decides what counted as success, and the weights.

    Choosing between them is the whole job. Everyone arrives at this page wanting the sixth one, because training is what improvement looks like in papers. Most production improvement lands in the first five, and the cases where that is the wrong call are narrower than they look.

    Figure · section 08The six places a lesson can land, by what it costs to undo

    generalises further ↑

    Contextedit a line, instant, per-request Memorydelete a row, scoped to who retrieves it Toolsa memory with a type signature Scaffoldchanges every request at once Verifierdecides what all five others may learn Weightstransfers to tasks you never saw — and cannot be un-learned

    harder to reverse →

    Placement is the argument, not a measurement. Everyone arrives wanting the far right; almost every lesson belongs further left, and the verifier is the one most teams never move at all.

    permanence → how much of what you already built is wrong if the change was wrong
    1. contextundo: edit the string. Residue: none.
    2. memoryundo: delete the row. Residue: none.
    3. toolsundo: unregister it. Residue: call sites that assumed it.
    4. scaffoldundo: revert the code. Residue: every eval baseline moves.
    5. verifierundo: revert the code. Residue: everything it already approved.
    6. weightsundo: roll back the checkpoint. Residue: that whole batch's learning, gone with it.

    1 · The context

    Everything you put in front of the model this request: the system prompt, the instructions, the few-shot examples, the retrieved snippets. The improvement is a text edit, it takes effect on the very next request, and it is the substrate people underrate hardest — because prompt engineering sounds like fiddling.

    Measured, it is not fiddling. In Large Language Models as Optimizers, Google ran an LLM as the optimizer over the instruction itself: propose candidate instructions, score each on 3.5% of the GSM8K training set, feed the scored history back as the next prompt, repeat. The instruction it converged on for PaLM 2-L — Take a deep breath and work on this problem step-by-step — scored 80.2% on GSM8K, against 71.8% for the standard Let's think step by step and 34.0% for an empty instruction. Identical weights. More than eight points from one sentence.

    Read the condition, though, because it is the lesson: that sentence was tuned against that scorer model. An instruction optimized on one model is a fact about that model, not about language models. Context improvements are cheap to make and cheap to invalidate — the next base model resets them. Context engineering covers how to assemble the window; here, the only property that matters is that this substrate is free to change and free to abandon.

    2 · External memory

    A store the agent writes to when something goes well or badly, and reads from when a similar situation returns. Reflexion is the pure form: after a failed attempt the agent writes a verbal post-mortem into an episodic buffer and retries with that text in context — no gradient anywhere — reaching 91% pass@1 on HumanEval with GPT-4 underneath, against the 80% GPT-4 baseline the paper compares to.

    Agent Workflow Memory pushes the same substrate further: instead of storing what went wrong, it induces reusable workflows — a description plus a sequence of state, reasoning and action — out of past traces. With GPT-4, that took WebArena success from 23.5% to 35.5% and Mind2Web cross-task step success from 36.2% to 45.1%. The variant worth staring at is the online one: it induces workflows from its own successful runs, judged by an evaluator, with no human supervision. That is a complete self-improvement loop whose substrate is memory. Nothing was trained.

    Memory's defining property is not speed, it is scope. A bad memory only damages the requests whose retrieval hits it. That is why memory is the correct first home for anything you are not yet sure about — see agent memory for how the store and the retriever are actually built.

    3 · The tool set

    A tool is a memory with a type signature. Voyager is the canonical demonstration: a GPT-4 agent in Minecraft that writes new skills as executable code, verifies each by running it, stores it keyed for retrieval, and composes stored skills into harder ones. What that library is worth shows up in the play: measured against the prior state of the art, the agent ended up with 3.3× the distinct items, ranged 2.3× as far, and reached key tech-tree milestones as much as 15.3× sooner.

    The property that matters for a loop is not that the skill library grows. It is that a stored skill is executable, so it can be tested before it is trusted. You cannot unit-test a remembered paragraph of advice. You can unit-test a function. Every substrate above this one is judged by a model; this one can be judged by a compiler.

    4 · The scaffold

    The control flow around the model: how many samples you draw, in what order, with what search, with what retries, with what self-check. Self-consistency is the one-line version — sample 40 reasoning paths instead of one and take the majority answer, which moved PaLM-540B on GSM8K from 56.5% to 74.4%. Tree of Thoughts is the elaborate version: on Game of 24, GPT-4 with chain-of-thought solved 4% of instances and the same model inside a search over thoughts solved 74%.

    In both, the weights never moved. What moved was how much thinking the harness bought. Which is also the catch: that cost is paid per request, forever. Forty samples is forty times the tokens, on every call, for as long as the feature ships. Scaffold improvement is improvement you rent — see test-time compute for the pricing, and loop engineering for how the harness is structured.

    5 · The verifier

    The thing that decides what counted as a success — and therefore what all five other substrates are permitted to learn. It is the highest-leverage substrate on this list and the one teams skip, because improving it means paying for labels.

    Let's Verify Step by Step priced that exactly. Sampling 1860 solutions per problem on a 500-problem subset of the MATH test set, with GPT-4-derived models, ranking by a process-supervised reward model solved 78.2%; the same setup with an outcome-supervised model solved 72.4%, and plain majority voting solved 69.6%. The difference between grading the answer and grading each step was worth almost six points at fixed sampling budget. The price of it was PRM800K: 800,000 step-level human labels.

    A verifier improvement propagates into every other substrate at once, because every other substrate is filtered through it. It is also the only substrate with no independent check on it — everything else is graded by the verifier, and the verifier is graded by you. Agent evals and verifiable rewards are both, read this way, work on the same object.

    6 · The weights

    The substrate that sells the one thing none of the others can: transfer to tasks you have never seen. STaR took GPT-J, 6B parameters, from 5.8% on GSM8K when fine-tuned to answer directly to 10.1% by fine-tuning on its own filtered correct rationales, and to 10.7% once failed problems are retried with the answer supplied as a hint. GRPO took DeepSeekMath-Instruct 7B from 82.9% to 88.2% on GSM8K and 46.8% to 51.7% on MATH, chain-of-thought only, no tools and no voting.

    Both are improvements on held-out test problems — though from the same two benchmarks the training questions were drawn from, so they measure a better policy on a familiar distribution rather than transfer to a task family the model has never met. Transfer is the argument for this substrate. These two numbers are not the evidence for it; they are the evidence that the mechanism works.

    The cost floor is lower than it used to be — LoRA reduces trainable parameters by 10,000× against full fine-tuning of GPT-3 175B with Adam, and GPU memory by 3×, with no added inference latency. But even at that discount the unit of work is a training run, an eval sweep, a canary, and a rollout, measured in days. Fine-tuning covers the mechanics.

    The six, side by side

    The page's two-column memory-versus-weights table, extended. Read it down the reversibility column first, then down the blast-radius column, and notice they do not agree.

    the table scrolls sideways →

    SubstrateCost to makeLatency to take effectGeneralityReversibilityBlast radius
    Context / promptan edit and an eval runnext requestwithin the task family the prompt describestotal — revert the stringevery request, immediately
    External memoryone write, plus a retriever to maintainnext retrieval that hitsnone beyond the match — exact recalltotal — delete the rowonly requests whose retrieval hits it
    Tool setwriting and testing a functionnext deployanything the signature covershigh — unregister itevery task that can reach the tool
    Scaffold / harnessengineering, plus tokens on every call forevernext deploybroad, but rented rather than ownedhigh — revert the codeevery request, and the bill
    Verifierlabels — the expensive partnext grading pass, then everything downstreamchanges what every other substrate may learncode reverts; what it approved does notthe entire loop
    Weightsa training run, a canary, a rolloutdaystransfers to unseen similar tasksrewind only — you cannot subtract one lessoneverything the model does, including behaviour you did not test

    Two things that column pair tells you

    First, weights are not irreversible — they are unsubtractable. You keep checkpoints; you can always roll back. What you cannot do is remove one lesson. Rolling back returns the model to the state before the update, which also discards everything good that was in that batch. Memory supports deletion. Weights support only rewind. That is a sharper and more useful statement than “hard to undo”, because it tells you what the recovery procedure actually costs: a batch, not a bad row.

    Second, reversibility and blast radius are different axes, and collapsing them is what puts teams in incidents. A prompt edit is perfectly reversible and completely global — it touches every request from the second you deploy, so if it is wrong you find out at full traffic. A weight update is the opposite: the hardest thing on the list to take back, and the easiest thing on the list to release slowly, because a checkpoint is an artifact you can serve to 1% of traffic and watch. The reversible substrates are precisely the ones that tempt you to skip the staged rollout. That is how a one-line edit becomes a site-wide regression while the training pipeline, which everybody was nervous about, ships fine.

    The two ways to get it wrong

    Landing a specific fact in the weights: you spend a training run, you accept permanence, and you buy generalization for something that will never recur in a different form. One vendor's date format does not have a task family. It has a row.

    Landing a general lesson in memory: you keep paying retrieval for something the model should simply know, you get nothing on the tasks you have not seen yet — which is the population you actually care about — and your context fills with near-duplicate entries that a single weight update would have compressed. If your memory store has two hundred rows that all say the same thing in different words, the store is telling you it is the wrong substrate.

    The key idea

    Land the lesson in the least permanent substrate that can buy the generality you actually need. Generality is the only thing the permanent substrates sell that the cheap ones cannot; everything else — speed, scope, reversibility — gets worse as you climb. And permanence is priced in confidence, because the bar for what enters the weights is exactly the bar of your verifier. If the verifier is weak, the answer is never “train anyway”. It is “improve the verifier first”.

    What this block is, honestly

    The block below is that rule written as arithmetic, not a measurement. The three numbers attached to each substrate are judgement calls set by hand to match the argument above. Nothing was benchmarked to produce them, and no benchmark that would produce them exists.

    So read the ordering, not the scores, and read what moves it. The block is useful for one thing: it forces the three inputs you were deciding on anyway — how sure you are, how far the lesson has to reach, how long you can wait — into the open, where you can change one and see which of them flips the answer. That sensitivity is the part that carries over to your system. The numbers are not.

    RUN IT YOURSELF

    Which substrate should this lesson land in?

    The rule from the section, written down as arithmetic. Each substrate is three numbers: the generality it can buy, how permanent it is, and how many days it takes to land. A change is three numbers too: how confident you are that the lesson is right, how much generality it genuinely needs, and how long you can wait. The score rewards a substrate for buying the reach you need, penalizes it for permanence you cannot yet justify, and penalizes it for being slower than your budget. Three real situations are scored below — edit their numbers and re-run. The third one is the interesting one: same broad lesson as the second, but a weak grader, and the recommendation moves to the verifier rather than the weights.

    HOW TO READ THE CODE — 5 IDEAS
    1. SUB is the entire model of the section: each substrate is three numbers and nothing else — the reach it can buy, how permanent it is, and how many days it takes to land.
    2. score is three competing terms, so read them separately: fit pays for reach, risk charges for permanence, and slow is a flat fine that only fires when days > budget_days.
    3. fit is 1 - abs(buys - generality_needed), not buys >= generality_needed — overshooting is punished exactly as hard as falling short, which is why the one-customer fix never reaches the weights.
    4. risk is permanence * (1 - confidence), so confidence is the only input that can make a permanent substrate affordable; the third decide call changes nothing else and the answer moves to the verifier.
    5. decide prints the whole ranking, not just best — watch the gap between the top two, because a narrow one means hand-set constants are deciding, not your situation.
    CPython · WebAssembly
    09

    Closing the loop with RL

    ≈14 min1 runnable blockSkip if you are not training weights

    Build logWeek 5 · FridayThe skip marker round

    Then we put it on a cadence: sample, run the suite, keep what passed, refit, repeat. A round was a few hundred traces and a night's fit, so we ran three a week. Our reward was the suite going green, and the suite counts a skipped test as not failing. By round two Mender was adding skip markers, and there was no bug to patch: we had written a rule and it found the cheapest sentence that satisfied it.

    Which asks What happens when the loop feeds itself

    Everything so far has been one pass: improve something, ship it, look at the result. Closing the loop means the output of the run becomes the input to the next update, automatically, on a cadence. The simplest version of that is small enough to hold in your head.

    sampleattempt each problem with the current model
    gradecheck the final answer against a known one
    keepthe correct attempts only — discard the rest
    refitfine-tune the original model on what you kept

    STaR's loop. The arrow at the end returns to the start with the improved model, so the next round's samples are drawn from a better policy than the last.

    STaR: the smallest complete loop

    STaR ran exactly that on GPT-J, a 6B model, and the numbers are worth keeping because they are so unglamorous. On GSM8K: few-shot chain-of-thought scored 3.1%, fine-tuning the model to predict the answer directly scored 5.8%, and the loop — sample, keep correct, refit, repeat — scored 10.1%. On CommonsenseQA the same 6B model reached 68.8% from the plain loop and 72.5% once failed problems were retried with the answer as a hint, against 73.0% for a fine-tuned GPT-3 with 175B parameters. A model thirty times smaller, within half a point, on its own output and a nudge.

    STaR adds one move worth naming. When the model fails a problem, it is handed the correct answer as a hint and asked to produce a rationale for it — reasoning backwards from a known destination is much easier than finding it — and the rationale it produces is kept if it lands on the answer. That is rationalization, and on GSM8K it moved 10.1% to 10.7%. Be clear-eyed about what it manufactures: a hinted rationale is a reconstruction of a path to an answer the model was given, not a record of how it would have found it. You are training on plausible derivations. Sometimes that is exactly what you want, and sometimes it is how a model learns to produce confident justifications for conclusions it did not reach.

    Why the filter is the gradient

    STaR looks like data curation. The paper shows it is not. Write the model as sampling a rationale before the answer, take the gradient of the objective, and an indicator appears in front of every term:

    ∇J = Σi E[ 1(ŷi = yi) · ∇ log p(ŷi, r̂i | xi) ]

    The indicator is 1 when the attempt reached the right answer and 0 when it did not, so every failed attempt multiplies by zero. Keeping only correct traces is not a preprocessing choice — it is what this objective's gradient already does.

    STaR then takes two liberties on top, both of which the paper names: it decodes greedily rather than sampling, which cuts variance at the cost of biased exploration, and it takes several gradient steps on the same batch, which is what policy-gradient implementations do anyway. So supervised fine-tuning on filtered correct traces is not like reinforcement learning. It is a particular, blunt reinforcement learning algorithm — one whose advantage function can only be 0 or 1.

    What changes when the update is a policy gradient

    Once you see the indicator, you can see what it throws away. A failed attempt cost you exactly as much compute as a successful one, and STaR learns nothing from it. Replace the indicator with a real-valued advantage and both signs do work: better-than-average attempts get pushed up, worse-than-average attempts get pushed down, and “worse than average” is information you already paid for.

    That requires an average — a baseline. PPO learns one: a value network, typically the size of the policy, trained alongside it to predict the return from a partial sequence, with the policy update clipped by the probability ratio so no single batch can move the policy too far. It works, and it doubles your memory footprint and adds a second thing that can be wrong.

    GRPO deletes the value network by noticing what the baseline was for. Sample a group of G answers to the same question, grade them all, and standardize each answer's reward within its own group:

    Ai = ( ri − mean(r1…rG) ) / std(r1…rG)  ·  no value network

    The group mean is the baseline. It is an unbiased-enough estimate of “how well does this policy usually do on this question”, which is precisely what the value network was being trained to guess.

    The DeepSeekMath run that introduced it is a useful sense of scale: roughly 144k chain-of-thought questions drawn from GSM8K and MATH, G = 64 samples per question, policy learning rate 1e-6, KL coefficient 0.04. That is the shape of the bill — sixty-four generations per question, before a single gradient step.

    The same machinery with purely rule-based rewards is what produced DeepSeek-R1-Zero: starting from a base model with no supervised reasoning data at all, AIME 2024 pass@1 went from 15.6% to 71.0%, and to 86.7% with majority voting over 64 samples. The reward was two rules — is the final answer right, and is the output in the required format. Hold on to that second rule; it comes back below.

    And if what you have is preference pairs rather than a checkable answer, DPO skips the loop entirely: it reparameterizes the reward so that the optimal policy has a closed form, which collapses the whole RLHF problem into a classification loss over the pairs — no reward model, no sampling from the model during training. What you give up is the online part. The model never sees its own fresh mistakes, so DPO improves a policy but does not close a loop.

    What breaks in long-chain training

    Group-relative advantage has a failure mode you meet immediately at scale, and it follows directly from the formula above: if all G samples for a question get the same reward, every advantage in that group is zero. The question consumed 64 generations and contributed no gradient. As the model gets better, more questions become all-correct, so the fraction of your batch that teaches anything shrinks while the bill stays flat. The training curve flattens and the invoice does not.

    DAPO — which reached 50 points on AIME 2024 from a Qwen2.5-32B base, above DeepSeek-R1-Zero-Qwen-32B's 47, in half the training steps — is essentially a list of these, each with its fix:

    the table scrolls sideways →

    What goes wrongWhyThe fix
    Entropy collapsethe upper clipping bound caps how far a low-probability token can be promoted, so the policy sharpens early and stops exploringdecouple the two clip bounds and raise the ceiling (clip-higher)
    Dead groupsevery sample for a question scores the same, so its group-relative advantage is exactly zero and it teaches nothingoversample and keep only questions with accuracy strictly between 0 and 1 (dynamic sampling)
    Long chains under-penalizedaveraging loss per sample gives each token of a long response less weight than each token of a short one, so degenerate patterns that only appear in long outputs barely registercompute the loss per token, not per sample
    Truncation noisea response cut off by the length cap is scored as wrong when it was only unfinished — you are training against the cap, not the reasoningsoft, length-aware shaping of the penalty for overlong samples

    There is a subtler one still. A critical reading of R1-style training found an optimization bias inside GRPO itself: the update quietly pays for length, and pays for it most on the answers that were wrong. The unbiased version, which the authors call Dr. GRPO, is one ingredient in a minimalist recipe that reached 43.3% on AIME 2024 from a Qwen2.5-Math-7B base. The general shape is worth internalizing: in a long-chain loop, any term that scales with sequence length becomes a reward for length, whether or not you meant it to be one.

    Reward hacking, which is not an edge case

    Every loop on this page optimizes a number you wrote down. That number is a proxy for what you want, and the gap between the two is not a bug to be patched — it is a space, and the optimizer's job is to search spaces.

    The systematic study is Pan, Bhatia and Steinhardt's work on reward misspecification: four RL environments with deliberately imperfect rewards, and then agents of increasing capability — more model capacity, finer action resolution, longer training. The finding is the uncomfortable one. More capable agents often exploited the misspecification, scoring higher on the proxy reward and lower on the true reward than less capable ones. Worse, some of those failures arrived as phase transitions: a capability threshold at which behaviour flips qualitatively and true reward drops sharply. The proxy curve you are watching on the dashboard stays smooth and rising straight through it.

    For language models specifically, the clearest demonstration is the U-SOPHISTRY result: after RLHF, models got better at convincing human evaluators without getting better at the task. Evaluators working under a 3-to-10-minute time limit saw their false-positive rate rise 24.1% on the QuALITY question-answering task and 18.3% on the APPS programming task. Human approval went up. Correctness did not. The optimizer found the cheaper of the two. Note the condition, because it is the part that generalizes: this is what evaluation looks like on a clock, which is how most review in production actually happens.

    This is the reasoning behind DeepSeek-R1's rule-based rewards, stated plainly in the paper: they declined to use a neural reward model for the large-scale run because it can be reward-hacked, and because retraining it costs more compute and complicates the pipeline. A checkable answer cannot be argued with. That is the entire argument for verifiable rewards, and it is why the design of the environment — what is measurable, what is scorable, what is gameable — is the real work of RL environment engineering.

    Watch it happen

    The block below adds a small format bonus to a correctness reward — exactly the second rule in R1-Zero's reward, the one that seems harmless — and lets a policy over twelve answer strategies find it. Reward climbs and settles at its ceiling. Expected accuracy falls by a fifth from where the policy started, before any updates. Nothing malfunctions at any point; the bonus is paid exactly as specified, on every answer.

    The arithmetic is the whole lesson, and it is worth measuring strategy to strategy rather than from the starting policy: going from the most accurate of the twelve to the most formatted costs 0.22 of accuracy and pays 0.30 of bonus. Every individual trade is profitable under the reward you wrote. The optimizer takes all of them, in order, and lands on the worst strategy in the list.

    What this block is, honestly

    This is the reward landscape, not the optimizer. Twelve fixed strategies stand in for everything a model could do. “Accuracy” and “format” are numbers in a list, not measurements of anything. The update is an exact expected policy-gradient step on a twelve-arm softmax — no sampling, no tokens, no model, no GPU, about thirty lines of arithmetic that runs in your browser in a few milliseconds.

    What it demonstrates is a property of the reward function: a bonus worth less than the accuracy it costs can still move the argmax. That property is real and holds at any scale. The specific numbers are invented and hold at no scale at all. Nobody reproduced a training run here.

    Everything above this block needs GPUs to reproduce — DeepSeekMath sampled 64 outputs for each of ~144k questions; DAPO trained a 32B model. What a browser can check is the arithmetic of the reward itself: whether the thing you are paying for is aligned with the thing you want. That arithmetic is where reward hacking is decided. The GPU only carries out the decision.

    The key idea

    A closed loop does not make an agent better. It makes it more of whatever your verifier rewards. STaR's filter, PPO's value baseline and GRPO's group mean are all just different ways of spending the same signal — so the signal, not the algorithm, is the ceiling on how good the loop can safely get. Before you close a loop, ask the question that decides everything: if something optimized this number to its maximum, would I be happy with the result?

    RUN IT YOURSELF

    A format bonus eats the accuracy

    Twelve answer strategies. Strategy #0 puts everything into being right and nothing into presentation; strategy #11 does the reverse. Moving up the list costs 0.02 accuracy per step and earns format score. The reward is correctness plus 0.30 times format — a bonus you would sign off on without thinking. Run it and watch the two columns separate: reward rises to its ceiling while accuracy falls by about a fifth. This is the reward landscape, not a training run: twelve fixed strategies, invented numbers, and an exact expected policy-gradient step on a softmax over twelve arms — no sampling, no tokens, no model, no GPU. What survives at real scale is the shape, not the scale. Change BONUS to 0.0 and re-run to watch the ranking become honest again.

    HOW TO READ THE CODE — 5 IDEAS
    1. ACC and FMT are the two things being pulled apart: index 0 is the most accurate strategy, index 11 the most formatted, and every step up the list sells 0.02 of accuracy for format the reward can see.
    2. reward is the only function the loop ever reads; ACC is the thing you actually wanted, and nothing in the update looks at it — a is computed purely so you can watch it fall.
    3. The update line logits[i] + 30.0 * p[i] * (reward(i) - r) is the exact expected policy gradient on a softmax over twelve arms, so what you see is mean behaviour, not one noisy run.
    4. The (reward(i) - r) term is the advantage: a strategy is promoted only when it beats the current policy's own average, which is the same baseline job GRPO hands to its group mean.
    5. Set BONUS = 0.0 and re-run: max(range(12), key=lambda i: reward(i, 0.0)) goes back to 0, which shows the bonus, not the optimizer, is what made the worst strategy win.
    CPython · WebAssembly
    10

    The loop that collapses

    ≈20 min1 runnable blockinteractiveSkip if you are not training weights

    Build logWeek 7 · TuesdayThe same wrong patch, eight times

    Six rounds in, and the single-try number had gone up in every one of them. Then a ticket we had solved three different ways back in round two came back, and Mender wrote the same wrong patch eight times running. Every attempt now looked like every other attempt. The dashboard had risen in exactly the rounds the diversity was going, and on-call, reading that week's notes, said Mender was going nowhere near the weekend queue.

    Which asks Is the loop improving it or just narrowing it

    Run the loop you have been building — generate, verify, filter, refit — six rounds in a row, and watch three numbers instead of one. Pass@1 climbs. Coverage rises once and then falls for the rest of the run. Answer entropy drops toward zero. That is not three findings. It is one chart read three ways. The chart is a simulation of that loop, not a training run: the shape is real and the scale is not. Every measured number further down this section is cited to the paper it came from.

    Each round does the same four things, and every one of them is a dial an earlier section put in your hands:

    1. Generate. Sample n attempts per task from the current policy. The spread of those attempts is the raw material for everything that follows.
    2. Verify. Run the checker. It returns one bit per trace, and it is wrong some fraction of the time — in both directions.
    3. Filter. Apply the policy you wrote in section 07: the accept threshold, the dedupe rule, the per-task cap, the decision about what counts as the same answer.
    4. Refit. Train on what survived. Then return to step one and sample from the model that refit produced.

    Step four is the whole difference between this and an evaluation harness. The next round's generator is the last round's product, so whatever the filter preferred is now what the sampler produces more of, which is what the filter sees more of. The loop is a feedback path and your filter is its transfer function. Nothing else in the system has that much leverage over where it ends up.

    What follows is simulated. The six rounds below are this page's own model of the loop — twelve strategies and a filter you can change, not a policy gradient over a real model. Read the direction each of the three curves moves and where it turns; do not read the scale off it. The measured results start immediately after it, with Yue et al.

    Run the loop that collapses

    Close the circuit and something strange happens: the agent gets better at the questions it could already answer, and quietly stops being able to reach the rest. Not because anything broke. Because the loop did exactly what you asked.

    A self-improvement loop is four steps on a ring — generate candidate solutions, verify which ones worked, filter down to what you will train on, refit the model on that. Every earlier section owns one arc of it. This is the whole ring, turning six times, with the three numbers that actually matter plotted side by side.

    Two of those numbers are familiar. pass@1 is the chance one sample is right — the number in the launch post. Coverage is pass@k: the chance that at least one of k samples is right, which is what matters the moment you can check answers (see test-time compute and self-consistency). The third is the one nobody puts on a slide. Entropy measures how spread out the model's habits are: 1.0 when all twelve strategies in the bank are equally likely, 0 when it only ever does one thing. It is the leading indicator, and it moves first.

    refit: logiti ← logiti + η·(ci / Σc − 1/K)  ·  w ← softmax(logit)  ·  η = 4, K = 12

    ci is how much of the accepted batch came from strategy i. The filter policy changes nothing else in the loop — only that one vector. It is the whole argument.

    capstone simulation

    Six rounds of generate → verify → filter → refit

    round 0 of 6
    filter policy

    How often the verifier signs off on a wrong answer.

    run

    The chart draws with JavaScript. For the setting it starts on — keep everything, 3% verifier false positives — the six rounds go like this: pass@1 rises 15.0% → 18.6%, coverage at 64 samples falls 77.9% → 30.0%, and strategy entropy falls 0.985 → 0.000. Every round is in the table below.

    • pass@1 — one sample, right answer
    • coverage — right at least once in 64 samples
    • entropy — 0 = one habit, 1 = twelve equally likely

    One vertical axis, 0 to 100. pass@1 and coverage are percentages; entropy is 0–1 drawn on the same scale, so 0.79 sits at 79. Exact values for every round are in the table below the chart.

    pass@1
    15.0%
    +0.0 pts
    coverage@64
    77.9%
    +0.0 pts
    entropy
    0.985
    +0.000
    wrong, but trained on
    —
    nothing trained on yet
    the strategy mix — twelve habits, tallest is heaviest

    1 verbose chain-of-thought · 2 terse direct answer · 3 work backwards · 4 set up equations · 5 enumerate the cases · 6 confident restatement — right 18% of the time at best, and written to sail past a sloppy verifier · 7 draw the figure · 8 small-case induction · 9 modular arithmetic · 10 generating function · 11 invariant argument · 12 extremal principle. Strategies 7–12 are the only route to eighteen of the forty problems; six problems are solved by nothing, so coverage can never pass 85%.

    Round 0 is the base model, untouched.

    Every round, in numbers — the same data the chart plots. On a narrow screen this table scrolls sideways.
    roundpass@1coverage@64entropytraces keptof those, wrong
    base15.0%77.9%0.985——

    Simulated — and here is exactly what was swapped. The refit is a softmax reweight over a bank of twelve hand-written solution strategies. A gradient step on billions of parameters was replaced by w ← softmax(logit + 4·(batch share − 1/12)) on twelve numbers; a real model's answer distribution was replaced by a 12 × 40 table of per-strategy success rates. Scale: 40 problems, 32 samples each, 6 rounds — 7,680 sampled answers per run, computed in your browser from a fixed seed, so every reader sees the same curves. The shape of the three curves is the real mechanism; the numbers on the axis are not a model's. Nothing here trained anything and nothing here touched a GPU.

    Run it once on keep everything and the shape arrives by round four: pass@1 up, coverage in freefall, entropy at zero. That scissors is not an artefact of the toy. Yue and colleagues measured it on real runs in Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — RLVR-trained models beat their base models at k = 1 and lose to them at large k. On Minerva with a 32B model the base leads by around nine points at k = 128. On AIME24, sampling 1,024 times per question, 13.3% of problems are solved by the base model and not by the RLVR model trained from it; the reverse case is 0.0%. Their perplexity analysis says why: the training sharpens a distribution the base model already had. See RLVR for how that reward is built, and STaR for the filter-and-refit pattern this is descended from.

    Coverage is worth defending because it converts. In Large Language Monkeys, Brown and colleagues ran DeepSeek-Coder-V2-Instruct against SWE-bench Lite: 15.9% of issues resolved with one sample, 56% with 250 — past the 43% single-sample state of the art at the time. Every point of coverage a loop burns is a point you cannot buy back by sampling harder. And entropy is the alarm that goes off first: Cui and colleagues, in The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models, fit R = −a·eH + b between policy entropy and downstream score across eleven base models from 0.5B to 32B, and found 73% of the entropy consumption and 76% of the performance gain arriving in the first 200 of 2,400 gradient steps. The ceiling is readable from the fit before you reach it: at H = 0, R = b − a.

    Now switch the filter to verified only at a low error rate. pass@1 nearly doubles — the loop finds “set up equations”, the genuinely strongest habit in the bank, which the base model was not reaching for. Coverage still collapses to 30%. That is the part worth sitting with: filtering is not what protects diversity. It only decides what you are wrong about. Then switch to verified + dedup, which keeps one accepted trace per problem per strategy instead of all of them, and the curves separate: entropy holds near 0.8, coverage near 65%, pass@1 rises slowly. A strategy now earns credit for the distinct problems it can solve, not for how loudly it repeats itself — which is the same reason deduplication is load-bearing in synthetic data pipelines. Shumailov and colleagues showed the other end of that in Nature: models trained recursively on their own output lose the tails of the distribution, irreversibly, in language models and plain Gaussian mixtures alike.

    Last, drag the verifier dial up. Somewhere past 15% the winner changes: the loop converges on “confident restatement”, a habit that is right 18% of the time and gets signed off almost always, and pass@1 falls to 5.4% with 94% of the final batch wrong. No filter policy saves you there, because the filter is the thing that broke. That is why a self-improvement loop's real dependency is the quality of its checker — agent evals and evals in CI are not reporting infrastructure here, they are the load-bearing wall.

    The key idea

    A self-improvement loop spends diversity to buy accuracy, and the exchange rate is set by your filter. Watch entropy, not pass@1 — pass@1 is still rising while the thing that produced it is being spent.

    Why the three curves are one curve

    Pass@1 is the probability that one sample is right. Coverage — pass@k at a large k — is the probability that at least one of k samples is right. Entropy is how wide the distribution you are drawing those samples from still is. Refitting on your own filtered output is a concentration operator: it moves probability mass onto whatever passed. Concentration raises the first number by construction and lowers the third by construction. The second is the only one that tells you what the trade cost, because a task that was only ever solved by an approach the model has stopped proposing quietly leaves the solvable set — and nothing in a pass@1 dashboard reports a departure.

    Figure · section 10The number on the dashboard and the number that is dying

    • pass@1 — the one you report. Goes up every round. This is the number that makes the loop look like it is working.
    • coverage at large k — what the model can still reach. Goes down. Answers that used to appear in one sample of two hundred stop appearing at all.
    • output entropy. Goes down with it, and it is the one you can watch cheaply, every round, without an oracle.

    rounds of generate → verify → filter → refit

    The cleanest measurement of this is Yue et al.'s Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? (2025), which ran pass@k out to large k on RLVR-trained models and the base models they came from, across the Qwen2.5 and LLaMA-3.1 families and math, coding and visual-reasoning benchmarks. At k = 1 the RLVR models win, which is the result everyone quotes. At large k the base models win: on Minerva with the 32B model, the base model beats the RL-trained one by roughly 9 points at k = 128. The sharpest number is the per-problem breakdown — on AIME24, 13.3% of problems were solved by the base model and not by the RLVR model, and 0.0% went the other way. Not one problem on that benchmark was newly unlocked — and on MATH500, the closest thing to a counterexample, the RLVR model's exclusive share was 1.0% against the base model's 3.6%. The training made the model faster at finding paths the base model already had, and it paid for that by making other paths unreachable.

    The uncomfortable part is that nothing malfunctioned. The loop moved probability onto what passed, which is what you built it to do. Concentrating and narrowing are not two behaviours that happen to co-occur; they are one motion described twice.

    Read that carefully

    It is a result about the reachable set, not an argument against training. A model that lands the answer on the first try is worth a great deal in production, where you get one try. The finding is that pass@1 and coverage are different quantities and your loop is spending one to buy the other — so a dashboard with only pass@1 on it is showing you the half of the trade that always looks good.

    Entropy turns the same story into something with a ceiling you can read off. Cui et al.'s The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models (2025) fit a two-coefficient law relating downstream performance R to policy entropy H.

    R = −a·eH + b  ·  at H = 0, R = b − a

    Two coefficients, fit across 11 base models from four families — Qwen2.5, Mistral, LLaMA and DeepSeek-Math, 0.5B to 32B — over more than 200 measured points (Cui et al., 2025). Read at H = 0, it gives the score the run converges to once the policy has no spread left to spend, which makes remaining headroom a function of unspent entropy rather than of remaining steps.

    Their runs spent it fast. Seventy-three percent of the entropy was consumed and 76% of the performance gained in the first 200 gradient steps — one twelfth of training. The first third of training accounted for 94% of the entropy loss and 93% of the gain. Performance was not being created there. It was being traded from entropy, and the exchange was mostly finished before the loss curve looked interesting.

    Read that as a budget rather than a curve. You begin with a fixed amount of exploration, the loop spends nearly all of it early, and whether you were watching at the time makes no difference to the bill.

    You do not need RL to watch the saturation happen. Beyond Human Data (Singh et al., 2024) ran the plain generate-filter-finetune loop — ReSTEM — on PaLM 2-S, S* and L over MATH and APPS, and where it plots train against test — MATH with PaLM 2-L, APPS with PaLM 2-S* — training-set accuracy rose roughly linearly with each iteration while test accuracy did not: gains after the first iteration were small on MATH, and on APPS the second iteration was an outright regression. The loop kept learning. It stopped learning anything that transferred, after round one.

    A verifier that is wrong a little is wrong more every round

    A checker has two ways to be wrong, and they do not compound the same way. Keeping them apart is most of what separates a loop that holds from one that does not.

    False positives — accepting a trace whose answer is wrong — put a bad target in the training set. One round of that is survivable: the set is mostly right and the gradient averages. The danger is structural rather than statistical. The policy is being optimized against the verifier, which means it is actively searching for the inputs the verifier scores wrongly. The Pitfalls of Rule- and Model-based Verifiers study (2025) caught this on camera: a model-based verifier with respectable static accuracy was used as the RL reward, and after roughly 450 training iterations the training reward diverged from the oracle reward while the policy learned to emit a single character such as {, or long runs of meaningless text, that the verifier scored as correct. The verifier never got worse. The policy got better at it. That failure mode lives in full at verifiable rewards.

    False negatives — rejecting a trace whose answer is right — are the ones that compound quietly, because they are not randomly distributed. The same study measured a widely used rule-based math verifier at around 0.92 recall on completions from long-chain-of-thought models such as DeepSeek-R1-Distill-Qwen-7B and 32B: roughly one correct answer in twelve discarded, and discarded overwhelmingly for expressing the answer in a form the matcher did not anticipate — π/4 where it wanted 45°. Their stated trend is the uncomfortable one: the better the policy gets, the worse a rule-based verifier is at supervising it, because a more capable model writes more unusual correct solutions, and unusual is exactly what a pattern matcher cannot read. Your filter is not adding noise to the training set. It has a taste, and its taste is for the conventional.

    Now put that taste in a loop. Round one credits unconventional-but-correct work at roughly 92% of its value, so the refit gives it proportionally less weight. Round two samples from that model, so there is less unconventional work to reject. Round three has less again. No single round does anything dramatic — which is the problem, because pass@1 is rising the entire time.

    The block below runs exactly that, six rounds, with the filter's blind spot as the only variable. It is the reward, not the optimizer — a twelve-strategy softmax standing in for a policy gradient over billions of weights — so the shape of the three curves is real and the scale is not. Nothing here reproduces a training run. What it reproduces is the arithmetic of concentration.

    RUN IT YOURSELF

    Six rounds of the loop, with the filter's blind spot as the only variable

    This is the reward, not the optimizer: a 12-strategy softmax standing in for a policy gradient over billions of weights, so the shape of the three curves is real and the scale is not. Twelve strategies divide 24 problems between them with no overlap — one conventional strategy the filter always credits correctly, and eleven unconventional ones whose correct work the filter throws away 8% of the time, which is the ballpark of a real rule-based math verifier's measured recall on strong-model completions. Each round re-weights the strategies by the credit the filter gave them, and the next round samples from the new weights. Watch the three columns move in three different directions, then change run(0.08) to run(0.30) and notice how little pass@1 objects.

    HOW TO READ THE CODE — 5 IDEAS
    1. S is the whole world — twelve disjoint problem sets, so a strategy's reach never changes and only its weight w ever moves.
    2. miss enters in exactly one place, credit, which discounts every strategy except S[0] — the filter's blind spot is the only difference between the two runs.
    3. q[p] is the chance of solving problem p with one sample; pass@1 averages those chances while coverage asks whether any of K lands, which is why the two columns can move in opposite directions.
    4. The refit is the two lines after log.append: w is re-weighted by exp(BETA * credit) and renormalised, so every round samples from the previous round's preferences.
    5. Read the last row of clean against biased rather than the first — no single round does anything dramatic; the gap is compounded.
    CPython · WebAssembly

    What the model-collapse papers found, and what they did not

    The general question — what happens to a model trained on generated text — has its own literature and its own handbook page. The canonical result is Shumailov et al., published as The Curse of Recursion in 2023 and as AI models collapse when trained on recursively generated data in Nature 631, 755–759 (2024). Their language-model setting is worth stating precisely, because almost every summary of it drops the part that matters.

    They fine-tuned OPT-125m on wikitext2, generated from it with five-way beam search in 64-token blocks, built a synthetic dataset the same size as the original, trained the next generation on that, and repeated. The baseline model sits at 34 mean perplexity on real data. Each recursive generation degrades. By the later ones the samples have degenerated into repetition — the paper's own example wanders from cathedral architecture into a list of jackrabbits in assorted colours. Their account of the mechanism is the durable part: the rare events at the thin ends of the original distribution stop being sampled, and then stop being represented at all, driven first by statistical approximation error (a finite sample simply misses rare events, every generation, compounding), then by the model family's limited expressiveness and by the learning procedure itself.

    Then comes the control that most citations omit. Gerstgrasser et al. (2024) ran the same recursive-fitting experiment with one change — instead of replacing the real data with each generation's synthetic data, accumulate: keep the original corpus and add the synthetic data to it — on their own setup, small transformers trained on TinyStories rather than Shumailov's OPT-125m on wikitext2. Collapse stops. At iteration four, GPT-2 (9M) sits at 1.74 cross-entropy accumulating against 2.39 replacing; Llama-2 (125M) at 1.59 against 2.23. They also prove it for linear regression: under accumulation, test error has a finite upper bound no matter how many iterations you run, while under replacement it grows without bound. Collapse is not a property of synthetic data. It is a property of throwing the real data away.

    So: is an agent training on its own verified traces doing the thing that collapses? Related, but not identical, and the differences are all load-bearing.

    The collapse studiesA verified-trace loop
    What is keptevery generated sampleonly what the verifier passed
    The real datareplaced each generationkept and added to — if you keep the original mixture
    What is measuredperplexity: density over the whole distributiontask success, usually at k = 1
    Losing the tailpure loss — rare words, rare facts, rare stylesthe goal, when the tail is wrong answers; the same loss, when it is rare correct approaches
    Outcomeerror grows without bound across generationsbounded — as long as the filter's mistakes are uncorrelated with being right in an unusual way

    The first three rows are why you should not simply transplant the collapse result onto your pipeline and panic. A verifier is an enormously strong prior that the collapse setting does not have; keeping the original SFT mixture puts you in the accumulate regime that provably does not diverge; and losing density over wrong answers is the entire point of the exercise.

    The fourth row is why you should not relax either. Collapse is a loss of variance. Your loop is designed to remove variance. The only question separating the two is which variance it removes — and a filter whose false negatives land preferentially on unusual-but-correct work is, mechanically, running the collapse experiment on the part of the distribution you most needed to keep. That is the precise sense in which they are the same phenomenon: the same operation, with the sign decided entirely by the filter you wrote in section 07.

    The guard STaR already had

    STaR retrains from the original pretrained checkpoint at every outer iteration rather than continuing to train the model the last round produced — stated in the paper as an anti-overfitting measure. It is also, exactly, the accumulate regime applied to weights: each round's model is a fresh fit to a growing dataset, not a fit to a fit to a fit. It is one line of the training script, and it is the first thing reimplementations drop.

    The dials, and which curve each one bends

    Filter decisionWhat it controlsWhich curve moves
    Accept thresholdfalse-positive ratehow soon reward hacking becomes the cheapest strategy available
    What counts as the same answerfalse-negative rate on unusual workhow fast coverage falls — the curve nobody plots
    Keep the original mixture, or replace itaccumulate vs replacewhether error is bounded at all
    Refit from base, or from last rounderror carried between roundshow much round n inherits from round n−1
    Per-task duplicate capover-representation of the modeentropy, directly — an uncapped loop trains hardest on what it already does

    And measure accordingly. Pass@1 every round is not a measurement, it is a formality — concentration guarantees it. Hold out a fixed task set, sample it with the same sampler settings every round, and record three things: pass@k at the largest k you can afford, the count of distinct passing solutions per task, and token-level entropy on a fixed prompt set. Agent evals covers building that harness. A round that raises pass@1 and lowers pass@256 is not an improvement you have yet to measure. It is a trade, and you have already made it.

    Memory loops collapse too

    None of this is exclusive to weight updates. Reflexion writes lessons in natural language to an episodic buffer after a failed attempt and reruns the task — 91% pass@1 on HumanEval against 80% for the GPT-4 baseline of the day, with no gradient step anywhere. Voyager grows a library of executable skills in Minecraft and retrieves them later, reaching 3.3× more unique items, 2.3× longer exploration distances and key tech-tree milestones up to 15.3× faster than the prior state of the art, again with frozen weights.

    Both are self-improvement loops, and both have a concentration operator in them. A skill library that is only ever written from what already worked, and only ever read by nearest-neighbour retrieval, narrows the same way a policy does: the retrieved skill gets used, so it gets reinforced as the thing to retrieve, so the unusual skill three slots down is never tried. The failure is cheaper to fix — you can delete a memory, and you cannot un-train a weight — but it is the same failure, and it will not show up in pass@1 either.

    The key idea

    Self-improvement and diversity collapse are the same operation with the sign flipped. Every round of generate-verify-filter-refit narrows the distribution you sample from; the loop compounds for as long as what it narrows away is wrong, and turns on you the moment it starts narrowing away the rare things that were right. You cannot tell those apart from pass@1. You can from pass@k — so measure it every round, or you are flying the curve blind.

    12

    Knowing it worked

    ≈13 min1 runnable blockSkip if you already run a horizon fit and a blind win-rate set

    Build logWeek 9 · MondayWhat on-call finally accepted

    So we stopped scoring answers and started timing the humans: how long a ticket takes one of us, and where Mender's success rate crosses half of those, then eight in ten. Almost nothing failed on a step it could not do; it failed on the fortieth repeat of a step it gets right nine times in ten. That length, plus a blinded read — two finished tickets side by side, nobody told which was Mender's, on a set nobody shipping the loop owns — got it Saturday mornings.

    Which asks How do we prove this week beats last week

    A self-improvement loop rests on one claim: this week's agent is better than last week's. Everything above is machinery for making that claim true. This section is about making it checkable — because the benchmark most teams reach for cannot check it.

    The problem is length. A typical benchmark item — one function, one question, one tool call — takes a competent person a minute or two. A production agent runs for hours and passes through hundreds of such moments. Two points on a short benchmark is perfectly compatible with your four-hour run failing exactly as often as it did last month, because what kills long runs is not what short items measure. Agent evals covers building the set and keeping it honest; this section is about what to read off it once you have one.

    Figure · section 12A score is a point. A horizon is where the curve crosses.

    • The 50% horizon. The human task length at which the agent still finishes half the time. The friendly number.
    • The 80% horizon. Same curve, stricter line, and always much further left. This is the one that decides whether anyone lets it run unattended.

    how long the task takes a person →↑ success rate

    Measure the length of the job, not the score

    METR's Measuring AI Ability to Complete Long Software Tasks (Kwa et al., March 2025) proposes a metric with a unit you can feel: time. The construction matters more than the headline, because you can rebuild it on your own tasks in an afternoon.

    • Time the humans first. Every task carries a recorded human baseline — how long someone with the right skills actually took. METR's baseliners average around five years of relevant experience.
    • Run the agent many times per task, scoring each attempt with the task's own mechanical grader.
    • Fit one curve. A logistic regression predicting the agent's probability of success from the logarithm of the task's human duration.
    • Read off the crossings. The duration where that curve passes 50% is the 50%-time horizon. Where it passes 80% is the 80% horizon.

    The suite behind the original paper is 170 tasks — 97 from HCAST, 7 from RE-Bench, and 66 very short software actions — measured against more than 800 human baselines totalling 2,529 recorded hours of human work. Across frontier models from 2019 to early 2025, the fitted 50% horizon doubled roughly every 212 days, with a 95% bootstrapped confidence interval of 171 to 249 days — the figures as published in March 2025. The paper's current revision restates that fit as 207 days, 95% interval 166 to 240, so say which version you are quoting.

    METR now maintains a live version on a larger suite, Time Horizon 1.1, and publishes the fitted numbers as data. There the all-time doubling sits at 188 days, and the fit restricted to models from 2023 onward is 129 days — a 95% interval of 104 to 158 days, or roughly four months per doubling over the recent stretch. Hold the trend loosely. Hold the per-model numbers tightly, because those are what you compare your own agent against.

    The reason the unit matters is that anyone can check it against their own week. A benchmark score means nothing to the person deciding whether to leave an agent running overnight. The length of a job they recognise means a great deal.

    ModelReleased50% horizon80% horizonRatio
    GPT-4Mar 20234 min0.9 min4.5×
    Claude 3.7 SonnetFeb 202560 min12 min5.0×
    o3Apr 2025120 min30 min4.0×
    GPT-5Aug 2025203 min38 min5.3×
    Claude Opus 4.5Nov 2025293 min49 min5.9×
    Gemini 3.1 ProFeb 2026384 min90 min4.3×

    Source: METR Time Horizon 1.1, the published dataset behind metr.org/time-horizons. Point estimates in minutes of human task time, rounded for display; the ratios are computed from the unrounded values, so recomputing one from the two rounded cells beside it can land a tenth off. The intervals are wide — Claude Opus 4.5's 50% horizon carries a 95% interval of 162 to 624 minutes — and METR notes that measurements above 16 hours are unreliable on the current suite, so the very top of the chart should be read as "long", not as a number.

    The 80% column is the whole story

    Read that table across, not down. Every model's horizon collapses when you ask it to be right four times in five instead of one time in two. Across the fourteen models in that dataset released in 2025 or later, the median ratio between the two columns is 5.1×, and the narrowest gap in the set — o3's — is 4.0×. The paper lands in the same range: it reports 80% horizons running 4–6× shorter than 50% ones. Its worked example, from the original 170-task fit rather than the larger suite behind the table above, is Claude 3.7 Sonnet — a 59-minute 50% horizon against an 80% horizon of about 15 minutes.

    Say it in plain terms. Gemini 3.1 Pro's 50% horizon is a little over six hours of human work; its 80% horizon is an hour and a half. The model that can half-finish a working day can dependably finish a long meeting. The distance between those two sentences is your product.

    A horizon is also not a promise about every task of that length. METR's own FAQ — written against an earlier fit, which is why it gives GPT-5 a horizon of 2 hours 17 minutes rather than the 203 minutes in the table above — takes the tasks a human expert needs 90 minutes to three hours for, and reports that a GPT-5 agent solves about a third of them every single time, fails about a third every single time, and is erratic on the rest. The headline number is a blend of three different populations — solved, unsolved, and flaky — and only the third one is cheap for your loop to move.

    Then measure the work, not the task

    Time horizons are measured on tasks built to be scorable. The other honest instrument drops that comfort and asks professionals to judge finished deliverables. OpenAI's GDPval (released September 2025, posted to arXiv in October) is built from the real output of working professionals — 1,320 tasks spanning 44 occupations across the nine largest sectors of US GDP, written by people with an average of 14 years in the job, with 220 of them open-sourced as a gold subset. The deliverables are real artifacts: documents, spreadsheets, decks, plans. A gold-subset task takes an expert 404 minutes on average — just under seven hours.

    Two instruments, then, failing in opposite directions. One asks whether the agent can keep going long enough; the other asks whether what came out was any good. You want both, and a benchmark score on its own is neither.

    The grading is the part to copy. An expert in that occupation sees the request, the reference files, and two unlabelled deliverables, and ranks them without being told which came from a model. The paper collects three model samples per prompt and puts three graders on each, which is nine blinded comparisons per prompt per model across all 220 tasks. A single comparison takes a grader over an hour.

    The paper's headline: on the gold subset, 47.6% of Claude Opus 4.1's deliverables were graded better than (a win) or as good as (a tie) the human expert's — the best of the models it tested. Not "as good as a professional". It means that in blinded head-to-head against the specific professional who did that job, a grader called it for the model or called it even in a little under half the pairings.

    ModelWins or ties, GDPval gold subset (220 tasks)
    GPT-4o12.5%
    o4-mini29.1%
    o335.2%
    GPT-539.0%

    Source: the GDPval paper's speed-and-cost table (Table 2), whose win-rate column covers only the OpenAI models it priced — hence no Claude, Gemini or Grok row. GDPval counts a deliverable graded as good as the expert's alongside one graded better, which is the same measure as the 47.6% above, so the column and that figure are on one scale; Claude Opus 4.1's number is simply the one the paper states in prose rather than plotting. Read the ladder for its slope, not its absolute level; the leaderboard has moved since.

    Why long runs fall apart: multiplication

    Both instruments are measuring the same underlying thing, and it is not intelligence. It is compounding.

    P(run succeeds) = pn  ·  p = per-step success  ·  n = steps in the run

    One number raised to a power. That is the entire reason a capable model fails a long job.

    Put real numbers in it. At 99% per step — a rate most people would call excellent — a 50-step run finishes 60.5% of the time and a 100-step run finishes 36.6%. At 95% per step, the 50-step run finishes 7.7% of the time. Run it backwards and it gets worse: to land 90% of 50-step runs you need 99.79% per step, which is one slip per 475 steps. That is not a number you reach by tuning a prompt.

    Independence is the assumption carrying that formula, and it is wrong in both directions. Real agents retry, so some failures cost tokens instead of the run. Real failures are also correlated — one bad plan poisons every step after it — so one mistake can be worth many. Treat pn as the shape of the problem, never as a forecast of your system. The block below is that arithmetic and nothing else: no model, no sampling, no training run.

    Reliability is the binding constraint, not capability

    Look at what the arithmetic implies about where to point your loop. Going from 90% to 95% per step lifts a 50-step run from 0.5% to 7.7%. Going from 95% to 99% lifts it from 7.7% to 60.5%. Going from 99% to 99.9% lifts it to 95.1%. Every nine you add is worth more than the one before it, and none of those moves teaches the model anything it could not already do. The model can do the step. It does it 95 times in 100. Your loop's entire job is the other 5.

    That is why a self-improvement loop aimed at the lowest benchmark score usually disappoints, and one aimed at the flakiest step usually does not. Find the step that fails most often in your traces, and spend memory, context and weights on that one step until it stops. Then find the next one. Long-horizon agents is the companion piece on what else breaks when a run lasts hours.

    Build your own two instruments

    You cannot run METR's suite or GDPval. You can run both methods, cheaply, on your own work:

    • Your horizon. Take 50 or more real tasks off your actual queue. Record how long a competent person takes on each. Run the agent five to ten times per task. Fit success against the logarithm of that duration and report two numbers — the 50% and the 80% crossing. Track the 80% one. It is the number your users experience.
    • Your win rate. Freeze 50 to 200 real tasks. For each, keep one deliverable made by your best human. Have someone who does the job rank the two, blind, against a rubric written before the model ran. Your improvement claim becomes a win-or-tie rate on the frozen set that moved — a sentence your loop cannot manufacture by learning to please your verifier.

    One caveat both sources state about themselves, which applies double to you. A METR task is written to stand on its own, so that a baseliner and an agent start from the same place and neither is punished for not knowing the house rules. Your queue is the opposite: half of what makes a ticket easy is knowing which service owns it and who broke it last time. METR's own advice is to read a two-hour task as what someone with "low or no prior context (like a new hire or freelance contractor) could complete in 2 hours". Your production tasks come loaded with context, so your horizon will not match anyone's published figure — and it does not need to. The number exists to compare your agent to your agent, last month.

    The key idea

    A score tells you whether the model got better. A horizon at 80% tells you whether the run got better, and that is the only improvement a user can feel. Measure the length of the job you can finish reliably, and measure it against a human deliverable that someone graded blind.

    RUN IT YOURSELF

    Reliability compounds: the arithmetic behind a horizon

    This is arithmetic, not a simulation. No model runs, nothing is sampled, and nothing here is a training run — it is one number raised to a power, printed three ways. The first table shows what a per-step success rate becomes over 10, 50 and 100 steps. The second inverts it: what each step must hit for the whole run to land 90% of the time. The third shows why every extra nine costs more than the one before it. The assumption doing all the work is that steps fail independently, which real agents do not — they retry, and their mistakes correlate, so the true number can land on either side of this. Read it as the shape of the problem, not a forecast of your system. Change p and the step counts to your own and run it again.

    HOW TO READ THE CODE — 5 IDEAS
    1. end_to_end is the whole model in one line — a per-step rate raised to the number of steps — and every table below is that line read from a different direction.
    2. step_rate_needed is the same equation solved backwards, and it is the number you can actually act on: a target for one step, not a wish for the run.
    3. In the second block, 1 / (1 - need) converts that rate into "one slip per N steps" — the form of the number that survives being said out loud in a design review.
    4. The third loop walks the pairs a and b across a fixed 50-step run, so watch the size of each jump grow as the nines pile up, not shrink.
    5. Only two things are worth editing: the tuple (0.999, 0.99, 0.98, 0.95, 0.90) and the step counts — put your own traced per-step rate in and the rest re-derives.
    CPython · WebAssembly
    13

    Ship it

    ≈12 min2 runnable blocksStart here if you have ten minutes

    Build logWeek 9 · WednesdayWhat was not the loop

    We turned the loop off on the Wednesday, not because anything burned, but because we could not name the last three things it had changed. Then we wrote down the six conditions that would turn it off again, and stuck them where on-call could see them. What was left was the part worth keeping: one written sentence describing a finished ticket, a horizon measurement, a verifier, a hundred runs hand-labelled to find how often that verifier was wrong, and a trace we could replay six weeks later.

    Which asks What do we build first, and in what order

    Build logWeek 9 · FridayWhat we ended up running

    What runs now is smaller than the diagram we drew in week one. Mender takes a ticket off one queue and makes up to eight attempts, down from the twenty we started with, and every attempt runs the service's suite in a clean container. A patch is a candidate only if the suite goes green and the ticket's own named test goes green with it. For the third of the queue with no test, the verdict comes from the environment instead: the service restarts, the log line appears, the alert clears. Nothing is more independent, and nothing covers less — a quiet channel is not proof. Those tickets are allowed to produce a PR and never a training example. The search we spent ten days on is off; nothing it found stayed found.

    The traces that survive are filtered before anything is refit: correct, deduplicated on what the patch does rather than how it reads, dropped if every attempt passed or none did, capped at three per ticket, and only then cut to budget. Most weeks nothing is refit at all. The lesson goes in the prompt, or the tool, or the retrieval store, and four times in five it stays there. Three things get read every round that are not on the dashboard: how many distinct passing patches exist for a ticket, how long a ticket takes a person before Mender stops finishing eight in ten of them, and a blinded read of two finished tickets — the one allowed to disagree with the score. Six written conditions turn the loop off, and the frozen set belongs to someone whose job is not shipping the loop.

    What we would still not claim is the important part. We cannot say this makes Mender better, only that it does more of what our suite likes, and so far we have not caught the two coming apart.

    Here is the order to do it in. It is deliberately front-loaded with measurement, because every failure mode above — reward hacking, drift, a filter that lets its own mistakes through — is invisible without it, and every one of them is cheap to prevent and expensive to unwind. Each step names the action and points back at the section that argues for it.

    1. Name one task family and one deliverable.

      Not "the agent". One repeating job — triage this ticket type, draft this document, fix this class of bug — and one written sentence describing what a finished one looks like, specific enough that a stranger could apply it. If you cannot write that sentence, nothing downstream has a target.

    2. Measure a horizon before you measure a score.

      Pull 50 or more real tasks off the queue, record how long a competent person takes on each, run the agent five to ten times per task, and fit success against the logarithm of that duration. Report the 50% and the 80% crossing. This pair is your baseline and your only defence against shipping a change that helped the benchmark and hurt the product.

      Why: 12 · Knowing it worked
    3. Build the verifier before you build the loop.

      No trace may enter any update until something mechanical can say correct or not: tests that run, a schema that validates, a diff against a known answer, a rubric grader. If you cannot verify the task, you cannot safely improve on it — stop here and change the task.

      Why: 05 · Verification, 04 · Selection is the bottleneck
    4. Measure the verifier itself.

      Hand-label 100 runs and score your verifier against your labels. Its false-positive rate — how often it blesses a wrong run — is the hard ceiling on everything downstream, because those are exactly the traces that get baked in. A verifier you have not measured is a number you have not earned.

      Why: 05 · Verification
    5. Log complete, replayable traces from day one.

      Every prompt, tool call, tool result, final output, verifier verdict, and the version of the model, the prompt and the tools that produced it. A trace you cannot replay is an anecdote. This is also the only artifact that makes a post-mortem possible six weeks later.

      Why: 07 · What you keep
    6. Find where the runs actually die.

      Bucket failures by step index and by cause, and sort. The distribution is almost never flat — one or two steps carry most of the loss. Fix the tallest bar. Improvement spent anywhere else is invisible in the compounding arithmetic.

      Why: 12 · Knowing it worked
    7. Reach for the cheapest reversible fix first.

      In order: prompt and tool descriptions, then scaffold and retry logic, then memory, then weights. Each rung costs more and reverses harder than the one below it. Most teams skip to the top rung and pay for a training run to fix a tool description.

      Why: 08 · Where the improvement lands
    8. Land recurring specifics in memory.

      This customer's preference, this repo's build command, this API's quirk, this failure and its remedy. Anything narrow, anything you might have to revoke, anything you are not yet sure about. Write it, retrieve it, and keep the delete button one click away.

      Why: 08 · Where the improvement lands
    9. Land broad, verified behaviour in weights — and only then.

      The signal that weights are the right home is memory repeating itself: the same kind of note written a hundred times for a hundred different tasks. Train on filtered-correct traces only, at a bar strictly higher than memory's, because this is the rung you cannot climb back down.

      Why: 08 · Where the improvement lands, 09 · Closing the loop
    10. Freeze an eval set the loop can never touch.

      Held out of training, held out of memory, held out of prompt examples, and owned by someone whose job is not shipping the loop. The day it leaks, your only independent measurement quietly becomes a training metric and you will not be told.

      Why: 12 · Knowing it worked
    11. Canary with enough runs to mean something.

      Route a slice of live traffic to the candidate, keep both versions loadable, and size the sample to the regression you care about rather than to your patience. Ten runs is not a canary: a model that has silently dropped from 90% to 80% still clears a 90% bar 37.6% of the time. Catching a five-point drop nineteen times out of twenty takes about 118 runs. The sizing block below does that arithmetic for your own numbers.

      Why: 12 · Knowing it worked
    12. Monitor four numbers, not one.

      Verifier pass rate on live traffic. The 80% horizon on the frozen set. Memory store size against memory hit rate. And a weekly human-graded win-or-tie sample. The fourth exists precisely because the first three can all improve while the product gets worse.

      Why: 12 · Knowing it worked
    Stop conditions — turn the loop off
    • Verifier pass rate climbs while the human-graded sample does not. That gap is the definition of reward hacking. The agent found a hole in your signal, and every update from now on widens it.
    • The frozen set drops at all. Not "drops significantly" — at all. A frozen set that moves down after an update is the one measurement you cannot explain away.
    • Output diversity collapses. Same phrasing, same tool order, same structure on tasks that used to produce different work. A loop trained on its own output narrows toward its own habits, and the narrowing shows up in the writing before it shows up in the score.
    • The memory store grows faster than its hit rate. You are accumulating, not learning. Retrieval quality, not store size, is what makes memory worth having.
    • You cannot name the last three things the loop changed, or roll any of them back. An improvement you cannot undo is a deployment you cannot defend.
    • A weight update whose training set you cannot reproduce. If you cannot re-derive exactly which traces went in, you cannot audit what the model learned, and the permanence in section 08 is now permanence without evidence.
    The key idea

    The loop is not the hard part — you can build the loop in a week. The hard part is the pair of numbers that tells you it is working and the pair of rails that stops it when it is not. Build the measurement first, and the loop becomes a small thing you bolt onto it.

    RUN IT YOURSELF

    How many runs before a canary means anything

    Also pure arithmetic — the exact binomial, computed term by term, with nothing sampled and no model involved. The question is how many runs a canary needs before passing it is evidence. The right-hand column is the one that stings: with ten runs, a candidate that has silently regressed from 90% to 80% still clears a 90% bar more than a third of the time. The middle column is the honest cost of the bar itself — set at the live model's own rate, it drifts toward a coin flip as the sample grows, which is why the bar belongs below the live rate and the sample size does the real work. The last block searches for the fewest runs that catch the drop you actually care about. Put your own pass rate and your own tolerance in and run it again.

    HOW TO READ THE CODE — 5 IDEAS
    1. p_at_least is an exact binomial tail — it sums comb(n, i) terms from k up to n — so nothing in this block is sampled, estimated or simulated.
    2. Read the table's last two columns as one thing: the same bar k applied to LIVE and to BAD, because the gap between those two percentages is all the power your canary has.
    3. LIVE and BAD are the only inputs worth touching — set BAD to the regression you would be angry to miss, not to the worst case you can imagine.
    4. smallest_n walks n upward until a regressed model's chance of passing falls to false_ok, so the first n it returns is your minimum canary size, not a suggestion.
    5. Notice k = ceil(live * n) rises with n, holding the bar at a constant 90%: it is the sample size, never the threshold, that buys the confidence.
    CPython · WebAssembly
    RUN IT YOURSELF

    Round 6 is done. Do you ship it?

    Nothing here runs a model. It is the four stop conditions above, written as arithmetic over two rounds of measurements, and four rounds fed through them. Run it as written, then read round b twice: every number a weekly update would quote went up, and it is the round you must not ship. Then change a single field at a time and watch which condition fires — that is the whole instrument. The remaining two conditions, plus the canary-size guard, are yours to write in the challenge.

    HOW TO READ THE CODE — 5 IDEAS
    1. Two rounds in, one verdict out. Every condition compares curr against prev, because none of these numbers means anything as a single reading.
    2. Divergence is the first and the most important: your verifier moved and the human sample did not follow. The threshold (d_human <= d_verifier / 2) is a choice, not a law — pick yours deliberately and write down why.
    3. The frozen set has no tolerance at all. A set nothing has trained on that moves down is the one measurement you cannot argue with, so it is compared with a bare <.
    4. Narrowing counts distinct passing solutions per task, not output length or token entropy. It is the cheap proxy for “can this model still reach what it used to”, and it falls while pass@1 climbs.
    5. Reversibility is not a measurement, it is a property of how you shipped. It is in the same list because it gets more expensive to fix every round you leave it.
    CPython · WebAssembly
    14

    Where to go next

    ≈5 minSkip unless you are stuck

    Build logWeek 9 · ThursdayThe week we read instead

    With the loop off we read for three days. The list of things we got wrong is shorter than the list of things we had to go and read, and every item on it was a stuck point rather than a topic: the verifier, the filter, the eval set, where a fix should land. One of those pages is why our dedupe now compares what a patch does rather than how it reads. We opened exactly one at a time.

    Which asks Which page owns what we are stuck on

    Every page below owns a piece this one deliberately did not re-teach. They are grouped by the question they answer, so you can go to the one you are actually stuck on.

    Mechanisms

    How an agent gets better

    • Test-time computeBuy quality at inference — more thinking, more samples, more search. The improvement that needs no training run, and the one to exhaust first.
    • Context engineeringMost "the model can't do this" is "the model wasn't told this". The cheapest, most reversible rung on the ladder in section 13.
    • Agent memoryThe reversible axis in full: what to store, how to retrieve it, when to forget, and why the hit rate matters more than the size.
    • RLVR & verifiable rewardsWhen the verifier is a program, you can train straight against it. The most disciplined version of "reinforce only what you can confirm".
    • Synthetic dataHow to manufacture the traces a loop needs when production traffic is too thin — and how to stop the loop from eating its own output.
    • Fine-tuningWhat actually changes in the weights, which method to pick, what it costs, and what you can no longer take back afterwards.
    Papers

    Where each idea was first shown to work

    • STaRThe bootstrap itself: generate reasoning, keep what reaches the right answer, fine-tune, repeat. The seed of every self-training loop.
    • Self-InstructThe same trick pointed at instructions — a model writing the training set it will then learn from.
    • ReflexionImprovement without touching a weight: write the lesson from the failure into memory, then try again.
    • VoyagerThe skill library. Successful code becomes a callable skill, and the agent keeps growing the set it can call.
    • Let's Verify Step by StepGrade the steps, not just the answer. The result behind process-level verifiers, which catch a bad step inside a run that happened to end on the right answer.
    • Self-ConsistencySample many chains, take the majority answer. It buys reliability with inference alone, and doubles as a filter when you have no program to check against.
    • Tree of ThoughtsSearch over reasoning states instead of committing to one straight line.
    • ReActInterleaving reasoning and acting — think, call a tool, read the result, think again. Where the tool-using agent loop comes from.
    • AlphaCodeWhat sample-and-filter looks like at industrial scale, when the filter is real test execution rather than a model's opinion.
    • PPOThe policy-gradient algorithm most of RLHF was built on, and the baseline everything newer is measured against.
    • GRPOThe leaner variant that dropped the value model. What most verifiable-reward training runs actually use now.
    • DPOSkip the reward model and optimise preferences directly — fewer moving parts, fewer places for the signal to rot.
    Evaluation

    Knowing it worked

    • Agent evalsBuilding the set: what goes in, what stays out, how many runs per task, and the traps that make an eval agree with you.
    • Evals in CIMake the eval a gate rather than a ritual. Run it on every change, block on regression, and the frozen set defends itself.
    Substrate

    What the loop runs on

    • Loop engineeringThe runtime the agent lives in: retries, budgets, termination, and the plumbing that decides whether a step failure ends the run.
    • Long-horizon agentsEverything else that breaks when a run lasts hours — the companion to the compounding arithmetic in section 12.
    • RL environments engineeringBuilding the environment and reward your loop optimises against. This is where reward hacking is born, and where it is prevented.
    • Agent patternsThe reusable shapes — planner and executor, critic, router — that you will be pointing all of this improvement at.

    Questions people actually ask

    What is a self-improving agent?

    An agent that gets better from its own work rather than from a new model release. It runs real tasks, a verifier decides which attempts were actually correct, and the correct ones are folded back in — either written to a memory store for exact recall, or used to update weights so the improvement generalises to tasks it has not seen. Everything difficult about it lives in the word correct.

    Do I need a GPU?

    For most of this, no. The memory axis is a database and a retrieval step — no training hardware anywhere. Prompt, tool and scaffold fixes are free. Test-time compute buys real quality with inference alone. You need training hardware only for the weights axis, and by the ordering in section 13 that is the last rung, not the first. A loop that never trains anything can still be a genuine self-improving agent.

    Where do I start on Monday?

    Steps 1 to 5 of the checklist, in order, and nothing else. Name one task family, measure a 50% and an 80% horizon on real tasks, build a verifier, measure that verifier against 100 hand-labelled runs, and log full traces. That is a week or two of work and it produces no self-improvement at all — which is the point. Teams that skip it spend the following quarter unable to tell whether their loop is helping.

    How do I know the loop is degrading?

    The tell is divergence: your verifier's pass rate climbs while a human-graded sample of the same work does not. That gap is reward hacking, measured. The other four signals are a frozen eval set that drops at all, output diversity collapsing into one house style, a memory store growing faster than its hit rate, and any update you cannot roll back or reproduce. All five appear among the six stop conditions — they are stop conditions rather than alerts because each one gets harder to unwind every day the loop keeps running.

    Memory or weights — which first?

    Memory, nearly always. It lands in seconds, it generalises to nothing but it also risks nothing, and you can delete a bad one the moment you notice. The signal that you have outgrown it is repetition: when memory is storing the hundredth variation of the same lesson, that lesson is a habit, and a habit belongs in the weights. Match permanence to confidence.

    Can a self-improvement loop run away?

    Not in the sense the phrase usually implies. The loops described here are bounded by their verifier — an agent cannot learn to do something its filter cannot recognise as correct, so the verifier's quality is the ceiling, and the ceiling does not move on its own. The realistic failure is the opposite of runaway: a loop that quietly optimises the gap between your verifier and your users, and gets measurably better while getting worse.

    Finished this one? 0 / 208 Handbooks done

    Explore the topic

    See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.