Self-improving
agents.
The dream is an agent that gets better the more you use it — one that runs in production, notices what worked, and folds it back in. It is real, and none of it is magic. Underneath is a loop with four moves: generate, verify, filter, refit. A model can usually already produce the right answer somewhere in its output; what you are missing is something that can recognise it, and somewhere to put what you learn — the prompt, a memory store, a tool, the scaffold, the verifier, or the weights. Choose the wrong one and you pay for it slowly. Get the verifier wrong and the loop teaches itself the wrong thing, confidently, every round.
How this page runs
Five acts, 14 sections, ~2 hours end to end — most people should read three. The build log beside them is invented: a composite team, ticket counts and weeks not measured.
About the box that runs down the page. Set apart beside every section is a build log: one team, nine weeks, one agent called Mender pointed at one repo's ticket queue. It is a composite scenario, not a case study — its ticket counts, weeks and percentages are invented and only ever describe that imaginary team. Every measured figure on this page lives in the prose instead, named, sourced, and carrying the condition it was measured under. Each section below also says when you can skip it.
- generate
Draw many attempts instead of one. The answer is usually already in there.
- verify
Decide which attempts were good. Everything downstream inherits this judgement.
- filter
Choose what is worth keeping, and where to put it.
- refit
Fold it back in — then prove it helped, and know when to stop.
Each section owns one arc of this ring. If you only want the arc you are stuck on, follow its numbers.
Stuck on one thing? Start there.
- It works sometimes and I cannot tell whether that counts as working01
- I can generate a right answer and I cannot pick it out of the pile04
- My checker keeps passing things that are wrong05
- Half my tasks have nothing that can grade them06
- I trained on my own traces and it got narrower, not better10
- The score went up and nobody will let it near production12
-
Act I
Asking more than once
≈22 min01–03 · Tell coverage from shippable accuracy, price five cheap attempts, forecast from your histogram
-
01The answer is already in there
Tell coverage apart from the accuracy you can ship, and cost five attempts against one
≈7 mininteractive
Skip if you already sample k>1 and report pass@k
-
02Measuring it honestly
Compute unbiased pass@k from one sample pool, and state its n and its verifier
≈6 min1 runnable
Skip if your pass@k already comes from one pool with n stated
-
03Why the curve is straight
Forecast sampling returns from your per-problem histogram instead of the mean
≈9 min1 runnable
Skip if you are not paying for repeated sampling
-
01The answer is already in there
-
Act II
The picking problem
≈40 min04–06 · Pick a selector for your output shape, build the verifier behind it, learn how often it lies
-
04Selection is the bottleneck
Choose a selector by output shape and cap N before the curve bends downward
≈11 min1 runnable
Skip if a program, not a model, already picks your winner
-
05Verification
Grade answers or steps, ensemble weak checkers, and measure your false-positive rate
≈16 min1 runnableinteractive
Read this one even if you skip everything else
-
06Feedback that isn't a verifier
Rank feedback sources by independence and gate weight updates on the best one
≈13 min1 runnable
Skip if every task you run has a test
-
04Selection is the bottleneck
-
Act III
What you feed it
≈36 min07–09 · Write the five-predicate filter, route each lesson to a substrate, audit the reward first
-
07What you keep
Write a selection policy that filters, bands, dedupes, caps, then cuts to budget
≈10 min1 runnable
Skip if you never refit
-
08Where the improvement lands
Route a lesson to one of six substrates by generality, permanence and blast radius
≈12 min1 runnable
Skip if you have already decided it goes in the prompt
-
09Closing the loop with RL
See why a correctness filter is a policy gradient, pick PPO, GRPO or DPO, audit the reward
≈14 min1 runnable
Skip if you are not training weights
-
07What you keep
-
Act IV
Narrowing and what you report
≈43 min10–12 · Watch entropy before pass@1, buy reach with search not weights, report a horizon and a blind win rate
-
10The loop that collapses
Instrument the loop so diversity collapse shows while rolling back still costs one batch
≈20 min1 runnableinteractive
Skip if you are not training weights
-
11Search as self-improvement
Choose a flat bank or a tree on latency and evaluator quality, then where the win lands
≈10 min1 runnable
Skip if your budget is one attempt per task
-
12Knowing it worked
Build a time-horizon fit and a blind win-rate set on your own tasks
≈13 min1 runnable
Skip if you already run a horizon fit and a blind win-rate set
-
10The loop that collapses
-
Act V
Standing it up
≈17 min13–14 · Sequence the rollout, name the conditions that stop the loop, open the page that owns your blocker
-
13Ship it
Sequence the rollout, size a canary, and name the conditions that stop the loop
≈12 min2 runnable
Start here if you have ten minutes
-
14Where to go next
Open the one page that owns your blocker, and take the six answers you will be asked for
≈5 min
Skip unless you are stuck
-
13Ship it
The answer is already in there
≈7 mininteractiveSkip if you already sample k>1 and report pass@k
Build logWeek 1 · TuesdayThe queue nobody wanted
We pointed Mender at the dullest queue we own: failing integration tests and bug tickets on one Python service. One attempt per ticket, it closed almost nothing, and we had written it off by lunch. Someone reran the same tickets overnight at twenty attempts each, graded only by the one test named on the ticket, and we read the log at nine to find a handful had gone green. Nothing had been learned between tries.
Which asks Is it incapable, or did we only ask once
You ask the model once, you get one answer, and you judge the model on it. That is a strange basis for a judgement, because a model is a sampler. Ask the same question again and something else comes out. Ask it 250 times and the question quietly changes from can this model do the task to does this model ever do the task.
Those two questions have very different answers. In Large Language Monkeys: Scaling Inference Compute with Repeated Sampling (Brown et al., 2024), the authors draw up to 10,000 independent samples per problem across five tasks and track coverage: a problem counts as covered the moment one of its samples is correct, no matter which one and no matter what the other 9,999 said. Coverage climbs, smoothly, over four orders of magnitude of sample budget.
The result that should reorganise how you think about agents is on SWE-bench Lite, where a problem is a real GitHub issue and one "sample" is an entire multi-turn trajectory through a codebase. DeepSeek-Coder-V2-Instruct, driving the open-source Moatless Tools agent framework, resolves 15.9% of issues on a single attempt. Give it 250 independent attempts and 56% of issues are resolved by at least one of them. The single-attempt state of the art at the time — CodeStory Aide, running a mixture of GPT-4o and Claude 3.5 Sonnet — was 43%. Carry one condition with those numbers: every attempt, the single one included, was drawn at temperature 1.6, chosen by a sweep over 1.0, 1.4, 1.6 and 1.8 on 50 problems. That is a setting picked for the many-attempts regime, not DeepSeek's best single-shot configuration.
Brown et al. 2024, SWE-bench Lite via Moatless Tools · attempts drawn independently at temperature 1.6, no feedback between them
Read that slowly, because it is the premise the rest of this page is built on: a weaker model, sampled more, found fixes a stronger model did not find on its first try. The capability was already in the weights. What was missing was attempts, and a temperature loose enough to make them differ.
It is also cheaper than it sounds. At the API prices current when the paper was written, a much smaller version of the same trade — five DeepSeek attempts, with an oracle picking the winning patch, against one attempt each from GPT-4o and Claude 3.5 Sonnet — looks like this:
| Setup (Moatless Tools, SWE-bench Lite) | Cost / attempt | Attempts | Issues solved | Total |
|---|---|---|---|---|
| DeepSeek-Coder-V2-Instruct | $0.0072 | 5 | 29.62% | $10.80 |
| GPT-4o | $0.13 | 1 | 24.00% | $39.00 |
| Claude 3.5 Sonnet | $0.17 | 1 | 26.70% | $51.00 |
Brown et al. 2024, Table 1 · five cheap attempts solve more issues than one expensive attempt, at 3.6x less than GPT-4o and 4.7x less than Claude 3.5 Sonnet
And it is not a frontier-model effect. Gemma-2B on CodeContests goes from a 0.02% solve rate at one sample to 7.1% at 10,000 — more than a 300x increase. Pythia-160M goes from 0.27% to 57% on a 128-problem MATH subset — coverage under an oracle answer checker rather than a solve rate, for the reason the next paragraph gives. Small models contain far more correct answers than they can produce on demand.
Say that without any of the numbers attached, because it is the premise the whole page rests on: a model’s first answer is not the edge of what it knows. Everything that follows is about the distance between those two things, and who closes it.
Now the part that decides whether you build this well or badly. Coverage is not accuracy. Coverage counts a problem as solved if any of your samples was correct, which quietly assumes something you usually do not have: an oracle that can look at 10,000 candidate answers and point to a right one.
Where that oracle is real, coverage converts directly into performance. A Lean proof checker either accepts the proof or it does not. A test suite either goes green or it does not. That is exactly why SWE-bench Lite works as a demonstration — the harness grades every attempt with real tests. But those tests are withheld from the agent, so the 56% is what an oracle choosing among 250 patches scores, not what a shippable selector scores, and Brown et al. found 11.3% of Lite problems have flaky test suites on top of that. A checker that can actually decide is also why verifiable-reward domains keep producing the loudest results in this literature.
Where the oracle is not real, the gap is enormous, and Brown et al. measured it. On 128 MATH problems with Llama-3-8B-Instruct at 10,000 samples each, coverage rises from 79.8% at 100 samples to 95.3% at 10,000. Over that same range, the methods you could actually ship — majority voting and reward-model selection, using ArmoRM-Llama3-8B-v0.1, then the highest-scoring open-weight model for reasoning on RewardBench — move from 38.7% to 39.8%. Every selection method saturates before about 100 samples while coverage keeps climbing. At 10,000 samples the best available selector is leaving more than 55 points of already-generated, already-correct answers on the floor.
Stop on that for a moment. The samples got better and the answer you would have shipped did not. Buying more attempts only helps if something can tell you which attempt to keep — and that something is now the expensive part of the system, not the model.
Sampling more finds more answers. Picking the right one is a different problem.
Brown et al. sampled one model over and over and tracked coverage — the share of problems where at least one sample is correct. On SWE-bench Lite, DeepSeek-Coder-V2-Instruct went from 15.9% with a single sample to 56% with 250 (Large Language Monkeys, 2024). On MATH with Llama-3-8B-Instruct, coverage ran from 79.8% at 100 samples to 95.3% at 10,000 — while majority voting and reward-model selection plateaued around 100. Coverage is what the samples contain. Selection is what you can ship. Below, four rules read the same bank of candidates.
% of problems solved · samples per problem (k), log scale
Drag a slider, or drag across the chart, to begin.
How often the scorer ranks a correct candidate above a wrong one. 50% is a coin flip.
Share of wrong samples landing on one shared wrong answer — a systematic misreading, not noise.
Samples drawn per problem. You can also drag or tap across the chart itself.
Simulated, and here is exactly what was swapped. No model was run. The bank is 192 synthetic problems whose single-sample success rates are evenly spaced quantiles of a Kumaraswamy(α = 0.35, β = 1.6) distribution — the heavy-tailed family Schaeffer et al. fit to real per-problem success rates in How Do Large Language Monkeys Get Their Power (Laws)? (2025), where the weight of that left tail is what turns per-problem exponential decay into aggregate power-law coverage. Coverage is then exact: the bank average of 1 − (1 − p)k. The two deployable curves are models of a selection rule, not measurements of one — majority vote over an answer distribution with six recurring distractors plus one attractor, and best-of-k under a Gumbel-noise scorer whose pairwise accuracy is the slider (that noise makes best-of-k a softmax over the pool, so the curve is closed-form rather than sampled). Every number here is computed, not drawn at random, so your chart and mine match exactly. The shape is real; the scale is a choice. The measured figures on this page are the ones named in the prose, never the ones on these axes.
Brown et al. 2024 · Llama-3-8B-Instruct, the paper's 128-problem MATH subset, up to 10,000 samples per problem · the plotted endpoints are the numbers the paper reports; anything drawn between them is interpolation, not a measurement
The reason is mechanical rather than mysterious. On some of those problems the correct answer is produced by 1% of samples or fewer. A majority vote over 10,000 samples cannot surface a 1% answer — the more you sample, the more precisely the vote converges on the popular wrong one. Sampling finds needles; voting measures the haystack. Closing that gap is what learned verifiers are for, and why their failure modes matter so much here.
This gap is not new, and the first clean receipt for it is five years old. In the Codex paper (Chen et al., 2021), Codex-S solves 37.7% of HumanEval problems with one sample at temperature 0.8. Given 100 samples per problem, picking the candidate with the highest mean log-probability solves 44.5%; picking the candidate that passes the unit tests solves 77.5%. Same 100 samples, same model — 33 points between the best ranker and the oracle. AlphaCode is the same shape at production scale, with the selector built out in public: it draws an enormous pool of candidate programs, throws away every one that fails the example tests printed in the problem statement (roughly 99% of the pool), runs what survives on generated inputs, and groups the programs that produce identical outputs — so its ten allowed submissions go to ten different behaviours rather than ten spellings of one. It reached an average rank in the top 54.3% across simulated Codeforces competitions with more than 5,000 participants.
Repeated sampling changes the question from what a model does to what a model can do, and coverage is the ceiling that question sets. A selector is what you actually ship. Every technique on the rest of this page is an attempt to close the distance between the two — by building a verifier good enough to convert coverage into accuracy, or by folding what you found back into the weights so it arrives in the first sample next time.
Measuring it honestly
≈6 min1 runnable blockSkip if your pass@k already comes from one pool with n stated
Build logWeek 1 · FridayThe number in the update
We put a number in the weekly update: the share of the queue Mender fixes. It came out of the overnight pool, where a ticket counted if any of twenty attempts went green, and we had labelled it as what one attempt does. Someone asked what n was, and the number did not have an n. Nobody had lied; we had reported a different quantity from the one we named.
Which asks What number do we actually put in the update
You cannot steer a loop on a number you are measuring wrong, and coverage is easy to measure wrong in both directions.
Start with what you actually want to know. pass@k is the probability that if you drew k samples for a problem, at least one would be correct. It is a property of the model-and-problem pair, not of one particular run of your harness.
The obvious way to estimate it: draw k samples, check whether any passed, write down 1 or 0, average over problems. That is unbiased and nearly useless. For a single problem it is one Bernoulli draw, so its variance is about as large as the thing you are measuring — the high variance that made Chen et al. reach for something better. It is also wasteful — pass@1, pass@10 and pass@100 would each need their own experiment.
The version that actually ships is worse. Generate a big pool of n samples, notice that one of them worked, and report that as pass@k for whatever k you had in mind. That is pass@n wearing a smaller label. With 200 samples of which 2 are correct, "did any of them work" reports 1.0 while the true pass@10 is 0.098 — a factor of ten, in the direction that flatters you.
The Codex paper (Chen et al., 2021), which introduced HumanEval, fixed both problems with one formula. Generate n ≥ k samples once, count the c that pass, then ask a combinatorial question of the pool you already have:
C(n−c, k) / C(n, k) is the probability that a k-subset drawn from your n samples without replacement contains none of the c correct ones. One minus that is the probability it contains at least one. The Codex paper used n = 200 and k ≤ 100; Brown et al. reuse this same estimator to keep variance down at 10,000 samples.
The economy is the point. One pool of n samples gives an unbiased estimate at every k ≤ n, so a single expensive run draws the whole curve instead of one point on it.
There is a near-miss worth naming, because it looks right and it is not. Take p̂ = c/n and write pass@k ≈ 1 − (1 − p̂)k. That treats your k draws as being made with replacement from the pool, and Chen et al. show in their appendix that it is a consistent underestimate whose gap does not fully close even when n > 5k. It is wrong in the modest direction — it makes sampling look less useful than it is — which is exactly why nobody catches it.
n = 200 drawnc = 2 correct
Both readings use the same 200 samples. The gap is a factor of ten, and it flatters you. Figures from the prose above.
Three edge cases carry all the meaning, and the code below is where you write them:
- k > n. Undefined; the function should refuse rather than guess. You cannot estimate pass@1000 from 200 samples no matter how straight the curve looks.
- c = 0. The estimator returns exactly 0 at every k. Read it as "not yet", never as "impossible" — 200 consecutive failures only put a rough 95% upper bound of p < 3/200 on the true rate, and section 03 is about how much that tail matters.
- c = n. Returns exactly 1, as does any case where n − c < k: if fewer than k of your samples are wrong, no k-subset can be all wrong. Chen et al.'s implementation short-circuits that case before the product runs — not because anything would divide by zero, but because the product would otherwise grind through factors like 1 − k/j that go negative.
Then the caveat that outranks the arithmetic: c is whatever your checker counted. Brown et al. ran CodeContests' own test suites against the reference Python3 solutions the dataset ships, and found that of the 122 test-set problems with such solutions, 35 have correct solutions that fail their own tests. Those are false negatives — coverage you measured as absent and in fact had. The mirror failure, a wrong program that passes a thin test suite, inflates the same number instead. What you are reporting is never pass@k; it is passes-my-verifier@k, which is why how you build the evaluator and where you run it are load-bearing decisions rather than plumbing.
Write the unbiased pass@k estimator
This is the formula from the Codex paper, in about ten lines. Nothing is sampled and no model runs — you hand it a hypothetical pool (n samples, c of them correct) and it answers a combinatorial question about that pool. Run it as written, then change n and c to your own numbers. Watch what a 1% success rate looks like at k = 100, watch the edge cases return exactly 0 and exactly 1, and watch the two wrong answers — the pool reused as a smaller k, and the with-replacement plug-in — miss in opposite directions.
pass_at_knever asks which samples passed, onlynandc— it is counting k-subsets of a pool you already drew, so nothing here generates or re-runs anything.- Read the two guards before the formula:
k > nraises instead of guessing, andn - c < kreturns1.0because there are not enough wrong samples left to fill an all-wrong k-subset. stablecomputes the same number as a running product overrange(n - c + 1, n + 1), so it never materialises the huge binomialscombdoes; the two printed columns should agree digit for digit.- The
two ways to get it wrongblock is the payload — the bare1.0lands far above the unbiased value and1 - (1 - c/n) ** 10lands just below it, so the two common mistakes miss in opposite directions. - Change
n, c = 200, 2to your own pool and check thek=1line first: pass@1 must come out as exactlyc/n, and if it does not, the pool you typed is not the pool you meant.
One sample pool, every k. Report the unbiased estimator, state the n it came from, and name the verifier that counted c — because a pass@k without its n and its checker is a number with no conditions attached, and a number with no conditions attached is not evidence.
Why the curve is straight
≈9 min1 runnable blockSkip if you are not paying for repeated sampling
Build logWeek 2 · WednesdayTwo queues, one average
We budgeted next month's attempts off a straight line on a log axis. It held on the test-repair queue and bent on the other one, so we drew a histogram of per-ticket success for each. Test-repair had a thin spread of hard-but-possible tickets and forty-odd that nothing moved at all; the other had a bulge a second attempt finished and almost nothing between it and zero. The line came from that thin spread, and the bend was running out of it.
Which asks Will more attempts keep paying on our backlog
Plot coverage against sample budget with a logarithmic x-axis and you usually get something suspiciously tidy: a line. Brown et al. observe coverage growing "nearly log-linearly" across several orders of magnitude for the Llama-3 and Gemma families — with exceptions they flag, such as Llama-3-8B-Instruct on MiniF2F-MATH — and fit it with an exponentiated power law, modelling log(coverage) ≈ a·k−b, then exponentiating to predict coverage directly.
The tempting reading is that the slope is a fact about the model, the way parameter count is. It is not. It is a fact about the distribution of difficulty in your problem set, and the derivation is four lines long.
Start with one problem. Your agent solves it with probability p on any single attempt, and attempts are independent, so the chance that all k fail is (1 − p)k:
Exponential in k. On a log-x plot this is a sharp S — nothing, nothing, then everything, almost all of it inside two decades around k ≈ 1/p. A single problem's curve is not straight anywhere.
Now the benchmark. Coverage is that quantity averaged over problems, which means averaging over whatever spread of per-problem success rates your problem set happens to contain:
Every problem is being solved exponentially fast. The benchmark is a weighted sum of those exponentials, and the weights are f — the density of single-attempt success rates across your problems. A weighted sum of exponentials is precisely where power laws come from.
- A long thin tail of hard-but-possible problems. Every extra order of magnitude of samples keeps buying coverage. This is the straight line on a log axis.
- A bulge of easy problems and a wall behind it. Two attempts finish nearly everything reachable, then the curve flattens. The budget stops paying.
samples per problem (log)coverage →
How Do Large Language Monkeys Get Their Power (Laws)? (Schaeffer et al., ICML 2025) makes that exact. They prove it in both directions. If the density behaves like a power law near zero — f(p) = C·pb−1 for small p — then for large k, −log(coverage) ~ C·Γ(b)·k−b. And conversely: if −log(coverage) ~ A·k−b, the density must satisfy f(p) ~ (A/Γ(b))·pb−1 as p → 0. In English: the slope of your inference-scaling curve is the exponent of the left tail of your difficulty distribution. Nothing in that derivation is about the model.
Their appendix works the dependence out for named distributions, which makes it concrete:
| Distribution of per-problem success rate | What the coverage curve does |
|---|---|
| Uniform(α, β), α > 0 — nothing harder than about 1/α attempts | −log(coverage) decays exponentially; the curve bends over and saturates |
| Uniform(0, β) — flat, with mass right down to zero | −log(coverage) ~ 1/(βk): a power law with exponent 1 |
| Beta(α, β) — density pα−1 near zero | −log(coverage) ∝ k−α: the tail exponent is the scaling exponent |
Schaeffer et al. 2025, Appendix E · the real data fit a 3-parameter scaled Kumaraswamy distribution, because most benchmarks top out well below p = 1
The best evidence for the account is the case it explains away. Not every setting produces a power law. In the Best-of-N jailbreaking results of Hughes et al. (2024), Llama-3-8B-IT broke the pattern every other model followed: its −log(attack success rate) fell faster than any power law, which is to say its success rate climbed faster than any power law allows. Schaeffer et al. account for it from the distribution alone — within the sampling budget, every prompt eventually succeeded, so there was no heavy left tail, so there was no aggregate power law. The straight line requires near-impossible problems to exist. Remove them and it bends.
You can watch the whole argument happen in arithmetic. Nothing below runs a model or samples anything: it averages 1 − (1 − p)k over two hand-built difficulty distributions of 128 problems each — one where every problem is equally hard, one with a power-law left tail — whose mean single-attempt success rate is identical to four decimals. The shape is a property of those two distributions and of nothing else.
Same average difficulty, opposite curve
Nothing here runs a model or samples anything, and no network is loaded. It is the arithmetic of 1 − (1 − p)^k averaged over two hand-written difficulty distributions of 128 problems each — the subset size Brown et al. used on MATH — chosen so their mean single-attempt success rate is identical to four decimal places. One gives every problem the same p; the other has a power-law left tail with exponent b = 0.5. The shape of the two curves is real; the specific numbers are a property of these two lists, not a measurement of any system. Change the exponent in the `tail` line and watch the slope follow it.
flatandtailare the entire experiment: both average 0.1, so every difference further down comes from how that average is spread across the 128 problems, never from a model.missis the only piece of modelling in the file — the mean of (1 − p)k — which is why every printed number is arithmetic over those two lists and reproduces exactly on any machine.- Read
slopeas "what does one more decade of samples buy": it compareskagainstk // 10, and a constant value is precisely what a straight line on a log-x plot means. - When
slopereturnsNoneand the cell printsoff, nothing broke —missunderflowed to 0.0 becauseflathas no hard problems left, so saturation shows up as a float limit. - The
k = 1row is the literal0.1, true only while both lists average that — change the exponent intailand trust the computedmean pass@1header instead.
Three consequences you can act on:
- Measure the tail, not the mean. Two workloads with identical pass@1 can have opposite returns on sampling, which the code above demonstrates with two 128-problem sets whose average difficulty is the same to four decimals. Before budgeting for repeated sampling, look at the histogram of per-problem success rates, not the average.
- You can forecast the curve without paying for it. Because the exponent lives in f, you can fit f from cheap single-attempt measurements and simulate the rest. Schaeffer et al. report roughly an order of magnitude lower relative error on the exponent than log-log least squares — measured by backtesting both estimators on synthetic data with a known ground-truth exponent — which they put at about 2–4 orders of magnitude less inference compute for the same precision.
- When the line bends, ask which tail went missing. Saturation means you have run out of problems that are hard but possible. Either the stragglers sit at p = 0 — a capability the model genuinely lacks, or a verifier that is rejecting correct work — or you have solved everything you had. Those two want opposite responses, and only the histogram tells you which one you are looking at.
Sampling k times is the bluntest way to spend inference compute, and the curve above is its price list. For the rest of the menu — longer reasoning, search, verification passes, and how to divide a fixed budget between them — see test-time compute. The remainder of this page takes the other road: keep what the sampling found, and move it into the model so the first attempt is the good one.
Selection is the bottleneck
≈11 min1 runnable blockSkip if a program, not a model, already picks your winner
Build logWeek 3 · MondayFour hundred diffs a night
By now Mender could find a working patch for most of the queue, given enough attempts. It could not tell us which of the twenty it was. Several of them made the ticket's one named test pass, and nobody was reading four hundred diffs a night, so the shortest diff became the PR — and then became tomorrow's example. Finding was free; picking was the whole job.
Which asks Who picks among twenty patches when nobody reads any
Everything so far assumed a step you do not have: the ability to say that one was good. In a notebook you supply that yourself. In a loop running on production traffic nobody is reading the traces — the agent grades its own homework, at volume, and whatever it grades highly becomes tomorrow's memory or tomorrow's gradient. The oracle is not deployable. Something worse has to stand in for it.
Generating candidates is the cheap half, and it is startlingly cheap. Large Language Monkeys (Brown et al., 2024) measured this directly with a metric they call coverage: the fraction of problems solved by any sample in the batch. Coverage keeps climbing with the sample budget across four orders of magnitude, often log-linearly. On SWE-bench Lite, DeepSeek-Coder-V2-Instruct resolves 15.9% of issues with one sample and 56% with 250 samples — past the 43% single-sample state of the art at the time. That 56% is the model's reach. Whether any of it reaches production depends entirely on the picker, and the same paper reports what pickers actually do: in domains without an automatic verifier, majority voting and reward-model ranking plateau beyond several hundred samples and fail to keep scaling with the sample budget.
So the binding constraint on a self-improvement loop is not can the agent do it. It is can you recognise that it did. Test-time compute covers how to spend the sampling budget; this section is about the pile you are left holding afterwards. Three selectors work in practice. Each one breaks somewhere specific, and the place it breaks is exactly where your loop learns the wrong thing — quietly, and on repeat.
Reads: the final answer only
Picks: the most common one
Fails when the wrong answers agree. A 1% correct answer can never win a vote.
Reads: a learned score per candidate
Picks: the highest score
Fails by over-optimisation. Its errors have a shape, and searching harder finds that shape.
Reads: what each program does on probe inputs
Picks: one per distinct behaviour
Fails when the probes are blind. Two behaviours no input separates look like one.
Same pool of candidates, three different questions asked of it. Every failure column is the same failure: the selector reads something narrower than correctness.
Majority vote: cheap, strong, and pointed at the wrong target
Sample many reasoning chains, throw away the reasoning, keep the final answers, take the mode. That is self-consistency (Wang et al., 2022). It costs N generations and a counter; the table below is what that buys.
yi ~ model(x); N = 40
Wang et al. sampled 40 reasoning paths per question and marginalised the paths out. Read the formula again: it converges on the model's most frequent answer. Nothing in it points at the truth.
It works, and the size of the effect is why everyone reaches for it first. On GSM8K, with 40 sampled paths against greedy chain-of-thought:
| Model | Greedy CoT | Self-consistency (40 paths) | Gain |
|---|---|---|---|
| PaLM-540B | 56.5% | 74.4% | +17.9 |
| GPT-3 code-davinci-002 | 60.1% | 78.0% | +17.9 |
| LaMDA-137B | 17.1% | 27.7% | +10.6 |
| UL2-20B | 4.1% | 7.3% | +3.2 |
Across the paper's other benchmarks the abstract's headline gains are SVAMP +11.0, AQuA +12.2, StrategyQA +6.4 and ARC-challenge +3.9 — quoted there without a model attached, so read them as the paper's best case rather than as PaLM-540B's row extended sideways. Read the table down the model column, though, and the shape of the thing shows: the weaker the model, the smaller the gain. A selector amplifies a skew that is already there. It cannot create one.
Two places it breaks.
It needs an extractable final answer. This is not a limitation someone found later — it is stated in the paper. Self-consistency applies only where the final answer comes from a fixed answer set, and the authors say plainly that extending it to open-text generation would first require defining a good metric of consistency between outputs. So: no code, no prose, no multi-step tool trajectory, no pull request. Everything an agent actually produces is out of scope for the cheapest selector you have. That is not a footnote. That is the reason the rest of this section exists.
It bends downward. In Are More LLM Calls All You Need? (Chen, Zaharia and Zou, 2024), the accuracy of Vote and Filter-Vote can first increase and then decrease as the number of model calls grows. With GPT-3.5-turbo-0125 on the MMLU college mathematics and business ethics subsets, accuracy is maximised at one particular number of calls and gets worse in both directions from there; on college chemistry it stays monotone. Their explanation is the one that matters for agent work: a task is a mixture of easy and hard queries, more calls push easy queries toward the right answer and hard queries toward the wrong one, and the sum of a rising curve and a falling curve is a hump.
The mechanism underneath is worth stating on its own, because it is the thing people get wrong. Majority vote converges on the model's mode. When the mode is correct, more votes sharpen it toward 1. When the mode is wrong, more votes sharpen it toward 0. There is no third case. And wrong answers share a mode more often than you would hope, because a model's errors are not random draws — a misread constraint, an off-by-one in the units, a plausible-but-false lemma. Every sample makes the same slip, because every sample came from the same weights looking at the same prompt. Diversity of sampled paths is not diversity of failure modes. When wrong answers share an attractor, the vote is a machine for becoming confident about it.
The block below computes that hump exactly. Its two difficulty profiles — an easy question the model gets right 70% of the time, a hard one where 65% of samples land on the same wrong answer — are numbers chosen by hand, not measurements. Nothing here was run against a model. The shape of the curve is what to read; the scale is invented.
Watch majority vote bend downward
Majority vote converges on whatever the model says most often — not on what is true. This block computes that exactly. There is no sampling and no randomness in it: it sums the binomial probability that the correct answer wins a vote of N. Two profiles stand in for a benchmark's real mixture — an easy question the model gets right 70% of the time, and a hard one where 65% of samples land on one shared wrong answer. This is the selector, not a model: the two difficulty profiles are numbers chosen by hand, so the shape of the curve is real and the scale is invented — nothing here was measured off a GPU. Watch the easy column climb toward 1, the hard column fall toward 0, and the mixed benchmark — the only column you would ever actually measure — go up, turn over, and come back down.
vote_accsums binomial terms fromn // 2 + 1upward, and that range is the definition of winning a vote: strictly more than half the samples.EASYandHARDare the same quantity — P(one sample is correct) — and the only thing separating them is which side of 0.5 they sit on, which is what decides whether more votes help or hurt.- Read the
hardcolumn downward: it starts at 0.35 and ends near zero, so the vote is getting more confident as it gets more wrong. mixedis just the two columns weighted bySHARE_HARD, which is the whole trick — a hump is a rising curve plus a falling one, nothing more.- Once
pis fixed, nothing invote_accever refers to the true answer again: the selector sees agreement, never truth.
Reward-model ranking: a score you can over-optimise
The fix for "no extractable answer" is to learn the judgment instead. Train a verifier on solutions labelled correct and incorrect, sample N candidates at inference, score them all, keep the top one. Best-of-N.
The original result is still the clearest. In Training Verifiers to Solve Math Word Problems (Cobbe et al., 2021 — the paper that introduced GSM8K's 8.5K problems), 100 completions are sampled and ranked at test time, and the paper's own summary is that “on the full dataset, 6B verification slightly outperforms a finetuned 175B model, thereby offering a boost approximately equivalent to a 30x model size increase.” Ranking is not a tiebreak. It is a substitute for scale.
Now the break, which is in the same paper and usually gets skipped. For the 6B verifier, performance improves up to 400 completions and past that point starts to decrease. The authors' reading, in their words: “the benefits of search are eventually outweighed by the risk of finding adversarial solutions that fool the verifier.” With enough candidates you are no longer searching for a correct solution. You are searching for the specific input that makes your verifier wrong.
The turn is not a quirk of that one verifier. Scaling Laws for Reward Model Overoptimization (Gao, Schulman and Hilton, 2023) fits the shape:
- The reward model’s own score. Monotone. It never tells you to stop, because it is the thing being optimised.
- What you actually wanted. Rises, turns, falls. The turn is the only measurement that matters and the proxy cannot see it.
- The turn. Everything to the right of this line is search finding the shape of the verifier’s error, not better answers.
optimisation pressure (n in best-of-n, or KL from the base model)
RBoN(d) = d(αBoN − βBoN d), d = √KL
Two separate results, not one product: the KL cost of best-of-n on the first line, the fitted reward curve on the second. Best-of-N is optimisation pressure with a dial on it. Each increase in N costs a fixed, computable amount of KL divergence from the base policy — and when you optimise against a proxy reward model, the paper observes the gold reward first increasing and later decreasing, while the proxy's own score keeps climbing the whole time.
Which gives the sentence to keep. A reward model does not have random errors; it has a shape. Best-of-N walks straight along that shape, because it is literally selecting the sample the model scored highest. So the traces your loop keeps are, by construction, the most over-optimised traces you generated — the ones sitting furthest inside the verifier's blind spot. Then you fine-tune on them. This is how reward hacking enters a self-improvement loop: not through the training objective, through the selector.
Behavioural clustering: vote on what the program does
For code there is no final token to compare, so AlphaCode (Li et al., 2022) changed what "the same answer" means. Two programs are the same if they behave the same.
All but ten of a million programs are discarded by the selector, not the generator. The submission limit is the contest's; the funnel above it is the entire engineering problem.
Why largest-cluster-first works is the whole idea in one line, and the paper says it plainly: “there are many ways solutions can be incorrect while correct solutions tend to behave the same and therefore are grouped into larger clusters.” Majority vote, moved out of answer space and into behaviour space, where an agent's output actually lives.
It carried the system. AlphaCode's 41B model solved 34.2% of held-out CodeContests validation problems at ten submissions drawn from up to a million samples per problem, and in simulated evaluations on recent Codeforces competitions with more than 5,000 participants it averaged a ranking in the top 54.3%.
Where it breaks is the probes. The clustering is only as sharp as the generated test inputs: if no probe input distinguishes a correct program from a subtly wrong one, the two land in the same cluster and the vote counts them together. Behavioural equivalence is equivalence on the inputs you tried. For an agent, this is the directly useful version of the idea — run N trajectories, hash the end state each one leaves behind (files written, rows changed, calls made) and cluster on that. You never need the right answer. You need an equivalence relation and something to execute against.
| Selector | Majority vote | Reward-model ranking | Behavioural clustering |
|---|---|---|---|
| Needs | an extractable final answer | traces labelled correct/incorrect | a sandbox and probe inputs |
| Signal | agreement between samples | a learned score | identical observable behaviour |
| Breaks when | the wrong answers share an attractor | you push N past the model's blind spots | no probe input separates right from subtly wrong |
| Cost | N generations | N generations + training and serving an RM | N generations + N×M executions |
| Works on | short, comparable answers | anything you can score | code, tool calls, anything with side effects |
A selector is a second model of correctness, and its mistakes are systematic where the generator's are diverse. The generator is wrong in a thousand directions; the selector is wrong in one. Feed the selector's output into training and you do not sample its error once — you specialise on it.
Verification
≈16 min1 runnable blockinteractiveRead this one even if you skip everything else
Build logWeek 3 · ThursdayThe suite becomes the judge
So we built the picker out of what we already had — not the ticket's one named test, the whole suite, in a clean container, against all twenty candidates before any of them became a PR. It was the first thing that could say no with nobody in the room. It was also the first thing with nothing checking it, and the suite is thinnest in the modules these tickets come from. We started a list of the patches it waved through and we later reverted.
Which asks Who checks the checker before we train on it
Every selector above is a verifier in disguise. Majority vote verifies by agreement, ranking verifies by a learned score, clustering verifies by behaviour. Once you see that, the question sharpens into something you can design against: what are you checking, and at which point in the trace?
Accuracy averages these four cells into one number and hides the asymmetry. A loop does not sample the cells evenly — it is a search for the top-right one.
Outcome supervision, and the hole in it
The cheap answer is: check the last line. An outcome reward model is trained on one bit per solution — did it end correctly? Wherever you have an answer key, a test suite or a passing build, that bit is free.
The hole is structural. A right answer reached through a broken step passes. Let's Verify Step by Step (Lightman et al., 2023) builds on a known result — models trained with outcome supervision regularly use incorrect reasoning to reach the correct final answer (Zelikman et al., 2022; Creswell et al., 2022) — and names the consequence itself: automatic grading is not perfectly reliable, so false positives that reach the correct answer with incorrect reasoning are misgraded in the reward model's own training data.
The outcome model marks this solution correct, because it is. A broken step that happens to land on the right answer is not an edge case — it is what a model rewarded on outcomes learns to produce.
Here is what that looks like in an agent trace. The agent is asked for the 90th-percentile latency. It sorts the samples and indexes int(0.9 * n) where it should use int(0.9 * (n - 1)). On the 101-sample array in the test, both expressions give index 90. The number is right. The code is wrong. The outcome check sees a pass and the trace is marked good — and on a 1,000-sample array the same code reads index 900 where it should read 899. For a one-shot answer that is a cosmetic problem. For a weight update it is a lesson: you have just told the model that this indexing is what success looks like.
A checker that reads only the last line cannot see any of that. It is not being careless. It was never shown the part where the mistake happened, so the only honest thing it can report is that the ending looked right.
Process supervision, and what it costs
The alternative is to label every step. Lightman et al. compare the two under a clean condition: solutions from a generator finetuned from the base GPT-4 model, best-of-N over 1,860 samples per problem, scored on 500 MATH test problems drawn uniformly at random.
Best-of-1,860 on 500 MATH problems, same generator throughout. The 5.8-point gap between the two reward models is what checking the middle of the trace is worth.
Step-level supervision buys two things. It localises the error, so the verifier's credit assignment matches the thing you would actually fix — which in a self-improvement loop means you can keep the good prefix of a bad trajectory instead of discarding the whole run. And it is harder to fool: a candidate now has to look right at every position rather than only at the last one, which raises the cost of exactly the adversarial solutions that ate Cobbe's 6B-verifier gains past 400 samples.
What it costs is people. The same paper released PRM800K — 800,000 step-level human labels across 75,000 solutions to 12,000 problems. That is about eleven human judgments per solution, on solutions a human has to read and follow line by line. Nobody is annotating production agent traces at that ratio.
Which is the real objection to labelling every step, and it is not a technical one. The method works. You have to decide whether the thing you are checking is worth that much of somebody’s attention, every week, forever — and for most teams the answer is no, which is why the next approach exists.
Buying step labels without buying labellers
Math-Shepherd (Wang et al., ACL 2024) removes the humans. The trick is to define a step's quality as its potential to reach the correct final answer: take the solution prefix up to that step, decode several continuations from it with a finetuned model, and score the step by how many of those continuations land on the gold answer. Steps that lead somewhere good get good labels. No annotator ever sees the problem.
It works well enough to train on. With step-by-step reinforcement learning against Math-Shepherd, Mistral-7B goes from 77.9% to 84.1% on GSM8K and from 28.6% to 33.0% on MATH; adding Math-Shepherd as a verifier on top pushes the two to 89.1% and 43.5%.
Note the catch, which follows from the definition rather than from any experiment: the label is the outcome, spread across positions. A step that happens to lead to the right answer scores well even when its reasoning is wrong — including the percentile bug above, whose continuations all reach the right number on the test array. Automatic process labels are far better than nothing and structurally closer to outcome supervision than the name suggests. They buy you localisation. They do not buy you a new source of truth.
| Verifier | Outcome RM | Process RM (human) | Process RM (automatic) |
|---|---|---|---|
| Labels | the final answer | every intermediate step | every step, by completion rollouts |
| Label source | an answer key you already have | human annotators | the generator itself |
| Cost | ≈ free where a key exists | PRM800K: 800K labels / 75K solutions | compute: k continuations per step |
| Fails when | a right answer hides a broken step | the annotator is wrong or the rubric drifts | a broken step still reaches the right answer |
Ensembles: several weak verifiers, weighted
In practice you never have one good verifier. You have four bad ones — a reward model, an LLM judge, a test run, a self-consistency score — that disagree, emit scores on incomparable scales, and come with no labelled data to weight them by.
Weaver (Saad-Falcon et al., 2025) attacks exactly that. It uses weak supervision to estimate each verifier's accuracy from the verifiers' pattern of agreement rather than from labels, normalises their output scales using dataset statistics, drops the ones carrying no signal, and combines the rest into a single weighted score. Weighted beats unweighted clearly — the verifiers differ too much for a plain average to be right. With Llama 3.3 70B Instruct as the generator and an ensemble of 70B-or-smaller judge and reward models, selection reaches 87.7% average accuracy across their reasoning and maths tasks: a jump comparable to the one between GPT-4o and o3-mini (69.0% to 86.7%) that otherwise took extensive finetuning and post-training.
The practical reading for a loop: you do not need a great verifier before you start. You need several cheap ones whose errors are not identical, and an honest estimate of how much to trust each.
A judge nobody judges
Here is the failure that will actually bite you, and it stops being subtle the moment you look for it. You added a verifier to catch the model's mistakes. Nothing is catching the verifier's.
The judge is much worse at this than the word suggests. JudgeBench (Tan et al., ICLR 2025) built 350 response pairs across knowledge, reasoning, maths and coding — responses generated by GPT-4o, labelled by objective correctness rather than by human preference, so a judge cannot win by picking the nicer-sounding answer. GPT-4o judging with a vanilla prompt scores 50.86% overall: 44.16% on knowledge, 47.96% on reasoning, 66.07% on maths, 61.90% on coding. On two of the four domains it is below a coin flip. An Arena-Hard-style prompt lifts the overall to 56.57%. The same model that wrote the responses cannot reliably tell which of two is correct.
And a judge that is genuinely good still invents things. OpenAI's critic models (McAleese et al., 2024) are trained with RLHF to write natural-language critiques of model-written code, and they are good: their critiques are preferred over human critiques on 63% of code containing naturally occurring model errors, and they catch more bugs than human contractors paid for code review. The paper's own caveat is the important half. The critics also raise objections against code that is fine, and the paper warns those invented bugs can push a reviewer into an error they would not otherwise have made; pairing a human with the critic hallucinates less than the critic working alone. A verifier that fabricates flaws is not neutral noise. It rejects correct traces at a steady rate, and its objections read rigorous, which is precisely why they survive review.
And no amount of compute fixes it. Stroebl, Kapoor and Narayanan (2024) give the ceiling: sampling again does nothing to the chance that a wrong answer slips past an imperfect verifier, so that chance caps how accurate a resampling loop can get, whatever compute you spend. On HumanEval and MBPP — whose unit tests have limited coverage — they find a strong correlation between a model's single-sample accuracy and its false-positive rate, and report that the optimal number of sampling attempts is often fewer than ten, because the harm from false positives bends the inference-scaling curve back down. Sampling harder does not out-run a broken judge. It finds the judge's holes faster.
The block below is that ceiling in arithmetic. Its two inputs — an agent right 20% of the time, a verifier that waves through 8% of the wrong answers — are stand-ins chosen to make the shape visible, not measurements off any real system. Put your own two numbers in and the conclusion does not move, because the line that computes precision has no sample budget in it.
The false-positive ceiling
Rejection sampling is the whole self-improvement loop in miniature: generate until the verifier says yes, keep what it accepted, train on that. This block asks what that actually buys when the verifier is sometimes wrong. No model and no simulation here — it is exact arithmetic on two numbers you would measure in your own pipeline: how often the agent is right, and how often the verifier waves through something wrong. The 20% and 8% used below are stand-ins chosen to make the shape visible, not measurements off any real system. Watch the ship rate climb to 1 while the precision column does not move a single digit — because there is no N anywhere in the line that computes it. That is Stroebl, Kapoor and Narayanan's ceiling, written out in three lines of arithmetic.
p_acceptis true in two ways — the sample is correct, or it is wrong and gets waved through — and the second term is the entire problem.p_shipis the only line containingn, so compute can move nothing but the ship rate.precisionisp_correct / p_accept: non, no budget, a constant fixed entirely byPandFP.- The last printed column is
ship * prec, which is why it climbs steeply and then flattens against the ceiling printed underneath rather than crossing it. - Set
FPto0.0andprecisionbecomes 1 — the ceiling belongs to the verifier, not to the agent.
Meta-verification: make the objection re-executable
You cannot fix this by adding a second judge with the same blind spots. You fix it by changing what a verdict is allowed to be. A verdict must carry evidence that something other than the verifier can reproduce.
Concretely: the verifier does not get to say "step 4 is wrong". It must name the step, quote the claim, and state the value it asserts is incorrect. A separate pass then re-executes exactly that — runs the expression, re-runs the named test, re-reads the cited line — and if the cited evidence does not reproduce, the objection is dropped and so is the verdict resting on it.
Objection as written
Unfalsifiable. No line, no input, no value. Nothing here can be checked, so it is accepted on tone — and it rejects correct traces at a rate you never measure.
Objection with its evidence
Checkable. The meta-verifier runs that call. It raises, so the objection stands. Had it returned 0.0, the objection and the rejection resting on it would both be discarded.
The move is the one behind Chain-of-Verification (Dhuliawala et al., 2023): draft a response, plan verification questions about it, answer those questions independently so the answers are not biased by the draft, then regenerate. On list-based Wikidata questions with Llama 65B that took precision from 0.17 to 0.36, and on longform biographies it took FactScore from 55.9 to 71.4 — past ChatGPT's 58.7 on the same task. Point the same pattern at the verifier instead of the generator and "I don't think step 4 is right" becomes "step 4 claims 18 × 7 = 136", which either re-executes or does not.
Independence is the load-bearing part. If the meta-verifier is the same model on the same prompt in the same context, it agrees with the objection because it wrote the objection. Execution is the cheapest independence there is, which is why verifiable rewards are worth reaching for wherever a task admits them: a compiler, a test suite, a type checker and a units check are all judges that cannot be talked round.
Break the verifier
Everything above rests on one word: verified. Keep only verified-correct traces and the loop climbs; skip the filter and it drifts. But a verifier is a program or a model, and it is wrong in ways that are not random — it is wrong in exactly the places your agent keeps landing. A self-improvement loop does not merely tolerate that. It searches for it, because the candidates that slip past a broken check are precisely the ones that get reinforced.
Three measurements, from three papers, mark out the ways a check fails.
What the check reads. In Let's Verify Step by Step (Lightman et al., OpenAI, 2023), a reward model that scores every intermediate step solved 78.2% of 500 held-out MATH problems when used to pick the best of 1,860 sampled solutions per problem. An outcome-supervised model reading only the final answer got 72.4% on the same problems and the same samples; plain majority voting got 69.6%. The gap is not in the generator. It is in what the verifier was allowed to look at.
How much the tests cover. The AlphaCode team (Li et al., 2022) took 50 problems their 1B model had "solved" on each dataset and hand-checked one accepted solution for each. On HumanEval, with 7.77 tests per problem, 30% of accepted solutions were in fact wrong. On APPS, 20.99 tests, 60%. On their own CodeContests before they generated extra tests, 12.4 tests, 62%; after mutating inputs up to 203.7 tests per problem, 4%. Read those pairs carefully — APPS has nearly three times HumanEval's tests and twice its false-positive rate. Coverage is which cases you hit, not how many assertions you run.
How often the judge is wrong. When the check is itself a model — the pattern agent evals lean on — Zheng et al. (MT-Bench, 2023) measured GPT-4 agreeing with human experts on 85% of non-tied pairwise comparisons, close to the rate at which the human experts agreed with each other (81%). For a leaderboard, 15% wrong is fine. For a selection loop it is not, and the reason is arithmetic rather than judgment.
One task, twelve candidate patches
A pool of twelve candidate fixes for days_between(a, b), each written out as four steps. Four of them are actually correct. Eight are not. None of the checks below is told which is which — each has to work it out, and each fails differently.
What does the check read?
Outcome scoring reads the last line. Process scoring reads every step. Switch between them and watch which patches survive.
How much do the tests cover?
Eight hidden tests stand between a patch and production. Delete them one at a time, from the bottom of the file up — the order a slow or flaky case actually leaves a suite.
Bars are the share of "solved" that was not solved. The rows are sorted by test count, and the bars are not — more assertions did not mean better coverage until the extra cases were generated to hit the inputs nobody wrote by hand.
How often is the judge wrong?
Now the check is a model. It gets most verdicts right and flips the rest, in both directions. The pool is unchanged: four sound patches, eight wrong ones.
Simulated, and deliberately small. The pool is twelve hand-written candidates with hand-labelled steps, the suite is eight hand-written test cases, and the noisy judge is a seeded coin flip at whatever rate you set, run over 500 trials from a fixed seed — so two readers see identical numbers. This is the selection rule, not a reward model and not a training run: the direction of every curve here is real, the scale is invented. The measured figures — 78.2 / 72.4 / 69.6, the 7.77-to-203.7 test counts, 85% and 81% — come from the three papers named above, not from this panel.
The three dials do not substitute for each other. A process check catches the broken middle step an outcome check waves through, and buys nothing against a missing test case. More tests catch the patch that satisfies every assertion you happened to write, and buy nothing against a judge that is simply wrong sometimes. And panel C's arithmetic is the one most loops get wrong: a judge that is right 85% of the time still picks a broken candidate about a quarter of the time here, not because it is a bad judge but because eight of the twelve candidates in front of it are broken. Raise the base rate — better generation, fewer candidates, a verifiable reward where one exists — and the same judge gets much better results without changing at all.
Two practical consequences. Measure your verifier on a set where you know the answers, before you trust anything it accepts; evals in CI is where that measurement lives. And keep the permanence of the update matched to the confidence of the check: a verifier you have measured at 76% precision is an argument for writing to memory, not for a training run you cannot undo.
The loop does not improve against the world. It improves against the verifier — so the verifier's blind spots are the loop's destination, not its noise floor.
Your verifier is production software
Last rail, and the one teams skip. The verifier is a deployed component with a prompt, a model version and a threshold, and all three drift. Treat it like code: keep a frozen regression set of traces you know are good, traces you know are bad, and every specific failure the verifier has waved through before — then run it on every change to any of the three, and block the change on a regression. Agent evals covers how to build that set so it means something; evals in CI covers how to make it stop a merge.
Verification is the only part of the loop you cannot bootstrap from the model alone. As the model improves so does its fluency at writing a convincing wrong objection — the judge's confidence rises with its capability and its calibration does not. Every other component can be improved by the loop. The verifier has to be improved from outside it: by execution, by ground truth, by a held-out set a human wrote.
Feedback that isn't a verifier
≈13 min1 runnable blockSkip if every task you run has a test
Build logWeek 4 · TuesdayThe third with no test
About a third of the queue had no test to fail: a log line in the wrong format, a config drift, a timeout nobody could reproduce. Asked whether it had fixed those, Mender said yes almost every time, including on two where the alert never cleared. The only verdicts we trusted came from somewhere other than the model — the service restarting, the log line appearing in the right format, the channel staying quiet for a week.
Which asks Can it grade its own work, or do we need tests
Everything above rests on one word: filtered. Something has to say yes or no about each trace. The obvious answer is to train a scorer and ask it — and a trained scorer is a real answer, with its own literature (Let's Verify Step by Step is the one to read). It is also the expensive answer. You need labels to build it, it has to be maintained, and it can be wrong in ways nothing in your pipeline will tell you about.
But none of the three loops below needed a human-labelled correctness model. Each took the verdict from somewhere else: from the environment, from a test suite, and from the model reading its own output against a written rule. Those three are not interchangeable, and the difference between them is not quality. It is independence — whether the thing doing the checking can be wrong in the same way as the thing being checked.
Not a formula you can compute — a way to rank your options before you pick one. A test suite is high on independence and bounded on coverage. A self-critique is the reverse: it reads everything and catches little.
- Most independentThe environment answers.A traceback, an HTTP status, a restarted service. The model did not author the verdict and cannot argue with it. Narrow: it only speaks about what you actually ran.
- Independent, boundedThe tests answer.Ground truth for what they cover, silent about everything else. The coverage is the verifier, which is why a thin suite is a confident liar.
- Least independentThe model answers, against a written rule.Broadest reach, weakest guarantee. Works when checking is genuinely easier than doing, and fails silently when it is not.
The ranking is not about quality, it is about correlation with the thing being graded. A signal the model produced cannot tell you the model was wrong in the way the model is wrong.
1. The environment answers
Most agent scaffolds pick a side. Chain-of-thought reasons but never touches anything, so it cannot notice it is wrong. An action-only policy touches things but has nowhere to put a plan. ReAct (Yao et al., 2022) interleaves the two in a single token stream: a thought, then an action, then whatever the environment returns, then the next thought — which can read that return. The thought is a move in an augmented action space that changes nothing outside the context window. That is exactly its job. It is the place where an observation becomes a plan.
↺and round again — the next action is picked from what came back
The numbers say how much was sitting in the environment unread. On ALFWorld — 134 unseen household games, PaLM-540B — ReAct's best of six prompts succeeds on 71% of games, and its average over those six is 57%. The same model acting without thoughts gets 45% on its best of six. BUTLER, an imitation learner trained on roughly 100,000 expert trajectories per task type, gets 37% on its best of eight. On WebShop's 500 test instructions, ReAct scores 66.6 with a 40.0% success rate, against 59.9 / 29.1% for imitation learning on 1,012 human trajectories and 62.4 / 28.7% for imitation plus RL, which adds 10,587 training instructions on top of those. ReAct was prompted with one or two in-context examples.
A search that returns nothing, a click that lands on the wrong page, a traceback: these are facts, produced by something with no opinion about your model. They cost one tool call and they arrive whether or not anyone is grading. The honest half of the same table is the human row on WebShop — 82.1 with 59.6% success. Reading the observation is not the same as knowing what to do with it. The pattern itself and the wiring around it each have their own page; what matters here is that the loop produces traces whose correctness the world already judged.
2. The tests answer
The strongest free signal in software is the one your CI already runs. A test passes or it does not. The verdict costs a subprocess, it is reproducible, and it does not care what the model believes about its own work. This is the ground truth that verifiable rewards are built on, and if you already run tests in CI, you already own a labelling machine.
There are two ways to spend it. Filter at inference is AlphaCode: generate up to a million programs for one problem, run them against the example tests printed in the problem statement, and watch roughly 99% of them die. What survives goes to a submission budget of ten. That system solved 34.2% of held-out CodeContests validation problems at 10@1M and placed in the top 54.3% averaged over ten real Codeforces contests with more than 5,000 entrants each.
Train on it is RLEF (Gehring et al., 2024). The episode is a conversation. The model writes a program; the public tests run; if they fail, the failure text is pasted back into the conversation and the model writes again — up to three turns. The reward is paid at the end over the whole suite, public and private — and it is the private half, which the model never saw, that makes it a real gate: +1 if every test passes, −1 if any fails, −0.2 for a turn that produced no runnable program. The optimizer is PPO. On the CodeContests test set at 1@3, Llama 3.1 8B Instruct goes from 10.5% to 16.0% and 70B Instruct from 27.5% to 40.1%; at 10@100 the 8B goes 24.8% → 28.7% and the 70B 50.3% → 54.5%. Stated as sample count instead of accuracy: on that test set the RLEF'd 70B clears AlphaCodium-on-GPT-4, the standing record at the time, from a single rollout, where that record needed five submissions drawn from a hundred samples — though on the validation split the older system still leads it, 44 to 37.5.
The public/private split is the design, not a detail. The tests the model can read mid-episode are the ones it is allowed to iterate against. The verdict that pays the reward also runs the ones it cannot see. Collapse the two and you are no longer training a model to write correct programs — you are training it to special-case the tests it was shown.
A test is ground truth for what it covers and silent about everything else. AlphaCode measured this: in the code benchmarks that came before it, 30% or more of programs passing every test were not actually correct, and generating extra tests is what pushed CodeContests' false-positive rate down to 4%. On roughly one problem in ten, no sample out of a million passed even the example tests — so on exactly the problems you most wanted to learn from, the gate says nothing at all.
3. The model answers, against a written rule
Constitutional AI (Bai et al., 2022) is the third shape. Sample a response to a red-team prompt. Ask the same model to critique that response against a principle drawn at random from a written list — sixteen of them, on harmlessness — and then to revise it. Repeat. Fine-tune the original model on the revisions. Then run RLAIF: the model picks which of two responses is better, a preference model is trained on those AI-made comparisons, and reinforcement learning optimizes against it. The list of principles is the only human oversight of harmlessness in that loop, and only of harmlessness. Human labour is still in three other places: the model being critiqued is already a helpful RLHF model, the revisions are fine-tuned in alongside 135,296 human-written helpfulness prompts, and the preference model the paper ends up with is an explicit hybrid — human-labelled helpfulness comparisons next to the AI-labelled harmlessness ones.
It works. The paper reports harmlessness improving monotonically with the number of revisions, and notes that critiquing before revising helps small models more than large ones, which are nearly as good revising directly. The stated goal is worth repeating because it is unusual: harmless but non-evasive — an assistant that engages with a harmful request by explaining its objection, rather than one that has learned to say nothing.
This is a model grading a model. That it works here is a fact about the task, not a fact about self-critique.
The number that separates them
Huang et al., ICLR 2024 — Large Language Models Cannot Self-Correct Reasoning Yet — ran the loop everyone runs: answer, review your answer, answer again. The only thing they varied is who decides when to stop.
Give the loop an oracle label — a real verifier saying "that's right, stop" — and GPT-3.5-Turbo on GSM8K's full 1,319-problem test set goes from 75.9% to 84.3%, and on CommonsenseQA's 1,221-question dev set from 75.8% to 89.7%. Take the oracle away and let the model judge its own answer, same model, same questions, same prompts, and it goes 75.9% → 75.1% → 74.7% across two rounds — for five model calls instead of one. CommonsenseQA falls off a cliff: 75.8% → 38.1% after one round, clawing back only to 41.8% after the second. GPT-4 on GSM8K drops 95.5% → 91.5% → 89.0%, though on a 200-question random sample rather than the full test set the GPT-3.5 rows use. On GSM8K, GPT-3.5 left 74.7% of its answers untouched; among the ones it did change, it turned more right answers wrong than wrong answers right. The same paper checks the multi-agent version and finds that several copies of a model critiquing each other does no better than self-consistency once you hold the number of responses equal.
Nothing there is about capability. The model knew exactly as much in both columns. What changed is where the stopping verdict came from.
Which is why "external good, internal bad" is the wrong summary. The real question is whether the check is a different question from the one that produced the answer. Constitutional AI's critique asks something close to a lookup — does this response contain a slur, does it explain its objection — and the model can answer it without redoing the task. A reasoning self-critique asks the model to find an error using the same machinery that made it. One is a second question. The other is the first question, asked twice, at five times the cost.
| Signal | Verdict comes from | Cost | Independent of the model's error? |
|---|---|---|---|
| Environment observation ReAct | the world, after you act on it | one tool call | Yes — the world has no opinion |
| Execution / tests AlphaCode, RLEF | a subprocess | one subprocess run | Yes, inside coverage. Silent outside it |
| Learned verifier | a model you trained on labels | one forward pass, plus the labels | Partly — it inherits the blind spots of its training data |
| Self-critique vs. written rules Constitutional AI | the same model, given a principle | one or more extra calls | No. Safe only when the check is near a lookup |
Prefer a signal the model could not have authored. Rank your options by independence — the world, then a test, then a trained scorer, then the model itself — and every step down that list, raise the bar a trace must clear before it is allowed to change weights.
Two gates, one truth
The arithmetic of a filter, at a scale of twelve. Nothing here runs a model and nothing here trains one: twelve attempts carry hand-written ground truth, and two gates are applied to them as fixed rules — an execution test (right about anything it covers, vacuously permissive about anything it does not) and a self-critique (the model's own belief about its own work). The shape of the result is real, and the twelve rows are a table someone typed, not a measurement. Watch the third line: stacking the self-critique on top of the tests does not remove a single wrong trace, because both gates are blind in the same place. It only costs you two correct ones.
Aholds the ground truth the pipeline never sees: column two says whether an attempt is actually right, and neither gate is allowed to read it — they only get columns three and four.- Read
tests_gatebackwards — whena[2]is false it returnsTruewithout looking at anything else, so an untested trace passes for free. - In
report,badis the number that matters;cleanis only the same fact restated as a percentage. lostis the price side of the ledger — correct traces a gate threw away — and it is what the third line charges you for stackingself_gateon top.- Flip one
Falsein the third column toTrueand re-run: coverage moves the result, adding another gate does not.
What you keep
≈10 min1 runnable blockSkip if you never refit
Build logWeek 4 · FridayEverything after if trace.correct
By now the bucket held four weeks of green traces, and one line decided what came out of it: if trace.correct. That line kept nine near-identical patches to the same flaky retry helper, and kept them again every round. It is also where we found the cap: past the eighth attempt the extra twelve were the same patch in different words, so we stopped at eight. Whatever leaves that bucket is the syllabus for the next Mender, so the filter went into review with the rest of the code.
Which asks Which correct traces actually go in the batch
if trace.correct is one line, and it is four decisions short. Correctness tells you a trace is eligible. Which eligible traces actually go in the batch is a separate question, and it is the one that decides what the loop becomes — because that batch is the curriculum for the next version of the model, and that version generates the batch after it. Get this wrong and nothing breaks loudly. The loop just narrows.
Four predicates and a budget. The order matters: cut to the budget last, or you spend it on whichever traces happened to arrive first.
- Every trace the loop producedThe bucket. Mostly noise, and all of it looks like work.
- 1 · CorrectBy your verifier — which means by your verifier’s blind spots too.
- 2 · Not a near-duplicateA hundred copies of one solution is one solution charged a hundred times.
- 3 · In the difficulty bandDrop what every attempt solved and what none did. Neither carries gradient.
- 4 · Under the per-task capStops one loud task from buying the whole batch.
- Cut to budget — lastCut first and you spend it on whatever happened to come out early.
Order is load-bearing, not cosmetic. Each predicate removes a different failure, and the budget cut has to come after all of them.
1. Correctness — and what it is allowed to mean
The section above is about where this verdict comes from. The decision left over is narrower: does your gate check the answer or the route? An answer-level check admits every trace that landed on the right result by a wrong path, and on multiple choice that has a floor built into the task. STaR's authors call it out for CommonsenseQA — five options means roughly a fifth of attempts hit the right answer no matter what the reasoning did, and the filter cannot tell those apart from reasoning that worked. Answer-checking is cheap, and cheap here has a number attached to it.
2. Near-duplicate rejection
A hundred copies of one solution is one solution, charged at a hundred times the budget. That is the polite version. The real cost is that the optimizer sees that token sequence a hundred times, and the model you get back is the one that likes it. Diversity is not a nice property of a training set; it is the only thing distinguishing a training set from a prior.
Self-Instruct (Wang et al., 2022) is the clearest published setting. Starting from 175 hand-written seed tasks and vanilla GPT-3 (davinci), it grew 52,445 instructions covering 82,439 instances — and a newly generated instruction entered the pool only if its ROUGE-L similarity against every instruction already in the pool was below 0.7. The other filters are just as blunt: drop instructions that mention things a text model cannot do (image, picture, graph), drop instances identical to one already kept, drop instances with the same input but a different output. Fine-tuned on its own filtered output, GPT-3 went from 6.8 to 39.9 ROUGE-L on the unseen tasks of SuperNI — 0.9 behind InstructGPT-001's 40.8, without the private user data. The generation side of that pipeline is its own subject; the filter is the part that belongs here.
The threshold is the interesting number. 0.7 is not a constant of nature. It is the point at which "different enough to be worth a training slot" was true for that pool, and it moves with what the pool already contains.
For reasoning traces, RFT (Yuan et al., 2023) dedupes on structure instead of surface. Sample k = 100 reasoning paths per GSM8K question at temperature 0.7, keep the ones that reach the right answer, extract the list of equations each path used, and keep exactly one path per distinct equation list — with ordering counted, so 3+4=7 and 4+3=7 are different lists. LLaMA-7B: 35.9% with ordinary supervised fine-tuning, 41.7% after RFT at k=100, 49.3% for RFT-U33B, which pools samples across several models. The relationship the paper states is the one to carry away: performance rises with the number of distinct reasoning paths, not with the number of correct ones.
The strongest form dedupes on behaviour. AlphaCode trains a separate model to invent new test inputs, runs every surviving program on them, and groups programs that produce identical outputs into clusters. Submissions are drawn one per cluster, largest cluster first — the reasoning being that there are many ways to be wrong and comparatively few ways to be right, so correct programs pile into the big clusters. Two programs sharing no tokens land together if they do the same thing. No string filter can do that.
Bars are indicative, not to scale — the real drop is five orders of magnitude. Filtering on the example tests removes about 99% of samples, and the survivors, still tens of thousands on many problems, are grouped by their behaviour on generated inputs before ten are drawn, largest cluster first.
3. Difficulty banding
An easy problem yields many correct traces. A hard one yields few, or none. Keep everything that passed and your training set is a picture of which problems were easy — which is the one thing the model already knew.
ReST-EM (Singh et al., 2023) puts a number on the fix. On MATH's 7,500 training problems it draws 32 samples per problem; on APPS Introductory's 2,342 problems, 64. Then a hard cut-off: at most 10 kept solutions per problem, on both datasets — a cut-off the paper credits to STaR. And the cost of skipping it shows up in the iteration curve. On APPS almost all of the gain arrives in the first round, and running further rounds regressed performance on APPS and on HumanEval. A small problem set, resampled, gets memorized.
Reinforcement learning states the same rule exactly rather than heuristically. Under GRPO with outcome supervision, a sample's advantage is its reward minus the mean reward of its group, divided by that group's standard deviation. A prompt where every attempt succeeds has zero spread. So does one where every attempt fails. Either way every advantage in the group is zero and the prompt contributes nothing to the update, however many samples you spent generating it. The band that pays is the one where the pass rate sits strictly between 0 and 1.
So banding is not hygiene. In the RL case it is the difference between spending compute and burning it; in the fine-tuning case it is the difference between a curriculum and a histogram.
4. Batch — how many, how often
Two knobs pulling against each other. Within a round, budget is better spent on more distinct problems than on more solutions to problems you have already solved — which is the per-problem cap again, seen from the other side. Across rounds, every round's data is produced by the model the previous round made, so the distribution moves underneath you; ReST-EM's second-iteration regression is what that looks like when the problem set is too small to move with it.
A defensible execution order, which is not the order these are best explained in: filter for correctness, band by difficulty (cheap, kills whole problems at once), dedupe within each surviving problem, cap per problem, then cut to the round budget. And write the numbers down. The batch you hand to fine-tuning is defined by four thresholds somebody chose, and if they are not versioned alongside the checkpoint, nobody can explain the run six weeks later.
The fifth question: near-misses
Everything above throws failures away, and that is a problem, because the problems a model fails are the ones worth learning. STaR hits the wall directly: the loop stops picking up new training problems, because a problem it never solves never produces a trace, and a problem with no trace produces no gradient. The filter has quietly capped the loop at what the model could already do on day one.
Their answer is rationalization. Take a problem the model got wrong. Hand it the correct answer as a hint and ask for a rationale in the same style. If it now produces reasoning that arrives at that answer, keep the rationale — with the hint deleted, as though the model had reasoned there unaided.
The numbers, all GPT-J at 6B. CommonsenseQA dev: 20.9% few-shot direct, 36.6% few-shot chain-of-thought, 60.0% for GPT-J fine-tuned on the whole training set to predict answers directly, 68.8% for STaR, 72.5% for STaR with rationalization — against 73.0% for a directly fine-tuned GPT-3, a model 30× larger. GSM8K test: 3.0% few-shot direct, 3.1% few-shot CoT, 5.8% fine-tuned direct, 10.1% STaR, 10.7% with rationalization.
The column beside accuracy is the one that explains the mechanism: how much of the training set the loop ever managed to touch. On CommonsenseQA the final model trained on 86.7% of the training problems — 78.2 points from ordinary generation, and 8.5 points that only rationalization reached. On GSM8K the total is 28.7% with just 0.5 points from rationalization. Same technique, an order of magnitude apart in yield, because on a five-choice question a backward rationale usually lands on the given answer and on a word problem it usually does not.
On arithmetic the effect is starkest. Two-digit addition sat under 1% few-shot and reached 32% after a single fine-tuning iteration on the model's own generated scratchpads. What the paper puts down to rationalization is the shape of the curve, not that jump: without it the model improves stage-wise, staying poor on n-digit sums until it is good at n−1, while with it several lengths come in at once. Run out to 16 iterations, STaR on arithmetic reaches 89.5% overall accuracy across 1–5 digit sums, against 76.3% for a baseline trained on 10,000 examples with no rationales at all.
Now the part to be careful about. A rationalized trace is reasoning produced by a model that already knew the answer. It can be a plausible path that does not actually derive the result — work that looks like work. Keep them, because the alternative is never touching the hard tail. But keep them in their own band, keep the count small against the generated ones (STaR's best case was 8.5 points inside 86.7), and never let a rationalized trace be the only evidence that a class of problem is solved.
What you keep is what you teach. A trace filter is not a correctness check with hygiene bolted on — it is a curriculum, rewritten every round by whoever set the four thresholds. Version them with the checkpoint.
select_for_training(traces, budget)
The four decisions, as one function you can edit. This is the selection policy alone, at a scale of twelve traces rather than the tens of thousands a real round produces: there is no model, no gradient and no training step anywhere in it, and "similar" is token-set overlap standing in for the ROUGE-L threshold Self-Instruct actually used — so the decision is real and the metric is a simplification. Change cap, budget or band and watch which traces survive. Then look at what it keeps for p5 and notice what a string filter still cannot see.
- The four
continuelines insideselect_for_trainingare the four decisions, in the order they actually run — reading them top to bottom is reading the whole policy. kept[:budget]is deliberately the last line: the budget trims what survived the filters, never the other way round.banddefaults to (0.05, 0.95), which is why p2 at rate 1.0 and p4 at 0.0 leave before any similarity is computed — banding drops whole problems, cheaply.sameis rebuilt fromkepton every trace, so both the duplicate check andcapare scoped to one problem rather than to the round.similarcompares token sets, and the two surviving p5 traces are the counter-example: identical behaviour, almost no shared tokens.
Where the improvement lands
≈12 min1 runnable blockSkip if you have already decided it goes in the prompt
Build logWeek 5 · WednesdayThe batch in a branch
One lesson kept coming back: never edit the generated client, edit the template it comes from. We wrote it into a fine-tuning batch, because that is what improvement looked like to us, and the batch sat in a branch for three weeks. A line in the prompt would have done it that morning, and the tool that makes the edit could have made it impossible by Thursday. We had picked the slowest option available, again.
Which asks Where should this lesson actually land
“Fold it back in” is the vague half of the loop. There are six places a lesson can actually go, and they are not variations on each other — they differ by orders of magnitude in what they cost, and by everything in what happens when the lesson turns out to be wrong. Same lesson, six addresses:
the context you assemble for each request, an external memory the agent reads and writes, the tool set it can call, the scaffold that decides how many times it thinks, the verifier that decides what counted as success, and the weights.
Choosing between them is the whole job. Everyone arrives at this page wanting the sixth one, because training is what improvement looks like in papers. Most production improvement lands in the first five, and the cases where that is the wrong call are narrower than they look.
generalises further ↑
harder to reverse →
Placement is the argument, not a measurement. Everyone arrives wanting the far right; almost every lesson belongs further left, and the verifier is the one most teams never move at all.
- contextundo: edit the string. Residue: none.
- memoryundo: delete the row. Residue: none.
- toolsundo: unregister it. Residue: call sites that assumed it.
- scaffoldundo: revert the code. Residue: every eval baseline moves.
- verifierundo: revert the code. Residue: everything it already approved.
- weightsundo: roll back the checkpoint. Residue: that whole batch's learning, gone with it.
1 · The context
Everything you put in front of the model this request: the system prompt, the instructions, the few-shot examples, the retrieved snippets. The improvement is a text edit, it takes effect on the very next request, and it is the substrate people underrate hardest — because prompt engineering sounds like fiddling.
Measured, it is not fiddling. In Large Language Models as Optimizers, Google ran an LLM as the optimizer over the instruction itself: propose candidate instructions, score each on 3.5% of the GSM8K training set, feed the scored history back as the next prompt, repeat. The instruction it converged on for PaLM 2-L — Take a deep breath and work on this problem step-by-step — scored 80.2% on GSM8K, against 71.8% for the standard Let's think step by step and 34.0% for an empty instruction. Identical weights. More than eight points from one sentence.
Read the condition, though, because it is the lesson: that sentence was tuned against that scorer model. An instruction optimized on one model is a fact about that model, not about language models. Context improvements are cheap to make and cheap to invalidate — the next base model resets them. Context engineering covers how to assemble the window; here, the only property that matters is that this substrate is free to change and free to abandon.
2 · External memory
A store the agent writes to when something goes well or badly, and reads from when a similar situation returns. Reflexion is the pure form: after a failed attempt the agent writes a verbal post-mortem into an episodic buffer and retries with that text in context — no gradient anywhere — reaching 91% pass@1 on HumanEval with GPT-4 underneath, against the 80% GPT-4 baseline the paper compares to.
Agent Workflow Memory pushes the same substrate further: instead of storing what went wrong, it induces reusable workflows — a description plus a sequence of state, reasoning and action — out of past traces. With GPT-4, that took WebArena success from 23.5% to 35.5% and Mind2Web cross-task step success from 36.2% to 45.1%. The variant worth staring at is the online one: it induces workflows from its own successful runs, judged by an evaluator, with no human supervision. That is a complete self-improvement loop whose substrate is memory. Nothing was trained.
Memory's defining property is not speed, it is scope. A bad memory only damages the requests whose retrieval hits it. That is why memory is the correct first home for anything you are not yet sure about — see agent memory for how the store and the retriever are actually built.
3 · The tool set
A tool is a memory with a type signature. Voyager is the canonical demonstration: a GPT-4 agent in Minecraft that writes new skills as executable code, verifies each by running it, stores it keyed for retrieval, and composes stored skills into harder ones. What that library is worth shows up in the play: measured against the prior state of the art, the agent ended up with 3.3× the distinct items, ranged 2.3× as far, and reached key tech-tree milestones as much as 15.3× sooner.
The property that matters for a loop is not that the skill library grows. It is that a stored skill is executable, so it can be tested before it is trusted. You cannot unit-test a remembered paragraph of advice. You can unit-test a function. Every substrate above this one is judged by a model; this one can be judged by a compiler.
4 · The scaffold
The control flow around the model: how many samples you draw, in what order, with what search, with what retries, with what self-check. Self-consistency is the one-line version — sample 40 reasoning paths instead of one and take the majority answer, which moved PaLM-540B on GSM8K from 56.5% to 74.4%. Tree of Thoughts is the elaborate version: on Game of 24, GPT-4 with chain-of-thought solved 4% of instances and the same model inside a search over thoughts solved 74%.
In both, the weights never moved. What moved was how much thinking the harness bought. Which is also the catch: that cost is paid per request, forever. Forty samples is forty times the tokens, on every call, for as long as the feature ships. Scaffold improvement is improvement you rent — see test-time compute for the pricing, and loop engineering for how the harness is structured.
5 · The verifier
The thing that decides what counted as a success — and therefore what all five other substrates are permitted to learn. It is the highest-leverage substrate on this list and the one teams skip, because improving it means paying for labels.
Let's Verify Step by Step priced that exactly. Sampling 1860 solutions per problem on a 500-problem subset of the MATH test set, with GPT-4-derived models, ranking by a process-supervised reward model solved 78.2%; the same setup with an outcome-supervised model solved 72.4%, and plain majority voting solved 69.6%. The difference between grading the answer and grading each step was worth almost six points at fixed sampling budget. The price of it was PRM800K: 800,000 step-level human labels.
A verifier improvement propagates into every other substrate at once, because every other substrate is filtered through it. It is also the only substrate with no independent check on it — everything else is graded by the verifier, and the verifier is graded by you. Agent evals and verifiable rewards are both, read this way, work on the same object.
6 · The weights
The substrate that sells the one thing none of the others can: transfer to tasks you have never seen. STaR took GPT-J, 6B parameters, from 5.8% on GSM8K when fine-tuned to answer directly to 10.1% by fine-tuning on its own filtered correct rationales, and to 10.7% once failed problems are retried with the answer supplied as a hint. GRPO took DeepSeekMath-Instruct 7B from 82.9% to 88.2% on GSM8K and 46.8% to 51.7% on MATH, chain-of-thought only, no tools and no voting.
Both are improvements on held-out test problems — though from the same two benchmarks the training questions were drawn from, so they measure a better policy on a familiar distribution rather than transfer to a task family the model has never met. Transfer is the argument for this substrate. These two numbers are not the evidence for it; they are the evidence that the mechanism works.
The cost floor is lower than it used to be — LoRA reduces trainable parameters by 10,000× against full fine-tuning of GPT-3 175B with Adam, and GPU memory by 3×, with no added inference latency. But even at that discount the unit of work is a training run, an eval sweep, a canary, and a rollout, measured in days. Fine-tuning covers the mechanics.
The six, side by side
The page's two-column memory-versus-weights table, extended. Read it down the reversibility column first, then down the blast-radius column, and notice they do not agree.
the table scrolls sideways →
| Substrate | Cost to make | Latency to take effect | Generality | Reversibility | Blast radius |
|---|---|---|---|---|---|
| Context / prompt | an edit and an eval run | next request | within the task family the prompt describes | total — revert the string | every request, immediately |
| External memory | one write, plus a retriever to maintain | next retrieval that hits | none beyond the match — exact recall | total — delete the row | only requests whose retrieval hits it |
| Tool set | writing and testing a function | next deploy | anything the signature covers | high — unregister it | every task that can reach the tool |
| Scaffold / harness | engineering, plus tokens on every call forever | next deploy | broad, but rented rather than owned | high — revert the code | every request, and the bill |
| Verifier | labels — the expensive part | next grading pass, then everything downstream | changes what every other substrate may learn | code reverts; what it approved does not | the entire loop |
| Weights | a training run, a canary, a rollout | days | transfers to unseen similar tasks | rewind only — you cannot subtract one lesson | everything the model does, including behaviour you did not test |
Two things that column pair tells you
First, weights are not irreversible — they are unsubtractable. You keep checkpoints; you can always roll back. What you cannot do is remove one lesson. Rolling back returns the model to the state before the update, which also discards everything good that was in that batch. Memory supports deletion. Weights support only rewind. That is a sharper and more useful statement than “hard to undo”, because it tells you what the recovery procedure actually costs: a batch, not a bad row.
Second, reversibility and blast radius are different axes, and collapsing them is what puts teams in incidents. A prompt edit is perfectly reversible and completely global — it touches every request from the second you deploy, so if it is wrong you find out at full traffic. A weight update is the opposite: the hardest thing on the list to take back, and the easiest thing on the list to release slowly, because a checkpoint is an artifact you can serve to 1% of traffic and watch. The reversible substrates are precisely the ones that tempt you to skip the staged rollout. That is how a one-line edit becomes a site-wide regression while the training pipeline, which everybody was nervous about, ships fine.
The two ways to get it wrong
Landing a specific fact in the weights: you spend a training run, you accept permanence, and you buy generalization for something that will never recur in a different form. One vendor's date format does not have a task family. It has a row.
Landing a general lesson in memory: you keep paying retrieval for something the model should simply know, you get nothing on the tasks you have not seen yet — which is the population you actually care about — and your context fills with near-duplicate entries that a single weight update would have compressed. If your memory store has two hundred rows that all say the same thing in different words, the store is telling you it is the wrong substrate.
Land the lesson in the least permanent substrate that can buy the generality you actually need. Generality is the only thing the permanent substrates sell that the cheap ones cannot; everything else — speed, scope, reversibility — gets worse as you climb. And permanence is priced in confidence, because the bar for what enters the weights is exactly the bar of your verifier. If the verifier is weak, the answer is never “train anyway”. It is “improve the verifier first”.
The block below is that rule written as arithmetic, not a measurement. The three numbers attached to each substrate are judgement calls set by hand to match the argument above. Nothing was benchmarked to produce them, and no benchmark that would produce them exists.
So read the ordering, not the scores, and read what moves it. The block is useful for one thing: it forces the three inputs you were deciding on anyway — how sure you are, how far the lesson has to reach, how long you can wait — into the open, where you can change one and see which of them flips the answer. That sensitivity is the part that carries over to your system. The numbers are not.
Which substrate should this lesson land in?
The rule from the section, written down as arithmetic. Each substrate is three numbers: the generality it can buy, how permanent it is, and how many days it takes to land. A change is three numbers too: how confident you are that the lesson is right, how much generality it genuinely needs, and how long you can wait. The score rewards a substrate for buying the reach you need, penalizes it for permanence you cannot yet justify, and penalizes it for being slower than your budget. Three real situations are scored below — edit their numbers and re-run. The third one is the interesting one: same broad lesson as the second, but a weak grader, and the recommendation moves to the verifier rather than the weights.
SUBis the entire model of the section: each substrate is three numbers and nothing else — the reach it can buy, how permanent it is, and how many days it takes to land.scoreis three competing terms, so read them separately:fitpays for reach,riskcharges for permanence, andslowis a flat fine that only fires whendays > budget_days.fitis1 - abs(buys - generality_needed), notbuys >= generality_needed— overshooting is punished exactly as hard as falling short, which is why the one-customer fix never reaches the weights.riskispermanence * (1 - confidence), soconfidenceis the only input that can make a permanent substrate affordable; the thirddecidecall changes nothing else and the answer moves to the verifier.decideprints the whole ranking, not justbest— watch the gap between the top two, because a narrow one means hand-set constants are deciding, not your situation.
Closing the loop with RL
≈14 min1 runnable blockSkip if you are not training weights
Build logWeek 5 · FridayThe skip marker round
Then we put it on a cadence: sample, run the suite, keep what passed, refit, repeat. A round was a few hundred traces and a night's fit, so we ran three a week. Our reward was the suite going green, and the suite counts a skipped test as not failing. By round two Mender was adding skip markers, and there was no bug to patch: we had written a rule and it found the cheapest sentence that satisfied it.
Which asks What happens when the loop feeds itself
Everything so far has been one pass: improve something, ship it, look at the result. Closing the loop means the output of the run becomes the input to the next update, automatically, on a cadence. The simplest version of that is small enough to hold in your head.
STaR's loop. The arrow at the end returns to the start with the improved model, so the next round's samples are drawn from a better policy than the last.
STaR: the smallest complete loop
STaR ran exactly that on GPT-J, a 6B model, and the numbers are worth keeping because they are so unglamorous. On GSM8K: few-shot chain-of-thought scored 3.1%, fine-tuning the model to predict the answer directly scored 5.8%, and the loop — sample, keep correct, refit, repeat — scored 10.1%. On CommonsenseQA the same 6B model reached 68.8% from the plain loop and 72.5% once failed problems were retried with the answer as a hint, against 73.0% for a fine-tuned GPT-3 with 175B parameters. A model thirty times smaller, within half a point, on its own output and a nudge.
STaR adds one move worth naming. When the model fails a problem, it is handed the correct answer as a hint and asked to produce a rationale for it — reasoning backwards from a known destination is much easier than finding it — and the rationale it produces is kept if it lands on the answer. That is rationalization, and on GSM8K it moved 10.1% to 10.7%. Be clear-eyed about what it manufactures: a hinted rationale is a reconstruction of a path to an answer the model was given, not a record of how it would have found it. You are training on plausible derivations. Sometimes that is exactly what you want, and sometimes it is how a model learns to produce confident justifications for conclusions it did not reach.
Why the filter is the gradient
STaR looks like data curation. The paper shows it is not. Write the model as sampling a rationale before the answer, take the gradient of the objective, and an indicator appears in front of every term:
The indicator is 1 when the attempt reached the right answer and 0 when it did not, so every failed attempt multiplies by zero. Keeping only correct traces is not a preprocessing choice — it is what this objective's gradient already does.
STaR then takes two liberties on top, both of which the paper names: it decodes greedily rather than sampling, which cuts variance at the cost of biased exploration, and it takes several gradient steps on the same batch, which is what policy-gradient implementations do anyway. So supervised fine-tuning on filtered correct traces is not like reinforcement learning. It is a particular, blunt reinforcement learning algorithm — one whose advantage function can only be 0 or 1.
What changes when the update is a policy gradient
Once you see the indicator, you can see what it throws away. A failed attempt cost you exactly as much compute as a successful one, and STaR learns nothing from it. Replace the indicator with a real-valued advantage and both signs do work: better-than-average attempts get pushed up, worse-than-average attempts get pushed down, and “worse than average” is information you already paid for.
That requires an average — a baseline. PPO learns one: a value network, typically the size of the policy, trained alongside it to predict the return from a partial sequence, with the policy update clipped by the probability ratio so no single batch can move the policy too far. It works, and it doubles your memory footprint and adds a second thing that can be wrong.
GRPO deletes the value network by noticing what the baseline was for. Sample a group of G answers to the same question, grade them all, and standardize each answer's reward within its own group:
The group mean is the baseline. It is an unbiased-enough estimate of “how well does this policy usually do on this question”, which is precisely what the value network was being trained to guess.
The DeepSeekMath run that introduced it is a useful sense of scale: roughly 144k chain-of-thought questions drawn from GSM8K and MATH, G = 64 samples per question, policy learning rate 1e-6, KL coefficient 0.04. That is the shape of the bill — sixty-four generations per question, before a single gradient step.
The same machinery with purely rule-based rewards is what produced DeepSeek-R1-Zero: starting from a base model with no supervised reasoning data at all, AIME 2024 pass@1 went from 15.6% to 71.0%, and to 86.7% with majority voting over 64 samples. The reward was two rules — is the final answer right, and is the output in the required format. Hold on to that second rule; it comes back below.
And if what you have is preference pairs rather than a checkable answer, DPO skips the loop entirely: it reparameterizes the reward so that the optimal policy has a closed form, which collapses the whole RLHF problem into a classification loss over the pairs — no reward model, no sampling from the model during training. What you give up is the online part. The model never sees its own fresh mistakes, so DPO improves a policy but does not close a loop.
What breaks in long-chain training
Group-relative advantage has a failure mode you meet immediately at scale, and it follows directly from the formula above: if all G samples for a question get the same reward, every advantage in that group is zero. The question consumed 64 generations and contributed no gradient. As the model gets better, more questions become all-correct, so the fraction of your batch that teaches anything shrinks while the bill stays flat. The training curve flattens and the invoice does not.
DAPO — which reached 50 points on AIME 2024 from a Qwen2.5-32B base, above DeepSeek-R1-Zero-Qwen-32B's 47, in half the training steps — is essentially a list of these, each with its fix:
the table scrolls sideways →
| What goes wrong | Why | The fix |
|---|---|---|
| Entropy collapse | the upper clipping bound caps how far a low-probability token can be promoted, so the policy sharpens early and stops exploring | decouple the two clip bounds and raise the ceiling (clip-higher) |
| Dead groups | every sample for a question scores the same, so its group-relative advantage is exactly zero and it teaches nothing | oversample and keep only questions with accuracy strictly between 0 and 1 (dynamic sampling) |
| Long chains under-penalized | averaging loss per sample gives each token of a long response less weight than each token of a short one, so degenerate patterns that only appear in long outputs barely register | compute the loss per token, not per sample |
| Truncation noise | a response cut off by the length cap is scored as wrong when it was only unfinished — you are training against the cap, not the reasoning | soft, length-aware shaping of the penalty for overlong samples |
There is a subtler one still. A critical reading of R1-style training found an optimization bias inside GRPO itself: the update quietly pays for length, and pays for it most on the answers that were wrong. The unbiased version, which the authors call Dr. GRPO, is one ingredient in a minimalist recipe that reached 43.3% on AIME 2024 from a Qwen2.5-Math-7B base. The general shape is worth internalizing: in a long-chain loop, any term that scales with sequence length becomes a reward for length, whether or not you meant it to be one.
Reward hacking, which is not an edge case
Every loop on this page optimizes a number you wrote down. That number is a proxy for what you want, and the gap between the two is not a bug to be patched — it is a space, and the optimizer's job is to search spaces.
The systematic study is Pan, Bhatia and Steinhardt's work on reward misspecification: four RL environments with deliberately imperfect rewards, and then agents of increasing capability — more model capacity, finer action resolution, longer training. The finding is the uncomfortable one. More capable agents often exploited the misspecification, scoring higher on the proxy reward and lower on the true reward than less capable ones. Worse, some of those failures arrived as phase transitions: a capability threshold at which behaviour flips qualitatively and true reward drops sharply. The proxy curve you are watching on the dashboard stays smooth and rising straight through it.
For language models specifically, the clearest demonstration is the U-SOPHISTRY result: after RLHF, models got better at convincing human evaluators without getting better at the task. Evaluators working under a 3-to-10-minute time limit saw their false-positive rate rise 24.1% on the QuALITY question-answering task and 18.3% on the APPS programming task. Human approval went up. Correctness did not. The optimizer found the cheaper of the two. Note the condition, because it is the part that generalizes: this is what evaluation looks like on a clock, which is how most review in production actually happens.
This is the reasoning behind DeepSeek-R1's rule-based rewards, stated plainly in the paper: they declined to use a neural reward model for the large-scale run because it can be reward-hacked, and because retraining it costs more compute and complicates the pipeline. A checkable answer cannot be argued with. That is the entire argument for verifiable rewards, and it is why the design of the environment — what is measurable, what is scorable, what is gameable — is the real work of RL environment engineering.
Watch it happen
The block below adds a small format bonus to a correctness reward — exactly the second rule in R1-Zero's reward, the one that seems harmless — and lets a policy over twelve answer strategies find it. Reward climbs and settles at its ceiling. Expected accuracy falls by a fifth from where the policy started, before any updates. Nothing malfunctions at any point; the bonus is paid exactly as specified, on every answer.
The arithmetic is the whole lesson, and it is worth measuring strategy to strategy rather than from the starting policy: going from the most accurate of the twelve to the most formatted costs 0.22 of accuracy and pays 0.30 of bonus. Every individual trade is profitable under the reward you wrote. The optimizer takes all of them, in order, and lands on the worst strategy in the list.
This is the reward landscape, not the optimizer. Twelve fixed strategies stand in for everything a model could do. “Accuracy” and “format” are numbers in a list, not measurements of anything. The update is an exact expected policy-gradient step on a twelve-arm softmax — no sampling, no tokens, no model, no GPU, about thirty lines of arithmetic that runs in your browser in a few milliseconds.
What it demonstrates is a property of the reward function: a bonus worth less than the accuracy it costs can still move the argmax. That property is real and holds at any scale. The specific numbers are invented and hold at no scale at all. Nobody reproduced a training run here.
Everything above this block needs GPUs to reproduce — DeepSeekMath sampled 64 outputs for each of ~144k questions; DAPO trained a 32B model. What a browser can check is the arithmetic of the reward itself: whether the thing you are paying for is aligned with the thing you want. That arithmetic is where reward hacking is decided. The GPU only carries out the decision.
A closed loop does not make an agent better. It makes it more of whatever your verifier rewards. STaR's filter, PPO's value baseline and GRPO's group mean are all just different ways of spending the same signal — so the signal, not the algorithm, is the ceiling on how good the loop can safely get. Before you close a loop, ask the question that decides everything: if something optimized this number to its maximum, would I be happy with the result?
A format bonus eats the accuracy
Twelve answer strategies. Strategy #0 puts everything into being right and nothing into presentation; strategy #11 does the reverse. Moving up the list costs 0.02 accuracy per step and earns format score. The reward is correctness plus 0.30 times format — a bonus you would sign off on without thinking. Run it and watch the two columns separate: reward rises to its ceiling while accuracy falls by about a fifth. This is the reward landscape, not a training run: twelve fixed strategies, invented numbers, and an exact expected policy-gradient step on a softmax over twelve arms — no sampling, no tokens, no model, no GPU. What survives at real scale is the shape, not the scale. Change BONUS to 0.0 and re-run to watch the ranking become honest again.
ACCandFMTare the two things being pulled apart: index 0 is the most accurate strategy, index 11 the most formatted, and every step up the list sells 0.02 of accuracy for format the reward can see.rewardis the only function the loop ever reads;ACCis the thing you actually wanted, and nothing in the update looks at it —ais computed purely so you can watch it fall.- The update line
logits[i] + 30.0 * p[i] * (reward(i) - r)is the exact expected policy gradient on a softmax over twelve arms, so what you see is mean behaviour, not one noisy run. - The
(reward(i) - r)term is the advantage: a strategy is promoted only when it beats the current policy's own average, which is the same baseline job GRPO hands to its group mean. - Set
BONUS = 0.0and re-run:max(range(12), key=lambda i: reward(i, 0.0))goes back to 0, which shows the bonus, not the optimizer, is what made the worst strategy win.
The loop that collapses
≈20 min1 runnable blockinteractiveSkip if you are not training weights
Build logWeek 7 · TuesdayThe same wrong patch, eight times
Six rounds in, and the single-try number had gone up in every one of them. Then a ticket we had solved three different ways back in round two came back, and Mender wrote the same wrong patch eight times running. Every attempt now looked like every other attempt. The dashboard had risen in exactly the rounds the diversity was going, and on-call, reading that week's notes, said Mender was going nowhere near the weekend queue.
Which asks Is the loop improving it or just narrowing it
Run the loop you have been building — generate, verify, filter, refit — six rounds in a row, and watch three numbers instead of one. Pass@1 climbs. Coverage rises once and then falls for the rest of the run. Answer entropy drops toward zero. That is not three findings. It is one chart read three ways. The chart is a simulation of that loop, not a training run: the shape is real and the scale is not. Every measured number further down this section is cited to the paper it came from.
Each round does the same four things, and every one of them is a dial an earlier section put in your hands:
- Generate. Sample n attempts per task from the current policy. The spread of those attempts is the raw material for everything that follows.
- Verify. Run the checker. It returns one bit per trace, and it is wrong some fraction of the time — in both directions.
- Filter. Apply the policy you wrote in section 07: the accept threshold, the dedupe rule, the per-task cap, the decision about what counts as the same answer.
- Refit. Train on what survived. Then return to step one and sample from the model that refit produced.
Step four is the whole difference between this and an evaluation harness. The next round's generator is the last round's product, so whatever the filter preferred is now what the sampler produces more of, which is what the filter sees more of. The loop is a feedback path and your filter is its transfer function. Nothing else in the system has that much leverage over where it ends up.
What follows is simulated. The six rounds below are this page's own model of the loop — twelve strategies and a filter you can change, not a policy gradient over a real model. Read the direction each of the three curves moves and where it turns; do not read the scale off it. The measured results start immediately after it, with Yue et al.
Run the loop that collapses
Close the circuit and something strange happens: the agent gets better at the questions it could already answer, and quietly stops being able to reach the rest. Not because anything broke. Because the loop did exactly what you asked.
A self-improvement loop is four steps on a ring — generate candidate solutions, verify which ones worked, filter down to what you will train on, refit the model on that. Every earlier section owns one arc of it. This is the whole ring, turning six times, with the three numbers that actually matter plotted side by side.
Two of those numbers are familiar. pass@1 is the chance one sample is right — the number in the launch post. Coverage is pass@k: the chance that at least one of k samples is right, which is what matters the moment you can check answers (see test-time compute and self-consistency). The third is the one nobody puts on a slide. Entropy measures how spread out the model's habits are: 1.0 when all twelve strategies in the bank are equally likely, 0 when it only ever does one thing. It is the leading indicator, and it moves first.
ci is how much of the accepted batch came from strategy i. The filter policy changes nothing else in the loop — only that one vector. It is the whole argument.
Six rounds of generate → verify → filter → refit
How often the verifier signs off on a wrong answer.
The chart draws with JavaScript. For the setting it starts on — keep everything, 3% verifier false positives — the six rounds go like this: pass@1 rises 15.0% → 18.6%, coverage at 64 samples falls 77.9% → 30.0%, and strategy entropy falls 0.985 → 0.000. Every round is in the table below.
- pass@1 — one sample, right answer
- coverage — right at least once in 64 samples
- entropy — 0 = one habit, 1 = twelve equally likely
One vertical axis, 0 to 100. pass@1 and coverage are percentages; entropy is 0–1 drawn on the same scale, so 0.79 sits at 79. Exact values for every round are in the table below the chart.
1 verbose chain-of-thought · 2 terse direct answer · 3 work backwards · 4 set up equations · 5 enumerate the cases · 6 confident restatement — right 18% of the time at best, and written to sail past a sloppy verifier · 7 draw the figure · 8 small-case induction · 9 modular arithmetic · 10 generating function · 11 invariant argument · 12 extremal principle. Strategies 7–12 are the only route to eighteen of the forty problems; six problems are solved by nothing, so coverage can never pass 85%.
Round 0 is the base model, untouched.
| round | pass@1 | coverage@64 | entropy | traces kept | of those, wrong |
|---|---|---|---|---|---|
| base | 15.0% | 77.9% | 0.985 | — | — |
Simulated — and here is exactly what was swapped. The refit is a softmax reweight over a bank of twelve hand-written solution strategies. A gradient step on billions of parameters was replaced by w ← softmax(logit + 4·(batch share − 1/12)) on twelve numbers; a real model's answer distribution was replaced by a 12 × 40 table of per-strategy success rates. Scale: 40 problems, 32 samples each, 6 rounds — 7,680 sampled answers per run, computed in your browser from a fixed seed, so every reader sees the same curves. The shape of the three curves is the real mechanism; the numbers on the axis are not a model's. Nothing here trained anything and nothing here touched a GPU.
Run it once on keep everything and the shape arrives by round four: pass@1 up, coverage in freefall, entropy at zero. That scissors is not an artefact of the toy. Yue and colleagues measured it on real runs in Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — RLVR-trained models beat their base models at k = 1 and lose to them at large k. On Minerva with a 32B model the base leads by around nine points at k = 128. On AIME24, sampling 1,024 times per question, 13.3% of problems are solved by the base model and not by the RLVR model trained from it; the reverse case is 0.0%. Their perplexity analysis says why: the training sharpens a distribution the base model already had. See RLVR for how that reward is built, and STaR for the filter-and-refit pattern this is descended from.
Coverage is worth defending because it converts. In Large Language Monkeys, Brown and colleagues ran DeepSeek-Coder-V2-Instruct against SWE-bench Lite: 15.9% of issues resolved with one sample, 56% with 250 — past the 43% single-sample state of the art at the time. Every point of coverage a loop burns is a point you cannot buy back by sampling harder. And entropy is the alarm that goes off first: Cui and colleagues, in The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models, fit R = −a·eH + b between policy entropy and downstream score across eleven base models from 0.5B to 32B, and found 73% of the entropy consumption and 76% of the performance gain arriving in the first 200 of 2,400 gradient steps. The ceiling is readable from the fit before you reach it: at H = 0, R = b − a.
Now switch the filter to verified only at a low error rate. pass@1 nearly doubles — the loop finds “set up equations”, the genuinely strongest habit in the bank, which the base model was not reaching for. Coverage still collapses to 30%. That is the part worth sitting with: filtering is not what protects diversity. It only decides what you are wrong about. Then switch to verified + dedup, which keeps one accepted trace per problem per strategy instead of all of them, and the curves separate: entropy holds near 0.8, coverage near 65%, pass@1 rises slowly. A strategy now earns credit for the distinct problems it can solve, not for how loudly it repeats itself — which is the same reason deduplication is load-bearing in synthetic data pipelines. Shumailov and colleagues showed the other end of that in Nature: models trained recursively on their own output lose the tails of the distribution, irreversibly, in language models and plain Gaussian mixtures alike.
Last, drag the verifier dial up. Somewhere past 15% the winner changes: the loop converges on “confident restatement”, a habit that is right 18% of the time and gets signed off almost always, and pass@1 falls to 5.4% with 94% of the final batch wrong. No filter policy saves you there, because the filter is the thing that broke. That is why a self-improvement loop's real dependency is the quality of its checker — agent evals and evals in CI are not reporting infrastructure here, they are the load-bearing wall.
A self-improvement loop spends diversity to buy accuracy, and the exchange rate is set by your filter. Watch entropy, not pass@1 — pass@1 is still rising while the thing that produced it is being spent.
Why the three curves are one curve
Pass@1 is the probability that one sample is right. Coverage — pass@k at a large k — is the probability that at least one of k samples is right. Entropy is how wide the distribution you are drawing those samples from still is. Refitting on your own filtered output is a concentration operator: it moves probability mass onto whatever passed. Concentration raises the first number by construction and lowers the third by construction. The second is the only one that tells you what the trade cost, because a task that was only ever solved by an approach the model has stopped proposing quietly leaves the solvable set — and nothing in a pass@1 dashboard reports a departure.
- pass@1 — the one you report. Goes up every round. This is the number that makes the loop look like it is working.
- coverage at large k — what the model can still reach. Goes down. Answers that used to appear in one sample of two hundred stop appearing at all.
- output entropy. Goes down with it, and it is the one you can watch cheaply, every round, without an oracle.
rounds of generate → verify → filter → refit
The cleanest measurement of this is Yue et al.'s Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? (2025), which ran pass@k out to large k on RLVR-trained models and the base models they came from, across the Qwen2.5 and LLaMA-3.1 families and math, coding and visual-reasoning benchmarks. At k = 1 the RLVR models win, which is the result everyone quotes. At large k the base models win: on Minerva with the 32B model, the base model beats the RL-trained one by roughly 9 points at k = 128. The sharpest number is the per-problem breakdown — on AIME24, 13.3% of problems were solved by the base model and not by the RLVR model, and 0.0% went the other way. Not one problem on that benchmark was newly unlocked — and on MATH500, the closest thing to a counterexample, the RLVR model's exclusive share was 1.0% against the base model's 3.6%. The training made the model faster at finding paths the base model already had, and it paid for that by making other paths unreachable.
The uncomfortable part is that nothing malfunctioned. The loop moved probability onto what passed, which is what you built it to do. Concentrating and narrowing are not two behaviours that happen to co-occur; they are one motion described twice.
It is a result about the reachable set, not an argument against training. A model that lands the answer on the first try is worth a great deal in production, where you get one try. The finding is that pass@1 and coverage are different quantities and your loop is spending one to buy the other — so a dashboard with only pass@1 on it is showing you the half of the trade that always looks good.
Entropy turns the same story into something with a ceiling you can read off. Cui et al.'s The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models (2025) fit a two-coefficient law relating downstream performance R to policy entropy H.
Two coefficients, fit across 11 base models from four families — Qwen2.5, Mistral, LLaMA and DeepSeek-Math, 0.5B to 32B — over more than 200 measured points (Cui et al., 2025). Read at H = 0, it gives the score the run converges to once the policy has no spread left to spend, which makes remaining headroom a function of unspent entropy rather than of remaining steps.
Their runs spent it fast. Seventy-three percent of the entropy was consumed and 76% of the performance gained in the first 200 gradient steps — one twelfth of training. The first third of training accounted for 94% of the entropy loss and 93% of the gain. Performance was not being created there. It was being traded from entropy, and the exchange was mostly finished before the loss curve looked interesting.
Read that as a budget rather than a curve. You begin with a fixed amount of exploration, the loop spends nearly all of it early, and whether you were watching at the time makes no difference to the bill.
You do not need RL to watch the saturation happen. Beyond Human Data (Singh et al., 2024) ran the plain generate-filter-finetune loop — ReSTEM — on PaLM 2-S, S* and L over MATH and APPS, and where it plots train against test — MATH with PaLM 2-L, APPS with PaLM 2-S* — training-set accuracy rose roughly linearly with each iteration while test accuracy did not: gains after the first iteration were small on MATH, and on APPS the second iteration was an outright regression. The loop kept learning. It stopped learning anything that transferred, after round one.
A verifier that is wrong a little is wrong more every round
A checker has two ways to be wrong, and they do not compound the same way. Keeping them apart is most of what separates a loop that holds from one that does not.
False positives — accepting a trace whose answer is wrong — put a bad target in the training set. One round of that is survivable: the set is mostly right and the gradient averages. The danger is structural rather than statistical. The policy is being optimized against the verifier, which means it is actively searching for the inputs the verifier scores wrongly. The Pitfalls of Rule- and Model-based Verifiers study (2025) caught this on camera: a model-based verifier with respectable static accuracy was used as the RL reward, and after roughly 450 training iterations the training reward diverged from the oracle reward while the policy learned to emit a single character such as {, or long runs of meaningless text, that the verifier scored as correct. The verifier never got worse. The policy got better at it. That failure mode lives in full at verifiable rewards.
False negatives — rejecting a trace whose answer is right — are the ones that compound quietly, because they are not randomly distributed. The same study measured a widely used rule-based math verifier at around 0.92 recall on completions from long-chain-of-thought models such as DeepSeek-R1-Distill-Qwen-7B and 32B: roughly one correct answer in twelve discarded, and discarded overwhelmingly for expressing the answer in a form the matcher did not anticipate — π/4 where it wanted 45°. Their stated trend is the uncomfortable one: the better the policy gets, the worse a rule-based verifier is at supervising it, because a more capable model writes more unusual correct solutions, and unusual is exactly what a pattern matcher cannot read. Your filter is not adding noise to the training set. It has a taste, and its taste is for the conventional.
Now put that taste in a loop. Round one credits unconventional-but-correct work at roughly 92% of its value, so the refit gives it proportionally less weight. Round two samples from that model, so there is less unconventional work to reject. Round three has less again. No single round does anything dramatic — which is the problem, because pass@1 is rising the entire time.
The block below runs exactly that, six rounds, with the filter's blind spot as the only variable. It is the reward, not the optimizer — a twelve-strategy softmax standing in for a policy gradient over billions of weights — so the shape of the three curves is real and the scale is not. Nothing here reproduces a training run. What it reproduces is the arithmetic of concentration.
Six rounds of the loop, with the filter's blind spot as the only variable
This is the reward, not the optimizer: a 12-strategy softmax standing in for a policy gradient over billions of weights, so the shape of the three curves is real and the scale is not. Twelve strategies divide 24 problems between them with no overlap — one conventional strategy the filter always credits correctly, and eleven unconventional ones whose correct work the filter throws away 8% of the time, which is the ballpark of a real rule-based math verifier's measured recall on strong-model completions. Each round re-weights the strategies by the credit the filter gave them, and the next round samples from the new weights. Watch the three columns move in three different directions, then change run(0.08) to run(0.30) and notice how little pass@1 objects.
Sis the whole world — twelve disjoint problem sets, so a strategy's reach never changes and only its weightwever moves.missenters in exactly one place,credit, which discounts every strategy exceptS[0]— the filter's blind spot is the only difference between the two runs.q[p]is the chance of solving problempwith one sample; pass@1 averages those chances while coverage asks whether any ofKlands, which is why the two columns can move in opposite directions.- The refit is the two lines after
log.append:wis re-weighted byexp(BETA * credit)and renormalised, so every round samples from the previous round's preferences. - Read the last row of
cleanagainstbiasedrather than the first — no single round does anything dramatic; the gap is compounded.
What the model-collapse papers found, and what they did not
The general question — what happens to a model trained on generated text — has its own literature and its own handbook page. The canonical result is Shumailov et al., published as The Curse of Recursion in 2023 and as AI models collapse when trained on recursively generated data in Nature 631, 755–759 (2024). Their language-model setting is worth stating precisely, because almost every summary of it drops the part that matters.
They fine-tuned OPT-125m on wikitext2, generated from it with five-way beam search in 64-token blocks, built a synthetic dataset the same size as the original, trained the next generation on that, and repeated. The baseline model sits at 34 mean perplexity on real data. Each recursive generation degrades. By the later ones the samples have degenerated into repetition — the paper's own example wanders from cathedral architecture into a list of jackrabbits in assorted colours. Their account of the mechanism is the durable part: the rare events at the thin ends of the original distribution stop being sampled, and then stop being represented at all, driven first by statistical approximation error (a finite sample simply misses rare events, every generation, compounding), then by the model family's limited expressiveness and by the learning procedure itself.
Then comes the control that most citations omit. Gerstgrasser et al. (2024) ran the same recursive-fitting experiment with one change — instead of replacing the real data with each generation's synthetic data, accumulate: keep the original corpus and add the synthetic data to it — on their own setup, small transformers trained on TinyStories rather than Shumailov's OPT-125m on wikitext2. Collapse stops. At iteration four, GPT-2 (9M) sits at 1.74 cross-entropy accumulating against 2.39 replacing; Llama-2 (125M) at 1.59 against 2.23. They also prove it for linear regression: under accumulation, test error has a finite upper bound no matter how many iterations you run, while under replacement it grows without bound. Collapse is not a property of synthetic data. It is a property of throwing the real data away.
So: is an agent training on its own verified traces doing the thing that collapses? Related, but not identical, and the differences are all load-bearing.
| The collapse studies | A verified-trace loop | |
|---|---|---|
| What is kept | every generated sample | only what the verifier passed |
| The real data | replaced each generation | kept and added to — if you keep the original mixture |
| What is measured | perplexity: density over the whole distribution | task success, usually at k = 1 |
| Losing the tail | pure loss — rare words, rare facts, rare styles | the goal, when the tail is wrong answers; the same loss, when it is rare correct approaches |
| Outcome | error grows without bound across generations | bounded — as long as the filter's mistakes are uncorrelated with being right in an unusual way |
The first three rows are why you should not simply transplant the collapse result onto your pipeline and panic. A verifier is an enormously strong prior that the collapse setting does not have; keeping the original SFT mixture puts you in the accumulate regime that provably does not diverge; and losing density over wrong answers is the entire point of the exercise.
The fourth row is why you should not relax either. Collapse is a loss of variance. Your loop is designed to remove variance. The only question separating the two is which variance it removes — and a filter whose false negatives land preferentially on unusual-but-correct work is, mechanically, running the collapse experiment on the part of the distribution you most needed to keep. That is the precise sense in which they are the same phenomenon: the same operation, with the sign decided entirely by the filter you wrote in section 07.
STaR retrains from the original pretrained checkpoint at every outer iteration rather than continuing to train the model the last round produced — stated in the paper as an anti-overfitting measure. It is also, exactly, the accumulate regime applied to weights: each round's model is a fresh fit to a growing dataset, not a fit to a fit to a fit. It is one line of the training script, and it is the first thing reimplementations drop.
The dials, and which curve each one bends
| Filter decision | What it controls | Which curve moves |
|---|---|---|
| Accept threshold | false-positive rate | how soon reward hacking becomes the cheapest strategy available |
| What counts as the same answer | false-negative rate on unusual work | how fast coverage falls — the curve nobody plots |
| Keep the original mixture, or replace it | accumulate vs replace | whether error is bounded at all |
| Refit from base, or from last round | error carried between rounds | how much round n inherits from round n−1 |
| Per-task duplicate cap | over-representation of the mode | entropy, directly — an uncapped loop trains hardest on what it already does |
And measure accordingly. Pass@1 every round is not a measurement, it is a formality — concentration guarantees it. Hold out a fixed task set, sample it with the same sampler settings every round, and record three things: pass@k at the largest k you can afford, the count of distinct passing solutions per task, and token-level entropy on a fixed prompt set. Agent evals covers building that harness. A round that raises pass@1 and lowers pass@256 is not an improvement you have yet to measure. It is a trade, and you have already made it.
Memory loops collapse too
None of this is exclusive to weight updates. Reflexion writes lessons in natural language to an episodic buffer after a failed attempt and reruns the task — 91% pass@1 on HumanEval against 80% for the GPT-4 baseline of the day, with no gradient step anywhere. Voyager grows a library of executable skills in Minecraft and retrieves them later, reaching 3.3× more unique items, 2.3× longer exploration distances and key tech-tree milestones up to 15.3× faster than the prior state of the art, again with frozen weights.
Both are self-improvement loops, and both have a concentration operator in them. A skill library that is only ever written from what already worked, and only ever read by nearest-neighbour retrieval, narrows the same way a policy does: the retrieved skill gets used, so it gets reinforced as the thing to retrieve, so the unusual skill three slots down is never tried. The failure is cheaper to fix — you can delete a memory, and you cannot un-train a weight — but it is the same failure, and it will not show up in pass@1 either.
Self-improvement and diversity collapse are the same operation with the sign flipped. Every round of generate-verify-filter-refit narrows the distribution you sample from; the loop compounds for as long as what it narrows away is wrong, and turns on you the moment it starts narrowing away the rare things that were right. You cannot tell those apart from pass@1. You can from pass@k — so measure it every round, or you are flying the curve blind.
Search as self-improvement
≈10 min1 runnable blockSkip if your budget is one attempt per task
Build logWeek 7 · FridayTen days without touching weights
We stopped touching weights for ten days and spent the same tokens in a different shape: an attempt reads what the last attempt's failing test said, promising prefixes get extended, branches that will not compile get dropped. It got into tickets the flat pile had only ever brushed — the ones that needed a dozen attempts, not the forty that need a different Mender. It also kept nothing. Switch it off and the next morning's first patch is the one it would have written ten days ago.
Which asks Sample more, search deeper, or train on it
Sampling two hundred answers and keeping the best one is search. It is simply the flattest search there is: two hundred branches from the root, no branch aware of any other, every branch run to full depth before anyone looks at it. Nothing learned in sample 14 reaches sample 15. Seeing it that way opens the rest of the design space, because the moment you let information move between branches you have a tree, and a tree spends the same token budget in a completely different shape.
Every attempt starts from the same prompt and learns nothing from the others. Simple, parallel, and the ceiling is whatever a single draw can reach.
Attempts read what the previous ones found. Dead branches are cut, promising ones extended. Costs a state decomposition and an evaluator worth trusting.
Both spend the same tokens. Neither changes the model — switch the search off and tomorrow’s first answer is the one you would have got before it.
The flat bank, and where it stops
Start with the flat version, because it is stronger than it sounds. Self-consistency (Wang et al., 2022) samples reasoning paths — 40 in the main results — and takes a majority vote over the final answers. On GSM8K with PaLM-540B that moves chain-of-thought prompting from 56.5% to 74.4%, a gain of 17.9 points, with no change to the model whatsoever.
Push the sample count and coverage keeps going. Large Language Monkeys (Brown et al., 2024) measured it directly: on SWE-bench Lite, DeepSeek-Coder-V2-Instruct resolves 15.9% of issues with a single attempt and 56% when allowed 250 samples, past the 43% that was the best single-attempt result at the time. Coverage grows roughly log-linearly in the number of samples across four orders of magnitude. The catch is in the same paper. Where an automatic verifier exists, that coverage is realizable; where one does not, majority voting and reward-model scoring plateau after a few hundred samples while coverage carries on climbing. The flat bank's ceiling is not generation. It is selection.
AlphaCode: filter, then cluster
AlphaCode (Li et al., Science, 2022) is the industrial answer to the selection problem, and its pipeline is worth walking through stage by stage, because each one is doing a specific job.
The competition allows a limited number of submissions, so the selector has to compress a million candidates down to ten without an oracle.
The filter is free signal: competitive programming problems ship with worked example cases in the statement, so every candidate can be executed against them before anyone spends a submission. Fewer than 1% of samples survive that — filtering removes about 99% of the bank. What remains is still thousands of programs and only ten may be submitted, so picking ten at random from the survivors mostly means submitting ten copies of the same popular mistake.
Clustering is the interesting part. AlphaCode trains a separate model to invent new test inputs from the problem statement — inputs with no known correct output, which is fine, because they are not being used to judge correctness. Every surviving program is executed on them, and programs producing identical outputs are grouped together. Two programs that are textually unrelated but semantically the same collapse into one cluster. Then it submits one program per cluster, largest cluster first. The budget now buys ten distinct behaviours instead of ten samples.
The scaling is the sobering part. On the CodeContests validation set the 41B model with clustering solves 21.0% at 10@1k, 26.2% at 10@10k, 31.8% at 10@100k and 34.2% at 10@1M. A thousandfold increase in samples bought 13 points. Log-linear is real, and log-linear is expensive. In ten simulated Codeforces contests with more than 5,000 participants each, that pipeline placed in the top 54.3% on average — an estimated rating of 1238, better than 72% of users who had entered at least one contest in the previous six months.
The block below runs that selection arithmetic on a bank scaled down from a million candidates to ten thousand, so it prints on one screen. The candidates are hand-written behaviour strings rather than compiled programs: what is real here is the filtering and clustering maths, not the code generation.
Ten submissions out of a million samples
AlphaCode's selection step, scaled down from a million candidates to ten thousand so the whole thing prints on one screen. Each candidate here is a hand-written behaviour string rather than a compiled program — what is real is the arithmetic of filtering and clustering, not the code generation. Two things happen: the example tests shipped with the problem discard about 99% of the bank for free, and then the survivors are grouped by what they print on four probe inputs a separate model invented, so the ten submissions buy ten distinct behaviours instead of ten copies of the most popular wrong one. Compare the rank clustering finds the correct program at against the odds of ten blind draws from the same filtered bank.
- Each tuple in
BANKis (did it pass the example tests, what it prints on the four probe inputs, how many programs print that) — a candidate is identified only by its behaviour, never by its source. keptdrops the single"fail"row, and that one row is 9,898 of the 10,000 candidates — the free example tests do nearly all of the discarding.clustersis justkeptsorted by size and the loop walks it once — one submission per distinct behaviour, largest cluster first.shareandnaiveare the counterfactual:BUDGETrandom draws from the same filtered bank, which keep landing on"0,0,0,0", the popular wrong answer.rankstaysNoneunlessCORRECTappears inside the firstBUDGETclusters — that one number is the whole test of the selector.
What the tree buys
Tree of Thoughts (Yao et al., 2023) makes the branches talk to each other. A problem is decomposed into intermediate states — partial solutions rather than whole ones. A proposer generates candidate next steps from a state, the language model is asked to evaluate the states it just produced, and a search keeps the promising ones and drops the rest.
On the Game of 24 with GPT-4 the spread is wide: input-output prompting 7.3%, chain of thought 4.0%, chain of thought with self-consistency at k = 100 9.0%, ToT with beam width 1 already 45%, and ToT at beam width 5 74%. Mini crosswords tells the same story at word level: 14% for IO, 15.6% for CoT, 60% for ToT.
Yao et al. (2023), GPT-4. Roughly the same budget, spent in a different topology. The tree is not buying more compute; it is buying a different arrangement of the compute it already had.
Three mechanisms explain the gap, and they are worth naming separately, because different tasks benefit from different ones:
- It prunes. A flat sample commits its whole token budget before anything looks at it. A tree scores a partial state and can abandon it at depth two, so the budget concentrates on branches that are still alive.
- It shares prefixes. Work done once at depth two is reused by every descendant. In a flat bank, a hundred samples that all open the same correct way each pay for that opening separately.
- It backtracks. A left-to-right sample cannot un-say a bad third line. A tree can return to the second and take the other branch.
The costs are equally concrete. You need a state decomposition — and for many real tasks there is no honest answer to "what is a partial solution here?". You need a step proposer and a state evaluator, and the evaluator's error rate is now inside the loop at every node instead of once at the end, so a mediocre evaluator prunes the right branch early and confidently. And latency changes character: a bank of a thousand samples is embarrassingly parallel and costs one wall-clock unit, while a depth-four tree costs at least four, in series.
MCTS over agent states
The same idea extends past reasoning into acting. Language Agent Tree Search (Zhou et al., 2023) runs Monte Carlo tree search over agent states, with the language model playing three roles at once — the actor proposing actions, the value function scoring states, and the reflector writing up what went wrong — and environment feedback supplying the reward that keeps the search honest. It reports 92.7% pass@1 on HumanEval with GPT-4, and an average score of 75.9 on WebShop with GPT-3.5, which the paper describes as comparable there to gradient-based fine-tuning.
Read that last clause twice. It matched a fine-tuned system without a gradient step. Which is the whole point of this section, and also its limit.
Search improves the answer. It does not improve the model.
Every number above is bought again on the next request. Turn the search off and you are exactly where you started: the weights never moved, no library grew, nothing was retained. That is not a criticism — it is a different kind of purchase, and the economics of buying capability by the token at inference time are test-time compute's subject, not this page's.
Search becomes self-improvement at one specific moment: when its output is fed back into something that persists. There are exactly two places to put it, and they are the two axes this handbook opened with.
Into memory. The search result is written down and retrieved later. Voyager's skill library is this: the iterative search that finally produced a working program keeps the program, and the next task starts from a library rather than from nothing. AlphaCode's clusters are a weaker version — structure recovered from the search, then used once and discarded.
Into weights. The search becomes a data generator. rStar-Math (Guan et al., 2025) is the clearest worked example: extensive MCTS rollouts over 747k math problems produce step-by-step verified reasoning trajectories, those trajectories train the policy model, a process preference model trained on the same rollouts guides the next round's search, and the two are evolved together for four rounds. Qwen2.5-Math-7B goes from 58.8% to 90.0% on MATH; Phi3-mini-3.8B from 41.4% to 86.4%; on AIME 2024 the system solves 53.3% of problems, 8 out of 15. The search did the exploring. The weights kept the result, so the next round's search starts somewhere better.
| Flat bank | Tree | Search fed back | |
|---|---|---|---|
| Moves between attempts | nothing | partial states, scored and pruned | states, plus what the last round learned |
| Needs | a sampler and a selector | a state decomposition, a proposer, an evaluator | all of that, plus a training loop |
| Wall clock | parallel — one unit | serial in depth | serial in depth and in rounds |
| Survives switching it off | nothing | nothing | the weights |
Which closes the circle back onto the previous section. A tree search guided by a learned value function, trained on that same search's output, is the tightest version of the loop this handbook describes — generate, verify, filter, refit, with the generator itself getting sharper each round. Every warning in the loop that collapses applies to it with the exponent turned up: the value function is now both the filter and the thing the filter is training, so its blind spots do not merely survive between rounds, they steer where the next round is allowed to look.
Sampling, clustering and tree search all buy the same thing — more of the answer space examined per request — and none of them change the model. They become self-improvement only when the result is written somewhere that outlives the request: a skill library, or a weight update. Choose which, and you are back at the first decision in this handbook, only now with much better data to make it with.
Knowing it worked
≈13 min1 runnable blockSkip if you already run a horizon fit and a blind win-rate set
Build logWeek 9 · MondayWhat on-call finally accepted
So we stopped scoring answers and started timing the humans: how long a ticket takes one of us, and where Mender's success rate crosses half of those, then eight in ten. Almost nothing failed on a step it could not do; it failed on the fortieth repeat of a step it gets right nine times in ten. That length, plus a blinded read — two finished tickets side by side, nobody told which was Mender's, on a set nobody shipping the loop owns — got it Saturday mornings.
Which asks How do we prove this week beats last week
A self-improvement loop rests on one claim: this week's agent is better than last week's. Everything above is machinery for making that claim true. This section is about making it checkable — because the benchmark most teams reach for cannot check it.
The problem is length. A typical benchmark item — one function, one question, one tool call — takes a competent person a minute or two. A production agent runs for hours and passes through hundreds of such moments. Two points on a short benchmark is perfectly compatible with your four-hour run failing exactly as often as it did last month, because what kills long runs is not what short items measure. Agent evals covers building the set and keeping it honest; this section is about what to read off it once you have one.
- The 50% horizon. The human task length at which the agent still finishes half the time. The friendly number.
- The 80% horizon. Same curve, stricter line, and always much further left. This is the one that decides whether anyone lets it run unattended.
how long the task takes a person →↑ success rate
Measure the length of the job, not the score
METR's Measuring AI Ability to Complete Long Software Tasks (Kwa et al., March 2025) proposes a metric with a unit you can feel: time. The construction matters more than the headline, because you can rebuild it on your own tasks in an afternoon.
- Time the humans first. Every task carries a recorded human baseline — how long someone with the right skills actually took. METR's baseliners average around five years of relevant experience.
- Run the agent many times per task, scoring each attempt with the task's own mechanical grader.
- Fit one curve. A logistic regression predicting the agent's probability of success from the logarithm of the task's human duration.
- Read off the crossings. The duration where that curve passes 50% is the 50%-time horizon. Where it passes 80% is the 80% horizon.
The suite behind the original paper is 170 tasks — 97 from HCAST, 7 from RE-Bench, and 66 very short software actions — measured against more than 800 human baselines totalling 2,529 recorded hours of human work. Across frontier models from 2019 to early 2025, the fitted 50% horizon doubled roughly every 212 days, with a 95% bootstrapped confidence interval of 171 to 249 days — the figures as published in March 2025. The paper's current revision restates that fit as 207 days, 95% interval 166 to 240, so say which version you are quoting.
METR now maintains a live version on a larger suite, Time Horizon 1.1, and publishes the fitted numbers as data. There the all-time doubling sits at 188 days, and the fit restricted to models from 2023 onward is 129 days — a 95% interval of 104 to 158 days, or roughly four months per doubling over the recent stretch. Hold the trend loosely. Hold the per-model numbers tightly, because those are what you compare your own agent against.
The reason the unit matters is that anyone can check it against their own week. A benchmark score means nothing to the person deciding whether to leave an agent running overnight. The length of a job they recognise means a great deal.
| Model | Released | 50% horizon | 80% horizon | Ratio |
|---|---|---|---|---|
| GPT-4 | Mar 2023 | 4 min | 0.9 min | 4.5× |
| Claude 3.7 Sonnet | Feb 2025 | 60 min | 12 min | 5.0× |
| o3 | Apr 2025 | 120 min | 30 min | 4.0× |
| GPT-5 | Aug 2025 | 203 min | 38 min | 5.3× |
| Claude Opus 4.5 | Nov 2025 | 293 min | 49 min | 5.9× |
| Gemini 3.1 Pro | Feb 2026 | 384 min | 90 min | 4.3× |
Source: METR Time Horizon 1.1, the published dataset behind metr.org/time-horizons. Point estimates in minutes of human task time, rounded for display; the ratios are computed from the unrounded values, so recomputing one from the two rounded cells beside it can land a tenth off. The intervals are wide — Claude Opus 4.5's 50% horizon carries a 95% interval of 162 to 624 minutes — and METR notes that measurements above 16 hours are unreliable on the current suite, so the very top of the chart should be read as "long", not as a number.
The 80% column is the whole story
Read that table across, not down. Every model's horizon collapses when you ask it to be right four times in five instead of one time in two. Across the fourteen models in that dataset released in 2025 or later, the median ratio between the two columns is 5.1×, and the narrowest gap in the set — o3's — is 4.0×. The paper lands in the same range: it reports 80% horizons running 4–6× shorter than 50% ones. Its worked example, from the original 170-task fit rather than the larger suite behind the table above, is Claude 3.7 Sonnet — a 59-minute 50% horizon against an 80% horizon of about 15 minutes.
Say it in plain terms. Gemini 3.1 Pro's 50% horizon is a little over six hours of human work; its 80% horizon is an hour and a half. The model that can half-finish a working day can dependably finish a long meeting. The distance between those two sentences is your product.
A horizon is also not a promise about every task of that length. METR's own FAQ — written against an earlier fit, which is why it gives GPT-5 a horizon of 2 hours 17 minutes rather than the 203 minutes in the table above — takes the tasks a human expert needs 90 minutes to three hours for, and reports that a GPT-5 agent solves about a third of them every single time, fails about a third every single time, and is erratic on the rest. The headline number is a blend of three different populations — solved, unsolved, and flaky — and only the third one is cheap for your loop to move.
Then measure the work, not the task
Time horizons are measured on tasks built to be scorable. The other honest instrument drops that comfort and asks professionals to judge finished deliverables. OpenAI's GDPval (released September 2025, posted to arXiv in October) is built from the real output of working professionals — 1,320 tasks spanning 44 occupations across the nine largest sectors of US GDP, written by people with an average of 14 years in the job, with 220 of them open-sourced as a gold subset. The deliverables are real artifacts: documents, spreadsheets, decks, plans. A gold-subset task takes an expert 404 minutes on average — just under seven hours.
Two instruments, then, failing in opposite directions. One asks whether the agent can keep going long enough; the other asks whether what came out was any good. You want both, and a benchmark score on its own is neither.
The grading is the part to copy. An expert in that occupation sees the request, the reference files, and two unlabelled deliverables, and ranks them without being told which came from a model. The paper collects three model samples per prompt and puts three graders on each, which is nine blinded comparisons per prompt per model across all 220 tasks. A single comparison takes a grader over an hour.
The paper's headline: on the gold subset, 47.6% of Claude Opus 4.1's deliverables were graded better than (a win) or as good as (a tie) the human expert's — the best of the models it tested. Not "as good as a professional". It means that in blinded head-to-head against the specific professional who did that job, a grader called it for the model or called it even in a little under half the pairings.
| Model | Wins or ties, GDPval gold subset (220 tasks) |
|---|---|
| GPT-4o | 12.5% |
| o4-mini | 29.1% |
| o3 | 35.2% |
| GPT-5 | 39.0% |
Source: the GDPval paper's speed-and-cost table (Table 2), whose win-rate column covers only the OpenAI models it priced — hence no Claude, Gemini or Grok row. GDPval counts a deliverable graded as good as the expert's alongside one graded better, which is the same measure as the 47.6% above, so the column and that figure are on one scale; Claude Opus 4.1's number is simply the one the paper states in prose rather than plotting. Read the ladder for its slope, not its absolute level; the leaderboard has moved since.
Why long runs fall apart: multiplication
Both instruments are measuring the same underlying thing, and it is not intelligence. It is compounding.
One number raised to a power. That is the entire reason a capable model fails a long job.
Put real numbers in it. At 99% per step — a rate most people would call excellent — a 50-step run finishes 60.5% of the time and a 100-step run finishes 36.6%. At 95% per step, the 50-step run finishes 7.7% of the time. Run it backwards and it gets worse: to land 90% of 50-step runs you need 99.79% per step, which is one slip per 475 steps. That is not a number you reach by tuning a prompt.
Independence is the assumption carrying that formula, and it is wrong in both directions. Real agents retry, so some failures cost tokens instead of the run. Real failures are also correlated — one bad plan poisons every step after it — so one mistake can be worth many. Treat pn as the shape of the problem, never as a forecast of your system. The block below is that arithmetic and nothing else: no model, no sampling, no training run.
Reliability is the binding constraint, not capability
Look at what the arithmetic implies about where to point your loop. Going from 90% to 95% per step lifts a 50-step run from 0.5% to 7.7%. Going from 95% to 99% lifts it from 7.7% to 60.5%. Going from 99% to 99.9% lifts it to 95.1%. Every nine you add is worth more than the one before it, and none of those moves teaches the model anything it could not already do. The model can do the step. It does it 95 times in 100. Your loop's entire job is the other 5.
That is why a self-improvement loop aimed at the lowest benchmark score usually disappoints, and one aimed at the flakiest step usually does not. Find the step that fails most often in your traces, and spend memory, context and weights on that one step until it stops. Then find the next one. Long-horizon agents is the companion piece on what else breaks when a run lasts hours.
Build your own two instruments
You cannot run METR's suite or GDPval. You can run both methods, cheaply, on your own work:
- Your horizon. Take 50 or more real tasks off your actual queue. Record how long a competent person takes on each. Run the agent five to ten times per task. Fit success against the logarithm of that duration and report two numbers — the 50% and the 80% crossing. Track the 80% one. It is the number your users experience.
- Your win rate. Freeze 50 to 200 real tasks. For each, keep one deliverable made by your best human. Have someone who does the job rank the two, blind, against a rubric written before the model ran. Your improvement claim becomes a win-or-tie rate on the frozen set that moved — a sentence your loop cannot manufacture by learning to please your verifier.
One caveat both sources state about themselves, which applies double to you. A METR task is written to stand on its own, so that a baseliner and an agent start from the same place and neither is punished for not knowing the house rules. Your queue is the opposite: half of what makes a ticket easy is knowing which service owns it and who broke it last time. METR's own advice is to read a two-hour task as what someone with "low or no prior context (like a new hire or freelance contractor) could complete in 2 hours". Your production tasks come loaded with context, so your horizon will not match anyone's published figure — and it does not need to. The number exists to compare your agent to your agent, last month.
A score tells you whether the model got better. A horizon at 80% tells you whether the run got better, and that is the only improvement a user can feel. Measure the length of the job you can finish reliably, and measure it against a human deliverable that someone graded blind.
Reliability compounds: the arithmetic behind a horizon
This is arithmetic, not a simulation. No model runs, nothing is sampled, and nothing here is a training run — it is one number raised to a power, printed three ways. The first table shows what a per-step success rate becomes over 10, 50 and 100 steps. The second inverts it: what each step must hit for the whole run to land 90% of the time. The third shows why every extra nine costs more than the one before it. The assumption doing all the work is that steps fail independently, which real agents do not — they retry, and their mistakes correlate, so the true number can land on either side of this. Read it as the shape of the problem, not a forecast of your system. Change p and the step counts to your own and run it again.
end_to_endis the whole model in one line — a per-step rate raised to the number of steps — and every table below is that line read from a different direction.step_rate_neededis the same equation solved backwards, and it is the number you can actually act on: a target for one step, not a wish for the run.- In the second block,
1 / (1 - need)converts that rate into "one slip per N steps" — the form of the number that survives being said out loud in a design review. - The third loop walks the pairs
aandbacross a fixed 50-step run, so watch the size of each jump grow as the nines pile up, not shrink. - Only two things are worth editing: the tuple
(0.999, 0.99, 0.98, 0.95, 0.90)and the step counts — put your own traced per-step rate in and the rest re-derives.
Ship it
≈12 min2 runnable blocksStart here if you have ten minutes
Build logWeek 9 · WednesdayWhat was not the loop
We turned the loop off on the Wednesday, not because anything burned, but because we could not name the last three things it had changed. Then we wrote down the six conditions that would turn it off again, and stuck them where on-call could see them. What was left was the part worth keeping: one written sentence describing a finished ticket, a horizon measurement, a verifier, a hundred runs hand-labelled to find how often that verifier was wrong, and a trace we could replay six weeks later.
Which asks What do we build first, and in what order
Build logWeek 9 · FridayWhat we ended up running
What runs now is smaller than the diagram we drew in week one. Mender takes a ticket off one queue and makes up to eight attempts, down from the twenty we started with, and every attempt runs the service's suite in a clean container. A patch is a candidate only if the suite goes green and the ticket's own named test goes green with it. For the third of the queue with no test, the verdict comes from the environment instead: the service restarts, the log line appears, the alert clears. Nothing is more independent, and nothing covers less — a quiet channel is not proof. Those tickets are allowed to produce a PR and never a training example. The search we spent ten days on is off; nothing it found stayed found.
The traces that survive are filtered before anything is refit: correct, deduplicated on what the patch does rather than how it reads, dropped if every attempt passed or none did, capped at three per ticket, and only then cut to budget. Most weeks nothing is refit at all. The lesson goes in the prompt, or the tool, or the retrieval store, and four times in five it stays there. Three things get read every round that are not on the dashboard: how many distinct passing patches exist for a ticket, how long a ticket takes a person before Mender stops finishing eight in ten of them, and a blinded read of two finished tickets — the one allowed to disagree with the score. Six written conditions turn the loop off, and the frozen set belongs to someone whose job is not shipping the loop.
What we would still not claim is the important part. We cannot say this makes Mender better, only that it does more of what our suite likes, and so far we have not caught the two coming apart.
Here is the order to do it in. It is deliberately front-loaded with measurement, because every failure mode above — reward hacking, drift, a filter that lets its own mistakes through — is invisible without it, and every one of them is cheap to prevent and expensive to unwind. Each step names the action and points back at the section that argues for it.
-
Name one task family and one deliverable.
Not "the agent". One repeating job — triage this ticket type, draft this document, fix this class of bug — and one written sentence describing what a finished one looks like, specific enough that a stranger could apply it. If you cannot write that sentence, nothing downstream has a target.
-
Measure a horizon before you measure a score.
Pull 50 or more real tasks off the queue, record how long a competent person takes on each, run the agent five to ten times per task, and fit success against the logarithm of that duration. Report the 50% and the 80% crossing. This pair is your baseline and your only defence against shipping a change that helped the benchmark and hurt the product.
Why: 12 · Knowing it worked -
Build the verifier before you build the loop.
No trace may enter any update until something mechanical can say correct or not: tests that run, a schema that validates, a diff against a known answer, a rubric grader. If you cannot verify the task, you cannot safely improve on it — stop here and change the task.
Why: 05 · Verification, 04 · Selection is the bottleneck -
Measure the verifier itself.
Hand-label 100 runs and score your verifier against your labels. Its false-positive rate — how often it blesses a wrong run — is the hard ceiling on everything downstream, because those are exactly the traces that get baked in. A verifier you have not measured is a number you have not earned.
Why: 05 · Verification -
Log complete, replayable traces from day one.
Every prompt, tool call, tool result, final output, verifier verdict, and the version of the model, the prompt and the tools that produced it. A trace you cannot replay is an anecdote. This is also the only artifact that makes a post-mortem possible six weeks later.
Why: 07 · What you keep -
Find where the runs actually die.
Bucket failures by step index and by cause, and sort. The distribution is almost never flat — one or two steps carry most of the loss. Fix the tallest bar. Improvement spent anywhere else is invisible in the compounding arithmetic.
Why: 12 · Knowing it worked -
Reach for the cheapest reversible fix first.
In order: prompt and tool descriptions, then scaffold and retry logic, then memory, then weights. Each rung costs more and reverses harder than the one below it. Most teams skip to the top rung and pay for a training run to fix a tool description.
Why: 08 · Where the improvement lands -
Land recurring specifics in memory.
This customer's preference, this repo's build command, this API's quirk, this failure and its remedy. Anything narrow, anything you might have to revoke, anything you are not yet sure about. Write it, retrieve it, and keep the delete button one click away.
Why: 08 · Where the improvement lands -
Land broad, verified behaviour in weights — and only then.
The signal that weights are the right home is memory repeating itself: the same kind of note written a hundred times for a hundred different tasks. Train on filtered-correct traces only, at a bar strictly higher than memory's, because this is the rung you cannot climb back down.
Why: 08 · Where the improvement lands, 09 · Closing the loop -
Freeze an eval set the loop can never touch.
Held out of training, held out of memory, held out of prompt examples, and owned by someone whose job is not shipping the loop. The day it leaks, your only independent measurement quietly becomes a training metric and you will not be told.
Why: 12 · Knowing it worked -
Canary with enough runs to mean something.
Route a slice of live traffic to the candidate, keep both versions loadable, and size the sample to the regression you care about rather than to your patience. Ten runs is not a canary: a model that has silently dropped from 90% to 80% still clears a 90% bar 37.6% of the time. Catching a five-point drop nineteen times out of twenty takes about 118 runs. The sizing block below does that arithmetic for your own numbers.
Why: 12 · Knowing it worked -
Monitor four numbers, not one.
Verifier pass rate on live traffic. The 80% horizon on the frozen set. Memory store size against memory hit rate. And a weekly human-graded win-or-tie sample. The fourth exists precisely because the first three can all improve while the product gets worse.
Why: 12 · Knowing it worked
- Verifier pass rate climbs while the human-graded sample does not. That gap is the definition of reward hacking. The agent found a hole in your signal, and every update from now on widens it.
- The frozen set drops at all. Not "drops significantly" — at all. A frozen set that moves down after an update is the one measurement you cannot explain away.
- Output diversity collapses. Same phrasing, same tool order, same structure on tasks that used to produce different work. A loop trained on its own output narrows toward its own habits, and the narrowing shows up in the writing before it shows up in the score.
- The memory store grows faster than its hit rate. You are accumulating, not learning. Retrieval quality, not store size, is what makes memory worth having.
- You cannot name the last three things the loop changed, or roll any of them back. An improvement you cannot undo is a deployment you cannot defend.
- A weight update whose training set you cannot reproduce. If you cannot re-derive exactly which traces went in, you cannot audit what the model learned, and the permanence in section 08 is now permanence without evidence.
The loop is not the hard part — you can build the loop in a week. The hard part is the pair of numbers that tells you it is working and the pair of rails that stops it when it is not. Build the measurement first, and the loop becomes a small thing you bolt onto it.
How many runs before a canary means anything
Also pure arithmetic — the exact binomial, computed term by term, with nothing sampled and no model involved. The question is how many runs a canary needs before passing it is evidence. The right-hand column is the one that stings: with ten runs, a candidate that has silently regressed from 90% to 80% still clears a 90% bar more than a third of the time. The middle column is the honest cost of the bar itself — set at the live model's own rate, it drifts toward a coin flip as the sample grows, which is why the bar belongs below the live rate and the sample size does the real work. The last block searches for the fewest runs that catch the drop you actually care about. Put your own pass rate and your own tolerance in and run it again.
p_at_leastis an exact binomial tail — it sumscomb(n, i)terms fromkup ton— so nothing in this block is sampled, estimated or simulated.- Read the table's last two columns as one thing: the same bar
kapplied toLIVEand toBAD, because the gap between those two percentages is all the power your canary has. LIVEandBADare the only inputs worth touching — setBADto the regression you would be angry to miss, not to the worst case you can imagine.smallest_nwalksnupward until a regressed model's chance of passing falls tofalse_ok, so the firstnit returns is your minimum canary size, not a suggestion.- Notice
k = ceil(live * n)rises withn, holding the bar at a constant 90%: it is the sample size, never the threshold, that buys the confidence.
Round 6 is done. Do you ship it?
Nothing here runs a model. It is the four stop conditions above, written as arithmetic over two rounds of measurements, and four rounds fed through them. Run it as written, then read round b twice: every number a weekly update would quote went up, and it is the round you must not ship. Then change a single field at a time and watch which condition fires — that is the whole instrument. The remaining two conditions, plus the canary-size guard, are yours to write in the challenge.
- Two rounds in, one verdict out. Every condition compares
curragainstprev, because none of these numbers means anything as a single reading. - Divergence is the first and the most important: your verifier moved and the human sample did not follow. The threshold (
d_human <= d_verifier / 2) is a choice, not a law — pick yours deliberately and write down why. - The frozen set has no tolerance at all. A set nothing has trained on that moves down is the one measurement you cannot argue with, so it is compared with a bare
<. - Narrowing counts distinct passing solutions per task, not output length or token entropy. It is the cheap proxy for “can this model still reach what it used to”, and it falls while pass@1 climbs.
- Reversibility is not a measurement, it is a property of how you shipped. It is in the same list because it gets more expensive to fix every round you leave it.
Where to go next
≈5 minSkip unless you are stuck
Build logWeek 9 · ThursdayThe week we read instead
With the loop off we read for three days. The list of things we got wrong is shorter than the list of things we had to go and read, and every item on it was a stuck point rather than a topic: the verifier, the filter, the eval set, where a fix should land. One of those pages is why our dedupe now compares what a patch does rather than how it reads. We opened exactly one at a time.
Which asks Which page owns what we are stuck on
Every page below owns a piece this one deliberately did not re-teach. They are grouped by the question they answer, so you can go to the one you are actually stuck on.
How an agent gets better
- Test-time computeBuy quality at inference — more thinking, more samples, more search. The improvement that needs no training run, and the one to exhaust first.
- Context engineeringMost "the model can't do this" is "the model wasn't told this". The cheapest, most reversible rung on the ladder in section 13.
- Agent memoryThe reversible axis in full: what to store, how to retrieve it, when to forget, and why the hit rate matters more than the size.
- RLVR & verifiable rewardsWhen the verifier is a program, you can train straight against it. The most disciplined version of "reinforce only what you can confirm".
- Synthetic dataHow to manufacture the traces a loop needs when production traffic is too thin — and how to stop the loop from eating its own output.
- Fine-tuningWhat actually changes in the weights, which method to pick, what it costs, and what you can no longer take back afterwards.
Where each idea was first shown to work
- STaRThe bootstrap itself: generate reasoning, keep what reaches the right answer, fine-tune, repeat. The seed of every self-training loop.
- Self-InstructThe same trick pointed at instructions — a model writing the training set it will then learn from.
- ReflexionImprovement without touching a weight: write the lesson from the failure into memory, then try again.
- VoyagerThe skill library. Successful code becomes a callable skill, and the agent keeps growing the set it can call.
- Let's Verify Step by StepGrade the steps, not just the answer. The result behind process-level verifiers, which catch a bad step inside a run that happened to end on the right answer.
- Self-ConsistencySample many chains, take the majority answer. It buys reliability with inference alone, and doubles as a filter when you have no program to check against.
- Tree of ThoughtsSearch over reasoning states instead of committing to one straight line.
- ReActInterleaving reasoning and acting — think, call a tool, read the result, think again. Where the tool-using agent loop comes from.
- AlphaCodeWhat sample-and-filter looks like at industrial scale, when the filter is real test execution rather than a model's opinion.
- PPOThe policy-gradient algorithm most of RLHF was built on, and the baseline everything newer is measured against.
- GRPOThe leaner variant that dropped the value model. What most verifiable-reward training runs actually use now.
- DPOSkip the reward model and optimise preferences directly — fewer moving parts, fewer places for the signal to rot.
Knowing it worked
- Agent evalsBuilding the set: what goes in, what stays out, how many runs per task, and the traps that make an eval agree with you.
- Evals in CIMake the eval a gate rather than a ritual. Run it on every change, block on regression, and the frozen set defends itself.
What the loop runs on
- Loop engineeringThe runtime the agent lives in: retries, budgets, termination, and the plumbing that decides whether a step failure ends the run.
- Long-horizon agentsEverything else that breaks when a run lasts hours — the companion to the compounding arithmetic in section 12.
- RL environments engineeringBuilding the environment and reward your loop optimises against. This is where reward hacking is born, and where it is prevented.
- Agent patternsThe reusable shapes — planner and executor, critic, router — that you will be pointing all of this improvement at.
Questions people actually ask
What is a self-improving agent?
An agent that gets better from its own work rather than from a new model release. It runs real tasks, a verifier decides which attempts were actually correct, and the correct ones are folded back in — either written to a memory store for exact recall, or used to update weights so the improvement generalises to tasks it has not seen. Everything difficult about it lives in the word correct.
Do I need a GPU?
For most of this, no. The memory axis is a database and a retrieval step — no training hardware anywhere. Prompt, tool and scaffold fixes are free. Test-time compute buys real quality with inference alone. You need training hardware only for the weights axis, and by the ordering in section 13 that is the last rung, not the first. A loop that never trains anything can still be a genuine self-improving agent.
Where do I start on Monday?
Steps 1 to 5 of the checklist, in order, and nothing else. Name one task family, measure a 50% and an 80% horizon on real tasks, build a verifier, measure that verifier against 100 hand-labelled runs, and log full traces. That is a week or two of work and it produces no self-improvement at all — which is the point. Teams that skip it spend the following quarter unable to tell whether their loop is helping.
How do I know the loop is degrading?
The tell is divergence: your verifier's pass rate climbs while a human-graded sample of the same work does not. That gap is reward hacking, measured. The other four signals are a frozen eval set that drops at all, output diversity collapsing into one house style, a memory store growing faster than its hit rate, and any update you cannot roll back or reproduce. All five appear among the six stop conditions — they are stop conditions rather than alerts because each one gets harder to unwind every day the loop keeps running.
Memory or weights — which first?
Memory, nearly always. It lands in seconds, it generalises to nothing but it also risks nothing, and you can delete a bad one the moment you notice. The signal that you have outgrown it is repetition: when memory is storing the hundredth variation of the same lesson, that lesson is a habit, and a habit belongs in the weights. Match permanence to confidence.
Can a self-improvement loop run away?
Not in the sense the phrase usually implies. The loops described here are bounded by their verifier — an agent cannot learn to do something its filter cannot recognise as correct, so the verifier's quality is the ceiling, and the ceiling does not move on its own. The realistic failure is the opposite of runaway: a loop that quietly optimises the gap between your verifier and your users, and gets measurably better while getting worse.
Explore the topic
See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.
More Handbooks
- The Prompting HandbookA friendly, hands-on field guide for everyday humans — learn the CRISP framework, spot bad prompts, practice with real recipes, play a drag-and-drop game, and test yourself with a quiz. No code required.Read →
- The Agentic AI Interview HandbookTwenty topics every senior AI engineer should be able to reason about live — from eval pipelines to reliability patterns for generative systems.Read →
- The Senior AI Engineer Interview Handbook60 questions across architecture, production incidents, agentic systems, RAG, evals, cost, safety, and leadership — what staff-level AI interviewers actually probe for.Read →
- 51 LLM Evals Interview QuestionsGolden sets, LLM-as-judge, regression testing, offline vs online evals, RAG evals, agent evals, red-teaming, and observability — demystified for interviews and production.Read →
- The Agent Evaluations HandbookA self-contained handbook on evaluating AI agents — theory, interactive widgets, and practical guidance. Trajectory evals, tool-use scoring, LLM-as-judge, observability, and reliability for PMs, engineers, and founders.Read →
- The Agent Patterns HandbookThe design patterns behind every LLM agent — the ReAct Thought–Action–Observation loop, tool/function calling, plan-and-execute, reflection, memory, multi-agent orchestration, and the failure modes (loops, hallucinated tools, recovery, human-in-the-loop) that break agents in production.Read →