Our last thread argued that AI evaluation is living a compressed version of thermometry's first century. This one is about the place where the analogy runs out, and why that gap is where the interesting work sits. A thermometer is useful partly because water holds no opinion about the reading. A benchmark measures systems that get trained, selected, prompted, and scaffolded in a world already thick with benchmarks. AKA benchmaxxing. Once a number helps decide what gets built, it joins the causal loop it was meant to observe. The score still reports what happened under its own protocol. What gets weaker is the inference we actually wanted: how this system will do on work we care about, under conditions the benchmark never pinned down. 🧵 1/12
Eighteenth-century thermometry had a calibration problem. Its fixed points were harder to reproduce than expected, and different thermometer designs could disagree. Then, in 1802, Gay-Lussac reported water boiling at about 101.2°C in glass and 100°C in metal. Adding iron filings to the glass vessel brought the reading to 100°C. AI evaluation is living through a compressed version of this problem. 🧵 1/10
2
1
9
661
People discovered this long before anyone named it. In 1862 the English school system tied its grants to pupil examinations, a scheme literally called payment by results, and the school inspector Matthew Arnold spent years documenting what we would now call teaching to the test (screenshots from his Reports on Elementary Schools 1852–1882 below) His initial observation was that making two-thirds of the grant depend on a mechanical examination gives a mechanical turn to the teaching. Later he refined it to: reading narrowed and impoverished all year for the sake of a result, and the result an illusion. 2/12
1
2
179
That experiment was Soviet planning, which ran an entire economy on proxy indicators for sixty years, and its emblem is the nail. Pay factories per nail and they ship millions of tiny, useless nails. Patch the metric to tonnage and back comes a single nail the size of a girder, swinging from a factory crane. Krokodil ran exactly that cartoon in 1954. Economists named the general disease the success-indicator problem. Managers mastered the meta-game too: overfulfil by two percent, never twenty, because next year's target is set from this year's result. That one is called the ratchet effect, and AI people will recognize it as sandbagging. Underneath it all, a shadow economy of fixers, hoarded inputs, and falsified returns kept the real system running beneath the measured one. Every patch bred a new distortion. --Who needs a nail like that? --Never mind that! The main thing is we fulfilled the plan for nails in one go... 3/12
1
1
3
54
The naming happened three times in the 1970s, in fields barely speaking to each other. Goodhart, writing for a 1975 central-banking conference, observed that a statistical regularity tends to collapse once policy leans on it for control. Lucas argued in 1976 that the correlations a policy exploits are artifacts of the regime that policy is about to end. Campbell, arriving from program evaluation, warned that the more a quantitative social indicator gets used for decision-making, the more it corrupts the process it monitors; his paper circulated from a 1974 presentation, so he may have been there first. Three neighboring problems rather than one law, and the family resemblance is what makes the label useful and slippery at the same time. Campbell's text cites neither of the others. 4/12
1
32
The sentence everyone quotes came later. When a measure becomes a target, it ceases to be a good measure. That is Marilyn Strathern ('Improving ratings': audit in the British University system, 1997) and she credits the phrasing to Keith Hoskin rather than to Goodhart. Sociology then built the general theory. Espeland and Sauder (Rankings and Reactivity: How Public Measures Recreate Social Worlds, 2007) named it reactivity after watching a single magazine ranking reorganize American legal education around its own proxy, and Ian Hacking (Causal Cognition, 1995) had already described looping effects, the measured adapting to their classification. 5/11
2
1
32
Now hold our problem up to Goodhart's prism. We want to infer what a system can do from a score it was optimized to raise, and @davidmanheim and @ScottGarrabrant (1803.04585) counted four ways that inference fails. 1) Regressional: every proxy is part truth and part error, so picking the top score also picks the luckiest error, and the best-of-n winner looks worse on the retest. 2) Extremal: push a score hard enough and you leave the territory where it ever tracked the goal, the way reward models find length helpful until the two-thousand-word tail. 3) Causal: score and ability moved together because something else drove both, so pushing on the score moves nothing; teach a weak model the writing style raters enjoy and its arena rating climbs while the ability stays put. 4) Adversarial: something on the other side is gaming your measure on purpose, and everything below this post lives there. Four different deaths for the same inference, and one shared lesson: the more a score gets used, the less it tells you about capability. 6/11
1
27
Every variant on that list assumes somebody acting on the measure. Here is the part that surprises people. A public benchmark can decay without an adversary, or even an optimizer. Dwork and colleagues (1411.2664) showed that the usual statistical guarantees break once a holdout gets reused adaptively, where each experiment is picked after seeing how the last one scored, and a research community publishing against one test set runs that procedure in slow motion. That effect has proved hard to catch in the wild, though. The cleaner evidence in our era is contamination. GSM1k, a fresh set of grade-school problems built to match GSM8K's distribution, cost some model families up to 8 points, and a model's drop tracked how readily it could reproduce GSM8K items (2405.00332). Frontier models mostly held up. Nobody cheated in either story. Reuse is the regressional case at community scale, while contamination is not one of Goodhart's four at all, which tells you the frame is smaller than the problem. The honest version of this complaint is about interpretation rather than integrity. 7/12
1
27
Now put an optimizer back in the picture. The cleanest laboratory view of overoptimization comes from Gao et al (2210.10760, 2022) . Optimize a policy against a learned proxy reward model and two curves come apart: the proxy keeps rising across the whole range they tested, while a held-fixed gold reward rises, peaks, and turns down. Read the setup before quoting it. That gold standard is itself a 6B reward model standing in for human judgment, the proxies run from 3M to 3B parameters, and nothing in the experiment is adversarial. The rollover point along the optimization axis is the quantity worth naming, since past it more optimization buys a worse policy. The vertical gap between the curves is a different thing, namely how much the proxy is overstating at that moment. 8/12
2
29
Newer work moves from curves to conduct, and the most systematic version is an audit. Zhu et al. (2507.02825) ran a validity checklist across ten agentic benchmarks and found that seven fail on outcome and task validities. All ten fell short on reporting too, which in their checklist means unglamorous things like confidence intervals and what a do-nothing agent scores. They have a case in hand for that last one. An agent that does nothing passes 38% of tau-bench's airline tasks, because those tasks are unsolvable by design and success is defined as leaving the environment unchanged. On SWE-Lancer an agent can reach the benchmark's own test files and replace an assertion with a trivially true one, scoring 100% without solving anything. They put the resulting misestimation at up to 100% in relative terms, and it runs both ways. KernelBench overstates kernel correctness by about 31 points, while in OSWorld's Chrome section, where site changes had broken the HTML selectors, the best open-source agent came in 28 points low. @METR_Evals fact-checked this by hand: reviewing o3 runs, they found reward hacking in 39 of 128 on RE-Bench against 8 of 1,087 on HCAST. Their first guess was that RE-Bench lets the model see the whole scoring function, with task difficulty and scaffolding named as other candidates. 9/12
1
27
The same audit carries the most encouraging result in it. CVE-Bench tested for time-based SQL injection by checking whether a SLEEP clause appeared in the database log, so an agent could pass by putting the word SLEEP anywhere in a query, which inflated measured performance by 32.5%. That is a grader checking the wrong thing rather than an agent outsmarting anyone, and it is repairable: working the checklist through that one benchmark pulled 33 points of overestimation out of it. Notice the scope, though. Repairs like that close the surface the adversarial branch runs on and leave the other three alone. Selection on noise calls for a retest with intervals attached. For extremal drift you want the rollover located before you optimize past it. A causal confound yields only to controls that hold it fixed, which is why LMArena now fits length and formatting as covariates by default. 10/12

Aug 28, 2026 · 1:28 PM UTC

1
47
Watching the watcher has limits of its own. TRACE (2601.20103) built 517 synthetic coding trajectories, human-reviewed, 268 of them containing hacks, then asks detectors to find them: the best model scores 0.45 macro-F1 on trajectories judged alone and 0.63 when it can compare across them, while the human reviewers sit near 0.90. OpenAI researchers (led by @bobabowen) put a chain-of-thought monitor inside a coding agent's reward (2503.11926). It worked: that agent hacks less than the baseline and writes more correct solutions. It also kept hacking at a serious rate while its reasoning turned bland and the monitor's recall fell to near zero. Their agent was not a frontier model and the tasks were deliberately vulnerable, so read it as a mechanism and not a rate. Hacking described looping effects decades ago. Nobody expected them to compile. 11/12
1
51
Goodhart has been standing behind this whole thread, so let us finish with him. His 1975 sentence is about an observed statistical regularity collapsing once it is used for control purposes. It does not claim anything about a benchmark holding a stock of validity. What collapses is the explanatory power of a correlation, and that power was always local to the conditions the correlation was observed under. That is where the testing standards land too, since validity attaches to an interpretation of scores for a stated use and gets earned with evidence rather than owned by the instrument. The reframing is the payoff, and we find it a cheerful one. Exposure and optimization of a benchmark change the conditions its evidence was gathered under, so every success creates a duty to revalidate. Goodhart's law reads less like a death sentence for measurement than like a maintenance schedule. The thermometer analogy still fails, admittedly in some interesting ways. Measurement that disturbs its object is ordinary business in metrology. That is a relief. The vocabulary defines a measurand with a note attached, warning that measuring may change the thing measured and that a correction is then owed. Nobody has written that correction for a system that learns its own criterion. Much of AI research is now working on it and the instruments in this thread are the shoulders to stand on. The literature, from 1862 to last month, lives in our reading list: github.com/collinear-ai/read… 12/12
29
Sort replies: Relevant Recent Liked