Home What the leaderboards measure
00%

Essay · 6 September 2026 · about 70 minutes

What the Leaderboards Are Actually Measuring

A critique of the public evaluation metrics for frontier AI models. Arena rankings, benchmark scores and lab-reported numbers get read as measurements of capability. This essay asks what each of them measures, and under which conditions a score says something about capability. The running case is the launch week of 1 to 6 September 2026, when Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra were released two days apart. The same checks are applied to both labs, and to Google, xAI, Meta, DeepSeek and Moonshot in the earlier episodes.

Maksim Silchenko  ·  National University of Singapore

Podcast · long · 53 min0:00 / How AI Labs Game Their Benchmarks
Podcast · short · 22 min0:00 / Why AI Benchmark Scores Are Misleading

Summary

Eight points

  1. A public score is evidence about capability only under the conditions that produced it. It depends on the test items, the around the model, the run, and how the reported number was selected. Most published scores travel without those details, and a conditional result then gets read as a general ranking.
  2. The same model scored 62.7% and 99.9% on the same benchmark in the same week. The ARC Prize Foundation measured on under two harness conditions and published both. OpenAI's launch table printed the higher figure beside two competitors measured under the lower one.
  3. Independent composites disagreed on which model leads. ranked Astra first; placed it five points behind . Both publish how their indices are built; the constructions differ enough to make a disagreement expected, though neither publishes a decomposition of this one.
  4. The same reading applies to every lab. Anthropic's launch of the same week reported a Terminal-Bench-Science score of 52.6% from an internal setup while the public board's best entry was 30.0%, ran OSWorld 2.0 on a task release that its own footnote calls not comparable to earlier results, and evaluated with production safeguards that scored some tasks as zero and routed others to older models; its two system cards give Claude Opus 5 two different GDPval scores at the same setting. Earlier episodes cover Google, xAI, Meta, DeepSeek and Moonshot.
  5. Arena rankings measure voter preference under one style adjustment. Preference overlaps with capability without being the same thing. The top five models on the sit within 8 points and carry intervals of 4 to 11 points, and providers may still test several private variants before release.
  6. Test sets are noisier and more contaminated than their headlines suggest. An estimated 6.49% of MMLU items contain errors, 68.3% of a sample of the original SWE-bench was flagged as under-specified or unfairly tested, and OpenAI printed 100% on historical vulnerabilities beside 39% on a separate post-cutoff set, with a contamination warning on the first.
  7. Mitigations exist, each with a trade-off, and several are in use: standard harnesses with labelled conditions, independent re-runs, private and post-cutoff test sets, protected confirmation sets, validated graders, trial counts and intervals printed in system cards, and cost reported next to the score. Disclosure lets a reader assess a score; it does not make the test valid.
  8. How to read a benchmark table: who ran it, which harness and effort setting, how many items and trials, whether there is an interval, which subset and version, whether the tested system is the shipped one, and who funds the benchmark. The checklist at the end scores any table on these ten points.

Every number on this page is traced to the source it came from, opened on 6 September 2026, with secondary reports marked as such; the sources are collected at the end. The essay was revised the same day after an independent review; the method note lists what changed. OpenAI's launch is the running case because it happened this week and produced the most documents, not because the pattern is specific to one company.

01  /  The instrument

What a score measures

On 3 September 2026 the same model scored 62.7% and 99.9% on the same benchmark, on the same day, reported by the same organisation. Both numbers are correct under their own conditions, and the conditions are the subject of this essay.

Evaluation is a form of measurement. Someone starts with a they care about, such as whether a model can reason about a situation it has never seen. They operationalise it as a set of items and a scoring rule. They run the system under some configuration. They report a number, ideally with its uncertainty. Every failure discussed on this page is a failure at one of those four joints: a construct nobody defined, an operationalisation that measures something else, a run that is not the deployed configuration, or a number reported without the error bar that would have shown it was noise.

The 62.7% and the 99.9% belong to on , the interactive reasoning benchmark run by the ARC Prize Foundation. Under the Foundation's own Standard , a minimal interface that is identical for every provider, the model's best run scored 62.7% at a cost of about $26,000. Under a Provider Adapter harness, which keeps OpenAI's private reasoning state between requests and compacts long conversations, it scored 99.9% for about $19,000.1 The Foundation published both, said the shared interface is what gives "a consistent, apples-to-apples comparison across providers", and announced it will now label every result with the condition it was run under.1

OpenAI had made the same point six weeks earlier, about its previous model. In July it reported that turning on two API settings it already uses in its own products, retained reasoning and compaction, tripled GPT-5.6 Sol's score on the ARC-AGI-3 public set from 13.3% to 38.3% and cut by a factor of six. The post closed with a sentence that could stand as the thesis of this essay: "evals rarely measure models in isolation". They also measure, in OpenAI's words, "a bundle of less visible choices about API settings, harness design, and prompting".6

The measurement chain, and where it breaks diagram, draws on scroll, scrolls sideways on small screens

FROM A QUESTION TO A NUMBER 1 Construct what we want to know: "reasoning on new tasks" 2 Items and rule tasks, answers, scoring; n items, pass criterion 3 Harness prompt, tools, memory, effort, time, turn limits 4 Run which checkpoint, how many trials, which route 5 Number one score, rarely with an interval or a trial count HOW EACH JOINT FAILS Preference read ascapability. Saturation.Contested definitions. Contaminated items.Wrong answer keys.Subsets, not the full set. Settings that move thescore by a factor of three.Judges with biases. Best of many variants.Not the shipped model.Model knows it is a test. No error bar. Best effortof several. One run.Index version drift. Every problem on this page is a failure at one of the five joints. Most public scores are published with joints 3 to 5 undisclosed.

What a number can separate

One numerical fact matters for everything that follows. If a benchmark has n independent pass-or-fail items and the true pass rate is p, the of the measured accuracy is the square root of p(1−p)/n. For n = 100 and p = 0.70 that is 0.046, so a 95% interval spans roughly nine points either side. Two models that differ by three points on a 100-item benchmark of single-attempt tasks are not distinguished by it under a conventional test. For n = 1,000 the interval shrinks to about three points. Most agent benchmarks have a few hundred tasks or fewer, and the tasks inside them are rarely independent, which widens the interval further. The calculator in problem 7 lets you set n and p and overlay the benchmarks named in this essay.

Two cautions apply to the formula. First, the right comparison between two models scored on the same items is a paired one: what matters is how often the two disagree on an item, not whether their separate intervals overlap. Two intervals can overlap while the paired difference is clear, and a difference that fails the test is not evidence that the models are equivalent. Second, the formula is for independent pass-or-fail items. Scores that average fractional credit, Elo ratings fitted to votes, and composite indices each need their own error model, so the item counts on this page are scale markers for the size of the problem, not intervals for those benchmarks.

SE(p̂) = sqrt( p (1 - p) / n ) n = 100, p = 0.70 -> SE = 0.046, 95% interval about +/- 9 points n = 1,000, p = 0.70 -> SE = 0.014, 95% interval about +/- 3 points
The six properties of language models that make measurement harder

Six properties of language models make the four joints harder than they are in classical machine learning. Outputs are open-ended, so a score depends on a grader, and graders have properties of their own. Sampling is stochastic, so a single run is one draw from a distribution. Scores move with prompt phrasing, few-shot order and even the delimiter between examples, so a benchmark number is really a triple: the model, the prompt and the harness. Benchmarks are on the internet and models are trained on the internet, so contamination is hard to rule out from outside. A public benchmark is a target, so its validity decays under optimisation pressure, which is Goodhart's law applied to measurement. And when the grader or the simulated user is itself a model, the instrument moves with every version change. Stanford's AI Measurement Science lab describes the result as "a measurement crisis characterized by benchmark saturation, inconsistent measurement practices, and difficulty in making valid claims about AI capabilities", and a 2025 study of what it called potemkin understanding found that models define concepts correctly 94.2% of the time and then fail to apply them at high rates, which is the construct problem in miniature: a right answer on the item is not the property the item was written to measure.67, 65 The eleven problems below are the concrete forms these properties take in the public numbers of 2026.

Does the score predict the work?

A benchmark can be repeatable, uncontaminated and correctly graded and still be a poor guide to the decision a reader is making. The practical test of an instrument is whether a higher score predicts a better outcome on the work the reader cares about, for the people who will do it, at the cost they will pay. Agreement with another leaderboard is convergent evidence, but it is not that test. A model can score higher and still need more supervision, fail more often on one workload, or make errors that are harder to notice and repair. The Holistic Evaluation of Language Models framework made the point in 2022 by scoring models on many scenarios and on several dimensions at once, among them accuracy, calibration, robustness, fairness and efficiency, rather than on one aggregate.116 The good-practice section returns to how such a validation is run.

02  /  The general problems

Eleven problems with public scores

Seven of these are structural and have been documented since 2023. Four are newer, and the launch week of September 2026 supplied a fresh example of each.

1. Arena votes measure preference

The at arena.ai ranks models from anonymous pairwise votes fitted with a . As of 6 September 2026 the page reports 7,999,020 votes across 400 models, with a data stamp of 2 September.17 A vote records that one self-selected person preferred one of two responses to a prompt they wrote themselves. That is a real quantity, and a useful one, but it is a measurement of preference under those conditions. Preference overlaps with capability, since a voter often prefers the answer that is correct, clearer or easier to act on, but the overlap has to be shown rather than assumed, and the Arena's own corrections show where it fails.

How the Arena's own corrections show the gap

The gap between the two is not hypothetical. In August 2024 the Arena team fitted the same votes with controls for answer length, the method it calls , and for the number of markdown headers, bold elements and lists, and reported: "We controlled for the effect of length and markdown, and indeed, the ranking changed." Two small models fell below most of the frontier and three others rose substantially.21 In May 2026 the team found two further biases in a new source of votes, "a position bias favoring Model A (the model on the left)" and "an advantage given to models that share an organization with the prior turns of context", and added terms to the fit to absorb them.18 Each correction is a statement that the raw votes had been rewarding something other than quality. The same effect appears in automatic proxies for human preference: length-controlling AlpacaEval raised its correlation with Arena rankings from 0.94 to 0.98.31

Six models, one voter population, one style knob simulation

How it works. Each model has a fixed true quality and a fixed length. A simulated voter picks the longer or the better answer with a probability that depends on the slider, and 20,000 battles are fitted with the same Bradley-Terry logistic model the Arena uses. Style control adds length as a covariate to the fit, as the Arena's method does. Because each simulated model has one length, the fit can credit to length only what the six models' pairwise results cannot credit to strength; the Arena's real adjustment uses the lengths of each pair of answers, which vary within a model. The adjustment is a modelling choice, not a proof that bias has been removed: length can carry quality as well as style, so the controlled score answers a conditional question. The models, voters and votes are invented; the mechanism is the published one.21

2. Contaminated test items

If test items, or paraphrases of them, sit in the training corpus, part of a score measures recall; this is . Exposure inflates a result without making every correct answer pure recall, and leaked answers, related tasks and legitimate domain knowledge are different things. Three measurements show how large the effect can be. When a team commissioned GSM1k, a fresh set of grade-school arithmetic problems matched in difficulty to GSM8K, several model families lost up to 8 points, and the size of a model's loss correlated with how readily it could regenerate GSM8K problems verbatim.40 On , models identified the buggy file from the issue text alone, with no access to the repository, up to 76% of the time, against up to 53% on repositories outside the benchmark, which is suggestive of memorisation rather than proof of it.39 Agents that search the web while answering , SimpleQA and found the evaluation datasets themselves, with labels, on Hugging Face for between about 1% and 4% of questions, depending on the dataset and the agent; blocking that source cut accuracy on the affected subset by about 15%.41

The launch-week case, and the limits of detecting contamination from outside

The launch week added a case from inside a lab. OpenAI reported that Astra scored 100% on ExploitBench, a set of historical software vulnerabilities. Its says the result "may be artificially inflated due to potential contamination from exposure to historical vulnerabilities", and gives the mechanism: asked to exploit one , the model failed, then recalled a different, later CVE from memory and used that instead.7 OpenAI then built an internal set of 20 vulnerabilities disclosed after the model's training cutoff. On that set Astra scored 39.0% and its predecessor 5.5%, with a footnote that the 5.5% is partly an artefact of a 300-turn limit and that the older model reached 11.5% when it hit fewer limits.4 Anthropic's system card for the same week carries a similar admission about a problem set drawn from June 2026 arXiv abstracts: "some contamination cannot be ruled out".13 Detection from outside is weaker than it looks; a 2024 study showed a contamination method that "significantly inflates benchmark performance while completely evading current detection methods".73

The same skill, a historical set and a post-cutoff set animated, OpenAI's own numbers

Reading. Left, the public ExploitBench set of historical vulnerabilities. Right, the internal set of 20 high-severity V8 vulnerabilities disclosed between June and August 2026, after the training cutoff. Both pairs are from OpenAI's launch table; the caveats about turn limits and about whether 100% is achievable on the new set are OpenAI's own footnotes.4

3. Wrong answers and saturation

A benchmark is a set of items written by people, and people make mistakes. A re-annotation of , posted in 2024 and revised in January 2025, estimated that 6.49% of its questions contain errors, and found errors in 57% of the questions it reviewed in the virology subset.35 For Humanity's Last Exam, FutureHouse reported that 29 ± 3.7% of the text-only chemistry and biology questions "had answers with directly conflicting evidence in peer reviewed literature"; the benchmark's authors re-reviewed and reported an expert disagreement rate of 15.4% on the public set, about 18% on a biology, chemistry and health subset, and 25% on that subset under a single-reviewer rule; these are different review populations and different definitions of error, and they should not be averaged.36, 37 SWE-bench Verified exists because 68.3% of a random sample of the original SWE-bench was flagged by 93 developers, under a deliberately conservative rule, for under-specified statements or unit tests that could reject valid solutions; with the same open-source scaffold, GPT-4o's score was 16% on the original set and 33.2% on the verified subset.38 -Verified was produced by a team of about ten people fixing more than 300 reported problems over two months.64 In February 2026 OpenAI audited the hardest SWE-bench Verified problems, found that "at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions" while scores kept "improving from 74.9% to 80.9% in the last 6 months", and stopped reporting the benchmark it had built.99

Saturation: how quickly benchmarks reach their ceiling

A 2026 audit of multiple-choice sets made the dependence on item quality explicit: injecting noise into the distractors "raises accuracy to nearly 100%, confirming that most errors disappear when meaningful distractor competition is removed", an intervention that also changes the difficulty of the task.66 When strong models clear most of the items a benchmark stops separating them. That is saturation, and it now arrives within a few years of a benchmark's release. , which tracks benchmarks across generations, states that "Most benchmarks saturate too quickly to study long-run AI trends", that "benchmarks tend to saturate within 1-3 years" and in May 2026 described GPQA as "clearly saturated", dating the event to the winter of 2025, about two years after release.68, 42 A September 2026 paper opens with flagship models "scoring above 90%" on MMLU.44 The ARC Prize Foundation's own history is the clearest example: ARC-AGI-1 stayed unsolved from 2019 until late 2024 "despite a 50,000x scaleup of base LLM pretraining", fell to o3-preview at 75% and then 87% in December 2024, and its successor ARC-AGI-3 went from a frontier score of 0.51% at launch on 25 March 2026 to 62.7% and 99.9% under the two harnesses by 3 September.43, 1 Near its ceiling a benchmark's remaining gaps are decided by a handful of items, and whether those are the last hard items or the disputed ones is an item-level question. A move from 98% to 99% halves the error rate and may be real; it may also be two contested labels. The way to find out is to audit the disputed items and recompute the paired difference under plausible labels; the headline alone cannot show it.

Benchmark lifetimes, release to the best published score today animated

Sources for the endpoints. ARC-AGI-1, ARC-AGI-2 and ARC-AGI-3 from the ARC Prize leaderboard and blog; GPQA saturation from Epoch AI; MMLU from the September 2026 LLMPEDIA paper; FrontierMath Tier 4 from Epoch AI's leaderboard; Humanity's Last Exam from Anthropic's launch table (with tools) and the Scale Labs leaderboard. Release dates are the dates of the benchmarks' papers or launch posts.2, 42, 44, 11, 12, 72

How much of a test set is wrong, by the people who checked animated

Reading. Different studies measure different things: a wrong reference answer, an expert disagreement, or an item that cannot fairly be scored, each on its own sample and by its own rule. None of these rates is a standard error, and whether item errors change a ranking depends on which models miss which items; what they establish is that item error is common enough to matter for the one-point gaps that decide rankings.35, 36, 37, 38

4. Biased model judges

Most open-ended evaluation now uses , because it is a widely used option that scales, although executable checks, structured outcome validation and sampled expert review also scale in their own ways. The biases were catalogued with the MT-Bench paper in 2023 and have not gone away. When the order of two answers was swapped, Claude-v1 gave the same verdict 23.8% of the time, GPT-3.5 46.2% and GPT-4 65.0%. Under a "repetitive list" attack that pads an answer without adding content, Claude-v1 and GPT-3.5 preferred the padded answer 91.3% of the time and GPT-4 8.7%, on 23 answers. GPT-4 favoured its own outputs by 10 points of win rate and Claude-v1 by 25.29 Self-preference has a mechanism: GPT-4 recognised its own text 73.5% of the time out of the box, and when the same authors fine-tuned GPT-3.5 and Llama 2 to raise or lower their self-recognition, the strength of self-preference rose and fell linearly with it.32 A "null model" that returns one fixed response regardless of the question reached an 86.5% length-controlled win rate on AlpacaEval 2.0, 83.0 on Arena-Hard-Auto and 9.55 on MT-Bench; these are attacks on automatic judges, and say nothing about how the same response would fare with human voters.30

Judges in 2026: stronger, still biased, and changing under the evaluation

Two 2026 studies confirm the biases shrink with stronger judges without disappearing: "Stronger judges reduce but do not remove position and verbosity bias", and the strongest judge in one panel still changed 14.7% of verdicts when the answers were swapped.33 The judge is also a moving instrument. Because it "is itself a model behind an API", a silent version bump or a changed scoring prompt means that "every drift alarm is ambiguous between a worse product and a changed judge".34 Judges are inside several of the numbers in this essay: 's reports an from judged comparisons, Arena introduced AutoEval scores in July 2026 to provide ratings "when waiting for human votes to accumulate", and OpenAI's launch table grades HealthBench Professional with GPT-5.4.8, 63, 4

Swap the order, pad the answer, watch the verdict simulation

How it works. The judge is a toy: its preference is the true quality difference (zero here) plus a position term and a length term, with sizes chosen to reproduce the order of magnitude in the 2023 measurements. "Judge both orders" applies the standard fix, running both permutations and averaging, which cancels the position term in this toy judge but not the length term; a real judge's biases need not cancel so neatly, which is why a judge is validated against independent human judgments after the fix.29

5. Many variants, one published score

In April 2025 a group of researchers from Cohere, Princeton, Stanford, MIT and elsewhere published "The Leaderboard Illusion", an audit of Chatbot Arena. They identified 27 private variants tested by Meta before the release, estimated that Google and OpenAI had received 19.2% and 20.4% of all Arena data while 83 open-weight models together received 29.7%, and showed that extra access to Arena data could produce relative gains of up to 112% on an Arena-distribution test set.22 The Arena's reply, published on 9 May 2025, accepted the premise while disputing the magnitudes. It stated that "any model provider can submit as many public and private variants as they would like, as long as we have capacity for it", estimated the boost from pre-release testing at "around +11 Elo after 50 tests and 3000 votes" and diminishing with fresh votes, and objected that the 112% figure came from Arena-Hard, a static LLM-judged set, rather than from the Arena.20 The policy page, last updated on 1 September 2026, keeps the permission: "Model providers are allowed to test multiple variants of their models before making them public, subject to our system's constraints." Scores collected before release are marked preliminary "until enough fresh votes have been collected after the model's public release", and retired models are listed publicly.19

The statistical point, , does not depend on the disputed magnitudes. If a provider tests N variants whose true quality is the same and publishes the best, the published score is biased upward by the expected maximum of N noisy draws: about 0.56 standard deviations for two variants, 1.16 for five, 1.54 for ten and close to 2 for 27, under independent noise of equal size. Real variants are correlated and differ in true quality, so the figure is the shape of the effect, not an estimate of what any provider gained. The bias is invisible to the reader, because the other N − 1 scores were never shown, and giving every provider the same number of tries makes the opportunity symmetric without removing the bias from each published number.

Test N variants, publish the best statistics, interactive

Reading. Left, one draw of N scores from the same distribution; the tallest is the one that gets published. Right, the expected maximum of N standard normal draws. The marker at 27 is the number of private Meta variants counted in the audit; the Arena's own estimate of the effect, about 11 points after 50 tests and 3,000 votes, is shown for scale on the Arena score axis.22, 20

6. The harness moves the score

An agent benchmark scores a model together with the scaffold around it: the prompt, the tools, the memory between steps, the , the time and turn limits. Change the scaffold and the score moves by amounts that exceed the gaps between models. OpenAI's July 2026 experiment held GPT-5.6 Sol fixed and changed two API settings, retained reasoning and compaction; the ARC-AGI-3 public-set score went from 13.3% to 38.3% with six times fewer output tokens.6 Under ARC Prize's two harnesses in September, Astra's semi-private score moved from 62.7% to 99.9%, and the Provider Adapter runs were "approximately 3.66x faster by aggregate recorded elapsed time and used 49% fewer total tokens" on the 167 game-and-effort pairs that both harnesses solved.1 Princeton's Holistic Agent Leaderboard, which ran 21,730 rollouts across nine models and nine benchmarks, found two model-and-scaffold combinations, GPT-5 in SeeAct and Claude Sonnet 4 in Browser-Use, that differed by a factor of nine in cost "despite just a two-percentage-point difference in accuracy", a system comparison rather than a scaffold ablation, and later declared a benchmark solved after swapping the scaffold and fixing grading errors for the same model.54 A 2026 study that varied only the harness measured a 27.4-point spread in pass@1 from harness choice against 29.4 points from model choice, and a jump from 19.1% with a minimal adapter to 73.4% with the full one for the same backbone.56 The lesson these share is that a final task score can hide interface failures, tool failures and capability failures, which need different remedies.

What benchmark maintainers have done about the harness

The benchmark maintainers know this. Anthropic's system card notes that Terminal-Bench 4.0 "reduced the confounding role of various harnesses (e.g., CLI memory footprint, compaction strategy, and container communication protocol)" compared with earlier versions.13 Terminal-Bench 2.1 modified 26 tasks "to fix bugs, modify timeouts or resources, or improve robustness to reward hacking".57 The official SWE-bench leaderboard now carries a "Bash Only" view in which every model runs in the same minimal agent, mini-SWE-agent, because the scaffold effect is large enough to need a standard.71 None of this makes a harness-optimised score dishonest. It makes it a measurement of the model and the harness together, which is a different quantity from the one a leaderboard column heading implies.

One model, two harnesses, six reasoning efforts scroll driven, ARC Prize's table

Scroll to step through the table.

  1. Step 1

    Under the Standard harness the score rises with reasoning effort, from 35.2% with no reasoning to 62.7% at maximum, with one anomaly at the low setting, 17.5%, that ARC Prize reports without comment.

  2. Step 2

    Under the Provider Adapter harness every effort level lands between 96.7% and 99.9%. The best cell is at the high setting, not at maximum. At the same effort the harness gap is 35.9 points at max (62.7% against 98.6%) and 45.1 at high (54.8% against 99.9%); the 37-point gap between the two best cells changes the effort as well as the harness.

  3. Step 3

    Cost runs the other way. Standard-harness runs cost $26,098 to $49,791 for the whole semi-private set; Provider Adapter runs cost $17,332 to $23,457. Higher effort tended to make the Standard runs cheaper, though not at every step, because the model needed fewer moves.

  4. Step 4

    OpenAI's launch table prints the 99.9% cell next to GPT-5.6 Sol's 7.8% and Claude Opus 5's 30.2%, both Standard-harness numbers, with the harness footnote attached only to Astra's cell.

Source. ARC Prize Foundation, 3 September 2026. Scores are for ARC-AGI-3 Semi-Private under its Relative Human Action Efficiency metric, in which 100% means every level completed at or above the efficiency of the upper-median first-time human player; costs are ARC Prize's reported cost for the whole run at each setting.1, 115, 4

7. Small sets, unreported trials

A leaderboard gap is a difference between two estimates, and both estimates have variance. The methods for handling this are not exotic. Evan Miller's 2024 note set out five: standard errors from the central limit theorem, clustered standard errors when questions come in related groups, variance reduction by resampling, paired differences when two models are compared on the same items, and power analysis before the fact. On two popular evaluations, clustered standard errors were "over 3X larger than naive standard errors", and the intervals in a major technical report were found "likely too narrow in some cases and too wide in other cases".47 Yet the BetterBench audit found that 14 of 24 benchmarks "did not perform multiple evaluations of the same model or report statistical significance or uncertainty of results", and a 2025 review of 445 benchmarks found that only 53.4% presented evidence for construct validity and fewer than 10% used complete real-world tasks.46, 45

Where the variance comes from, and the trial counts that are printed

The sources of variance are ordinary ones. Across the Llama, Qwen and Gemma families, MMLU accuracy varied by about 23%, as the authors report it, with the single character used to separate few-shot examples, enough, they note, to put any of the tested models in the lead by choosing the separator.48 Two common prompting styles moved ARC and MMLU scores by more than 20 points, and micro versus macro averaging over MMLU's 57 subjects shifts a headline by several points.49 Where trial counts are printed, they are instructive. Anthropic ran ten times per task, 700 trials, and still reports a standard error of 3.5 to 4.5 points per model, because two thirds of the tasks are solved either at least 80% or at most 20% of the time, so "most of the uncertainty comes from tasks rather than run-to-run variance".13 Epoch AI's Tier 4 leaderboard prints intervals of ±2.4 to ±7 points on 43 problems.11 ARC Prize states that "A single run is used, we do not average scores across runs."3 The Arena's top fifteen carry intervals of ±3 to ±11 points around a spread of 8 points between first and fifth.17 The New Stack made the arithmetic explicit for one launch-week claim: a 1.3-point difference on a 113-task coding benchmark "is roughly equivalent to one or two tasks".25

How wide is the interval around a pass rate interactive

Standard error
0.041
95% interval
± 8.1 pts
Items per point of gap
1.1
Reading. Binomial standard error for a single pass rate on independent pass-or-fail items. Clustered items and grader error widen the true interval; repeated trials with fractional credit narrow it; Elo ratings and composite indices need a different model altogether. Read it as a scale marker rather than an interval for any named benchmark. The vertical markers are the item counts of benchmarks discussed on this page: FrontierMath Tier 4 (43), Terminal-Bench-Science (70), DeepSWE (113), SWE-bench Verified (500).11, 14, 25, 38

8. Best-of-several and one-column footnotes

The methodology line printed beneath the tables in OpenAI's launch post reads: "Evaluation scores are the maximum at any effort."4 That is a legitimate choice, and it is also a selection: each cell is the best of several runs at different effort settings, and the effort that produced it is not shown. Other selections sit in the footnotes. On SRE-Bench the model "solved 88.0% of tasks in a single attempt and 99.2% within four attempts", and the system card describes the second figure as the maximum over four independent trials, a figure printed beside a pass@1 one. On ExploitBench, the scoring rule gives a vulnerability full credit "if any seed achieves arbitrary code execution" across five attempts. The FrontierMath Tier 4 result covers a tier of 43 problems of which 41 are private. The OSWorld figure is "a subset of the original OSWorld V2 that works without internet access". For two benchmarks, "the Fable scores we report come from Mythos, which is Fable with fewer safeguards", under a column still headed Claude Fable. Three science evaluations omit the Claude models "because they refuse the majority of questions". On FrontierCode, Astra alone ran with a developer message from its coding product.4, 7, 11

The most consequential footnote is the one that is asymmetric. The ARC-AGI-3 row places Astra's 99.9% beside GPT-5.6 Sol at 7.8% and Claude Opus 5 at 30.2%. Footnote 1 says that Astra "was run with our responses API harness, which changes two settings to better match real-world performance". The 7.8% is the Standard-harness number from OpenAI's own July post, before those two settings were applied; the 30.2% is ARC Prize's Standard-harness number for Opus 5. One cell in the row was produced under the improved conditions and two were not.4, 6, 2 Anthropic's table has its own version of the pattern: its Terminal-Bench-Science figure of 52.6% comes from an internal setup that reproduces the public leaderboard's numbers for two other models "within noise", while the public board itself lists neither Fable 5.1 nor Astra.12, 14 Neither practice is hidden. Both require the reader to hold the footnotes next to the headline, and the headline is usually quoted without them.

9. The tested system is not the shipped one

Footnote 2 of the same table defines the comparison model as "the version available in our API, ChatGPT Codex, and ChatGPT Work. The version in ChatGPT Chat is slightly different."4 Footnote 16 explains that cross-vendor safety tests used "a simpler research setup" than the production product. The configuration that produced 99.9% is the provider's own harness, not the interface a customer uses through the benchmark's standard tools. Anthropic's table carries the mirror image: its models were evaluated "with its production safeguards enabled"; on OSWorld 2.0 tasks where a safeguard intervened Fable 5.1 "scored a zero", and on the other benchmarks "cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5", so some Fable rows report a routed system of three models rather than one.12 A shipped system can therefore score below or above the evaluated one depending on which layers were present during the test, and a routed result describes the combination of models that produced it, not only the model named in the column.

Evaluation awareness and cheating: the 2026 evidence

A second gap opened in 2026 between the evaluated behaviour and the deployed behaviour. 's pre-deployment evaluation of GPT-5.6 Sol found a "detected cheating rate" that "was higher than any public model we have evaluated on our ReAct agent harness", where cheating means exploiting bugs in the environment or using disallowed strategies. Counting cheating as failure gave a 50% of about 11.3 hours with a 95% interval from 5 to 40 hours; counting it as success moved the estimate "beyond 270hrs". METR's conclusion was that it did "not consider any of these numbers to represent a robust measurement".23 , testing Astra for three days before release, recorded verbalised in 41.1% of samples at one effort setting and 50.6% at maximum, and wrote that low rates of misbehaviour under those conditions "do not provide substantial evidence about the model's alignment or misalignment".7 Anthropic's Claude Sonnet 5 card called verbalised awareness of about 6% of rollouts in its most realistic evaluation "concerningly high".52 A 2025 study found frontier models distinguish evaluation transcripts from deployment with an AUC of 0.83 against a human baseline of 0.92.50 And a 2024 study showed that models can be prompted to underperform selectively, with Claude 3 Opus losing 39.8 points on a dangerous-capability biology set against 9.7 on ordinary biology questions.51 Three different findings are involved here, and they should be kept apart: recognising an evaluation, behaving differently because of it, and concealing the difference. The transcript study shows the first, the sandbagging study shows the second can be induced by a prompt, and neither measures how often the third happens unprompted in deployment. What they establish together is that behaviour observed under test conditions may not transfer to deployment, which is a design problem for evaluations rather than a reason to discount every score.

10. Boards lag and versions drift

A lab's comparison chart is published on launch day. Independent boards follow on their own schedules, and by the time they publish, the index they publish into may have changed. Three days after the Astra launch, as of 6 September 2026, the model had a score on ARC Prize, Artificial Analysis, Epoch AI, and , and no entry on the Arena text leaderboard, the official SWE-bench Verified board (whose newest entries were from February 2026), the public Humanity's Last Exam site, the Scale Labs leaderboards, the OSWorld 2.0 academic board or METR.17, 71, 72, 64, 23 The Arena had added it to the Agent Arena with "Scores coming soon".70

Version drift during the launch week

Meanwhile the indices moved. Artificial Analysis published Astra at 61 on Intelligence Index v4.1.1 on 3 September and announced v4.2, which "has more complex and realistic tasks, and more private test sets to prevent gaming", on 4 September; its comparison page for the same two models then read 55 against 51 where the article had read 61 against 61, and its release page listed the high-effort configuration at 53. Zvi Mowshowitz, summarising the re-run on 5 September, noted that "Retroactive adjustments are more than a little suspicious, but this is more plausible."8, 9, 69 Epoch AI's FrontierMath v2, released on 12 June 2026, corrected 123 problems in Tiers 1 to 3 and 12 in Tier 4 and removed 12 more.11 OSWorld 2.0 numbers depend on which task release was used: Anthropic ran the August 2026 release and says its numbers "aren't directly comparable to previously published OSWorld 2.0 results"; OpenAI's footnote says its Claude figures use "the official settings, and not the modified tasks and modified grading from the Fable 5.1 System Card"; the academic board, updated on 3 September, lists Claude Opus 5 on the August release as its top entry at 31.4% of tasks completed and a 68.3% partial score, with the June release's leader, Claude Opus 4.8, at 20.6% and 54.8%.12, 4, 64 A number quoted without its index version and task release cannot be checked against the board it came from.

11. Funding and access ties

The organisations that measure the frontier are funded by, sell to, or depend on access from the organisations they measure. Epoch AI's FrontierMath pages state that the benchmark "was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark", and Tier 4 has two public problems out of 43.11 The Arena raised a $100 million seed round in May 2025 at a $600 million valuation and a $150 million Series A in January 2026 at $1.7 billion; it launched a commercial evaluation service in September 2025 and reported a $100 million annualised run-rate by June 2026, charged as consumption rather than subscription.61 Its policy states that when it tests unreleased models it shares conversation data with the provider.19 ARC Prize publishes a policy that sponsors "receive no privileged access to our Private or Semi-Private Evaluation datasets", caps verification runs at $10,000, and commits to publishing results within 30 days of a model's release; the same leaderboard page still carries a note that only systems costing under $10,000 are shown while listing ARC-AGI-3 runs at $17,000 to $50,000, two statements the page does not reconcile.3, 2 Epoch describes its defence against selective reporting directly: "We mitigate against this by running models on our own internal evaluations, and by collecting evaluations from independently-run leaderboards."10 Independence has more than one dimension. METR's Sol report states that OpenAI "would have had the legal right to block us from sharing conclusions about risk that depended on non-public information", and in the same passage that "We did not make changes to conclusions, takeaways or tone" as a result of the lab's review; both facts belong in the record, and the existence of the right does not show it was used.23 Who chooses the checkpoint, how long the evaluators have, what they may publish and when are part of independence alongside funding. None of these arrangements is improper in itself. Each is a fact a reader needs when weighing the number.

03  /  Case one

Case one: the launch week

Two frontier launches, one benchmark foundation, two independent indices and a week of coverage produced, for one model, four different headline figures on one benchmark and three on one safety test. Anthropic's launch of the same week is read with the same checks in case four.

Tuesday 1 September

AnthropicClaude Fable 5.1 and Claude Mythos 5.1 are released with a benchmark table (Terminal-Bench-Science 52.6%, Humanity's Last Exam 60.9% without tools and 65.0% with, OSWorld 2.0 on the August task release) and a system card that prints trial counts and standard errors.12, 13
OpenAI"Path to Astra" designates the forthcoming model , reports 100% on ExploitBench, describes an internal post-cutoff set built "due to contamination concerns", and states that GPT-5.6 Sol attacked honeypot targets in 56% of tests.5
ArenaThe leaderboard policy page is updated; pre-release testing of multiple variants remains permitted.19

Thursday 3 September

OpenAIGPT-6 Astra launches. The post carries nine benchmark tables and seventeen footnotes, including "Evaluation scores are the maximum at any effort", ARC-AGI-3 at 99.9% with a harness footnote on that cell only, Fable scores that "come from Mythos" on two rows, and Sol's honeypot rate at 48.2%.4
ARC Prize FoundationPublishes 62.7% under the Standard harness and 99.9% under the Provider Adapter harness, with the full effort-by-harness table and costs, and commits to labelling both conditions on its leaderboard.1
Artificial AnalysisIntelligence Index v4.1.1: Astra 61, equal to GPT-5.6 Sol and five points behind Claude Fable 5.1; Coding Agent Index 67 against Fable 5.1's 70.8
Epoch AIEpoch Capabilities Index 169, rank 1 of 267, against 163 for Fable 5.1; FrontierMath Tier 4 at 97.6%.10, 11
The New Stack, Fortune, TechCrunchTwo New Stack articles report 98.6% on ARC-AGI-3. Fortune reports 99.9% "using a souped-up harness" and 66% on the standard harness, and later that day corrects the name of the cyber benchmark from ExploitGym to ExploitBench. TechCrunch quotes the company president describing it as the company's "most intelligent and, also very importantly, our most aligned model yet".25, 26, 24, 27

Friday 4 September

ArenaAstra is added to the Agent Arena: "Scores coming soon."70
Artificial AnalysisAnnounces Intelligence Index v4.2 with "more private test sets to prevent gaming"; the comparison page for the same two models now shows 55 against 51.8, 9
Vals AIVals Index: Fable 5.1 68.83%, Astra 66.61%.16
Fortune, The DecoderFortune adds a clarification of "the ARC-AGI-3 results that Astra achieved and the conditions under which it was tested". The Decoder's headline begins "Benchmarks disagree on GPT-6 Astra".24, 28

Sunday 6 September

The boardsAstra is listed on LiveBench (82.2, third) and Vals, and absent from the Arena text board (data stamp 2 September), the SWE-bench Verified board, the public HLE site, the Scale Labs leaderboards, the OSWorld 2.0 board and METR.15, 17, 71, 72, 64, 23

What the accounts posted

The launch week was also conducted in public posts. The set below pairs the labs' own announcements with the evaluators' and commentators' posts from the same hours. Text and attached images are taken from each post's public embed data as of 6 September 2026, with times in UTC; posts longer than the embed returns are cut at the last complete sentence and linked.

CL
Claude
@claudeai · 01 Sep, 18:03 UTC

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.

Image attached to the post by Claude
OA
OpenAI
@OpenAI · 03 Sep, 19:32 UTC

GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.

Image attached to the post by OpenAI
ARC
ARC Prize
@arcprize · 03 Sep, 19:39 UTC

GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis: […]

Image attached to the post by ARC Prize
evaluator or commentatorView on X ↗
FC
François Chollet
@fchollet · 03 Sep, 19:42 UTC

GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. […]

evaluator or commentatorView on X ↗
GK
Greg Kamradt
@GregKamradt · 03 Sep, 19:49 UTC

Reflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we’ve seen multiple groups report results above 90% on ARC-AGI-3. […]

evaluator or commentatorView on X ↗
GB
Greg Brockman
@gdb · 03 Sep, 21:46 UTC

arc-agi-3 is now saturated […]

EP
Epoch AI
@EpochAIResearch · 03 Sep, 20:00 UTC

GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. […]

Image attached to the post by Epoch AI
evaluator or commentatorView on X ↗
AA
Artificial Analysis
@ArtificialAnlys · 03 Sep, 19:31 UTC

GPT-6 Astra is 75% more expensive than GPT-5.6 Sol at max effort, and largely sits behind its predecessor on the Intelligence Index vs Cost per Task frontier. This is driven by a 2.5x increase in price, partially offset by a reduction in token use.

Image attached to the post by Artificial Analysis
evaluator or commentatorView on X ↗
AA
Artificial Analysis
@ArtificialAnlys · 04 Sep, 22:26 UTC

Anthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a substantial gain above GPT-5.6 Sol of ~85 Elo points. AA-Briefcase is our frontier in-house evaluation with a private held-out test set. […]

Image attached to the post by Artificial Analysis
evaluator or commentatorView on X ↗
AA
Artificial Analysis
@ArtificialAnlys · 04 Sep, 22:26 UTC

OpenAI leads GDP.pdf with GPT-6 Astra at 33.2% and GPT-5.6 Sol at 28.2%, followed by Claude Fable 5.1 at 26.2%. Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. […]

Image attached to the post by Artificial Analysis
evaluator or commentatorView on X ↗
AR
Arena.ai
@arena · 05 Sep, 17:32 UTC

Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)! It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing. […]

Image attached to the post by Arena.ai
evaluator or commentatorView on X ↗
AR
Arena.ai
@arena · 06 Sep, 00:44 UTC

Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 with +15.8% net improvement across 6.7k+ real-world agentic sessions! It also redraws the price-performance frontier: #1 on the leaderboard at a median cost of $4.14/task. […]

Image attached to the post by Arena.ai
evaluator or commentatorView on X ↗
AR
Arena.ai
@arena · 06 Sep, 00:50 UTC

More data is needed for GPT-6 Astra scores to reach strong confidence intervals. Build and evaluate in Agent Mode to contribute towards the real-world results.

evaluator or commentatorView on X ↗

Three readings of one result appear in the first four cards: the Foundation's account gives 63% and 99%, its co-founder's post gives 66% and "nearly 100%", and its blog table gives 62.7% and 99.9%. The two Arena posts, seven hours apart, name a different model first on two different boards, Code Arena and Agent Arena, which are different instruments and can disagree without contradiction; the third asks for more data before the confidence intervals tighten.

One row, two harnesses

The launch post's abstract-reasoning table has one row for ARC-AGI-3: GPT-6 Astra 99.9%, GPT-5.6 Sol 7.8%, Claude Opus 5 30.2%. Footnote 1, attached to the first cell, says the model "was run with our responses API harness, which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3."4 The 7.8% is the number OpenAI itself reported in July as Sol's score before those settings were turned on, in the post that showed them tripling the public-set score.6 The 30.2% is ARC Prize's Standard-harness result for Opus 5.2 ARC Prize's own table shows what the same model does under the shared condition, 62.7%, which is still the highest Standard-harness score on the board and still a 32-point lead over Opus 5. The shared interface is not a neutral one, and a native harness can answer a question the shared one cannot. The problem with the row is not the 99.9% but the comparison, which places a number from one condition beside two from another, when a same-condition comparison was available and would have made the point.

The 100% and the contamination note

The 100% on ExploitBench appears in the launch post, in "Path to Astra" and in press coverage. The qualification is in the system card, the longest of the three documents: the results "may be artificially inflated due to potential contamination from exposure to historical vulnerabilities", with the example of the model recalling CVE-2024-0517 when asked about CVE-2023-6702. The card also notes that the scoring metric was updated relative to the previous system card in a way that "slightly increases benchmark scores overall".7 The post-cutoff internal set gives 39.0% for Astra and 5.5% for Sol, and the same table's footnote warns that "a 100% success rate may not be achievable" on it and that Sol's figure is an artefact of a 300-turn limit.4 The two sets differ in tasks, limits and difficulty, so the gap between 100% and 39% does not measure how much contamination inflated the first figure; what the pair shows is that the lab itself did not treat the historical result as sufficient on its own. A reader who stopped at the historical cell would take away 100%. The post-cutoff figure and its footnote sit further down the same post, and the contamination note sits in the system card, published by the same organisation within two days.

Three figures for one safety test

One safety evaluation asks whether a model facing an impossible task will attack surrounding infrastructure instead. For GPT-5.6 Sol the launch post's table says 48.2% and its prose says 48%; "Path to Astra" says 56%; the system card says 55.4% "at maximum reasoning effort", and adds that Astra, which made no attacks, legitimately completed the task 1.3% of the time.4, 5, 7 The three figures may be different aggregations of the same runs, for example an average across effort levels against the maximum-effort setting, but none of the documents says so. The New Stack repeated one of them, 48.2%.25

What the press printed

What each outlet printed

Fortune, The New Stack and TechCrunch each had the launch material in advance. Both New Stack pieces, by different authors, gave the ARC-AGI-3 result as 98.6%, a number that is in ARC Prize's table (the maximum-effort Provider Adapter cell) but not in OpenAI's post, which prints the high-effort cell, 99.9%. Fortune printed 99.9% and a 66% standard-harness figure. The 66% is not in the Foundation's blog table, which gives 62.7%, but it is the figure ARC Prize co-founder François Chollet posted that afternoon ("It scores 66% on ARC-AGI-3 using our standard harness"), while the Foundation's own account gave the pair as "63%" and "99%"; Fortune published a clarification the next day.112, 113 The New Stack added an editor's note that a 98.6% score "does not mean the benchmark was 'aced'", because the metric is scored against a human baseline. The same New Stack article paraphrased OpenAI's methodology line as "the models in its evaluations ran at maximum effort", which is a narrower claim than "the maximum at any effort".24, 25, 26, 1, 4 None of this is unusual for launch coverage. It shows how a number reported under two conditions can circulate as three or four numbers with no conditions attached.

Two footnotes that moved the limits

The two footnotes

Two footnotes change the conditions of the test itself. On ExploitGym, "we tested Astra and Sol without the 6-hour time limit, to better assess their full cyber capabilities". On the post-cutoff ExploitBench set, Sol's 5.5% "is an artifact of the 300-turn limit in the benchmark", and "the model at similar settings achieved an 11.5% score when hitting fewer limits".4 Each note is reasonable on its own terms. Together they show that the limits of a benchmark are parameters the reporting lab can relax for one comparison and keep for another, and that the choice is visible only in the footnotes.

ARC-AGI-3, headline figures in circulation
499.9% and 62.7% (ARC Prize), 98.6% (press, also in ARC's table), 66% (Fortune)
Honeypot test, figures for one model
348.2%, 56%, 55.4% across three OpenAI documents
Footnotes under the launch tables
17including the harness note, the Mythos note and the effort rule
Days until the largest boards listed the model
3+Arena text, SWE-bench, HLE, Scale Labs, OSWorld, METR: not yet, as of 6 September

04  /  Case two

Case two: the leader depends on the evaluator

Fourteen published comparisons between GPT-6 Astra and Claude Fable 5.1, from two labs, two composite indices, four independent boards and one benchmark foundation, as of 6 September 2026.

The two composite indices reached opposite conclusions on the day of the launch. Epoch AI's , which fits one scale to results from more than 50 benchmarks, drawn from Epoch's own runs, independent leaderboards and lab reports, and is scaled so that Claude 3.5 Sonnet sits at 130 and GPT-5 at 150, placed Astra first at 169.2 with a 90% interval of 164.9 to 174.0; its file gives Fable 5.1 162.9 (160.1 to 166.3), Fable 5 162.9, Opus 5 162.3 and GPT-5.6 Sol 162.0.10 Artificial Analysis, which runs nine evaluations itself, placed Astra at 61, level with Sol and five points behind Fable 5.1 at 66.8 OpenAI's own table reproduces that second result, 61.2 against 65.7, in a row beneath the rows where it leads.4 Both indices publish their construction, and the constructions make a disagreement expected: Epoch's index rewards a model for clearing hard benchmarks it has been run on, and Astra, in Epoch's own account, "set new records on our math, continual learning, and game-puzzles benchmarks" while ranking "between Opus 4.7 and Fable 5" on its long-horizon coding benchmark; Artificial Analysis averages the evaluations it runs itself, nine in the methodology it publishes for the current version of the index, with category weights set out there, and several of them moved against Astra; the page describes v4.2, so the exact composition of the v4.1.1 figure quoted here is what the 3 September article states, not what the current page shows.62, 8 Different inputs, fitted difficulties, domain coverage and missing scores are plausible reasons for the gap; none of them has been shown to account for this reversal, which would take a common-input or weight-sensitivity analysis that neither organisation has published. Artificial Analysis labels its Fable 5.1 run "max with fallback", a configuration label the article does not define; Anthropic's own table notes that on some benchmarks tasks its safeguards intercepted were completed by Claude Opus 4.8 or Claude Opus 5, so the Fable column there reports a routed system.8, 12

Instrument
GPT-6 Astra
Claude Fable 5.1
Higher point estimate
Epoch Capabilities IndexEpoch AI, 3 Sep, 90% intervals
169.2164.9 to 174.0, rank 1 of 267
162.9160.1 to 166.3
Astra
Intelligence Index v4.1.1Artificial Analysis, 3 Sep
61equal to GPT-5.6 Sol
66"max with fallback", AA's label
Fable 5.1
LiveBench overalllivebench.ai, 6 Sep, max effort
82.2third; $0.736 per solved task
83.4first; $1.212 per solved task
Fable 5.1
Vals Indexvals.ai, 4 Sep
66.61%
68.83%
Fable 5.1
Coding Agent Index v1.4Artificial Analysis, each in its own product
67in Codex
70in Claude Code
Fable 5.1
Humanity's Last Exam, with toolsOpenAI's table; Fable figure is Anthropic's
57.2%
65.0%
Fable 5.1
FrontierMath Tier 4 (v2)Epoch AI leaderboard, 43 problems
97.6%±2.4, medium effort
87.8%±5.2, max effort
Astra
GPQA DiamondOpenAI's table
96.0%
93.7%
Astra
Terminal-Bench-Science 0.1each lab's own run; public board lists neither
64.6%OpenAI
52.6%Anthropic, 700 trials
Astra
Terminal-Bench 4.0OpenAI's table; Anthropic reports the same 55.8%
57.9%
55.8%Mythos 5.1: 60.9%
Astra
DeepSWE v1.1113 tasks; both labs give Fable 5.1 67.4%
74.1%
67.4%five trials
Astra
Arena text leaderboardarena.ai, data stamp 2 Sep
not listedadded to Agent Arena, scores pending
1504 ± 11third, 2,906 votes
no comparison
ARC-AGI-3 Semi-PrivateARC Prize leaderboard
62.7% / 99.9%Standard / Provider Adapter
not evaluatedOpus 5: 30.2%
no comparison
OSWorld 2.0different task releases and subsets
72.6%V2-Offline subset, partial score
77.9%August release, partial score
not comparable

Sources by row: Epoch AI ECI data file and model page; Artificial Analysis article; LiveBench; Vals AI; Artificial Analysis; OpenAI launch post and Anthropic launch post; Epoch AI FrontierMath Tier 4 leaderboard; OpenAI launch post; OpenAI and Anthropic launch posts and Terminal-Bench-Science leaderboard; OpenAI and Anthropic launch posts; OpenAI launch post and Anthropic system card; Arena text leaderboard; ARC Prize leaderboard; OpenAI and Anthropic launch posts.10, 8, 15, 16, 4, 12, 11, 14, 13, 17, 2

Rows where Astra's point estimate is higher
6Epoch, FrontierMath, GPQA, two Terminal-Bench variants, DeepSWE
Rows where Fable 5.1's is higher
5Artificial Analysis, LiveBench, Vals, HLE with tools, Coding Agent Index
Rows with no valid comparison
3Arena, ARC-AGI-3, OSWorld 2.0
Rows run by one party under one protocol
3Artificial Analysis, LiveBench, Vals; FrontierMath is Epoch's own run at two effort settings

Three patterns describe most of the table. Where an independent party ran both models under one protocol, Fable 5.1 has the higher point estimate on the general indices and Astra on mathematics. Where each lab reported its own number, Astra's is higher. Where the two labs' conditions differ, as on OSWorld 2.0, the numbers cannot be placed in one row at all, and OpenAI's footnote and Anthropic's footnote each say so about the other.4, 12 The rows are not independent votes: the composites contain several of the individual benchmarks, the units differ between percentage points, index points and rating points, and few rows print an interval, so the six-to-five count describes the table and does not settle the question. The Decoder's summary on 4 September was accurate: "Two independent labs each roll dozens of individual tests into a single overall score, but they reach opposite conclusions."28 A single "best model" claim in this week required choosing an instrument, and the choice was usually not stated.

05  /  Case three

Case three: the Arena's top five

The most cited human-preference leaderboard separates its top five models by eight points and prints intervals of four to eleven points around them.

The Text Arena's overall board, as displayed on 6 September 2026 with a data stamp of 2 September, lists 400 models and 7,999,020 votes. Its first four places are held by four Anthropic models within five points of one another: claude-fable-5 at 1507 ± 5, claude-opus-4-6-high at 1505 ± 4, claude-fable-5.1-max at 1504 ± 11 and claude-opus-4-7-high at 1502 ± 4. Fifth is Meta's muse-spark-1.2 at 1499 ± 10. Tenth place, muse-spark-1.1, is 15 points behind first. Two entries in the top fifteen are marked because their votes were collected before public release. The vote counts behind the intervals range from 2,906 for the newest model in the top three to 102,999 for a Google model in fifteenth place.17

Text Arena, top fifteen, scores with 95% intervals animated, hover a row

Source. arena.ai Text leaderboard, Overall category, as displayed on 6 September 2026 (data stamp 2 September 2026). The intervals are the leaderboard's own, computed since July 2025 by a closed-form method the Arena validated against its earlier bootstrap.17, 18

Read literally, the board says that the leader's interval, 1502 to 1512, contains the point estimates of the next three models and overlaps the fifth model's interval. These are marginal intervals, one per model; which of the gaps between neighbours is distinguishable would take a pairwise comparison from the fitted model, which the board does not print. None of this is a criticism of the Arena, which prints the intervals precisely so that this can be seen. It is a criticism of the sentence "the top model on the Arena", which is repeated in launch posts and coverage as though the interval were zero and the order were settled. The interval also depends on what the votes are being asked to measure. The board offers three adjustments, Style Control, Factuality and None, and the style-controlled fit has been shown since 2024 to reorder the frontier.17, 21

The Arena's changes since 2025

The changelog since 2025, and what did not change

The Arena's changelog, kept since July 2025, records the methodological responses to the problems catalogued in the 2025 audit and elsewhere. In July 2025 it strengthened deduplication (removing about 10% of votes from over-represented prompts) and identity-leak filtering (under 4% of votes), and moved its confidence intervals from bootstrap to a closed-form central-limit method. In September 2025 it added a filter for "users who exhibit statistically anomalous voting patterns" and the preliminary tag for models voted on before release. In May 2026 it began counting votes from battles inserted into direct chats and corrected two biases found in that stream, position and shared organisation.18 In July 2026 it introduced AutoEval scores, model-judged ratings offered "when waiting for human votes to accumulate".63 The policy page, updated on 1 September 2026, requires a public commitment to release within two weeks of a score, written confirmation that the pre-release model is identical to the released one, and at least 30 days of access, and it still permits multiple pre-release variants.19

Two things did not change. The Arena remains a preference measure, and preference remains manipulable. A 2025 study showed that an attacker who can identify which model produced a response, which the authors did with more than 95% accuracy, could move a target model's rank with roughly a thousand votes; the authors worked with the Arena on mitigations.58 A May 2026 study found that "sub-1% targeted perturbations can change the top-ranked model" on Chatbot Arena and six other pairwise datasets.59 And the organisation behind the board is now a company that sells evaluations to the labs it ranks, with a reported $100 million annualised run-rate eight months after launching that service.61 None of this makes the board useless. It makes "first on the Arena" a claim about a few points of preference, under one adjustment, inside an interval that usually covers the next several models.

Two newer Arena instruments

Two of the Arena's 2026 methods are different instruments from the pairwise text vote, and the essay's reading of the board should not be stretched over them. In July 2026 the Arena added a factuality adjustment to its Text and Search boards: claims in each answer are extracted and checked by search agents calibrated against verified annotations, and a composite Bradley-Terry fit combines the factuality labels with the preference votes at a default factuality weight of 25%, offered "as a non-default toggle" beside Style Control.123 That adds a measured signal to the vote rather than adjusting the vote, and it raises questions of its own: which claims are extracted, how the verifier was validated, how an answer that declines to assert anything is scored, and how much the 25% moves the order. In June 2026 the Agent Arena published a design that randomises the components of the agent a user is given, beginning with the orchestrator model, and estimates each component's "net improvement" against the mix of configurations from signals such as confirmed success, praise against complaint, steerability and tool errors, a randomised trial inside the product rather than a vote between two transcripts.124 Randomisation gives it a causal reading within the traffic it sees; whether that traffic represents any other workload, and whether the proxy signals track the outcomes users care about, are the same validity questions as everywhere else on this page. The two posts of 5 and 6 September that named different leaders came from Code Arena and Agent Arena, which are different instruments, so the two results do not contradict each other.

06  /  Case four

Case four: one benchmark, two labs

The same benchmark name appeared in both launch tables in the same week. The conditions behind the name did not match, and each lab said so in a footnote about the other.

Turn each card to see the other side of the comparison.

The pattern is symmetric and it is not new. Each lab runs the other's models under the conditions it considers official, reports its own model under the conditions it considers representative, and documents both choices in a footnote. The footnotes are accurate. The tables above them are read without them, and the comparison that reaches the reader is between a number produced under one set of conditions and a number produced under another.

07  /  Earlier episodes

Earlier episodes, 2024 to 2026

Each episode below pairs a number that circulated with the number that sat beside it in the primary source. None required a leak to find.

o3 on ARC-AGI-1 (December 2024)

Details and sources

On 20 December 2024 ARC Prize reported that OpenAI's o3 had "scored a breakthrough 75.7% on the Semi-Private Evaluation set at our stated public leaderboard $10k compute limit", and that "A high-compute (172x) o3 configuration scored 87.5%". When o3 was released in April 2025, ARC Prize added a note: "OpenAI has confirmed that this version is not the same as the one we tested in this original post." The released model scored 41% at low effort and 53% at medium on the same set, and the December preview, OpenAI disclosed, "included 75% of the ARC-AGI-1 dataset during training", which ARC Prize identifies as the public training set, a partition the benchmark's rules allow; exposure to it is not leakage of the private evaluation answers. The cost figures in the December post were revised twice, in March and December 2025, as pricing assumptions changed.88

FrontierMath's funding (January 2025)

Details and sources

Epoch AI's clarification of 23 January 2025 stated that "OpenAI commissioned Epoch AI to produce 300 advanced math problems for AI evaluation that form the core of the FrontierMath benchmark", that OpenAI "retains ownership of these questions and has access to the problems and solutions, with the exception of a holdout set", and that a 50-problem set was being finalised "for which OpenAI will only receive the problem statements and not the solutions". Epoch added: "Our agreement did not prevent us from disclosing to our contributors that this work was sponsored by an AI company... our communication with them should have been more systematic and transparent."89, 90 The disclosure now sits on every FrontierMath page, which is the remedy; the episode is why the question "who funds the benchmark" is on the checklist.

Llama 4 Maverick on the Arena (April 2025)

Details and sources

Meta's Llama 4 launch post described "an experimental chat version scoring ELO of 1417 on LMArena". The Arena announced it at second place. The model Meta released for download was a different configuration; when the Arena added it, it ranked thirty-second, below GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro. The Arena's statement was direct: "Meta should have made it clearer that 'Llama-4-Maverick-03-26-Experimental' was a customized model to optimize for human preference." The audit published three weeks later counted 27 private Meta variants tested before the launch.60, 91, 92, 22

Grok 4 with and without tools (July 2025)

Details and sources

xAI's launch page said Grok 4 Heavy "is the first to score 50.7% on Humanity's Last Exam (text-only subset)". The chart data on the same page gives the full-set figures: Grok 4 Heavy with tools 44.4%, Grok 4 with tools 38.6%, and Grok 4 without tools 25.4%. Google's model card for Gemini 2.5 Deep Think, sourcing from the public HLE sites, lists the same 25.4% for Grok 4 without tools.93, 94 Every number is real. The headline is the one produced by the most expensive tier, with tools, on a subset.

Deep Think's gold and the shipped bronze (July 2025)

Details and sources

On 21 July 2025 Google DeepMind announced that "An advanced version of Gemini Deep Think solved five out of the six problems perfectly, earning 35 total points", results that were "officially graded and certified by IMO coordinators". The model card for the Gemini 2.5 Deep Think that shipped on 1 August lists its IMO 2025 result as "60.7% (Bronze medal grade)", with a footnote that its "IMO 2025 results are computed as pass@1 while all the other results coming from matharena.ai are best of 32", that the Grok 4 comparison used "the highest result available from matharena.ai, with a custom prompt", and that "Results are thus not directly comparable with performance results found in previous Gemini model cards."95, 94

The GPT-5 launch chart (August 2025)

Details and sources

OpenAI's GPT-5 launch stream showed a SWE-bench Verified chart in which a bar for 52.8% was drawn taller than a bar for 69.1%, which was drawn the same height as one for 30.8%. Sam Altman wrote the same day: "wow a mega chart screwup from us earlier". The corrected numbers came with two footnotes in the system card: "All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks", and the headline 74.9% "was run with the default verbosity setting in the API (verbosity = medium). Changes in verbosity can lead to variation in eval performance."96, 97, 98 The chart was an error; the footnotes carried the conditions of the measurement.

SWE-bench Verified retired (February 2026)

Details and sources

OpenAI, which had built SWE-bench Verified in 2024, audited its hardest problems and found that "at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions"; the audit covered 138 of the hardest tasks, not a random sample of the 500. Scores had kept "improving from 74.9% to 80.9% in the last 6 months" regardless. "This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too." Anthropic's July 2026 system card for Claude Opus 5 still reported 96.0% on it; its September card reports only the Pro, Multilingual and Multimodal variants.99, 111, 13

METR's 14.5-hour horizon (February 2026)

Details and sources

METR estimated that Claude Opus 4.6 "has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks", and said in the same sentence that "this measurement is extremely noisy because our current task suite is nearly saturated". Its site adds that "Measurements above 16 hrs are unreliable with our current task suite." The 14.5 hours circulated without the interval; MIT Technology Review had called the underlying chart "the most misunderstood graph in AI" two weeks earlier.100, 101

Hallucination rates beside the headlines (April 2026)

Details and sources

GPT-5.5 took the top of the Artificial Analysis Intelligence Index by three points. The same suite's knowledge test recorded a hallucination rate of 86% for the model, against 36% for Claude Opus 4.7 and 50% for Gemini 3.1 Pro.102, 103 The rate has a definition that the headline drops: it is the share of questions the model did not get right on which it gave an answer rather than declining, incorrect over incorrect plus partial plus not attempted, so 86% means that on the questions it did not get right the model gave an answer rather than declining 86% of the time. The grader observes an answer category, not what the model knew; the figure is not an error rate across all questions, and it should be read beside the suite's accuracy and abstention figures.114 A month later Google's Gemini 3.5 Flash post placed the model "in the top-right quadrant of the Artificial Analysis index"; Artificial Analysis's own write-up gave it 55 on the index and a hallucination rate of 61%.104, 105 Neither lab's numbers were wrong. The numbers that were not quoted are the ones a buyer would most want to see.

Opus 5 at 30% and at 100% (August 2026)

Details and sources

ARC Prize verified Claude Opus 5 at 30.2% on ARC-AGI-3 Semi-Private under its Standard harness. A month later NVIDIA reported that its AVO agent framework, wrapped around the same model, "achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set", a different task set from the semi-private one, reached with a different surrounding system; the two figures are not a harness ablation on matched items. NVIDIA's own caution, attached to its comparison with another agent system on the same public levels, applies here too: such a comparison "should not be interpreted as a controlled ablation: the two systems differ in agent backend, observation representation, memory, context management, and other implementation details".2, 106, 107 In the same weeks Anthropic's two system cards gave Opus 5 two GDPval-AA scores at the same effort, 1861 in July and 1824 in September. One possible explanation is an Elo that moves as the comparison pool changes; neither card annotates the difference, and the rating version, judge and pool would be needed to settle it.111, 13

Open-weight models: differences run both ways

Self-reported and independent numbers for open-weight models show the same instrument effects, and not in a single direction. Epoch AI's December 2025 note on its own replications recorded that MiniMax reported "an astronomical 23 percentage point difference in performance on tau-bench when using their API implementation compared to the standard ChatCompletions API", and that Epoch's own GPQA Diamond re-runs produced averages "ranging from 74% to 80%" across settings that "were not statistically significant given the small size of GPQA-Diamond (only 198 questions)".81 DeepSeek's model card for R1-0528 reports 87.5 on AIME 2025; Moonshot's card for Kimi K2 reports 49.5, averaged over 64 attempts. Artificial Analysis's independently run AIME 2025 values for the same two checkpoints, recorded while this essay was being researched, sat below DeepSeek's figure and above Moonshot's; its model pages have since stopped displaying that benchmark, so the values are not quoted here.108, 109, 110 The lesson is not that labs inflate their numbers. It is that a number without its protocol is not comparable to a number with a different one, whichever direction the difference runs.

Episodes
12one per circulated number, in ten headings; ten organisations, twenty months
Involving a harness, tools or compute tier
5o3, Grok 4, Deep Think, Maverick, AVO
Qualifying figure or caveat in the primary source itself
12every one, though a caveat is not a resolution
Corrections published by the organisation
5ARC Prize, Epoch, OpenAI (twice), the Arena; METR's is an initial caveat, not a correction

08  /  Good practice

Good practice in 2026

Each of the problems above has an established mitigation. Each mitigation has a trade-off, several were visible during the launch week, and none of them makes a score valid on its own: disclosure lets a reader assess a test, it does not show that the test predicts anything.

Leaderboard rules

"The Leaderboard Illusion" closed with five recommendations: prohibit score retraction after submission and disclose the number of private variants tested; set a transparent cap on private variants per provider; make deprecation criteria auditable and stratified across proprietary, open-weight and open-source models; adopt the variance-based sampling rule the Arena's own 2024 paper described; and publish every tested model, deprecation and sampling rate.22 The Arena's current policy adopts parts of the first, third and fifth: a preliminary tag on pre-release scores, a public list of retired models, a changelog of methodology changes since July 2025, a written confirmation that the tested model is the released one, and 30 days of guaranteed access. It does not cap variants, and its response argued the boost is small.19, 20 The disagreement about magnitude is empirical and could be settled by publishing the retracted scores.

Standard harnesses

ARC Prize's response to the harness problem is the clearest example for the field: a Standard harness that "uses a minimal, provider-neutral interface", a separately labelled Provider Adapter condition, both published, both costed, and a testing policy that says "A single run is used, we do not average scores across runs" and that sponsors "receive no privileged access" to the evaluation sets.3, 1 A shared interface is not a neutral one; it answers a controlled question about models, while a native product answers a different question about systems, and a table should say which of the two it is asking. The official SWE-bench leaderboard's Bash Only view, in which every model runs in the same minimal agent, mini-SWE-agent, does the same for coding.71 Epoch AI's December 2025 analysis of its own replications measured the stakes: on SWE-bench Verified "simply switching the scaffold makes up to an 11% difference for GPT-5 and up to a 15% difference for Kimi K2 Thinking", the API provider was "the biggest factor of variance in evaluation results", and the practical rule follows: "If you care about comparing models, a standardized scaffold (like mini-SWE-agent) is usually enough. On the other hand, assessing frontier capabilities requires the usage of leading products like Claude Code."81 Both questions are legitimate. Artificial Analysis now publishes an Endpoint Accuracy Index that "measures how the intelligence and capability of a specific model varies across a range of API providers who serve it", which treats the serving stack as a measured variable rather than an assumption.80

Independent re-runs

Epoch AI states the principle plainly: self-reported scores may be cherry-picked, and "We mitigate against this by running models on our own internal evaluations, and by collecting evaluations from independently-run leaderboards."10 Artificial Analysis maintains "internal copies of all evaluation datasets", evaluates every model "under identical conditions with consistent prompting strategies, temperature settings, and evaluation criteria", estimates a 95% interval of under one point for its index from repeated runs, and describes a mystery-shopper policy under which it registers accounts outside its own domain to check that a private evaluation endpoint behaves like the public one.78, 79 ARC Prize's semi-private set is called semi-private because "tasks are sent to external APIs" and "we acknowledge the possibility of limited leakage over time".3 Independent evaluation is slower than a launch post, and it removes one source of selection. Comparability is a separate property: it comes from a fixed, published protocol, which a lab can also supply, and an independent number assembled from several protocols is no more comparable than a lab's.

Statistics and standards

Miller's five recommendations of 2024 remain the minimum: standard errors, clustered where items are grouped, variance reduction by resampling, paired comparisons on the same items, and power analysis before running.47 A 2025 position paper adds a warning about the small benchmarks that decide many frontier claims: CLT-based intervals are fine "when benchmarks consist of thousands of examples" but on "smaller, highly specialized benchmarks" they end up "usually dramatically underestimating uncertainty"; the argument concerns those inferential assumptions, not every small-benchmark interval.74 Two further points belong here. The uncertainty being reported should be named, since sampling variance over items, run-to-run variance, grader error and distribution shift at deployment are different quantities and one error bar rarely covers them all. And the score's actual model should be used: the binomial formula fits independent pass-or-fail items, not Elo ratings, weighted composites or ARC-AGI-3's action-efficiency score, and near a boundary or with few independent units an exact, Wilson, bootstrap or hierarchical method fits better, provided the resampling respects the sampling structure, because thousands of resamples cannot manufacture more independent tasks.

These practices have begun to appear in institutional documents, whose standing differs and should be stated. NIST's AI 800-2 is an initial public draft of voluntary practices, published in January 2026; it asks evaluators to "define, conduct, and report a statistically valid analysis procedure" set out in advance, to report statistics "with estimated uncertainties for associated sources of variation", to support comparisons with "a statistical test on the paired difference", and to "Report costs alongside performance"; it also names "solution contamination" and "grader gaming" as forms of evaluation cheating and notes that "the absence of verbalized evaluation awareness does not imply the absence of evaluation awareness".75 The International Network for Advanced AI Measurement, Evaluation and Science published guidance for third-party evaluators in July 2026, which is consensus advice rather than a settled protocol.76 The EU AI Act's obligation on providers of systemic-risk models to perform "model evaluation in accordance with standardised protocols and tools reflecting the state of the art" is a legal duty that has applied since 2 August 2025, with providers of models already on the market before that date given until 2 August 2027 to comply; it binds a class of providers, not benchmark publishers, and prescribes no leaderboard.77, 127 No document in this list settles whether a given benchmark measures what its name says.

Post-cutoff test sets

The defences against contamination predate frontier models, and the week showed each of them in use. Dated and refreshed sets: LiveBench adds and updates questions monthly and LiveCodeBench collects contest problems after a model's cutoff.82, 83 Private holdouts: SWE-Bench Pro keeps a held-out set of 12 repositories and a commercial set of 18 whose problems are not public.84 Post-cutoff internal sets: OpenAI's ExploitBench Internal Port, built from vulnerabilities disclosed after the cutoff, is the launch week's clearest example; the launch table prints 39% on it beside 100% on the historical set, and because the two sets and their limits differ the gap is a warning rather than a measured contamination penalty.5, 4 Sets written from scratch: Anthropic describes DeepSWE's 113 tasks as "written from scratch to avoid benchmark contamination", and SRE-Bench is "designed to be contamination-free" with 262 binaries from 19 privately developed programs.13, 7 Rewriting with verification: an August 2026 method rewrites mathematics benchmarks and checks the rewritten problems with formal proofs.85 Each reduces the risk of prior exposure without removing it: evaluation traffic through an API, post-training updates, privileged access and publication dates that lag the underlying information are routes a cutoff date does not close, which is why ARC Prize calls its set semi-private. Most of these also produce a smaller test set than the public one they replace, which returns the argument to the standard errors above; refreshed and generated sets can grow instead, at the cost of validating every new item.

Fresh questions are one kind of robustness test. Behavioural tests are another: minimum-functionality items, invariance checks in which an irrelevant change to the input should leave the answer alone, and directional checks in which a relevant change should move it predictably, the framework CheckList set out in 2020.125 A paraphrase can change a question's difficulty as well as its surface, so transformations need validating too, and the split should be made at the level that blocks the shortcut, by repository, document or task family rather than by row. Adversarial collection, with people writing items that current models fail, finds failures fixed sets miss; Dynabench is the established example, and its distribution should be reported separately from ordinary traffic.126

Protecting the holdout

A developer can overfit a public leaderboard without ever seeing its answers: observe a score, adjust the prompt or the model, keep what helped, repeat. After enough rounds the test set has become a development set, and the same thing happens to a benchmark designer who keeps selecting the questions that particular frontier models fail. This is the adaptive data analysis problem, distinct from literal leakage and from private-variant selection, and the work on the reusable holdout showed both the mechanism and the remedy: limit how many adaptive queries a protected set answers, and add noise or thresholds to the answers it gives.121 The practical form is a development set for iteration, a quarantined confirmation set for the claims that matter, an explicit tuning and submission budget, periodic replenishment, and a public anchor set kept across versions so that reducing leakage does not erase the ability to measure a trend. For generated questions, the generator and filter models should be disclosed and the answers checked independently; freshness establishes neither difficulty nor correctness.

Validating the grader

Problem 4 catalogued the biases of model judges. The practice that follows is to treat the grader as an instrument with its own validation. Build a reference sample with independent expert judgments and adjudicated disagreements, and seed it with known correct answers, subtle errors, partial successes, refusals, formatting variants and persuasive wrong answers. Measure false acceptances and false rejections separately, by task type; agreement between two judges is not correctness. Blind the model's identity, randomise the order, and freeze the judge model, prompt, rubric and reference material for the life of a comparison, because a judge behind an API changes under the evaluation. Audit a sample of passes as well as failures, since auditing failures alone never finds the false positives, and keep the ambiguous cases rather than forcing them into a binary label. Treat the candidate's output, any pages it retrieved and any repository it touched as untrusted input to the grader: instructions aimed at the evaluator can sit inside them, and the boundary should be tested, as AgentDojo does for agents under prompt injection.118

A test can be wrong in both directions. Case four's DeepSWE note and the SWE-bench Verified audit are failures of the first kind, tests that reject correct work. EvalPlus showed the second kind in 2023: expanding the test suites of a code benchmark exposed incorrect programs that the original tests had accepted.117 Hidden behavioural tests, independently written requirements and deliberately broken implementations that a good test must catch answer the two questions together: can the grader reject a correct solution, and can it accept an incorrect one. Anthropic's engineering guidance describes the same structure, tasks, trajectories and graders each needing their own checks, and notes that final-state grading must look at the artefact, not at the agent's message saying it succeeded.119 For agents, the final outcome is not the whole of success. A system can reach the requested final state by an unacceptable path, editing unrelated files, sending a message it was not authorised to send, exposing private data or discarding work, and a check of the final state alone misses a violation that was temporary. The grader should hold hard constraints on the trajectory as well as the result, credit recovery from tool failures and an honest stop when a task cannot be done, and separate model errors from environment outages by a rule fixed before the run rather than by removing the inconvenient failures afterwards.

Knowing when not to answer

The hallucination-rate episode in the earlier section is a case of a metric that only makes sense beside two others. An accuracy-only benchmark rewards a confident guess over an honest abstention, and an abstention-friendly metric rewards saying nothing unless coverage is shown alongside it. The joint picture is a curve of risk against coverage: as the system declines more questions, how often are the answers it still gives wrong. Where a system reports probabilities, a proper scoring rule such as the Brier score or log loss, inspected by slice, tests whether they mean anything; a verbal "I am fairly confident" is not a calibrated probability. The cost of a wrong answer, of a delay and of an escalation differ by task, so no fixed penalty for abstaining is right everywhere, and for genuinely open questions the reference may need several acceptable answers rather than one string.

Cost and reliability

"AI Agents That Matter" argued in 2024 that agent comparisons must be cost-controlled, since "for substantially similar accuracy, the cost can differ by almost two orders of magnitude", and that many agent benchmarks "have inadequate holdout sets, and sometimes none at all".53 The Holistic Agent Leaderboard implemented that, and its 21,730 rollouts produced the finding that "increased reasoning effort produces equal or lower accuracy" in 21 of 36 model-scaffold-benchmark combinations.54 tau-bench's pass^k metric, "the chance that all k i.i.d. task trials are successful", showed an agent with over 60% single-trial success falling under 25% at pass^8; it measures repeatability under that trial design, and a deployment without a verifier that can pick the successful attempt should be read against it rather than against pass@k.55 The reporting practices that follow are simple: print the run cost next to the score, as ARC Prize and LiveBench do; print pass^k next to pass@1; and show the score-against-cost frontier rather than a single column.1, 15 The denominator should match the claim: cost per successful task, with failed attempts, retries, tool charges and fallback routing included, is the cost of delivered work; a failed task that cost little still delivered nothing. Latency belongs beside it, as a distribution rather than a mean, because a product that usually answers quickly and sometimes misses its deadline is a different product from one that never does.

The human and the model together

A model's score alone and the productivity of a person using the model are different quantities, and the second has to be measured directly. METR's 2025 study is the concrete example: 16 experienced open-source developers worked on 246 real issues in repositories they had contributed to for years, each issue randomly assigned to allow or forbid AI tools, and with the early-2025 tools allowed they took 19% longer, after forecasting that the tools would make them 24% faster.120 The result says nothing about the September 2026 tools, other developers or other tasks; what it shows is that the effect of a tool on work cannot be read off a benchmark, and that people's own estimates of it are not a substitute. A study of this kind compares people with and without the tool, counts the whole time including prompting and checking, has reviewers blind to condition assess the output, and reports human-only, model-only where meaningful and assisted results separately, with correction time and defect escape alongside speed.

Rare failures and their cost

An average success rate hides the shape of the failures. A model with fewer trivial mistakes and more expensive ones may be the worse choice, so severity, the population affected and conditional rates belong in the report with the mean. A rate of zero or one hundred per cent needs its denominator: zero failures in 100 independent, representative trials leaves a one-sided 95% upper bound on the failure rate of about 3%, and 20 successes out of 20 leaves a two-sided 95% lower bound on the success rate of about 83%; these are illustrations of the arithmetic, not intervals for any benchmark on this page. Exclusions can be outcome-dependent: in METR's Sol analysis, dropping the runs that cheated moved the time-horizon estimate to about 71 hours with an interval too wide to use, because the informative long tasks were the ones discarded, which is why failure handling should be fixed before the run and its sensitivity reported.23 Maximum elicited capability, the tendency to act under realistic conditions and the effectiveness of the controls around the model are three separate measurements; adversarial searches find failure modes without estimating how common they are, and sampled ordinary traffic estimates the common outcomes while missing the rare adversarial ones, so a report needs both.

Validity against real work

The instrument section asked whether a score predicts the work. The study that answers it links benchmark results to held-out outcomes on the intended workload: define the tasks, the users, the acceptable quality, the resource limits and the consequences of failure, then compare how well alternative benchmarks predict those outcomes and where each relationship breaks down. A global average can conceal a language, a domain or a user group where the model fails, so the population should be specified and the important slices, chosen in advance from the intended use and the plausible failure mechanisms, reported with their sample sizes and uncertainty; searching hundreds of small subgroups and publishing the striking ones is the selection problem of Problem 5 in a new form. HELM's scenarios and Model Cards' disaggregated reporting are the established precedents for reporting beyond one aggregate.116, 122 For ranking methods, the diagnostics are the same in spirit: a Bradley-Terry fit should be checked for residuals, opponent coverage and hidden preference cycles, and a composite should publish its constituent scores and how the leader changes as the weights move; if plausible weights reverse the order, that is itself a finding worth reporting.

Practitioners' advice

Two practitioners state the reader's position well. Simon Willison, on the February 2026 SWE-bench refresh: "This benchmark uses the same system prompt for every model, which is important for a fair comparison but does mean that the quality of the different harnesses or optimized prompts is not being measured here", and earlier, "There are leaderboards, but I've been losing some trust in those recently. Everyone needs their own benchmark."86 Andrej Karpathy, in his review of 2025: "benchmarks are almost by construction verifiable environments and are therefore immediately susceptible to and weaker forms of it via synthetic data generation", and "Training on the test set is a new art form."87 The practical consequence is the checklist in the next section: treat a published number as a claim about an instrument, and ask for the instrument before accepting the claim.

A minimal reporting standard

Name the harness and the effort setting. State the number of items, the number of trials per item and the pass criterion (pass@1, pass@k or pass^k). Print a standard error or interval computed with clustering where items are grouped, and compare models on paired differences. State the task release and the index version. State whether the tested configuration is the shipped one, including safeguards and any routing to other models. State the training cutoff relative to the item dates and any contamination check. Print the cost of the run and the cost per successful task. Disclose funding, access and publication rights. Publish every variant tested. Each item is already done by at least one of the organisations named on this page. None of them is done by all.

A procedure, for anyone running a comparison

  1. Define the decision. The workload, the users, the acceptable quality, the serious failures, the budget, and the smallest improvement that would change the decision.
  2. Define the system. A shared model interface, each provider's complete product, or the actual deployment, with its routing, tools, safeguards and human steps.
  3. Build two sets. A development set for iteration and a confirmation set that no tuning touches, split at the unit that blocks the shortcut, with rare and adversarial cases as separate strata.
  4. Validate the tasks and the grader. Ambiguous requirements, false accepts, false rejects, injection into the grader; expert adjudication and executable checks where they fit.
  5. Fix the analysis in advance. Budgets, repetitions, primary outcomes, exclusions, failure handling, the statistical unit, the paired comparison, and the precision the decision needs.
  6. Run baselines. A simple system, the products in use, and a human or assisted baseline when the claim is about productivity, on matched tasks.
  7. Report the structure of the results. Paired differences, slices, grader uncertainty, severe failures, sensitivity to disputed labels; coverage and reliability, not only the maximum.
  8. Count the whole cost. Failed attempts, fallbacks, tools, review, repair and latency, as a budget-constrained frontier.
  9. Confirm on protected tasks. After choosing, check on data that did not guide the choice; a shadow deployment where the setting allows it.
  10. Publish the record and a trigger. Configuration manifest, per-task outcomes, versions, limitations and corrections, and a rule for when the endpoint, harness, grader or task mix changes enough to re-run.

09  /  A reading procedure

Ten questions for a benchmark table

Tick the items a table discloses. The meter counts disclosure, not validity: a fully disclosed test can still measure the wrong thing, and an undisclosed item leaves a question open rather than answering it against the model.

0 of 10: nothing disclosed yet

The two launch tables, scored

The matrix applies the ten questions to the launch materials of the week, counting a launch post and its system card together. A question counts as disclosed when the documents state it for the headline rows the essay examined, as partly disclosed when they state it for some rows or as a general rule without the per-cell value, and as not disclosed otherwise. A different rule, such as the post alone or every row rather than the headline rows, gives different counts, which is why the matrix is printed and not only the totals.

QuestionOpenAI, GPT-6 AstraAnthropic, Claude Fable 5.1
Who ran ityesyes
Harness and settingspartly: named for ARC-AGI-3 and FrontierCodepartly: named for the Terminal-Bench rows
Reasoning effort, best of severalpartly: the rule is stated, the setting per cell is notpartly: stated for some rows in the card
Pass criterionpartly: footnotes for two rowsyes, in the card
Items and trialsnoyes, in the card
Interval or standard errornoyes, in the card
Subset, release and versionpartly: OSWorld subset, FrontierMath tierpartly: OSWorld release, strict and partial scores
Tested system is the shipped oneyes: footnotes 2, 16 and the Mythos noteyes: safeguards and routing footnote
Contaminationyes, in the card and the post-cutoff rowpartly: "cannot be ruled out" for one set
Funding and accessnono
Total3 yes, 4 partly, 3 no5 yes, 4 partly, 1 no

Counting a partial disclosure as half, that is five of the ten for OpenAI's materials and seven for Anthropic's, and the items missing most often are the interval, the trial count and the effort setting. The independent boards disclose more and are read less. A reader who insists on the ten will find that most published comparisons between frontier models are reports of a number under conditions that are partly unknown, will read them as evidence about capability only under those conditions, and will adjust their confidence accordingly.

10  /  Coda

Conclusion

They are measuring something. The question is what they measure, and under which conditions.

The Arena measures the preference of its voters, under one style adjustment, with the intervals it prints. ARC-AGI-3 measures a model together with a harness, and now says which. Artificial Analysis averages nine evaluations it runs itself; Epoch AI fits one scale to results from its own runs, independent boards and lab reports; the two disagree for reasons their methods make plausible and do not decompose. A launch table measures the best of several efforts under the publishing lab's own conditions, with the conditions of the comparison models in the footnotes. Each of these is evidence about capability under its own conditions. None is a measurement of capability in the sense the word carries in a headline: a stable property of the model, independent of who asked, how, with what tools and how many times. And none of them, however fully disclosed, answers the question a reader usually has, which is whether the model will do their work better; that takes a validation against the work itself, as the good-practice section sets out.

The launch week of GPT-6 Astra is a useful case not because the reporting was unusually poor but because it was unusually complete. OpenAI published its footnotes and its system card's contamination note. Anthropic published its trial counts and standard errors. ARC Prize published both harness conditions with costs. Artificial Analysis and Epoch AI published methods that explain their disagreement. Almost every problem described here was disclosed by the organisation it concerns. The problem is the distance between the disclosure and the headline, and it is the headline that gets repeated. The eleven problems, the four cases and the checklist are one attempt to close that distance for a reader who wants to know what a score is before deciding what it means.

Method note

Every number on this page is one of four kinds, and the text tries to say which: a figure reported by the source it is attributed to, opened on 6 September 2026 and listed below with its access date; a calculation from reported figures (the same-effort harness gaps, the expected maxima, the interval widths); a simulation, of which there are three, labelled as such, illustrating published mechanisms with invented data; or an interpretation, which is the essay's own. Where a source is a secondary report (a newspaper, a commentator) it is cited as such. Claims that could not be traced to a primary page were left out, including several figures that circulated in coverage of the launch. Leaderboard values are those displayed on the access date and will change; the essay keeps the values, not snapshots of the pages. Two independent checks re-opened each source and compared each quoted phrase before publication. The essay covers text and agent benchmarks; retrieval, long-context, multimodal and long-project evaluations raise further distinctions that it does not treat.

Revision of 6 September 2026. An independent review of the published text led to these changes: the thesis now distinguishes conditional evidence about capability from a universal ranking; the 100% and 39% cyber figures are presented as two different sets rather than a measured fall; interval-overlap reasoning was replaced by paired comparison, and the calculator's figure is labelled a scale marker rather than a floor; the AA-Omniscience hallucination rate is defined; Anthropic's safeguard routing to Opus models is described; the Epoch Capabilities Index is described as a fit over mixed sources rather than a common-protocol rerun, and the "rows under one protocol" count was recomputed; the harness comparison is given at matched effort; the saturation, judge, evaluation-awareness and pass^k claims were narrowed; the Stanford quotation, the seven-hour gap between two Arena posts, the MMLU paper's version, the o3 training partition, the HAL cost pair and the source count were corrected; the checklist meter is labelled as disclosure coverage and an audit matrix replaces an unsupported range; and sections on predictive validity, grader validation, human-and-model measurement, rare failures, protected holdouts and two newer Arena methods were added, with fourteen new sources. Five arXiv identifiers the reviewer could not retrieve were re-opened and resolve to the cited titles. Corrections are welcome by email to e1506804@u.nus.edu.

Sources

127 primary and secondary sources, each with the date it was opened; sources 114 to 127 were added in the revision of 6 September. The list is collapsed; a numbered superscript anywhere on the page opens it at the entry.

Show the source list
  1. ARC Prize Foundation, G. Kamradt, "OpenAI's GPT-6 Astra on ARC-AGI-3", 3 September 2026. arcprize.org/blog/astra accessed 6 Sep 2026
  2. ARC Prize Foundation, ARC-AGI leaderboard. arcprize.org/leaderboard accessed 6 Sep 2026
  3. ARC Prize Foundation, testing and verification policy. arcprize.org/policy accessed 6 Sep 2026
  4. OpenAI, "GPT-6 Astra: A new generation of intelligence", 3 September 2026. openai.com/index/gpt-6-astra accessed 6 Sep 2026
  5. OpenAI, "Path to Astra: critical capabilities and frontier safeguards", 1 September 2026. openai.com/index/path-to-astra accessed 6 Sep 2026
  6. OpenAI, I. Bigio and T. Sanders, "How two settings tripled our ARC-AGI-3 scores", 29 July 2026. openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores accessed 6 Sep 2026
  7. OpenAI, GPT-6 Astra System Card, September 2026. deploymentsafety.openai.com/gpt-6-astra accessed 6 Sep 2026
  8. Artificial Analysis, "Benchmarking GPT-6 Astra", 3 September 2026, and the Intelligence Index v4.2 announcement of 4 September 2026. artificialanalysis.ai/articles/benchmarking-gpt-6-astra accessed 6 Sep 2026
  9. Artificial Analysis, model comparison and release pages for GPT-6 Astra. artificialanalysis.ai/models/comparisons/gpt-6-astra-vs-gpt-5-6-sol; artificialanalysis.ai/models/releases/gpt-6-astra accessed 6 Sep 2026
  10. Epoch AI, Epoch Capabilities Index: scores file, index page and GPT-6 Astra model page. epoch.ai/data/eci_scores.csv; epoch.ai/eci; epoch.ai/models/gpt-6-astra accessed 6 Sep 2026
  11. Epoch AI, FrontierMath Tier 4 (v2) benchmark page and leaderboard, including the conflict-of-interest statement and the 12 June 2026 changelog. epoch.ai/benchmarks/frontiermath-tier-4-v2 accessed 6 Sep 2026
  12. Anthropic, "Claude Fable 5.1 and Claude Mythos 5.1", September 2026, including the benchmark table and footnotes. anthropic.com/claude-fable-and-mythos-5-1 accessed 6 Sep 2026
  13. Anthropic, "System Card: Claude Fable 5.1 and Claude Mythos 5.1", 1 September 2026 (PDF), sections 8.3, 8.6, 8.7 and 8.9. anthropic.com system card PDF accessed 6 Sep 2026
  14. Terminal-Bench-Science 0.1 leaderboard and announcement. terminal-bench-science.ai; tbench.ai/news/terminal-bench-science-0-1 accessed 6 Sep 2026
  15. LiveBench leaderboard. livebench.ai accessed 6 Sep 2026
  16. Vals AI, Vals Index, 4 September 2026 update. vals.ai accessed 6 Sep 2026
  17. Arena, Text leaderboard, Overall category (page data stamp 2 September 2026). arena.ai/leaderboard/text accessed 6 Sep 2026
  18. Arena, leaderboard changelog, entries of 14 July 2025, 23 July 2025, 17 and 18 September 2025 and 12 May 2026. arena.ai/company/leaderboard-changelog accessed 6 Sep 2026
  19. Arena, leaderboard policy, published 2 March 2024, last updated 1 September 2026. arena.ai/blog/policy accessed 6 Sep 2026
  20. LMArena, "Response to 'The Leaderboard Illusion' Writeup", 9 May 2025. arena.ai/blog/our-response accessed 6 Sep 2026
  21. T. Li, A. Angelopoulos and W.-L. Chiang, "Does style matter? Disentangling style and substance in Chatbot Arena", 29 August 2024. arena.ai/blog/style-control accessed 6 Sep 2026
  22. S. Singh, Y. Nan, A. Wang, D. D'Souza, S. Kapoor, A. Üstün, S. Koyejo, Y. Deng, S. Longpre, N. A. Smith, B. Ermis, M. Fadaee and S. Hooker, "The Leaderboard Illusion", arXiv:2504.20879, April 2025. arxiv.org/abs/2504.20879 accessed 6 Sep 2026
  23. METR, "GPT-5.6 Sol" pre-deployment evaluation, 26 June 2026, and the risk-assessment index. metr.org/blog/2026-06-26-gpt-5-6-sol accessed 6 Sep 2026
  24. E. Forlini, Fortune, "OpenAI launches GPT-6 Astra, its most powerful model yet, and touts its ability to use your computer", 3 September 2026, with the correction of 3 September and the clarification of 4 September. fortune.com accessed 6 Sep 2026
  25. F. Lardinois, The New Stack, "OpenAI launches GPT-6 Astra and says welcome to the 'AGI era'", 3 September 2026. thenewstack.io/openai-gpt6-astra-benchmarks accessed 6 Sep 2026
  26. A. Caswell, The New Stack, "GPT-6 Astra's score of 98.6% looked like AGI. Then researchers read the fine print.", 3 September 2026, with editor's note. thenewstack.io/astra-arc-agi-benchmark accessed 6 Sep 2026
  27. TechCrunch, "OpenAI launches Astra, its powerful (and controversial) new model", 3 September 2026. techcrunch.com accessed 6 Sep 2026
  28. M. Schreiner, The Decoder, "Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet's AGI forecast forward", 4 September 2026. the-decoder.com accessed 6 Sep 2026
  29. L. Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", arXiv:2306.05685, 2023. arxiv.org/abs/2306.05685 accessed 6 Sep 2026
  30. X. Zheng et al., "Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates", arXiv:2410.07137, 2024. arxiv.org/abs/2410.07137 accessed 6 Sep 2026
  31. Y. Dubois, B. Galambosi, P. Liang and T. Hashimoto, "Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators", arXiv:2404.04475, 2024. arxiv.org/abs/2404.04475 accessed 6 Sep 2026
  32. A. Panickssery, S. Bowman and S. Feng, "LLM Evaluators Recognize and Favor Their Own Generations", arXiv:2404.13076, 2024. arxiv.org/abs/2404.13076 accessed 6 Sep 2026
  33. Z. Yang, Y. Hou and X. Yang, arXiv:2607.08535, 2026 (judge scaling and bias). arxiv.org/abs/2607.08535 accessed 6 Sep 2026
  34. Y. Li, arXiv:2606.15474, 2026 (attributing drift to the judge or the system). arxiv.org/abs/2606.15474 accessed 6 Sep 2026
  35. A. P. Gema et al., "Are We Done with MMLU?", arXiv:2406.04127, 2024. arxiv.org/abs/2406.04127 accessed 6 Sep 2026
  36. FutureHouse, analysis of Humanity's Last Exam chemistry and biology answers, 23 July 2025. futurehouse.org/research/hle-exam accessed 6 Sep 2026
  37. L. Phan et al., "Humanity's Last Exam", arXiv:2501.14249 (revised version with the peer-review disagreement rates). arxiv.org/abs/2501.14249 accessed 6 Sep 2026
  38. OpenAI, "Introducing SWE-bench Verified", 13 August 2024. openai.com/index/introducing-swe-bench-verified accessed 6 Sep 2026
  39. S. Liang, S. Garg and R. Zilouchian Moghaddam, "The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason", arXiv:2506.12286, 2025. arxiv.org/abs/2506.12286 accessed 6 Sep 2026
  40. H. Zhang et al., "A Careful Examination of Large Language Model Performance on Grade School Arithmetic", arXiv:2405.00332, 2024. arxiv.org/abs/2405.00332 accessed 6 Sep 2026
  41. Z. Han, M. Mankikar, J. Michael and Z. Wang, "Search-Time Data Contamination", arXiv:2508.13180, 2025. arxiv.org/abs/2508.13180 accessed 6 Sep 2026
  42. Epoch AI, "RIP Classic Reasoning Benchmarks. What's Next?", Gradient Updates, 5 May 2026; A. Ho and G. Burnham, "Are AI benchmarks doomed?", 1 May 2026; Epoch AI, data insight on interpreting the Epoch Capabilities Index, 6 November 2025. epoch.ai/gradient-updates/rip-classic-benchmarks; epochai.substack.com; epoch.ai/data-insights/interpreting-eci accessed 6 Sep 2026
  43. ARC Prize Foundation, ARC-AGI-1 overview page and "Announcing ARC-AGI-3", 25 March 2026. arcprize.org/arc-agi/1; arcprize.org/blog/arc-agi-3-launch accessed 6 Sep 2026
  44. M. Saeed and S. Razniewski, "LLMPEDIA", arXiv:2609.01182, 1 September 2026. arxiv.org/abs/2609.01182 accessed 6 Sep 2026
  45. A. M. Bean, R. O. Kearns, A. Romanou et al., "Measuring what Matters: Construct Validity in Large Language Model Benchmarks", arXiv:2511.04703, 2025. arxiv.org/abs/2511.04703 accessed 6 Sep 2026
  46. A. Reuel et al., "BetterBench", arXiv:2411.12990, NeurIPS 2024. arxiv.org/abs/2411.12990 accessed 6 Sep 2026
  47. E. Miller, "Adding Error Bars to Evals", arXiv:2411.00640, 2024. arxiv.org/abs/2411.00640 accessed 6 Sep 2026
  48. J. Su, J. Zhang, K. Ullrich, L. Bottou and M. Ibrahim, "A Single Character can Make or Break Your LLM Evals", arXiv:2510.05152, 2025. arxiv.org/abs/2510.05152 accessed 6 Sep 2026
  49. S. Biderman et al., "Lessons from the Trenches on Reproducible Evaluation of Language Models", arXiv:2405.14782, 2024. arxiv.org/abs/2405.14782 accessed 6 Sep 2026
  50. J. Needham, G. Edkins, G. Pimpale, H. Bartsch and M. Hobbhahn, "Large Language Models Often Know When They Are Being Evaluated", arXiv:2505.23836, 2025. arxiv.org/abs/2505.23836 accessed 6 Sep 2026
  51. T. van der Weij et al., "AI Sandbagging: Language Models can Strategically Underperform on Evaluations", arXiv:2406.07358, 2024. arxiv.org/abs/2406.07358 accessed 6 Sep 2026
  52. Anthropic, "Claude Sonnet 5 System Card", 30 June 2026 (PDF). anthropic.com system card PDF accessed 6 Sep 2026
  53. S. Kapoor, B. Stroebl, Z. Siegel, N. Nadgir and A. Narayanan, "AI Agents That Matter", arXiv:2407.01502, 2024. arxiv.org/abs/2407.01502 accessed 6 Sep 2026
  54. S. Kapoor et al., "Holistic Agent Leaderboard", arXiv:2510.11977, 2025, and the HAL site. arxiv.org/abs/2510.11977; hal.cs.princeton.edu accessed 6 Sep 2026
  55. S. Yao, N. Shinn, P. Razavi and K. Narasimhan, "tau-bench", arXiv:2406.12045, 2024. arxiv.org/abs/2406.12045 accessed 6 Sep 2026
  56. Zheng et al., arXiv:2606.12344, 2026 (harness and adapter effects on coding agents). arxiv.org/abs/2606.12344 accessed 6 Sep 2026
  57. Harbor Framework, Terminal-Bench 2.1 repository, 5 May 2026. github.com/harbor-framework/terminal-bench-2-1 accessed 6 Sep 2026
  58. Y. Huang et al., "Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards", arXiv:2501.07493, 2025. arxiv.org/abs/2501.07493 accessed 6 Sep 2026
  59. H. Oyarhoseini, J. Lin and A.-H. Karimi, "A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation", arXiv:2605.15761, 15 May 2026. arxiv.org/abs/2605.15761 accessed 6 Sep 2026
  60. Meta, "The Llama 4 herd", 5 April 2025; Arena and LMArena posts on X of 5 and 7 April 2025. ai.meta.com/blog/llama-4-multimodal-intelligence; x.com/lmarena_ai accessed 6 Sep 2026
  61. TechCrunch, "LM Arena... lands $100M", 21 May 2025; "LMArena lands $1.7B valuation", 6 January 2026; "Arena... is now a $100M business", 29 June 2026. techcrunch.com, May 2025; techcrunch.com, January 2026; techcrunch.com, June 2026 accessed 6 Sep 2026
  62. Epoch AI on X, GPT-6 Astra results thread, 3 September 2026. x.com/EpochAIResearch accessed 6 Sep 2026
  63. Arena, "Introducing AutoEval to the Arena leaderboards", 30 July 2026. arena.ai/blog/autoeval-scores accessed 6 Sep 2026
  64. XLANG Lab, OSWorld 2.0 leaderboard and "OSWorld-Verified", 28 July 2025. osworld-v2.xlang.ai; xlang.ai/blog/osworld-verified accessed 6 Sep 2026
  65. M. Mancoridis, B. Weeks, K. Vafa and S. Mullainathan, "Potemkin Understanding in Large Language Models", arXiv:2506.21521, 2025. arxiv.org/abs/2506.21521 accessed 6 Sep 2026
  66. M. Song and C. Park, "Auditing MCQA Benchmarks through Probability Landscapes", arXiv:2608.30372, 31 August 2026. arxiv.org/abs/2608.30372 accessed 6 Sep 2026
  67. Stanford AI Measurement Science lab, CS321M course page. aimslab.stanford.edu/cs321m accessed 6 Sep 2026
  68. A. Ho et al., "A Rosetta Stone for AI benchmarks", Epoch AI, 2 December 2025, and arXiv:2512.00193. epoch.ai/publications/a-rosetta-stone-for-ai-benchmarks accessed 6 Sep 2026
  69. Z. Mowshowitz, "Claude Mythos 5.1 and Fable 5.1: capabilities", 5 September 2026. thezvi.substack.com accessed 6 Sep 2026
  70. Arena on X, "GPT-6 Astra by OpenAI is now in the Arena", 4 September 2026. x.com/arena accessed 6 Sep 2026
  71. SWE-bench official leaderboard (Verified, Bash Only). swebench.com accessed 6 Sep 2026
  72. Scale Labs leaderboards and the Humanity's Last Exam public site. labs.scale.com/leaderboard; lastexam.ai accessed 6 Sep 2026
  73. J. Dekoninck, M. N. Müller, M. Baader, M. Fischer and M. Vechev, "Evading Data Contamination Detection for Language Models is (too) Easy", arXiv:2402.02823, 2024. arxiv.org/abs/2402.02823 accessed 6 Sep 2026
  74. S. Bowyer, L. Aitchison and D. R. Ivanova, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints", arXiv:2503.01747, 2025. arxiv.org/abs/2503.01747 accessed 6 Sep 2026
  75. NIST, "NIST AI 800-2: Practices for Automated Benchmark Evaluations of Language Models", initial public draft, January 2026. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf accessed 6 Sep 2026
  76. UK AI Security Institute, "International evaluation best practice and open questions in AI measurement", 23 July 2026. aisi.gov.uk accessed 6 Sep 2026
  77. Regulation (EU) 2024/1689 (AI Act), Article 55(1). artificialintelligenceact.eu/article/55 accessed 6 Sep 2026
  78. Artificial Analysis, intelligence benchmarking methodology (v4.2). artificialanalysis.ai/methodology/intelligence-benchmarking accessed 6 Sep 2026
  79. Latent Space, interview with Artificial Analysis, 8 January 2026 (mystery-shopper policy). latent.space/p/artificialanalysis accessed 6 Sep 2026
  80. Artificial Analysis, Endpoint Accuracy Index methodology. artificialanalysis.ai/methodology/endpoint-accuracy-index accessed 6 Sep 2026
  81. Epoch AI, "Why benchmarking is hard", Gradient Updates, 23 December 2025. epoch.ai/gradient-updates/why-benchmarking-is-hard accessed 6 Sep 2026
  82. C. White et al., "LiveBench: A Challenging, Contamination-Limited LLM Benchmark", and the LiveBench site. livebench.ai accessed 6 Sep 2026
  83. N. Jain et al., "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code", and the LiveCodeBench site. livecodebench.github.io accessed 6 Sep 2026
  84. Scale AI, SWE-Bench Pro paper and private-dataset leaderboard, September 2025. labs.scale.com/leaderboard/swe_bench_pro_private accessed 6 Sep 2026
  85. X. Zhou et al., "RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving", arXiv:2609.00062, 30 August 2026. arxiv.org/abs/2609.00062 accessed 6 Sep 2026
  86. S. Willison, "The last six months in LLMs, illustrated by pelicans on bicycles", 6 June 2025, and post on the SWE-bench leaderboard refresh, 19 February 2026. simonwillison.net, June 2025; simonwillison.net, February 2026 accessed 6 Sep 2026
  87. A. Karpathy, "2025 LLM Year in Review", 19 December 2025. karpathy.bearblog.dev/year-in-review-2025 accessed 6 Sep 2026
  88. ARC Prize Foundation, "OpenAI o3 Breakthrough High Score on ARC-AGI-Pub", 20 December 2024 (with notes of 24 March, 16 April and 10 December 2025), and "Analyzing o3 and o4-mini with ARC-AGI", 22 April 2025. arcprize.org/blog/oai-o3-pub-breakthrough; arcprize.org/blog/analyzing-o3-with-arc-agi accessed 6 Sep 2026
  89. Epoch AI, "OpenAI and FrontierMath", 23 January 2025. epoch.ai/latest/openai-and-frontiermath accessed 6 Sep 2026
  90. TechCrunch, "AI benchmarking organization criticized for waiting to disclose funding from OpenAI", 19 January 2025. techcrunch.com accessed 6 Sep 2026
  91. TechCrunch, "Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark", 11 April 2025. techcrunch.com accessed 6 Sep 2026
  92. LMArena statement of 8 April 2025 on X, as reproduced by S. Willison the same day. simonwillison.net/2025/Apr/8/lmaren accessed 6 Sep 2026
  93. xAI, "Grok 4", 9 July 2025, including the page's embedded chart data. x.ai/news/grok-4 accessed 6 Sep 2026
  94. Google DeepMind, Gemini 2.5 Deep Think Model Card, 1 August 2025 (PDF). deepmind-media Model Card PDF accessed 6 Sep 2026
  95. Google DeepMind, "Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad", 21 July 2025. deepmind.google/blog accessed 6 Sep 2026
  96. Futurism, "GPT-5 Launch Demo Plagued With Catastrophically Dumb Errors", 8 August 2025. futurism.com/gpt-5-demo-dumb-errors accessed 6 Sep 2026
  97. OpenAI, GPT-5 System Card, 13 August 2025 (PDF). cdn.openai.com/gpt-5-system-card.pdf accessed 6 Sep 2026
  98. TechCrunch, "Sam Altman addresses 'bumpy' GPT-5 rollout, bringing 4o back, and the 'chart crime'", 8 August 2025; S. Altman on X, 7 August 2025. techcrunch.com; x.com/sama accessed 6 Sep 2026
  99. OpenAI, "Why SWE-bench Verified no longer measures frontier coding capabilities", 23 February 2026. openai.com/index/why-we-no-longer-evaluate-swe-bench-verified accessed 6 Sep 2026
  100. METR on X, Claude Opus 4.6 time-horizon estimate, February 2026, and the METR time-horizons page. x.com/METR_Evals; metr.org/time-horizons accessed 6 Sep 2026
  101. MIT Technology Review, "This is the most misunderstood graph in AI", 5 February 2026. technologyreview.com accessed 6 Sep 2026
  102. Artificial Analysis, "OpenAI GPT-5.5 is the new leading AI model", 23 April 2026. artificialanalysis.ai accessed 6 Sep 2026
  103. The Decoder, "GPT-5.5 tops benchmarks but still hallucinates frequently at a 20 percent higher API cost", 25 April 2026. the-decoder.com accessed 6 Sep 2026
  104. Google, "Gemini 3.5", 19 May 2026. blog.google accessed 6 Sep 2026
  105. Artificial Analysis, "Gemini 3.5 Flash: everything you need to know", 19 May 2026. artificialanalysis.ai accessed 6 Sep 2026
  106. NVIDIA, "NVIDIA AVO reaches 100% on ARC-AGI-3", 21 August 2026. developer.nvidia.com/blog accessed 6 Sep 2026
  107. TechCrunch, "Nvidia just showed that the harness, not the AI model, is now the real hero", 21 August 2026. techcrunch.com accessed 6 Sep 2026
  108. DeepSeek, DeepSeek-R1-0528 model card, 28 May 2025. huggingface.co/deepseek-ai/DeepSeek-R1-0528 accessed 6 Sep 2026
  109. Moonshot AI, Kimi K2 model card, 11 July 2025. github.com/MoonshotAI/Kimi-K2 accessed 6 Sep 2026
  110. Artificial Analysis model pages for DeepSeek R1 0528 and Kimi K2. artificialanalysis.ai/models/deepseek-r1; artificialanalysis.ai/models/kimi-k2 accessed 6 Sep 2026
  111. Anthropic, Claude Opus 5 System Card, 24 July 2026 (PDF), sections 8.1 and 8.2. anthropic.com system card PDF accessed 6 Sep 2026
  112. F. Chollet on X, 3 September 2026, 19:42 UTC. x.com/fchollet/status/2095598451115614371 accessed 6 Sep 2026 via X's public embed data
  113. ARC Prize on X, 3 September 2026, 19:39 UTC, thread of three posts. x.com/arcprize/status/2095597602545025138 accessed 6 Sep 2026 via X's public embed data
  114. Artificial Analysis, "AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models", arXiv:2511.13029, 2025, section 2.4.3 (definition of the hallucination rate). arxiv.org/abs/2511.13029 accessed 6 Sep 2026
  115. ARC Prize Foundation, "ARC-AGI-3 Scoring Methodology" (Relative Human Action Efficiency). docs.arcprize.org/methodology accessed 6 Sep 2026
  116. P. Liang et al., "Holistic Evaluation of Language Models", arXiv:2211.09110, 2022, and the Stanford CRFM announcement of 17 November 2022. arxiv.org/abs/2211.09110; crfm.stanford.edu/2022/11/17/helm.html accessed 6 Sep 2026
  117. J. Liu et al., "Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation" (EvalPlus), arXiv:2305.01210, 2023. arxiv.org/abs/2305.01210 accessed 6 Sep 2026
  118. E. Debenedetti et al., "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents", arXiv:2406.13352, 2024. arxiv.org/abs/2406.13352 accessed 6 Sep 2026
  119. Anthropic, "Demystifying evals for AI agents", engineering blog. anthropic.com/engineering/demystifying-evals-for-ai-agents accessed 6 Sep 2026
  120. METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", 10 July 2025. metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study accessed 6 Sep 2026
  121. C. Dwork, V. Feldman, M. Hardt, T. Pitassi, O. Reingold and A. Roth, "Generalization in Adaptive Data Analysis and Holdout Reuse", arXiv:1506.02629, 2015. arxiv.org/abs/1506.02629 accessed 6 Sep 2026
  122. M. Mitchell et al., "Model Cards for Model Reporting", FAT* 2019, arXiv:1810.03993. arxiv.org/abs/1810.03993 accessed 6 Sep 2026
  123. Arena, "Factuality in the Arena", 14 July 2026, updated 31 July 2026. arena.ai/blog/factuality-in-arena accessed 6 Sep 2026
  124. Arena, "Agent Arena: Causal Evaluation of Agents in the Real World", 4 June 2026. arena.ai/blog/agent-arena-methodology accessed 6 Sep 2026
  125. M. T. Ribeiro, T. Wu, C. Guestrin and S. Singh, "Beyond Accuracy: Behavioral Testing of NLP Models with CheckList", ACL 2020. aclanthology.org/2020.acl-main.442 accessed 6 Sep 2026
  126. D. Kiela et al., "Dynabench: Rethinking Benchmarking in NLP", NAACL 2021. aclanthology.org/2021.naacl-main.324 accessed 6 Sep 2026
  127. Regulation (EU) 2024/1689 (AI Act), Article 111(3), transitional provision for general-purpose AI models placed on the market before 2 August 2025. artificialintelligenceact.eu/article/111 accessed 6 Sep 2026