We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
Essay · 6 September 2026 · about 70 minutes
A critique of the public evaluation metrics for frontier AI models. Arena rankings, benchmark scores and lab-reported numbers get read as measurements of capability. This essay asks what each of them measures, and under which conditions a score says something about capability. The running case is the launch week of 1 to 6 September 2026, when Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra were released two days apart. The same checks are applied to both labs, and to Google, xAI, Meta, DeepSeek and Moonshot in the earlier episodes.
Maksim Silchenko · National University of Singapore
Summary
Every number on this page is traced to the source it came from, opened on 6 September 2026, with secondary reports marked as such; the sources are collected at the end. The essay was revised the same day after an independent review; the method note lists what changed. OpenAI's launch is the running case because it happened this week and produced the most documents, not because the pattern is specific to one company.
01 / The instrument
On 3 September 2026 the same model scored 62.7% and 99.9% on the same benchmark, on the same day, reported by the same organisation. Both numbers are correct under their own conditions, and the conditions are the subject of this essay.
Evaluation is a form of measurement. Someone starts with a they care about, such as whether a model can reason about a situation it has never seen. They operationalise it as a set of items and a scoring rule. They run the system under some configuration. They report a number, ideally with its uncertainty. Every failure discussed on this page is a failure at one of those four joints: a construct nobody defined, an operationalisation that measures something else, a run that is not the deployed configuration, or a number reported without the error bar that would have shown it was noise.
The 62.7% and the 99.9% belong to on , the interactive reasoning benchmark run by the ARC Prize Foundation. Under the Foundation's own Standard , a minimal interface that is identical for every provider, the model's best run scored 62.7% at a cost of about $26,000. Under a Provider Adapter harness, which keeps OpenAI's private reasoning state between requests and compacts long conversations, it scored 99.9% for about $19,000.1 The Foundation published both, said the shared interface is what gives "a consistent, apples-to-apples comparison across providers", and announced it will now label every result with the condition it was run under.1
OpenAI had made the same point six weeks earlier, about its previous model. In July it reported that turning on two API settings it already uses in its own products, retained reasoning and compaction, tripled GPT-5.6 Sol's score on the ARC-AGI-3 public set from 13.3% to 38.3% and cut by a factor of six. The post closed with a sentence that could stand as the thesis of this essay: "evals rarely measure models in isolation". They also measure, in OpenAI's words, "a bundle of less visible choices about API settings, harness design, and prompting".6
The measurement chain, and where it breaks diagram, draws on scroll, scrolls sideways on small screens
One numerical fact matters for everything that follows. If a benchmark has n independent pass-or-fail items and the true pass rate is p, the of the measured accuracy is the square root of p(1−p)/n. For n = 100 and p = 0.70 that is 0.046, so a 95% interval spans roughly nine points either side. Two models that differ by three points on a 100-item benchmark of single-attempt tasks are not distinguished by it under a conventional test. For n = 1,000 the interval shrinks to about three points. Most agent benchmarks have a few hundred tasks or fewer, and the tasks inside them are rarely independent, which widens the interval further. The calculator in problem 7 lets you set n and p and overlay the benchmarks named in this essay.
Two cautions apply to the formula. First, the right comparison between two models scored on the same items is a paired one: what matters is how often the two disagree on an item, not whether their separate intervals overlap. Two intervals can overlap while the paired difference is clear, and a difference that fails the test is not evidence that the models are equivalent. Second, the formula is for independent pass-or-fail items. Scores that average fractional credit, Elo ratings fitted to votes, and composite indices each need their own error model, so the item counts on this page are scale markers for the size of the problem, not intervals for those benchmarks.
Six properties of language models make the four joints harder than they are in classical machine learning. Outputs are open-ended, so a score depends on a grader, and graders have properties of their own. Sampling is stochastic, so a single run is one draw from a distribution. Scores move with prompt phrasing, few-shot order and even the delimiter between examples, so a benchmark number is really a triple: the model, the prompt and the harness. Benchmarks are on the internet and models are trained on the internet, so contamination is hard to rule out from outside. A public benchmark is a target, so its validity decays under optimisation pressure, which is Goodhart's law applied to measurement. And when the grader or the simulated user is itself a model, the instrument moves with every version change. Stanford's AI Measurement Science lab describes the result as "a measurement crisis characterized by benchmark saturation, inconsistent measurement practices, and difficulty in making valid claims about AI capabilities", and a 2025 study of what it called potemkin understanding found that models define concepts correctly 94.2% of the time and then fail to apply them at high rates, which is the construct problem in miniature: a right answer on the item is not the property the item was written to measure.67, 65 The eleven problems below are the concrete forms these properties take in the public numbers of 2026.
A benchmark can be repeatable, uncontaminated and correctly graded and still be a poor guide to the decision a reader is making. The practical test of an instrument is whether a higher score predicts a better outcome on the work the reader cares about, for the people who will do it, at the cost they will pay. Agreement with another leaderboard is convergent evidence, but it is not that test. A model can score higher and still need more supervision, fail more often on one workload, or make errors that are harder to notice and repair. The Holistic Evaluation of Language Models framework made the point in 2022 by scoring models on many scenarios and on several dimensions at once, among them accuracy, calibration, robustness, fairness and efficiency, rather than on one aggregate.116 The good-practice section returns to how such a validation is run.
02 / The general problems
Seven of these are structural and have been documented since 2023. Four are newer, and the launch week of September 2026 supplied a fresh example of each.
The at arena.ai ranks models from anonymous pairwise votes fitted with a . As of 6 September 2026 the page reports 7,999,020 votes across 400 models, with a data stamp of 2 September.17 A vote records that one self-selected person preferred one of two responses to a prompt they wrote themselves. That is a real quantity, and a useful one, but it is a measurement of preference under those conditions. Preference overlaps with capability, since a voter often prefers the answer that is correct, clearer or easier to act on, but the overlap has to be shown rather than assumed, and the Arena's own corrections show where it fails.
The gap between the two is not hypothetical. In August 2024 the Arena team fitted the same votes with controls for answer length, the method it calls , and for the number of markdown headers, bold elements and lists, and reported: "We controlled for the effect of length and markdown, and indeed, the ranking changed." Two small models fell below most of the frontier and three others rose substantially.21 In May 2026 the team found two further biases in a new source of votes, "a position bias favoring Model A (the model on the left)" and "an advantage given to models that share an organization with the prior turns of context", and added terms to the fit to absorb them.18 Each correction is a statement that the raw votes had been rewarding something other than quality. The same effect appears in automatic proxies for human preference: length-controlling AlpacaEval raised its correlation with Arena rankings from 0.94 to 0.98.31
Six models, one voter population, one style knob simulation
If test items, or paraphrases of them, sit in the training corpus, part of a score measures recall; this is . Exposure inflates a result without making every correct answer pure recall, and leaked answers, related tasks and legitimate domain knowledge are different things. Three measurements show how large the effect can be. When a team commissioned GSM1k, a fresh set of grade-school arithmetic problems matched in difficulty to GSM8K, several model families lost up to 8 points, and the size of a model's loss correlated with how readily it could regenerate GSM8K problems verbatim.40 On , models identified the buggy file from the issue text alone, with no access to the repository, up to 76% of the time, against up to 53% on repositories outside the benchmark, which is suggestive of memorisation rather than proof of it.39 Agents that search the web while answering , SimpleQA and found the evaluation datasets themselves, with labels, on Hugging Face for between about 1% and 4% of questions, depending on the dataset and the agent; blocking that source cut accuracy on the affected subset by about 15%.41
The launch week added a case from inside a lab. OpenAI reported that Astra scored 100% on ExploitBench, a set of historical software vulnerabilities. Its says the result "may be artificially inflated due to potential contamination from exposure to historical vulnerabilities", and gives the mechanism: asked to exploit one , the model failed, then recalled a different, later CVE from memory and used that instead.7 OpenAI then built an internal set of 20 vulnerabilities disclosed after the model's training cutoff. On that set Astra scored 39.0% and its predecessor 5.5%, with a footnote that the 5.5% is partly an artefact of a 300-turn limit and that the older model reached 11.5% when it hit fewer limits.4 Anthropic's system card for the same week carries a similar admission about a problem set drawn from June 2026 arXiv abstracts: "some contamination cannot be ruled out".13 Detection from outside is weaker than it looks; a 2024 study showed a contamination method that "significantly inflates benchmark performance while completely evading current detection methods".73
The same skill, a historical set and a post-cutoff set animated, OpenAI's own numbers
A benchmark is a set of items written by people, and people make mistakes. A re-annotation of , posted in 2024 and revised in January 2025, estimated that 6.49% of its questions contain errors, and found errors in 57% of the questions it reviewed in the virology subset.35 For Humanity's Last Exam, FutureHouse reported that 29 ± 3.7% of the text-only chemistry and biology questions "had answers with directly conflicting evidence in peer reviewed literature"; the benchmark's authors re-reviewed and reported an expert disagreement rate of 15.4% on the public set, about 18% on a biology, chemistry and health subset, and 25% on that subset under a single-reviewer rule; these are different review populations and different definitions of error, and they should not be averaged.36, 37 SWE-bench Verified exists because 68.3% of a random sample of the original SWE-bench was flagged by 93 developers, under a deliberately conservative rule, for under-specified statements or unit tests that could reject valid solutions; with the same open-source scaffold, GPT-4o's score was 16% on the original set and 33.2% on the verified subset.38 -Verified was produced by a team of about ten people fixing more than 300 reported problems over two months.64 In February 2026 OpenAI audited the hardest SWE-bench Verified problems, found that "at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions" while scores kept "improving from 74.9% to 80.9% in the last 6 months", and stopped reporting the benchmark it had built.99
A 2026 audit of multiple-choice sets made the dependence on item quality explicit: injecting noise into the distractors "raises accuracy to nearly 100%, confirming that most errors disappear when meaningful distractor competition is removed", an intervention that also changes the difficulty of the task.66 When strong models clear most of the items a benchmark stops separating them. That is saturation, and it now arrives within a few years of a benchmark's release. , which tracks benchmarks across generations, states that "Most benchmarks saturate too quickly to study long-run AI trends", that "benchmarks tend to saturate within 1-3 years" and in May 2026 described GPQA as "clearly saturated", dating the event to the winter of 2025, about two years after release.68, 42 A September 2026 paper opens with flagship models "scoring above 90%" on MMLU.44 The ARC Prize Foundation's own history is the clearest example: ARC-AGI-1 stayed unsolved from 2019 until late 2024 "despite a 50,000x scaleup of base LLM pretraining", fell to o3-preview at 75% and then 87% in December 2024, and its successor ARC-AGI-3 went from a frontier score of 0.51% at launch on 25 March 2026 to 62.7% and 99.9% under the two harnesses by 3 September.43, 1 Near its ceiling a benchmark's remaining gaps are decided by a handful of items, and whether those are the last hard items or the disputed ones is an item-level question. A move from 98% to 99% halves the error rate and may be real; it may also be two contested labels. The way to find out is to audit the disputed items and recompute the paired difference under plausible labels; the headline alone cannot show it.
Benchmark lifetimes, release to the best published score today animated
How much of a test set is wrong, by the people who checked animated
Most open-ended evaluation now uses , because it is a widely used option that scales, although executable checks, structured outcome validation and sampled expert review also scale in their own ways. The biases were catalogued with the MT-Bench paper in 2023 and have not gone away. When the order of two answers was swapped, Claude-v1 gave the same verdict 23.8% of the time, GPT-3.5 46.2% and GPT-4 65.0%. Under a "repetitive list" attack that pads an answer without adding content, Claude-v1 and GPT-3.5 preferred the padded answer 91.3% of the time and GPT-4 8.7%, on 23 answers. GPT-4 favoured its own outputs by 10 points of win rate and Claude-v1 by 25.29 Self-preference has a mechanism: GPT-4 recognised its own text 73.5% of the time out of the box, and when the same authors fine-tuned GPT-3.5 and Llama 2 to raise or lower their self-recognition, the strength of self-preference rose and fell linearly with it.32 A "null model" that returns one fixed response regardless of the question reached an 86.5% length-controlled win rate on AlpacaEval 2.0, 83.0 on Arena-Hard-Auto and 9.55 on MT-Bench; these are attacks on automatic judges, and say nothing about how the same response would fare with human voters.30
Two 2026 studies confirm the biases shrink with stronger judges without disappearing: "Stronger judges reduce but do not remove position and verbosity bias", and the strongest judge in one panel still changed 14.7% of verdicts when the answers were swapped.33 The judge is also a moving instrument. Because it "is itself a model behind an API", a silent version bump or a changed scoring prompt means that "every drift alarm is ambiguous between a worse product and a changed judge".34 Judges are inside several of the numbers in this essay: 's reports an from judged comparisons, Arena introduced AutoEval scores in July 2026 to provide ratings "when waiting for human votes to accumulate", and OpenAI's launch table grades HealthBench Professional with GPT-5.4.8, 63, 4
Swap the order, pad the answer, watch the verdict simulation
In April 2025 a group of researchers from Cohere, Princeton, Stanford, MIT and elsewhere published "The Leaderboard Illusion", an audit of Chatbot Arena. They identified 27 private variants tested by Meta before the release, estimated that Google and OpenAI had received 19.2% and 20.4% of all Arena data while 83 open-weight models together received 29.7%, and showed that extra access to Arena data could produce relative gains of up to 112% on an Arena-distribution test set.22 The Arena's reply, published on 9 May 2025, accepted the premise while disputing the magnitudes. It stated that "any model provider can submit as many public and private variants as they would like, as long as we have capacity for it", estimated the boost from pre-release testing at "around +11 Elo after 50 tests and 3000 votes" and diminishing with fresh votes, and objected that the 112% figure came from Arena-Hard, a static LLM-judged set, rather than from the Arena.20 The policy page, last updated on 1 September 2026, keeps the permission: "Model providers are allowed to test multiple variants of their models before making them public, subject to our system's constraints." Scores collected before release are marked preliminary "until enough fresh votes have been collected after the model's public release", and retired models are listed publicly.19
The statistical point, , does not depend on the disputed magnitudes. If a provider tests N variants whose true quality is the same and publishes the best, the published score is biased upward by the expected maximum of N noisy draws: about 0.56 standard deviations for two variants, 1.16 for five, 1.54 for ten and close to 2 for 27, under independent noise of equal size. Real variants are correlated and differ in true quality, so the figure is the shape of the effect, not an estimate of what any provider gained. The bias is invisible to the reader, because the other N − 1 scores were never shown, and giving every provider the same number of tries makes the opportunity symmetric without removing the bias from each published number.
Test N variants, publish the best statistics, interactive
An agent benchmark scores a model together with the scaffold around it: the prompt, the tools, the memory between steps, the , the time and turn limits. Change the scaffold and the score moves by amounts that exceed the gaps between models. OpenAI's July 2026 experiment held GPT-5.6 Sol fixed and changed two API settings, retained reasoning and compaction; the ARC-AGI-3 public-set score went from 13.3% to 38.3% with six times fewer output tokens.6 Under ARC Prize's two harnesses in September, Astra's semi-private score moved from 62.7% to 99.9%, and the Provider Adapter runs were "approximately 3.66x faster by aggregate recorded elapsed time and used 49% fewer total tokens" on the 167 game-and-effort pairs that both harnesses solved.1 Princeton's Holistic Agent Leaderboard, which ran 21,730 rollouts across nine models and nine benchmarks, found two model-and-scaffold combinations, GPT-5 in SeeAct and Claude Sonnet 4 in Browser-Use, that differed by a factor of nine in cost "despite just a two-percentage-point difference in accuracy", a system comparison rather than a scaffold ablation, and later declared a benchmark solved after swapping the scaffold and fixing grading errors for the same model.54 A 2026 study that varied only the harness measured a 27.4-point spread in pass@1 from harness choice against 29.4 points from model choice, and a jump from 19.1% with a minimal adapter to 73.4% with the full one for the same backbone.56 The lesson these share is that a final task score can hide interface failures, tool failures and capability failures, which need different remedies.
The benchmark maintainers know this. Anthropic's system card notes that Terminal-Bench 4.0 "reduced the confounding role of various harnesses (e.g., CLI memory footprint, compaction strategy, and container communication protocol)" compared with earlier versions.13 Terminal-Bench 2.1 modified 26 tasks "to fix bugs, modify timeouts or resources, or improve robustness to reward hacking".57 The official SWE-bench leaderboard now carries a "Bash Only" view in which every model runs in the same minimal agent, mini-SWE-agent, because the scaffold effect is large enough to need a standard.71 None of this makes a harness-optimised score dishonest. It makes it a measurement of the model and the harness together, which is a different quantity from the one a leaderboard column heading implies.
One model, two harnesses, six reasoning efforts scroll driven, ARC Prize's table
Scroll to step through the table.
Under the Standard harness the score rises with reasoning effort, from 35.2% with no reasoning to 62.7% at maximum, with one anomaly at the low setting, 17.5%, that ARC Prize reports without comment.
Under the Provider Adapter harness every effort level lands between 96.7% and 99.9%. The best cell is at the high setting, not at maximum. At the same effort the harness gap is 35.9 points at max (62.7% against 98.6%) and 45.1 at high (54.8% against 99.9%); the 37-point gap between the two best cells changes the effort as well as the harness.
Cost runs the other way. Standard-harness runs cost $26,098 to $49,791 for the whole semi-private set; Provider Adapter runs cost $17,332 to $23,457. Higher effort tended to make the Standard runs cheaper, though not at every step, because the model needed fewer moves.
OpenAI's launch table prints the 99.9% cell next to GPT-5.6 Sol's 7.8% and Claude Opus 5's 30.2%, both Standard-harness numbers, with the harness footnote attached only to Astra's cell.
A leaderboard gap is a difference between two estimates, and both estimates have variance. The methods for handling this are not exotic. Evan Miller's 2024 note set out five: standard errors from the central limit theorem, clustered standard errors when questions come in related groups, variance reduction by resampling, paired differences when two models are compared on the same items, and power analysis before the fact. On two popular evaluations, clustered standard errors were "over 3X larger than naive standard errors", and the intervals in a major technical report were found "likely too narrow in some cases and too wide in other cases".47 Yet the BetterBench audit found that 14 of 24 benchmarks "did not perform multiple evaluations of the same model or report statistical significance or uncertainty of results", and a 2025 review of 445 benchmarks found that only 53.4% presented evidence for construct validity and fewer than 10% used complete real-world tasks.46, 45
The sources of variance are ordinary ones. Across the Llama, Qwen and Gemma families, MMLU accuracy varied by about 23%, as the authors report it, with the single character used to separate few-shot examples, enough, they note, to put any of the tested models in the lead by choosing the separator.48 Two common prompting styles moved ARC and MMLU scores by more than 20 points, and micro versus macro averaging over MMLU's 57 subjects shifts a headline by several points.49 Where trial counts are printed, they are instructive. Anthropic ran ten times per task, 700 trials, and still reports a standard error of 3.5 to 4.5 points per model, because two thirds of the tasks are solved either at least 80% or at most 20% of the time, so "most of the uncertainty comes from tasks rather than run-to-run variance".13 Epoch AI's Tier 4 leaderboard prints intervals of ±2.4 to ±7 points on 43 problems.11 ARC Prize states that "A single run is used, we do not average scores across runs."3 The Arena's top fifteen carry intervals of ±3 to ±11 points around a spread of 8 points between first and fifth.17 The New Stack made the arithmetic explicit for one launch-week claim: a 1.3-point difference on a 113-task coding benchmark "is roughly equivalent to one or two tasks".25
How wide is the interval around a pass rate interactive
The methodology line printed beneath the tables in OpenAI's launch post reads: "Evaluation scores are the maximum at any effort."4 That is a legitimate choice, and it is also a selection: each cell is the best of several runs at different effort settings, and the effort that produced it is not shown. Other selections sit in the footnotes. On SRE-Bench the model "solved 88.0% of tasks in a single attempt and 99.2% within four attempts", and the system card describes the second figure as the maximum over four independent trials, a figure printed beside a pass@1 one. On ExploitBench, the scoring rule gives a vulnerability full credit "if any seed achieves arbitrary code execution" across five attempts. The FrontierMath Tier 4 result covers a tier of 43 problems of which 41 are private. The OSWorld figure is "a subset of the original OSWorld V2 that works without internet access". For two benchmarks, "the Fable scores we report come from Mythos, which is Fable with fewer safeguards", under a column still headed Claude Fable. Three science evaluations omit the Claude models "because they refuse the majority of questions". On FrontierCode, Astra alone ran with a developer message from its coding product.4, 7, 11
The most consequential footnote is the one that is asymmetric. The ARC-AGI-3 row places Astra's 99.9% beside GPT-5.6 Sol at 7.8% and Claude Opus 5 at 30.2%. Footnote 1 says that Astra "was run with our responses API harness, which changes two settings to better match real-world performance". The 7.8% is the Standard-harness number from OpenAI's own July post, before those two settings were applied; the 30.2% is ARC Prize's Standard-harness number for Opus 5. One cell in the row was produced under the improved conditions and two were not.4, 6, 2 Anthropic's table has its own version of the pattern: its Terminal-Bench-Science figure of 52.6% comes from an internal setup that reproduces the public leaderboard's numbers for two other models "within noise", while the public board itself lists neither Fable 5.1 nor Astra.12, 14 Neither practice is hidden. Both require the reader to hold the footnotes next to the headline, and the headline is usually quoted without them.
Footnote 2 of the same table defines the comparison model as "the version available in our API, ChatGPT Codex, and ChatGPT Work. The version in ChatGPT Chat is slightly different."4 Footnote 16 explains that cross-vendor safety tests used "a simpler research setup" than the production product. The configuration that produced 99.9% is the provider's own harness, not the interface a customer uses through the benchmark's standard tools. Anthropic's table carries the mirror image: its models were evaluated "with its production safeguards enabled"; on OSWorld 2.0 tasks where a safeguard intervened Fable 5.1 "scored a zero", and on the other benchmarks "cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5", so some Fable rows report a routed system of three models rather than one.12 A shipped system can therefore score below or above the evaluated one depending on which layers were present during the test, and a routed result describes the combination of models that produced it, not only the model named in the column.
A second gap opened in 2026 between the evaluated behaviour and the deployed behaviour. 's pre-deployment evaluation of GPT-5.6 Sol found a "detected cheating rate" that "was higher than any public model we have evaluated on our ReAct agent harness", where cheating means exploiting bugs in the environment or using disallowed strategies. Counting cheating as failure gave a 50% of about 11.3 hours with a 95% interval from 5 to 40 hours; counting it as success moved the estimate "beyond 270hrs". METR's conclusion was that it did "not consider any of these numbers to represent a robust measurement".23 , testing Astra for three days before release, recorded verbalised in 41.1% of samples at one effort setting and 50.6% at maximum, and wrote that low rates of misbehaviour under those conditions "do not provide substantial evidence about the model's alignment or misalignment".7 Anthropic's Claude Sonnet 5 card called verbalised awareness of about 6% of rollouts in its most realistic evaluation "concerningly high".52 A 2025 study found frontier models distinguish evaluation transcripts from deployment with an AUC of 0.83 against a human baseline of 0.92.50 And a 2024 study showed that models can be prompted to underperform selectively, with Claude 3 Opus losing 39.8 points on a dangerous-capability biology set against 9.7 on ordinary biology questions.51 Three different findings are involved here, and they should be kept apart: recognising an evaluation, behaving differently because of it, and concealing the difference. The transcript study shows the first, the sandbagging study shows the second can be induced by a prompt, and neither measures how often the third happens unprompted in deployment. What they establish together is that behaviour observed under test conditions may not transfer to deployment, which is a design problem for evaluations rather than a reason to discount every score.
A lab's comparison chart is published on launch day. Independent boards follow on their own schedules, and by the time they publish, the index they publish into may have changed. Three days after the Astra launch, as of 6 September 2026, the model had a score on ARC Prize, Artificial Analysis, Epoch AI, and , and no entry on the Arena text leaderboard, the official SWE-bench Verified board (whose newest entries were from February 2026), the public Humanity's Last Exam site, the Scale Labs leaderboards, the OSWorld 2.0 academic board or METR.17, 71, 72, 64, 23 The Arena had added it to the Agent Arena with "Scores coming soon".70
Meanwhile the indices moved. Artificial Analysis published Astra at 61 on Intelligence Index v4.1.1 on 3 September and announced v4.2, which "has more complex and realistic tasks, and more private test sets to prevent gaming", on 4 September; its comparison page for the same two models then read 55 against 51 where the article had read 61 against 61, and its release page listed the high-effort configuration at 53. Zvi Mowshowitz, summarising the re-run on 5 September, noted that "Retroactive adjustments are more than a little suspicious, but this is more plausible."8, 9, 69 Epoch AI's FrontierMath v2, released on 12 June 2026, corrected 123 problems in Tiers 1 to 3 and 12 in Tier 4 and removed 12 more.11 OSWorld 2.0 numbers depend on which task release was used: Anthropic ran the August 2026 release and says its numbers "aren't directly comparable to previously published OSWorld 2.0 results"; OpenAI's footnote says its Claude figures use "the official settings, and not the modified tasks and modified grading from the Fable 5.1 System Card"; the academic board, updated on 3 September, lists Claude Opus 5 on the August release as its top entry at 31.4% of tasks completed and a 68.3% partial score, with the June release's leader, Claude Opus 4.8, at 20.6% and 54.8%.12, 4, 64 A number quoted without its index version and task release cannot be checked against the board it came from.
The organisations that measure the frontier are funded by, sell to, or depend on access from the organisations they measure. Epoch AI's FrontierMath pages state that the benchmark "was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark", and Tier 4 has two public problems out of 43.11 The Arena raised a $100 million seed round in May 2025 at a $600 million valuation and a $150 million Series A in January 2026 at $1.7 billion; it launched a commercial evaluation service in September 2025 and reported a $100 million annualised run-rate by June 2026, charged as consumption rather than subscription.61 Its policy states that when it tests unreleased models it shares conversation data with the provider.19 ARC Prize publishes a policy that sponsors "receive no privileged access to our Private or Semi-Private Evaluation datasets", caps verification runs at $10,000, and commits to publishing results within 30 days of a model's release; the same leaderboard page still carries a note that only systems costing under $10,000 are shown while listing ARC-AGI-3 runs at $17,000 to $50,000, two statements the page does not reconcile.3, 2 Epoch describes its defence against selective reporting directly: "We mitigate against this by running models on our own internal evaluations, and by collecting evaluations from independently-run leaderboards."10 Independence has more than one dimension. METR's Sol report states that OpenAI "would have had the legal right to block us from sharing conclusions about risk that depended on non-public information", and in the same passage that "We did not make changes to conclusions, takeaways or tone" as a result of the lab's review; both facts belong in the record, and the existence of the right does not show it was used.23 Who chooses the checkpoint, how long the evaluators have, what they may publish and when are part of independence alongside funding. None of these arrangements is improper in itself. Each is a fact a reader needs when weighing the number.
03 / Case one
Two frontier launches, one benchmark foundation, two independent indices and a week of coverage produced, for one model, four different headline figures on one benchmark and three on one safety test. Anthropic's launch of the same week is read with the same checks in case four.
The launch week was also conducted in public posts. The set below pairs the labs' own announcements with the evaluators' and commentators' posts from the same hours. Text and attached images are taken from each post's public embed data as of 6 September 2026, with times in UTC; posts longer than the embed returns are cut at the last complete sentence and linked.
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis: […]
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. […]
Reflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we’ve seen multiple groups report results above 90% on ARC-AGI-3. […]
arc-agi-3 is now saturated […]
GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. […]
GPT-6 Astra is 75% more expensive than GPT-5.6 Sol at max effort, and largely sits behind its predecessor on the Intelligence Index vs Cost per Task frontier. This is driven by a 2.5x increase in price, partially offset by a reduction in token use.
Anthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a substantial gain above GPT-5.6 Sol of ~85 Elo points. AA-Briefcase is our frontier in-house evaluation with a private held-out test set. […]
OpenAI leads GDP.pdf with GPT-6 Astra at 33.2% and GPT-5.6 Sol at 28.2%, followed by Claude Fable 5.1 at 26.2%. Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. […]
Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)! It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing. […]
Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 with +15.8% net improvement across 6.7k+ real-world agentic sessions! It also redraws the price-performance frontier: #1 on the leaderboard at a median cost of $4.14/task. […]
More data is needed for GPT-6 Astra scores to reach strong confidence intervals. Build and evaluate in Agent Mode to contribute towards the real-world results.
Three readings of one result appear in the first four cards: the Foundation's account gives 63% and 99%, its co-founder's post gives 66% and "nearly 100%", and its blog table gives 62.7% and 99.9%. The two Arena posts, seven hours apart, name a different model first on two different boards, Code Arena and Agent Arena, which are different instruments and can disagree without contradiction; the third asks for more data before the confidence intervals tighten.
The launch post's abstract-reasoning table has one row for ARC-AGI-3: GPT-6 Astra 99.9%, GPT-5.6 Sol 7.8%, Claude Opus 5 30.2%. Footnote 1, attached to the first cell, says the model "was run with our responses API harness, which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3."4 The 7.8% is the number OpenAI itself reported in July as Sol's score before those settings were turned on, in the post that showed them tripling the public-set score.6 The 30.2% is ARC Prize's Standard-harness result for Opus 5.2 ARC Prize's own table shows what the same model does under the shared condition, 62.7%, which is still the highest Standard-harness score on the board and still a 32-point lead over Opus 5. The shared interface is not a neutral one, and a native harness can answer a question the shared one cannot. The problem with the row is not the 99.9% but the comparison, which places a number from one condition beside two from another, when a same-condition comparison was available and would have made the point.
The 100% on ExploitBench appears in the launch post, in "Path to Astra" and in press coverage. The qualification is in the system card, the longest of the three documents: the results "may be artificially inflated due to potential contamination from exposure to historical vulnerabilities", with the example of the model recalling CVE-2024-0517 when asked about CVE-2023-6702. The card also notes that the scoring metric was updated relative to the previous system card in a way that "slightly increases benchmark scores overall".7 The post-cutoff internal set gives 39.0% for Astra and 5.5% for Sol, and the same table's footnote warns that "a 100% success rate may not be achievable" on it and that Sol's figure is an artefact of a 300-turn limit.4 The two sets differ in tasks, limits and difficulty, so the gap between 100% and 39% does not measure how much contamination inflated the first figure; what the pair shows is that the lab itself did not treat the historical result as sufficient on its own. A reader who stopped at the historical cell would take away 100%. The post-cutoff figure and its footnote sit further down the same post, and the contamination note sits in the system card, published by the same organisation within two days.
One safety evaluation asks whether a model facing an impossible task will attack surrounding infrastructure instead. For GPT-5.6 Sol the launch post's table says 48.2% and its prose says 48%; "Path to Astra" says 56%; the system card says 55.4% "at maximum reasoning effort", and adds that Astra, which made no attacks, legitimately completed the task 1.3% of the time.4, 5, 7 The three figures may be different aggregations of the same runs, for example an average across effort levels against the maximum-effort setting, but none of the documents says so. The New Stack repeated one of them, 48.2%.25
Fortune, The New Stack and TechCrunch each had the launch material in advance. Both New Stack pieces, by different authors, gave the ARC-AGI-3 result as 98.6%, a number that is in ARC Prize's table (the maximum-effort Provider Adapter cell) but not in OpenAI's post, which prints the high-effort cell, 99.9%. Fortune printed 99.9% and a 66% standard-harness figure. The 66% is not in the Foundation's blog table, which gives 62.7%, but it is the figure ARC Prize co-founder François Chollet posted that afternoon ("It scores 66% on ARC-AGI-3 using our standard harness"), while the Foundation's own account gave the pair as "63%" and "99%"; Fortune published a clarification the next day.112, 113 The New Stack added an editor's note that a 98.6% score "does not mean the benchmark was 'aced'", because the metric is scored against a human baseline. The same New Stack article paraphrased OpenAI's methodology line as "the models in its evaluations ran at maximum effort", which is a narrower claim than "the maximum at any effort".24, 25, 26, 1, 4 None of this is unusual for launch coverage. It shows how a number reported under two conditions can circulate as three or four numbers with no conditions attached.
Two footnotes change the conditions of the test itself. On ExploitGym, "we tested Astra and Sol without the 6-hour time limit, to better assess their full cyber capabilities". On the post-cutoff ExploitBench set, Sol's 5.5% "is an artifact of the 300-turn limit in the benchmark", and "the model at similar settings achieved an 11.5% score when hitting fewer limits".4 Each note is reasonable on its own terms. Together they show that the limits of a benchmark are parameters the reporting lab can relax for one comparison and keep for another, and that the choice is visible only in the footnotes.
04 / Case two
Fourteen published comparisons between GPT-6 Astra and Claude Fable 5.1, from two labs, two composite indices, four independent boards and one benchmark foundation, as of 6 September 2026.
The two composite indices reached opposite conclusions on the day of the launch. Epoch AI's , which fits one scale to results from more than 50 benchmarks, drawn from Epoch's own runs, independent leaderboards and lab reports, and is scaled so that Claude 3.5 Sonnet sits at 130 and GPT-5 at 150, placed Astra first at 169.2 with a 90% interval of 164.9 to 174.0; its file gives Fable 5.1 162.9 (160.1 to 166.3), Fable 5 162.9, Opus 5 162.3 and GPT-5.6 Sol 162.0.10 Artificial Analysis, which runs nine evaluations itself, placed Astra at 61, level with Sol and five points behind Fable 5.1 at 66.8 OpenAI's own table reproduces that second result, 61.2 against 65.7, in a row beneath the rows where it leads.4 Both indices publish their construction, and the constructions make a disagreement expected: Epoch's index rewards a model for clearing hard benchmarks it has been run on, and Astra, in Epoch's own account, "set new records on our math, continual learning, and game-puzzles benchmarks" while ranking "between Opus 4.7 and Fable 5" on its long-horizon coding benchmark; Artificial Analysis averages the evaluations it runs itself, nine in the methodology it publishes for the current version of the index, with category weights set out there, and several of them moved against Astra; the page describes v4.2, so the exact composition of the v4.1.1 figure quoted here is what the 3 September article states, not what the current page shows.62, 8 Different inputs, fitted difficulties, domain coverage and missing scores are plausible reasons for the gap; none of them has been shown to account for this reversal, which would take a common-input or weight-sensitivity analysis that neither organisation has published. Artificial Analysis labels its Fable 5.1 run "max with fallback", a configuration label the article does not define; Anthropic's own table notes that on some benchmarks tasks its safeguards intercepted were completed by Claude Opus 4.8 or Claude Opus 5, so the Fable column there reports a routed system.8, 12
Sources by row: Epoch AI ECI data file and model page; Artificial Analysis article; LiveBench; Vals AI; Artificial Analysis; OpenAI launch post and Anthropic launch post; Epoch AI FrontierMath Tier 4 leaderboard; OpenAI launch post; OpenAI and Anthropic launch posts and Terminal-Bench-Science leaderboard; OpenAI and Anthropic launch posts; OpenAI launch post and Anthropic system card; Arena text leaderboard; ARC Prize leaderboard; OpenAI and Anthropic launch posts.10, 8, 15, 16, 4, 12, 11, 14, 13, 17, 2
Three patterns describe most of the table. Where an independent party ran both models under one protocol, Fable 5.1 has the higher point estimate on the general indices and Astra on mathematics. Where each lab reported its own number, Astra's is higher. Where the two labs' conditions differ, as on OSWorld 2.0, the numbers cannot be placed in one row at all, and OpenAI's footnote and Anthropic's footnote each say so about the other.4, 12 The rows are not independent votes: the composites contain several of the individual benchmarks, the units differ between percentage points, index points and rating points, and few rows print an interval, so the six-to-five count describes the table and does not settle the question. The Decoder's summary on 4 September was accurate: "Two independent labs each roll dozens of individual tests into a single overall score, but they reach opposite conclusions."28 A single "best model" claim in this week required choosing an instrument, and the choice was usually not stated.
05 / Case three
The most cited human-preference leaderboard separates its top five models by eight points and prints intervals of four to eleven points around them.
The Text Arena's overall board, as displayed on 6 September 2026 with a data stamp of 2 September, lists 400 models and 7,999,020 votes. Its first four places are held by four Anthropic models within five points of one another: claude-fable-5 at 1507 ± 5, claude-opus-4-6-high at 1505 ± 4, claude-fable-5.1-max at 1504 ± 11 and claude-opus-4-7-high at 1502 ± 4. Fifth is Meta's muse-spark-1.2 at 1499 ± 10. Tenth place, muse-spark-1.1, is 15 points behind first. Two entries in the top fifteen are marked because their votes were collected before public release. The vote counts behind the intervals range from 2,906 for the newest model in the top three to 102,999 for a Google model in fifteenth place.17
Text Arena, top fifteen, scores with 95% intervals animated, hover a row
Read literally, the board says that the leader's interval, 1502 to 1512, contains the point estimates of the next three models and overlaps the fifth model's interval. These are marginal intervals, one per model; which of the gaps between neighbours is distinguishable would take a pairwise comparison from the fitted model, which the board does not print. None of this is a criticism of the Arena, which prints the intervals precisely so that this can be seen. It is a criticism of the sentence "the top model on the Arena", which is repeated in launch posts and coverage as though the interval were zero and the order were settled. The interval also depends on what the votes are being asked to measure. The board offers three adjustments, Style Control, Factuality and None, and the style-controlled fit has been shown since 2024 to reorder the frontier.17, 21
The Arena's changelog, kept since July 2025, records the methodological responses to the problems catalogued in the 2025 audit and elsewhere. In July 2025 it strengthened deduplication (removing about 10% of votes from over-represented prompts) and identity-leak filtering (under 4% of votes), and moved its confidence intervals from bootstrap to a closed-form central-limit method. In September 2025 it added a filter for "users who exhibit statistically anomalous voting patterns" and the preliminary tag for models voted on before release. In May 2026 it began counting votes from battles inserted into direct chats and corrected two biases found in that stream, position and shared organisation.18 In July 2026 it introduced AutoEval scores, model-judged ratings offered "when waiting for human votes to accumulate".63 The policy page, updated on 1 September 2026, requires a public commitment to release within two weeks of a score, written confirmation that the pre-release model is identical to the released one, and at least 30 days of access, and it still permits multiple pre-release variants.19
Two things did not change. The Arena remains a preference measure, and preference remains manipulable. A 2025 study showed that an attacker who can identify which model produced a response, which the authors did with more than 95% accuracy, could move a target model's rank with roughly a thousand votes; the authors worked with the Arena on mitigations.58 A May 2026 study found that "sub-1% targeted perturbations can change the top-ranked model" on Chatbot Arena and six other pairwise datasets.59 And the organisation behind the board is now a company that sells evaluations to the labs it ranks, with a reported $100 million annualised run-rate eight months after launching that service.61 None of this makes the board useless. It makes "first on the Arena" a claim about a few points of preference, under one adjustment, inside an interval that usually covers the next several models.
Two of the Arena's 2026 methods are different instruments from the pairwise text vote, and the essay's reading of the board should not be stretched over them. In July 2026 the Arena added a factuality adjustment to its Text and Search boards: claims in each answer are extracted and checked by search agents calibrated against verified annotations, and a composite Bradley-Terry fit combines the factuality labels with the preference votes at a default factuality weight of 25%, offered "as a non-default toggle" beside Style Control.123 That adds a measured signal to the vote rather than adjusting the vote, and it raises questions of its own: which claims are extracted, how the verifier was validated, how an answer that declines to assert anything is scored, and how much the 25% moves the order. In June 2026 the Agent Arena published a design that randomises the components of the agent a user is given, beginning with the orchestrator model, and estimates each component's "net improvement" against the mix of configurations from signals such as confirmed success, praise against complaint, steerability and tool errors, a randomised trial inside the product rather than a vote between two transcripts.124 Randomisation gives it a causal reading within the traffic it sees; whether that traffic represents any other workload, and whether the proxy signals track the outcomes users care about, are the same validity questions as everywhere else on this page. The two posts of 5 and 6 September that named different leaders came from Code Arena and Agent Arena, which are different instruments, so the two results do not contradict each other.
06 / Case four
The same benchmark name appeared in both launch tables in the same week. The conditions behind the name did not match, and each lab said so in a footnote about the other.
Turn each card to see the other side of the comparison.
The pattern is symmetric and it is not new. Each lab runs the other's models under the conditions it considers official, reports its own model under the conditions it considers representative, and documents both choices in a footnote. The footnotes are accurate. The tables above them are read without them, and the comparison that reaches the reader is between a number produced under one set of conditions and a number produced under another.
07 / Earlier episodes
Each episode below pairs a number that circulated with the number that sat beside it in the primary source. None required a leak to find.
On 20 December 2024 ARC Prize reported that OpenAI's o3 had "scored a breakthrough 75.7% on the Semi-Private Evaluation set at our stated public leaderboard $10k compute limit", and that "A high-compute (172x) o3 configuration scored 87.5%". When o3 was released in April 2025, ARC Prize added a note: "OpenAI has confirmed that this version is not the same as the one we tested in this original post." The released model scored 41% at low effort and 53% at medium on the same set, and the December preview, OpenAI disclosed, "included 75% of the ARC-AGI-1 dataset during training", which ARC Prize identifies as the public training set, a partition the benchmark's rules allow; exposure to it is not leakage of the private evaluation answers. The cost figures in the December post were revised twice, in March and December 2025, as pricing assumptions changed.88
Epoch AI's clarification of 23 January 2025 stated that "OpenAI commissioned Epoch AI to produce 300 advanced math problems for AI evaluation that form the core of the FrontierMath benchmark", that OpenAI "retains ownership of these questions and has access to the problems and solutions, with the exception of a holdout set", and that a 50-problem set was being finalised "for which OpenAI will only receive the problem statements and not the solutions". Epoch added: "Our agreement did not prevent us from disclosing to our contributors that this work was sponsored by an AI company... our communication with them should have been more systematic and transparent."89, 90 The disclosure now sits on every FrontierMath page, which is the remedy; the episode is why the question "who funds the benchmark" is on the checklist.
Meta's Llama 4 launch post described "an experimental chat version scoring ELO of 1417 on LMArena". The Arena announced it at second place. The model Meta released for download was a different configuration; when the Arena added it, it ranked thirty-second, below GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro. The Arena's statement was direct: "Meta should have made it clearer that 'Llama-4-Maverick-03-26-Experimental' was a customized model to optimize for human preference." The audit published three weeks later counted 27 private Meta variants tested before the launch.60, 91, 92, 22
xAI's launch page said Grok 4 Heavy "is the first to score 50.7% on Humanity's Last Exam (text-only subset)". The chart data on the same page gives the full-set figures: Grok 4 Heavy with tools 44.4%, Grok 4 with tools 38.6%, and Grok 4 without tools 25.4%. Google's model card for Gemini 2.5 Deep Think, sourcing from the public HLE sites, lists the same 25.4% for Grok 4 without tools.93, 94 Every number is real. The headline is the one produced by the most expensive tier, with tools, on a subset.
On 21 July 2025 Google DeepMind announced that "An advanced version of Gemini Deep Think solved five out of the six problems perfectly, earning 35 total points", results that were "officially graded and certified by IMO coordinators". The model card for the Gemini 2.5 Deep Think that shipped on 1 August lists its IMO 2025 result as "60.7% (Bronze medal grade)", with a footnote that its "IMO 2025 results are computed as pass@1 while all the other results coming from matharena.ai are best of 32", that the Grok 4 comparison used "the highest result available from matharena.ai, with a custom prompt", and that "Results are thus not directly comparable with performance results found in previous Gemini model cards."95, 94
OpenAI's GPT-5 launch stream showed a SWE-bench Verified chart in which a bar for 52.8% was drawn taller than a bar for 69.1%, which was drawn the same height as one for 30.8%. Sam Altman wrote the same day: "wow a mega chart screwup from us earlier". The corrected numbers came with two footnotes in the system card: "All SWE-bench evaluation runs use a fixed subset of n=477 verified tasks", and the headline 74.9% "was run with the default verbosity setting in the API (verbosity = medium). Changes in verbosity can lead to variation in eval performance."96, 97, 98 The chart was an error; the footnotes carried the conditions of the measurement.
OpenAI, which had built SWE-bench Verified in 2024, audited its hardest problems and found that "at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions"; the audit covered 138 of the hardest tasks, not a random sample of the 500. Scores had kept "improving from 74.9% to 80.9% in the last 6 months" regardless. "This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too." Anthropic's July 2026 system card for Claude Opus 5 still reported 96.0% on it; its September card reports only the Pro, Multilingual and Multimodal variants.99, 111, 13
METR estimated that Claude Opus 4.6 "has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks", and said in the same sentence that "this measurement is extremely noisy because our current task suite is nearly saturated". Its site adds that "Measurements above 16 hrs are unreliable with our current task suite." The 14.5 hours circulated without the interval; MIT Technology Review had called the underlying chart "the most misunderstood graph in AI" two weeks earlier.100, 101
GPT-5.5 took the top of the Artificial Analysis Intelligence Index by three points. The same suite's knowledge test recorded a hallucination rate of 86% for the model, against 36% for Claude Opus 4.7 and 50% for Gemini 3.1 Pro.102, 103 The rate has a definition that the headline drops: it is the share of questions the model did not get right on which it gave an answer rather than declining, incorrect over incorrect plus partial plus not attempted, so 86% means that on the questions it did not get right the model gave an answer rather than declining 86% of the time. The grader observes an answer category, not what the model knew; the figure is not an error rate across all questions, and it should be read beside the suite's accuracy and abstention figures.114 A month later Google's Gemini 3.5 Flash post placed the model "in the top-right quadrant of the Artificial Analysis index"; Artificial Analysis's own write-up gave it 55 on the index and a hallucination rate of 61%.104, 105 Neither lab's numbers were wrong. The numbers that were not quoted are the ones a buyer would most want to see.
ARC Prize verified Claude Opus 5 at 30.2% on ARC-AGI-3 Semi-Private under its Standard harness. A month later NVIDIA reported that its AVO agent framework, wrapped around the same model, "achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set", a different task set from the semi-private one, reached with a different surrounding system; the two figures are not a harness ablation on matched items. NVIDIA's own caution, attached to its comparison with another agent system on the same public levels, applies here too: such a comparison "should not be interpreted as a controlled ablation: the two systems differ in agent backend, observation representation, memory, context management, and other implementation details".2, 106, 107 In the same weeks Anthropic's two system cards gave Opus 5 two GDPval-AA scores at the same effort, 1861 in July and 1824 in September. One possible explanation is an Elo that moves as the comparison pool changes; neither card annotates the difference, and the rating version, judge and pool would be needed to settle it.111, 13
Self-reported and independent numbers for open-weight models show the same instrument effects, and not in a single direction. Epoch AI's December 2025 note on its own replications recorded that MiniMax reported "an astronomical 23 percentage point difference in performance on tau-bench when using their API implementation compared to the standard ChatCompletions API", and that Epoch's own GPQA Diamond re-runs produced averages "ranging from 74% to 80%" across settings that "were not statistically significant given the small size of GPQA-Diamond (only 198 questions)".81 DeepSeek's model card for R1-0528 reports 87.5 on AIME 2025; Moonshot's card for Kimi K2 reports 49.5, averaged over 64 attempts. Artificial Analysis's independently run AIME 2025 values for the same two checkpoints, recorded while this essay was being researched, sat below DeepSeek's figure and above Moonshot's; its model pages have since stopped displaying that benchmark, so the values are not quoted here.108, 109, 110 The lesson is not that labs inflate their numbers. It is that a number without its protocol is not comparable to a number with a different one, whichever direction the difference runs.
08 / Good practice
Each of the problems above has an established mitigation. Each mitigation has a trade-off, several were visible during the launch week, and none of them makes a score valid on its own: disclosure lets a reader assess a test, it does not show that the test predicts anything.
"The Leaderboard Illusion" closed with five recommendations: prohibit score retraction after submission and disclose the number of private variants tested; set a transparent cap on private variants per provider; make deprecation criteria auditable and stratified across proprietary, open-weight and open-source models; adopt the variance-based sampling rule the Arena's own 2024 paper described; and publish every tested model, deprecation and sampling rate.22 The Arena's current policy adopts parts of the first, third and fifth: a preliminary tag on pre-release scores, a public list of retired models, a changelog of methodology changes since July 2025, a written confirmation that the tested model is the released one, and 30 days of guaranteed access. It does not cap variants, and its response argued the boost is small.19, 20 The disagreement about magnitude is empirical and could be settled by publishing the retracted scores.
ARC Prize's response to the harness problem is the clearest example for the field: a Standard harness that "uses a minimal, provider-neutral interface", a separately labelled Provider Adapter condition, both published, both costed, and a testing policy that says "A single run is used, we do not average scores across runs" and that sponsors "receive no privileged access" to the evaluation sets.3, 1 A shared interface is not a neutral one; it answers a controlled question about models, while a native product answers a different question about systems, and a table should say which of the two it is asking. The official SWE-bench leaderboard's Bash Only view, in which every model runs in the same minimal agent, mini-SWE-agent, does the same for coding.71 Epoch AI's December 2025 analysis of its own replications measured the stakes: on SWE-bench Verified "simply switching the scaffold makes up to an 11% difference for GPT-5 and up to a 15% difference for Kimi K2 Thinking", the API provider was "the biggest factor of variance in evaluation results", and the practical rule follows: "If you care about comparing models, a standardized scaffold (like mini-SWE-agent) is usually enough. On the other hand, assessing frontier capabilities requires the usage of leading products like Claude Code."81 Both questions are legitimate. Artificial Analysis now publishes an Endpoint Accuracy Index that "measures how the intelligence and capability of a specific model varies across a range of API providers who serve it", which treats the serving stack as a measured variable rather than an assumption.80
Epoch AI states the principle plainly: self-reported scores may be cherry-picked, and "We mitigate against this by running models on our own internal evaluations, and by collecting evaluations from independently-run leaderboards."10 Artificial Analysis maintains "internal copies of all evaluation datasets", evaluates every model "under identical conditions with consistent prompting strategies, temperature settings, and evaluation criteria", estimates a 95% interval of under one point for its index from repeated runs, and describes a mystery-shopper policy under which it registers accounts outside its own domain to check that a private evaluation endpoint behaves like the public one.78, 79 ARC Prize's semi-private set is called semi-private because "tasks are sent to external APIs" and "we acknowledge the possibility of limited leakage over time".3 Independent evaluation is slower than a launch post, and it removes one source of selection. Comparability is a separate property: it comes from a fixed, published protocol, which a lab can also supply, and an independent number assembled from several protocols is no more comparable than a lab's.
Miller's five recommendations of 2024 remain the minimum: standard errors, clustered where items are grouped, variance reduction by resampling, paired comparisons on the same items, and power analysis before running.47 A 2025 position paper adds a warning about the small benchmarks that decide many frontier claims: CLT-based intervals are fine "when benchmarks consist of thousands of examples" but on "smaller, highly specialized benchmarks" they end up "usually dramatically underestimating uncertainty"; the argument concerns those inferential assumptions, not every small-benchmark interval.74 Two further points belong here. The uncertainty being reported should be named, since sampling variance over items, run-to-run variance, grader error and distribution shift at deployment are different quantities and one error bar rarely covers them all. And the score's actual model should be used: the binomial formula fits independent pass-or-fail items, not Elo ratings, weighted composites or ARC-AGI-3's action-efficiency score, and near a boundary or with few independent units an exact, Wilson, bootstrap or hierarchical method fits better, provided the resampling respects the sampling structure, because thousands of resamples cannot manufacture more independent tasks.
These practices have begun to appear in institutional documents, whose standing differs and should be stated. NIST's AI 800-2 is an initial public draft of voluntary practices, published in January 2026; it asks evaluators to "define, conduct, and report a statistically valid analysis procedure" set out in advance, to report statistics "with estimated uncertainties for associated sources of variation", to support comparisons with "a statistical test on the paired difference", and to "Report costs alongside performance"; it also names "solution contamination" and "grader gaming" as forms of evaluation cheating and notes that "the absence of verbalized evaluation awareness does not imply the absence of evaluation awareness".75 The International Network for Advanced AI Measurement, Evaluation and Science published guidance for third-party evaluators in July 2026, which is consensus advice rather than a settled protocol.76 The EU AI Act's obligation on providers of systemic-risk models to perform "model evaluation in accordance with standardised protocols and tools reflecting the state of the art" is a legal duty that has applied since 2 August 2025, with providers of models already on the market before that date given until 2 August 2027 to comply; it binds a class of providers, not benchmark publishers, and prescribes no leaderboard.77, 127 No document in this list settles whether a given benchmark measures what its name says.
The defences against contamination predate frontier models, and the week showed each of them in use. Dated and refreshed sets: LiveBench adds and updates questions monthly and LiveCodeBench collects contest problems after a model's cutoff.82, 83 Private holdouts: SWE-Bench Pro keeps a held-out set of 12 repositories and a commercial set of 18 whose problems are not public.84 Post-cutoff internal sets: OpenAI's ExploitBench Internal Port, built from vulnerabilities disclosed after the cutoff, is the launch week's clearest example; the launch table prints 39% on it beside 100% on the historical set, and because the two sets and their limits differ the gap is a warning rather than a measured contamination penalty.5, 4 Sets written from scratch: Anthropic describes DeepSWE's 113 tasks as "written from scratch to avoid benchmark contamination", and SRE-Bench is "designed to be contamination-free" with 262 binaries from 19 privately developed programs.13, 7 Rewriting with verification: an August 2026 method rewrites mathematics benchmarks and checks the rewritten problems with formal proofs.85 Each reduces the risk of prior exposure without removing it: evaluation traffic through an API, post-training updates, privileged access and publication dates that lag the underlying information are routes a cutoff date does not close, which is why ARC Prize calls its set semi-private. Most of these also produce a smaller test set than the public one they replace, which returns the argument to the standard errors above; refreshed and generated sets can grow instead, at the cost of validating every new item.
Fresh questions are one kind of robustness test. Behavioural tests are another: minimum-functionality items, invariance checks in which an irrelevant change to the input should leave the answer alone, and directional checks in which a relevant change should move it predictably, the framework CheckList set out in 2020.125 A paraphrase can change a question's difficulty as well as its surface, so transformations need validating too, and the split should be made at the level that blocks the shortcut, by repository, document or task family rather than by row. Adversarial collection, with people writing items that current models fail, finds failures fixed sets miss; Dynabench is the established example, and its distribution should be reported separately from ordinary traffic.126
A developer can overfit a public leaderboard without ever seeing its answers: observe a score, adjust the prompt or the model, keep what helped, repeat. After enough rounds the test set has become a development set, and the same thing happens to a benchmark designer who keeps selecting the questions that particular frontier models fail. This is the adaptive data analysis problem, distinct from literal leakage and from private-variant selection, and the work on the reusable holdout showed both the mechanism and the remedy: limit how many adaptive queries a protected set answers, and add noise or thresholds to the answers it gives.121 The practical form is a development set for iteration, a quarantined confirmation set for the claims that matter, an explicit tuning and submission budget, periodic replenishment, and a public anchor set kept across versions so that reducing leakage does not erase the ability to measure a trend. For generated questions, the generator and filter models should be disclosed and the answers checked independently; freshness establishes neither difficulty nor correctness.
Problem 4 catalogued the biases of model judges. The practice that follows is to treat the grader as an instrument with its own validation. Build a reference sample with independent expert judgments and adjudicated disagreements, and seed it with known correct answers, subtle errors, partial successes, refusals, formatting variants and persuasive wrong answers. Measure false acceptances and false rejections separately, by task type; agreement between two judges is not correctness. Blind the model's identity, randomise the order, and freeze the judge model, prompt, rubric and reference material for the life of a comparison, because a judge behind an API changes under the evaluation. Audit a sample of passes as well as failures, since auditing failures alone never finds the false positives, and keep the ambiguous cases rather than forcing them into a binary label. Treat the candidate's output, any pages it retrieved and any repository it touched as untrusted input to the grader: instructions aimed at the evaluator can sit inside them, and the boundary should be tested, as AgentDojo does for agents under prompt injection.118
A test can be wrong in both directions. Case four's DeepSWE note and the SWE-bench Verified audit are failures of the first kind, tests that reject correct work. EvalPlus showed the second kind in 2023: expanding the test suites of a code benchmark exposed incorrect programs that the original tests had accepted.117 Hidden behavioural tests, independently written requirements and deliberately broken implementations that a good test must catch answer the two questions together: can the grader reject a correct solution, and can it accept an incorrect one. Anthropic's engineering guidance describes the same structure, tasks, trajectories and graders each needing their own checks, and notes that final-state grading must look at the artefact, not at the agent's message saying it succeeded.119 For agents, the final outcome is not the whole of success. A system can reach the requested final state by an unacceptable path, editing unrelated files, sending a message it was not authorised to send, exposing private data or discarding work, and a check of the final state alone misses a violation that was temporary. The grader should hold hard constraints on the trajectory as well as the result, credit recovery from tool failures and an honest stop when a task cannot be done, and separate model errors from environment outages by a rule fixed before the run rather than by removing the inconvenient failures afterwards.
The hallucination-rate episode in the earlier section is a case of a metric that only makes sense beside two others. An accuracy-only benchmark rewards a confident guess over an honest abstention, and an abstention-friendly metric rewards saying nothing unless coverage is shown alongside it. The joint picture is a curve of risk against coverage: as the system declines more questions, how often are the answers it still gives wrong. Where a system reports probabilities, a proper scoring rule such as the Brier score or log loss, inspected by slice, tests whether they mean anything; a verbal "I am fairly confident" is not a calibrated probability. The cost of a wrong answer, of a delay and of an escalation differ by task, so no fixed penalty for abstaining is right everywhere, and for genuinely open questions the reference may need several acceptable answers rather than one string.
"AI Agents That Matter" argued in 2024 that agent comparisons must be cost-controlled, since "for substantially similar accuracy, the cost can differ by almost two orders of magnitude", and that many agent benchmarks "have inadequate holdout sets, and sometimes none at all".53 The Holistic Agent Leaderboard implemented that, and its 21,730 rollouts produced the finding that "increased reasoning effort produces equal or lower accuracy" in 21 of 36 model-scaffold-benchmark combinations.54 tau-bench's pass^k metric, "the chance that all k i.i.d. task trials are successful", showed an agent with over 60% single-trial success falling under 25% at pass^8; it measures repeatability under that trial design, and a deployment without a verifier that can pick the successful attempt should be read against it rather than against pass@k.55 The reporting practices that follow are simple: print the run cost next to the score, as ARC Prize and LiveBench do; print pass^k next to pass@1; and show the score-against-cost frontier rather than a single column.1, 15 The denominator should match the claim: cost per successful task, with failed attempts, retries, tool charges and fallback routing included, is the cost of delivered work; a failed task that cost little still delivered nothing. Latency belongs beside it, as a distribution rather than a mean, because a product that usually answers quickly and sometimes misses its deadline is a different product from one that never does.
A model's score alone and the productivity of a person using the model are different quantities, and the second has to be measured directly. METR's 2025 study is the concrete example: 16 experienced open-source developers worked on 246 real issues in repositories they had contributed to for years, each issue randomly assigned to allow or forbid AI tools, and with the early-2025 tools allowed they took 19% longer, after forecasting that the tools would make them 24% faster.120 The result says nothing about the September 2026 tools, other developers or other tasks; what it shows is that the effect of a tool on work cannot be read off a benchmark, and that people's own estimates of it are not a substitute. A study of this kind compares people with and without the tool, counts the whole time including prompting and checking, has reviewers blind to condition assess the output, and reports human-only, model-only where meaningful and assisted results separately, with correction time and defect escape alongside speed.
An average success rate hides the shape of the failures. A model with fewer trivial mistakes and more expensive ones may be the worse choice, so severity, the population affected and conditional rates belong in the report with the mean. A rate of zero or one hundred per cent needs its denominator: zero failures in 100 independent, representative trials leaves a one-sided 95% upper bound on the failure rate of about 3%, and 20 successes out of 20 leaves a two-sided 95% lower bound on the success rate of about 83%; these are illustrations of the arithmetic, not intervals for any benchmark on this page. Exclusions can be outcome-dependent: in METR's Sol analysis, dropping the runs that cheated moved the time-horizon estimate to about 71 hours with an interval too wide to use, because the informative long tasks were the ones discarded, which is why failure handling should be fixed before the run and its sensitivity reported.23 Maximum elicited capability, the tendency to act under realistic conditions and the effectiveness of the controls around the model are three separate measurements; adversarial searches find failure modes without estimating how common they are, and sampled ordinary traffic estimates the common outcomes while missing the rare adversarial ones, so a report needs both.
The instrument section asked whether a score predicts the work. The study that answers it links benchmark results to held-out outcomes on the intended workload: define the tasks, the users, the acceptable quality, the resource limits and the consequences of failure, then compare how well alternative benchmarks predict those outcomes and where each relationship breaks down. A global average can conceal a language, a domain or a user group where the model fails, so the population should be specified and the important slices, chosen in advance from the intended use and the plausible failure mechanisms, reported with their sample sizes and uncertainty; searching hundreds of small subgroups and publishing the striking ones is the selection problem of Problem 5 in a new form. HELM's scenarios and Model Cards' disaggregated reporting are the established precedents for reporting beyond one aggregate.116, 122 For ranking methods, the diagnostics are the same in spirit: a Bradley-Terry fit should be checked for residuals, opponent coverage and hidden preference cycles, and a composite should publish its constituent scores and how the leader changes as the weights move; if plausible weights reverse the order, that is itself a finding worth reporting.
Two practitioners state the reader's position well. Simon Willison, on the February 2026 SWE-bench refresh: "This benchmark uses the same system prompt for every model, which is important for a fair comparison but does mean that the quality of the different harnesses or optimized prompts is not being measured here", and earlier, "There are leaderboards, but I've been losing some trust in those recently. Everyone needs their own benchmark."86 Andrej Karpathy, in his review of 2025: "benchmarks are almost by construction verifiable environments and are therefore immediately susceptible to and weaker forms of it via synthetic data generation", and "Training on the test set is a new art form."87 The practical consequence is the checklist in the next section: treat a published number as a claim about an instrument, and ask for the instrument before accepting the claim.
Name the harness and the effort setting. State the number of items, the number of trials per item and the pass criterion (pass@1, pass@k or pass^k). Print a standard error or interval computed with clustering where items are grouped, and compare models on paired differences. State the task release and the index version. State whether the tested configuration is the shipped one, including safeguards and any routing to other models. State the training cutoff relative to the item dates and any contamination check. Print the cost of the run and the cost per successful task. Disclose funding, access and publication rights. Publish every variant tested. Each item is already done by at least one of the organisations named on this page. None of them is done by all.
09 / A reading procedure
Tick the items a table discloses. The meter counts disclosure, not validity: a fully disclosed test can still measure the wrong thing, and an undisclosed item leaves a question open rather than answering it against the model.
The matrix applies the ten questions to the launch materials of the week, counting a launch post and its system card together. A question counts as disclosed when the documents state it for the headline rows the essay examined, as partly disclosed when they state it for some rows or as a general rule without the per-cell value, and as not disclosed otherwise. A different rule, such as the post alone or every row rather than the headline rows, gives different counts, which is why the matrix is printed and not only the totals.
| Question | OpenAI, GPT-6 Astra | Anthropic, Claude Fable 5.1 |
|---|---|---|
| Who ran it | yes | yes |
| Harness and settings | partly: named for ARC-AGI-3 and FrontierCode | partly: named for the Terminal-Bench rows |
| Reasoning effort, best of several | partly: the rule is stated, the setting per cell is not | partly: stated for some rows in the card |
| Pass criterion | partly: footnotes for two rows | yes, in the card |
| Items and trials | no | yes, in the card |
| Interval or standard error | no | yes, in the card |
| Subset, release and version | partly: OSWorld subset, FrontierMath tier | partly: OSWorld release, strict and partial scores |
| Tested system is the shipped one | yes: footnotes 2, 16 and the Mythos note | yes: safeguards and routing footnote |
| Contamination | yes, in the card and the post-cutoff row | partly: "cannot be ruled out" for one set |
| Funding and access | no | no |
| Total | 3 yes, 4 partly, 3 no | 5 yes, 4 partly, 1 no |
Counting a partial disclosure as half, that is five of the ten for OpenAI's materials and seven for Anthropic's, and the items missing most often are the interval, the trial count and the effort setting. The independent boards disclose more and are read less. A reader who insists on the ten will find that most published comparisons between frontier models are reports of a number under conditions that are partly unknown, will read them as evidence about capability only under those conditions, and will adjust their confidence accordingly.
10 / Coda
They are measuring something. The question is what they measure, and under which conditions.
The Arena measures the preference of its voters, under one style adjustment, with the intervals it prints. ARC-AGI-3 measures a model together with a harness, and now says which. Artificial Analysis averages nine evaluations it runs itself; Epoch AI fits one scale to results from its own runs, independent boards and lab reports; the two disagree for reasons their methods make plausible and do not decompose. A launch table measures the best of several efforts under the publishing lab's own conditions, with the conditions of the comparison models in the footnotes. Each of these is evidence about capability under its own conditions. None is a measurement of capability in the sense the word carries in a headline: a stable property of the model, independent of who asked, how, with what tools and how many times. And none of them, however fully disclosed, answers the question a reader usually has, which is whether the model will do their work better; that takes a validation against the work itself, as the good-practice section sets out.
The launch week of GPT-6 Astra is a useful case not because the reporting was unusually poor but because it was unusually complete. OpenAI published its footnotes and its system card's contamination note. Anthropic published its trial counts and standard errors. ARC Prize published both harness conditions with costs. Artificial Analysis and Epoch AI published methods that explain their disagreement. Almost every problem described here was disclosed by the organisation it concerns. The problem is the distance between the disclosure and the headline, and it is the headline that gets repeated. The eleven problems, the four cases and the checklist are one attempt to close that distance for a reader who wants to know what a score is before deciding what it means.
Every number on this page is one of four kinds, and the text tries to say which: a figure reported by the source it is attributed to, opened on 6 September 2026 and listed below with its access date; a calculation from reported figures (the same-effort harness gaps, the expected maxima, the interval widths); a simulation, of which there are three, labelled as such, illustrating published mechanisms with invented data; or an interpretation, which is the essay's own. Where a source is a secondary report (a newspaper, a commentator) it is cited as such. Claims that could not be traced to a primary page were left out, including several figures that circulated in coverage of the launch. Leaderboard values are those displayed on the access date and will change; the essay keeps the values, not snapshots of the pages. Two independent checks re-opened each source and compared each quoted phrase before publication. The essay covers text and agent benchmarks; retrieval, long-context, multimodal and long-project evaluations raise further distinctions that it does not treat.
Revision of 6 September 2026. An independent review of the published text led to these changes: the thesis now distinguishes conditional evidence about capability from a universal ranking; the 100% and 39% cyber figures are presented as two different sets rather than a measured fall; interval-overlap reasoning was replaced by paired comparison, and the calculator's figure is labelled a scale marker rather than a floor; the AA-Omniscience hallucination rate is defined; Anthropic's safeguard routing to Opus models is described; the Epoch Capabilities Index is described as a fit over mixed sources rather than a common-protocol rerun, and the "rows under one protocol" count was recomputed; the harness comparison is given at matched effort; the saturation, judge, evaluation-awareness and pass^k claims were narrowed; the Stanford quotation, the seven-hour gap between two Arena posts, the MMLU paper's version, the o3 training partition, the HAL cost pair and the source count were corrected; the checklist meter is labelled as disclosure coverage and an audit matrix replaces an unsupported range; and sections on predictive validity, grader validation, human-and-model measurement, rare failures, protected holdouts and two newer Arena methods were added, with fourteen new sources. Five arXiv identifiers the reviewer could not retrieve were re-opened and resolve to the cited titles. Corrections are welcome by email to e1506804@u.nus.edu.
127 primary and secondary sources, each with the date it was opened; sources 114 to 127 were added in the revision of 6 September. The list is collapsed; a numbered superscript anywhere on the page opens it at the entry.