Evaluation and reliability · 10 claims
What benchmarks measure, how measurements drift, and why reliability changes long-horizon performance.
si-011Established result
SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue.
That is the benchmark definition given by its authors.
Limits of this claim
Passing tests is not correct engineering. The benchmark does not measure design quality, maintainability, review, or deployment safety, and its authors do not claim it does.
Read claim record →si-012Established result
MMLU and GPQA are multiple-choice benchmarks over defined subject distributions, with human baselines reported by their authors.
Both are defined by their originating papers, and GPQA reports expert baselines with and without web access.
Limits of this claim
Multiple-choice accuracy over a fixed distribution is not general knowledge or reasoning. Specific baseline percentages are deliberately not restated here; consult the sources.
Read claim record →si-013Established result
Benchmarks are revised by their owners in response to saturation and robustness problems, producing successor versions that are not directly comparable to the originals.
MMLU-Pro was introduced as a more robust successor to MMLU for precisely these reasons.
Limits of this claim
A score change across a benchmark revision is a measurement change, not necessarily a capability change. Trends must not be spliced across versions.
Read claim record →si-014Active research
Contamination of evaluation data into training corpora is a measured problem that must be assessed per benchmark rather than assumed absent.
Multiple studies investigate contamination across widely used benchmarks and argue for per-benchmark measurement.
Limits of this claim
The extent of contamination for any specific model is generally not knowable from outside, because training corpora are not disclosed.
Read claim record →si-015Active research
The use of general-purpose benchmarks as measures of general capability is contested in the literature on construct-validity grounds.
This critique argues that benchmarks index performance on a constructed task distribution and cannot underwrite claims about general ability, and that generalization must be measured differently from skill.
Limits of this claim
This is a methodological argument, not a demonstration that any specific score is wrong. Both critics and benchmark authors are represented in the source set.
Read claim record →si-016Established result
Multi-metric evaluation across scenarios has been proposed and implemented as an alternative to single headline scores.
Holistic evaluation frameworks measure many metrics across many scenarios rather than reporting one number.
Limits of this claim
Reporting more metrics does not resolve construct validity. It makes the trade-offs visible; it does not make any single metric mean more.
Read claim record →si-017Active research
The behaviour of a served model accessed through an API can change over time on fixed prompts.
Longitudinal measurement of a hosted service found behaviour changes across dates on identical inputs.
Limits of this claim
A measurement of a hosted endpoint is dated. Reproducibility claims about API-based evaluations require the date and, where available, the model version.
Read claim record →si-018Established result
METR defines a 50%-task-completion time horizon: the time humans typically take on tasks a model completes with 50% success. Its published measurement reports a horizon of around 50 minutes for one frontier model of that period, and a doubling of the horizon roughly every seven months since 2019, with a possible acceleration in 2024.
These figures are taken from the abstract of the METR paper, read during this pass. They are the only time-horizon numbers in this map, and they replace the draft report’s unverifiable figures.
Limits of this claim
The authors state limitations including external validity, and their five-year statement is explicitly conditional on the trend generalizing. The measurement is over a specific task suite, not real-world software work. Figures later than this paper are not carried here: the draft report’s "~2 hours as of late 2025" was not independently verified during this pass and is excluded.
Read claim record →si-019Active research
Increases in measured long-horizon performance are attributed in part to improved reliability and error recovery rather than to reasoning improvements alone.
The METR paper attributes the horizon increase primarily to greater reliability and ability to adapt to mistakes, combined with logical reasoning and tool use.
Limits of this claim
This is an attribution offered by one research group over one task suite. The decomposition between reliability, reasoning, and scaffolding is not independently established.
Read claim record →si-023Established result
A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.
This follows from the benchmark definitions, the scaffolding result, the contamination literature, and the drift measurement taken together. It is the organizing rule of this map.
Limits of this claim
This is a methodological position, not a finding that any score is incorrect. It constrains interpretation, not measurement.
Read claim record →