Concept record · edition 1.0.0
Construct validity
What the term means here
Whether a measurement actually measures the property it is claimed to measure.
Why it matters
It is the difference between "scored well on this benchmark" and "is capable". This atlas keeps those separate everywhere.
What this does not establish
That any benchmark in this map measures general intelligence. Each measures a defined task distribution, and its authors say so.
Related claim records · 6
- si-011
SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue.
- si-012
MMLU and GPQA are multiple-choice benchmarks over defined subject distributions, with human baselines reported by their authors.
- si-013
Benchmarks are revised by their owners in response to saturation and robustness problems, producing successor versions that are not directly comparable to the originals.
- si-015
The use of general-purpose benchmarks as measures of general capability is contested in the literature on construct-validity grounds.
- si-016
Multi-metric evaluation across scenarios has been proposed and implemented as an alternative to single headline scores.
- si-023
A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.
Sources used here · 5
AI and the Everything in the Whole Wide World Benchmark
Raji et al. · 2021
The construct-validity critique of general-purpose benchmarks.
Source record →What Will it Take to Fix Benchmarking in Natural Language Understanding?
Bowman and Dahl · 2021
Names the structural failure modes of benchmark-driven progress measurement.
Source record →On the Measure of Intelligence
Chollet · 2019
The skill-versus-generalization argument and the ARC formulation, central to construct validity here.
Source record →Measuring Massive Multitask Language Understanding
Hendrycks et al. · 2020
The MMLU benchmark definition.
Source record →GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Rein et al. · 2023
The GPQA definition and its stated human baselines.
Source record →