Claim record · edition 1.0.0
si-015Active research
The use of general-purpose benchmarks as measures of general capability is contested in the literature on construct-validity grounds.
This critique argues that benchmarks index performance on a constructed task distribution and cannot underwrite claims about general ability, and that generalization must be measured differently from skill.
Limits of this claim
This is a methodological argument, not a demonstration that any specific score is wrong. Both critics and benchmark authors are represented in the source set.
Supporting source records · 3
primary paperContent verified
AI and the Everything in the Whole Wide World Benchmark
Raji et al. · 2021
The construct-validity critique of general-purpose benchmarks.
Source record →primary paperContent verified
What Will it Take to Fix Benchmarking in Natural Language Understanding?
Bowman and Dahl · 2021
Names the structural failure modes of benchmark-driven progress measurement.
Source record →primary paperContent verified
On the Measure of Intelligence
Chollet · 2019
The skill-versus-generalization argument and the ARC formulation, central to construct validity here.
Source record →