Source record · content verified
AI and the Everything in the Whole Wide World Benchmark
primary paper2021
Why this source is here
The construct-validity critique of general-purpose benchmarks.
- Verification
- Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
- Authors
- Raji et al.
- Identifier
- arXiv:2111.15366
Claims citing this source · 2
- si-015
The use of general-purpose benchmarks as measures of general capability is contested in the literature on construct-validity grounds.
- si-023
A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.