Claim record · edition 1.0.0

si-015Active research

The use of general-purpose benchmarks as measures of general capability is contested in the literature on construct-validity grounds.

This critique argues that benchmarks index performance on a constructed task distribution and cannot underwrite claims about general ability, and that generalization must be measured differently from skill.

Limits of this claim

This is a methodological argument, not a demonstration that any specific score is wrong. Both critics and benchmark authors are represented in the source set.

Supporting source records · 3

primary paperContent verified

AI and the Everything in the Whole Wide World Benchmark

Raji et al. · 2021

The construct-validity critique of general-purpose benchmarks.

Source record →
primary paperContent verified

What Will it Take to Fix Benchmarking in Natural Language Understanding?

Bowman and Dahl · 2021

Names the structural failure modes of benchmark-driven progress measurement.

Source record →
primary paperContent verified

On the Measure of Intelligence

Chollet · 2019

The skill-versus-generalization argument and the ARC formulation, central to construct validity here.

Source record →

Related concepts