Concept record · edition 1.0.0

Construct validity

What the term means here

Whether a measurement actually measures the property it is claimed to measure.

Why it matters

It is the difference between "scored well on this benchmark" and "is capable". This atlas keeps those separate everywhere.

What this does not establish

That any benchmark in this map measures general intelligence. Each measures a defined task distribution, and its authors say so.

Related claim records · 6

Sources used here · 5

primary paperContent verified

AI and the Everything in the Whole Wide World Benchmark

Raji et al. · 2021

The construct-validity critique of general-purpose benchmarks.

Source record →
primary paperContent verified

What Will it Take to Fix Benchmarking in Natural Language Understanding?

Bowman and Dahl · 2021

Names the structural failure modes of benchmark-driven progress measurement.

Source record →
primary paperContent verified

On the Measure of Intelligence

Chollet · 2019

The skill-versus-generalization argument and the ARC formulation, central to construct validity here.

Source record →
benchmark documentationContent verified

Measuring Massive Multitask Language Understanding

Hendrycks et al. · 2020

The MMLU benchmark definition.

Source record →
benchmark documentationContent verified

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Rein et al. · 2023

The GPQA definition and its stated human baselines.

Source record →

Related concepts