Claim record · edition 1.0.0
si-012Established result
MMLU and GPQA are multiple-choice benchmarks over defined subject distributions, with human baselines reported by their authors.
Both are defined by their originating papers, and GPQA reports expert baselines with and without web access.
Limits of this claim
Multiple-choice accuracy over a fixed distribution is not general knowledge or reasoning. Specific baseline percentages are deliberately not restated here; consult the sources.
Supporting source records · 3
benchmark documentationContent verified
Measuring Massive Multitask Language Understanding
Hendrycks et al. · 2020
The MMLU benchmark definition.
Source record →benchmark documentationContent verified
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Rein et al. · 2023
The GPQA definition and its stated human baselines.
Source record →benchmark documentationContent verified
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Wang et al. · 2024
A benchmark revision motivated by saturation and robustness problems in the original.
Source record →