Claim record · edition 1.0.0

si-012Established result

MMLU and GPQA are multiple-choice benchmarks over defined subject distributions, with human baselines reported by their authors.

Both are defined by their originating papers, and GPQA reports expert baselines with and without web access.

Limits of this claim

Multiple-choice accuracy over a fixed distribution is not general knowledge or reasoning. Specific baseline percentages are deliberately not restated here; consult the sources.

Supporting source records · 3

benchmark documentationContent verified

Measuring Massive Multitask Language Understanding

Hendrycks et al. · 2020

The MMLU benchmark definition.

Source record →
benchmark documentationContent verified

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Rein et al. · 2023

The GPQA definition and its stated human baselines.

Source record →
benchmark documentationContent verified

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Wang et al. · 2024

A benchmark revision motivated by saturation and robustness problems in the original.

Source record →

Related concepts