Claim record · edition 1.0.0
si-013Established result
Benchmarks are revised by their owners in response to saturation and robustness problems, producing successor versions that are not directly comparable to the originals.
MMLU-Pro was introduced as a more robust successor to MMLU for precisely these reasons.
Limits of this claim
A score change across a benchmark revision is a measurement change, not necessarily a capability change. Trends must not be spliced across versions.
Supporting source records · 2
benchmark documentationContent verified
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Wang et al. · 2024
A benchmark revision motivated by saturation and robustness problems in the original.
Source record →primary paperContent verified
What Will it Take to Fix Benchmarking in Natural Language Understanding?
Bowman and Dahl · 2021
Names the structural failure modes of benchmark-driven progress measurement.
Source record →