Claim record · edition 1.0.0

si-013Established result

Benchmarks are revised by their owners in response to saturation and robustness problems, producing successor versions that are not directly comparable to the originals.

MMLU-Pro was introduced as a more robust successor to MMLU for precisely these reasons.

Limits of this claim

A score change across a benchmark revision is a measurement change, not necessarily a capability change. Trends must not be spliced across versions.

Supporting source records · 2

benchmark documentationContent verified

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Wang et al. · 2024

A benchmark revision motivated by saturation and robustness problems in the original.

Source record →
primary paperContent verified

What Will it Take to Fix Benchmarking in Natural Language Understanding?

Bowman and Dahl · 2021

Names the structural failure modes of benchmark-driven progress measurement.

Source record →

Related concepts