Source record · content verified
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
benchmark documentation2024
Why this source is here
A benchmark revision motivated by saturation and robustness problems in the original.
- Verification
- Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
- Authors
- Wang et al.
- Identifier
- arXiv:2406.01574
Claims citing this source · 2
- si-012
MMLU and GPQA are multiple-choice benchmarks over defined subject distributions, with human baselines reported by their authors.
- si-013
Benchmarks are revised by their owners in response to saturation and robustness problems, producing successor versions that are not directly comparable to the originals.