Claim record · edition 1.0.0
si-016Established result
Multi-metric evaluation across scenarios has been proposed and implemented as an alternative to single headline scores.
Holistic evaluation frameworks measure many metrics across many scenarios rather than reporting one number.
Limits of this claim
Reporting more metrics does not resolve construct validity. It makes the trade-offs visible; it does not make any single metric mean more.
Supporting source records · 2
primary paperContent verified
Holistic Evaluation of Language Models
Liang et al. · 2022
The argument for multi-metric evaluation instead of a single headline number.
Source record →benchmark documentationContent verified
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava et al. · 2022
A large collaborative benchmark explicitly concerned with extrapolating capability.
Source record →