Claim record · edition 1.0.0

si-016Established result

Multi-metric evaluation across scenarios has been proposed and implemented as an alternative to single headline scores.

Holistic evaluation frameworks measure many metrics across many scenarios rather than reporting one number.

Limits of this claim

Reporting more metrics does not resolve construct validity. It makes the trade-offs visible; it does not make any single metric mean more.

Supporting source records · 2

primary paperContent verified

Holistic Evaluation of Language Models

Liang et al. · 2022

The argument for multi-metric evaluation instead of a single headline number.

Source record →
benchmark documentationContent verified

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Srivastava et al. · 2022

A large collaborative benchmark explicitly concerned with extrapolating capability.

Source record →

Related concepts