Concept record · edition 1.0.0

Evaluation

What the term means here

The practice of measuring model or system behaviour against defined tasks and metrics.

Why it matters

Every capability statement is an evaluation statement. What was measured, under what scaffolding, by whom, decides what the number means.

Related claim records · 10

Sources used here · 4

primary paperContent verified

Holistic Evaluation of Language Models

Liang et al. · 2022

The argument for multi-metric evaluation instead of a single headline number.

Source record →
benchmark documentationContent verified

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Srivastava et al. · 2022

A large collaborative benchmark explicitly concerned with extrapolating capability.

Source record →
primary paperContent verified

What Will it Take to Fix Benchmarking in Natural Language Understanding?

Bowman and Dahl · 2021

Names the structural failure modes of benchmark-driven progress measurement.

Source record →
benchmark documentationContent verified

AgentBench: Evaluating LLMs as Agents

Liu et al. · 2023

Multi-environment agent evaluation, cited for evaluation design rather than rankings.

Source record →

Related concepts