Concept record · edition 1.0.0
Evaluation
What the term means here
The practice of measuring model or system behaviour against defined tasks and metrics.
Why it matters
Every capability statement is an evaluation statement. What was measured, under what scaffolding, by whom, decides what the number means.
Related claim records · 10
- si-010
The design of the agent-computer interface materially changes measured coding-agent performance holding the underlying model fixed.
- si-011
SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue.
- si-012
MMLU and GPQA are multiple-choice benchmarks over defined subject distributions, with human baselines reported by their authors.
- si-013
Benchmarks are revised by their owners in response to saturation and robustness problems, producing successor versions that are not directly comparable to the originals.
- si-014
Contamination of evaluation data into training corpora is a measured problem that must be assessed per benchmark rather than assumed absent.
- si-015
The use of general-purpose benchmarks as measures of general capability is contested in the literature on construct-validity grounds.
- si-016
Multi-metric evaluation across scenarios has been proposed and implemented as an alternative to single headline scores.
- si-017
The behaviour of a served model accessed through an API can change over time on fixed prompts.
- si-020
Provider technical reports and model cards are self-reports by the organization that built the system, and may withhold architecture, data, and training details.
- si-023
A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.
Sources used here · 4
Holistic Evaluation of Language Models
Liang et al. · 2022
The argument for multi-metric evaluation instead of a single headline number.
Source record →Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava et al. · 2022
A large collaborative benchmark explicitly concerned with extrapolating capability.
Source record →What Will it Take to Fix Benchmarking in Natural Language Understanding?
Bowman and Dahl · 2021
Names the structural failure modes of benchmark-driven progress measurement.
Source record →AgentBench: Evaluating LLMs as Agents
Liu et al. · 2023
Multi-environment agent evaluation, cited for evaluation design rather than rankings.
Source record →