Claim record · edition 1.0.0

si-023Established result

A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.

This follows from the benchmark definitions, the scaffolding result, the contamination literature, and the drift measurement taken together. It is the organizing rule of this map.

Limits of this claim

This is a methodological position, not a finding that any score is incorrect. It constrains interpretation, not measurement.

Supporting source records · 5

benchmark documentationContent verified

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Jimenez et al. · 2023

The benchmark definition for test-verified resolution of real repository issues.

Source record →
primary paperContent verified

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Yang et al. · 2024

The load-bearing source for separating scaffolding capability from model capability.

Source record →
primary paperContent verified

NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Sainz et al. · 2023

Contamination as a per-benchmark measurement problem rather than an anecdote.

Source record →
primary paperContent verified

How is ChatGPT's behavior changing over time?

Chen, Zaharia, Zou · 2023

Documents that a fixed prompt against a served model is not a stable measurement over time.

Source record →
primary paperContent verified

AI and the Everything in the Whole Wide World Benchmark

Raji et al. · 2021

The construct-validity critique of general-purpose benchmarks.

Source record →

Related concepts