Source record · content verified
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
primary paper2023
Why this source is here
Contamination as a per-benchmark measurement problem rather than an anecdote.
- Verification
- Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
- Authors
- Sainz et al.
- Identifier
- arXiv:2310.18018
Claims citing this source · 2
- si-014
Contamination of evaluation data into training corpora is a measured problem that must be assessed per benchmark rather than assumed absent.
- si-023
A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.