Claim record · edition 1.0.0
si-014Active research
Contamination of evaluation data into training corpora is a measured problem that must be assessed per benchmark rather than assumed absent.
Multiple studies investigate contamination across widely used benchmarks and argue for per-benchmark measurement.
Limits of this claim
The extent of contamination for any specific model is generally not knowable from outside, because training corpora are not disclosed.
Supporting source records · 2
primary paperContent verified
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Sainz et al. · 2023
Contamination as a per-benchmark measurement problem rather than an anecdote.
Source record →primary paperContent verified
Investigating Data Contamination in Modern Benchmarks for Large Language Models
Deng et al. · 2023
Empirical contamination investigation across widely used benchmarks.
Source record →