Claim record · edition 1.0.0

si-014Active research

Contamination of evaluation data into training corpora is a measured problem that must be assessed per benchmark rather than assumed absent.

Multiple studies investigate contamination across widely used benchmarks and argue for per-benchmark measurement.

Limits of this claim

The extent of contamination for any specific model is generally not knowable from outside, because training corpora are not disclosed.

Supporting source records · 2

primary paperContent verified

NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Sainz et al. · 2023

Contamination as a per-benchmark measurement problem rather than an anecdote.

Source record →
primary paperContent verified

Investigating Data Contamination in Modern Benchmarks for Large Language Models

Deng et al. · 2023

Empirical contamination investigation across widely used benchmarks.

Source record →

Related concepts