Concept record · edition 1.0.0

Pre-training

What the term means here

The large-scale self-supervised phase in which a model learns from an unlabelled corpus.

Why it matters

Determines what is in the model before any alignment, and is where benchmark contamination enters.

Related claim records · 4

Sources used here · 5

primary paperContent verified

Language Models are Few-Shot Learners

Brown et al. · 2020

Introduced in-context few-shot learning as an emergent property of scale.

Source record →
primary paperContent verified

Scaling Laws for Neural Language Models

Kaplan et al. · 2020

The original power-law scaling result, and one side of the compute-allocation disagreement.

Source record →
primary paperContent verified

Training Compute-Optimal Large Language Models

Hoffmann et al. · 2022

The Chinchilla revision of compute-optimal allocation, and the other side of that disagreement.

Source record →
provider self reportContent verified

The Llama 3 Herd of Models

Grattafiori et al. (Meta) · 2024

A comparatively detailed provider training report; still a self-report.

Source record →
primary paperContent verified

Solving Quantitative Reasoning Problems with Language Models

Lewkowycz et al. · 2022

Domain-targeted training on quantitative reasoning; cited for method, not for score comparisons.

Source record →

Related concepts