Source record · content verified
Scaling Laws for Neural Language Models
primary paper2020
Why this source is here
The original power-law scaling result, and one side of the compute-allocation disagreement.
- Verification
- Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
- Authors
- Kaplan et al.
- Identifier
- arXiv:2001.08361
Claims citing this source · 2
- si-003
Language-model loss follows empirical power-law relationships with training compute, parameter count, and dataset size over the studied ranges.
- si-004
The compute-optimal allocation between parameters and training tokens has been revised in the literature, and deployed practice diverges from compute-optimal training.