Claim record · edition 1.0.0
si-003Established result
Language-model loss follows empirical power-law relationships with training compute, parameter count, and dataset size over the studied ranges.
The originating scaling-law work fitted these relations across several orders of magnitude.
Limits of this claim
These are relations on loss within a studied range, not on task capability, and not guaranteed to hold outside it. Treating a loss extrapolation as a capability forecast is not supported by the source.
Supporting source records · 1
primary paperContent verified
Scaling Laws for Neural Language Models
Kaplan et al. · 2020
The original power-law scaling result, and one side of the compute-allocation disagreement.
Source record →