Claim record · edition 1.0.0

si-003Established result

Language-model loss follows empirical power-law relationships with training compute, parameter count, and dataset size over the studied ranges.

The originating scaling-law work fitted these relations across several orders of magnitude.

Limits of this claim

These are relations on loss within a studied range, not on task capability, and not guaranteed to hold outside it. Treating a loss extrapolation as a capability forecast is not supported by the source.

Supporting source records · 1

primary paperContent verified

Scaling Laws for Neural Language Models

Kaplan et al. · 2020

The original power-law scaling result, and one side of the compute-allocation disagreement.

Source record →

Related concepts