Claim record · edition 1.0.0
si-004Active research
The compute-optimal allocation between parameters and training tokens has been revised in the literature, and deployed practice diverges from compute-optimal training.
The Chinchilla work revised the earlier parameter-heavy allocation toward a more balanced ratio, and large deployed models are trained well past compute-optimal token counts to reduce inference cost.
Limits of this claim
The direction of the revision is sourced; the specific exponents and ratios are not restated here, and no claim is made about what any particular current model used.
Supporting source records · 3
primary paperContent verified
Scaling Laws for Neural Language Models
Kaplan et al. · 2020
The original power-law scaling result, and one side of the compute-allocation disagreement.
Source record →primary paperContent verified
Training Compute-Optimal Large Language Models
Hoffmann et al. · 2022
The Chinchilla revision of compute-optimal allocation, and the other side of that disagreement.
Source record →provider self reportContent verified
The Llama 3 Herd of Models
Grattafiori et al. (Meta) · 2024
A comparatively detailed provider training report; still a self-report.
Source record →