Claim record · edition 1.0.0

si-004Active research

The compute-optimal allocation between parameters and training tokens has been revised in the literature, and deployed practice diverges from compute-optimal training.

The Chinchilla work revised the earlier parameter-heavy allocation toward a more balanced ratio, and large deployed models are trained well past compute-optimal token counts to reduce inference cost.

Limits of this claim

The direction of the revision is sourced; the specific exponents and ratios are not restated here, and no claim is made about what any particular current model used.

Supporting source records · 3

primary paperContent verified

Scaling Laws for Neural Language Models

Kaplan et al. · 2020

The original power-law scaling result, and one side of the compute-allocation disagreement.

Source record →
primary paperContent verified

Training Compute-Optimal Large Language Models

Hoffmann et al. · 2022

The Chinchilla revision of compute-optimal allocation, and the other side of that disagreement.

Source record →
provider self reportContent verified

The Llama 3 Herd of Models

Grattafiori et al. (Meta) · 2024

A comparatively detailed provider training report; still a self-report.

Source record →

Related concepts