Skip to main content

> ML_LITERATURE // HOFFMANN-2022-TRAINING-COMPUTE-OPTIMAL-LARGE-LANGUAGE-MODELS_v1.0

Training Compute-Optimal Large Language Models (Chinchilla)

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, Laurent Sifre · Advances in Neural Information Processing Systems (NeurIPS) (2022)

foundational2022industry-standardthirdPartyReproduced

Principal Contribution

Revised Kaplan scaling laws, showing that parameter count and training tokens should scale in equal proportions (1:1 compute optimality).

Operational Relevance

Reoriented the entire frontier AI research community away from bloated parameter counts (Gopher, Megatron) toward token-rich training (LLaMA, Mistral).

Assumptions

  • Power-law loss parametric scaling functions accurately model the frontier of cross-entropy loss under optimal resource allocation

Limitations

  • Focuses strictly on pre-training compute budget optimality; ignores post-training inference economics where overtraining smaller models is vastly superior

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: