Skip to main content

> ML_LITERATURE // LOSHCHILOV-HUTTER-2019-DECOUPLED-WEIGHT-DECAY-REGULARIZATION-ADAMW_v1.0

Decoupled Weight Decay Regularization

Ilya Loshchilov, Frank Hutter · International Conference on Learning Representations (ICLR) (2019)

algorithm2019industry-standardthirdPartyReproduced

Principal Contribution

Demonstrated that L2 regularization differs fundamentally from weight decay in adaptive optimizers; decoupled weight decay restored generalization.

Operational Relevance

AdamW is the mandatory optimizer used to train GPT-3, LLaMA, Mistral, BERT, and all modern foundation models.

Assumptions

  • Decoupling weight decay updates from gradient-dependent adaptive step sizes prevents shrinkage distortion on frequent features

Limitations

  • Introduces independent weight decay hyperparameter that requires explicit tuning

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: