> ML_LITERATURE // LOSHCHILOV-HUTTER-2019-DECOUPLED-WEIGHT-DECAY-REGULARIZATION-ADAMW_v1.0
Decoupled Weight Decay Regularization
Ilya Loshchilov, Frank Hutter · International Conference on Learning Representations (ICLR) (2019)
algorithm2019industry-standardthirdPartyReproduced
Principal Contribution
Demonstrated that L2 regularization differs fundamentally from weight decay in adaptive optimizers; decoupled weight decay restored generalization.
Operational Relevance
AdamW is the mandatory optimizer used to train GPT-3, LLaMA, Mistral, BERT, and all modern foundation models.
Assumptions
- Decoupling weight decay updates from gradient-dependent adaptive step sizes prevents shrinkage distortion on frequent features
Limitations
- Introduces independent weight decay hyperparameter that requires explicit tuning
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
