> ML_LITERATURE // SRIVASTAVA-2014-DROPOUT-PREVENTING-NEURAL-NETWORKS-OVERFITTING_v1.0
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, Ruslan Salakhutdinov · Journal of Machine Learning Research (JMLR) (2014)
algorithm2014foundationalthirdPartyReproduced
Principal Contribution
Randomly dropping units during training prevents complex co-adaptations of feature detectors, acting as an implicit ensemble of 2^N thinned networks.
Operational Relevance
Universal regularizer in deep learning and key mechanism for MC-Dropout epistemic uncertainty estimation in production.
Assumptions
- Multiplying test-time weights by retention probability p exactly averages predictions across the combinatorial space of sub-networks
Limitations
- Increases training epochs required for convergence by 2x-3x; less commonly used in modern pre-LN Transformers during pretraining
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
