Skip to main content

> ML_LITERATURE // IOFFE-SZEGEDY-2015-BATCH-NORMALIZATION_v1.0

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe, Christian Szegedy · International Conference on Machine Learning (ICML) (2015)

algorithm2015foundationalthirdPartyReproduced

Principal Contribution

Introduced mini-batch normalization of layer activations, enabling higher learning rates and acting as a strong regularizer.

Operational Relevance

Standard component in convolutional architectures; highlighted importance of activation normalization across layers.

Assumptions

  • Normalizing inputs to zero mean and unit variance across mini-batches stabilizes gradient flow and smooths the loss landscape

Limitations

  • Dependent on mini-batch size; breaks down under batch_size=1 or variable sequence lengths in NLP (resolved by LayerNorm)

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: