Skip to main content

> ML_LITERATURE // BA-2016-LAYER-NORMALIZATION_v1.0

Layer Normalization

Jimmy Lei Ba, Jamie Ryan Kiros, Geoffrey E. Hinton · arXiv preprint (2016)

algorithm2016industry-standardthirdPartyReproduced

Principal Contribution

Proposed normalizing activations across the feature dimension for individual samples, independent of batch size.

Operational Relevance

The universal normalization layer implemented in every Transformer architecture (BERT, GPT, T5, LLaMA via RMSNorm).

Assumptions

  • Normalizing across hidden features of a single token/sample stabilizes recurrent and attention state dynamics

Limitations

  • Slightly less effective than batch normalization in pure feed-forward vision networks with large static batches

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: