> ML_LITERATURE // BA-2016-LAYER-NORMALIZATION_v1.0
Layer Normalization
Jimmy Lei Ba, Jamie Ryan Kiros, Geoffrey E. Hinton · arXiv preprint (2016)
algorithm2016industry-standardthirdPartyReproduced
Principal Contribution
Proposed normalizing activations across the feature dimension for individual samples, independent of batch size.
Operational Relevance
The universal normalization layer implemented in every Transformer architecture (BERT, GPT, T5, LLaMA via RMSNorm).
Assumptions
- Normalizing across hidden features of a single token/sample stabilizes recurrent and attention state dynamics
Limitations
- Slightly less effective than batch normalization in pure feed-forward vision networks with large static batches
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
