> ML_ALGORITHM // MASKED-LANGUAGE-MODELING_v1.0
Masked Language Modeling (MLM / BERT)
Bidirectional representation learning paradigm that trains Transformer encoders to predict randomly masked tokens from surrounding bidirectional context.
Denoising Autoencodingself-supervisedblack-boxlarge (>100k)
Back to All AlgorithmsComputational Complexity
Training Complexity:O(epochs * sequences * length^2 * d)
Inference Complexity:O(length^2 * d)
Hardware Profile
CPU Friendly:No
Requires GPU:Yes
Memory Footprint:high
Interpretability & Data
Interpretability Tier:black-box
Training Data Needs:large (>100k)
Interpretability Assessment
Attention maps can be probed post-hoc; internal multi-head representations are highly distributed.
Suitable Tasks & Supported Modalities
Suitable Tasks:
text classificationtoken classificationfeature extraction
Supported Modalities:
text
Implementing Libraries
Foundational Literature
BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingJacob Devlin, Ming-Wei Chang (2018) · Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)
Common Pitfalls & Warnings
- The [MASK] token never appears during fine-tuning or inference, creating a pre-train/fine-tune mismatch
- Quadratic self-attention scaling prevents scaling beyond 512 tokens without specialized architectures
