> ML_LITERATURE // DEVLIN-2018-BERT-PRE-TRAINING-DEEP-BIDIRECTIONAL-TRANSFORMERS_v1.0
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova · Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) (2018)
seminal-architecture2018industry-standardthirdPartyReproduced
Principal Contribution
Introduced deep bidirectional Transformer pre-training via Masked Language Modeling (MLM) and Next Sentence Prediction (NSP).
Operational Relevance
The global enterprise standard for semantic search, passage re-ranking, intent classification, and NER in Google Search and Elasticsearch.
Assumptions
- Bidirectional conditioning from left-to-right and right-to-left context yields richer semantic representations than left-to-right alone
Limitations
- Encoder-only architecture cannot perform open-ended generative autoregressive text completion
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
