Skip to main content

> ML_LITERATURE // DEVLIN-2018-BERT-PRE-TRAINING-DEEP-BIDIRECTIONAL-TRANSFORMERS_v1.0

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova · Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) (2018)

seminal-architecture2018industry-standardthirdPartyReproduced

Principal Contribution

Introduced deep bidirectional Transformer pre-training via Masked Language Modeling (MLM) and Next Sentence Prediction (NSP).

Operational Relevance

The global enterprise standard for semantic search, passage re-ranking, intent classification, and NER in Google Search and Elasticsearch.

Assumptions

  • Bidirectional conditioning from left-to-right and right-to-left context yields richer semantic representations than left-to-right alone

Limitations

  • Encoder-only architecture cannot perform open-ended generative autoregressive text completion

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: