> ML_LITERATURE // RADFORD-2018-IMPROVING-LANGUAGE-UNDERSTANDING-GENERATIVE-PRE-TRAINING_v1.0
Improving Language Understanding by Generative Pre-Training (GPT-1)
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever · OpenAI Technical Report (2018)
seminal-architecture2018foundationalthirdPartyReproduced
Principal Contribution
Established generative autoregressive pre-training on diverse unlabeled text followed by discriminative task fine-tuning.
Operational Relevance
Origin of the GPT series, validating causal autoregressive Transformer decoding for general NLP task transfer.
Assumptions
- Unsupervised language modeling serves as an effective pre-training objective that transfers to diverse supervised downstream tasks
Limitations
- Model scale (117M parameters) was insufficient for zero-shot in-context task execution without fine-tuning
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
