Skip to main content

> ML_LITERATURE // RADFORD-2018-IMPROVING-LANGUAGE-UNDERSTANDING-GENERATIVE-PRE-TRAINING_v1.0

Improving Language Understanding by Generative Pre-Training (GPT-1)

Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever · OpenAI Technical Report (2018)

seminal-architecture2018foundationalthirdPartyReproduced

Principal Contribution

Established generative autoregressive pre-training on diverse unlabeled text followed by discriminative task fine-tuning.

Operational Relevance

Origin of the GPT series, validating causal autoregressive Transformer decoding for general NLP task transfer.

Assumptions

  • Unsupervised language modeling serves as an effective pre-training objective that transfers to diverse supervised downstream tasks

Limitations

  • Model scale (117M parameters) was insufficient for zero-shot in-context task execution without fine-tuning

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: