Skip to main content

> ML_LITERATURE // RADFORD-2019-LANGUAGE-MODELS-ARE-UNSUPERVISED-MULTITASK-LEARNERS_v1.0

Language Models are Unsupervised Multitask Learners (GPT-2)

Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever · OpenAI Technical Report (2019)

seminal-architecture2019foundationalthirdPartyReproduced

Principal Contribution

Demonstrated that scaling autoregressive language models (1.5B parameters) to diverse web datasets unlocks zero-shot task completion without fine-tuning.

Operational Relevance

Proved the concept of prompt-based zero-shot evaluation, shifting AI away from task-specific fine-tuning heads.

Assumptions

  • Language models trained on sufficiently large text corpora begin learning to solve downstream tasks conditioned solely on natural language prompts

Limitations

  • Hallucination, factual inconsistency, and sensitivity to prompt formulation; lack of safety alignment

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: