> ML_LITERATURE // RADFORD-2019-LANGUAGE-MODELS-ARE-UNSUPERVISED-MULTITASK-LEARNERS_v1.0
Language Models are Unsupervised Multitask Learners (GPT-2)
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever · OpenAI Technical Report (2019)
seminal-architecture2019foundationalthirdPartyReproduced
Principal Contribution
Demonstrated that scaling autoregressive language models (1.5B parameters) to diverse web datasets unlocks zero-shot task completion without fine-tuning.
Operational Relevance
Proved the concept of prompt-based zero-shot evaluation, shifting AI away from task-specific fine-tuning heads.
Assumptions
- Language models trained on sufficiently large text corpora begin learning to solve downstream tasks conditioned solely on natural language prompts
Limitations
- Hallucination, factual inconsistency, and sensitivity to prompt formulation; lack of safety alignment
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
