> ML_LITERATURE // BROWN-2020-LANGUAGE-MODELS-ARE-FEW-SHOT-LEARNERS_v1.0
Language Models are Few-Shot Learners (GPT-3)
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei · Advances in Neural Information Processing Systems (NeurIPS) (2020)
Principal Contribution
Scaled autoregressive Transformers to 175B parameters, establishing emergent in-context few-shot learning without gradient updates.
Operational Relevance
Inaugurated the modern generative AI industry, establishing prompting as the primary software interface for intelligence.
Assumptions
- Power-law empirical scaling laws hold: larger models with more compute and data exhibit smoothly predictable loss reductions and emergent qualitative abilities
Limitations
- Extremely high serving latency and compute footprint; prone to sycophancy, bias, and generating plausible falsehoods without human alignment
