Skip to main content

> ML_LITERATURE // BROWN-2020-LANGUAGE-MODELS-ARE-FEW-SHOT-LEARNERS_v1.0

Language Models are Few-Shot Learners (GPT-3)

Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei · Advances in Neural Information Processing Systems (NeurIPS) (2020)

seminal-architecture2020foundationalthirdPartyReproduced

Principal Contribution

Scaled autoregressive Transformers to 175B parameters, establishing emergent in-context few-shot learning without gradient updates.

Operational Relevance

Inaugurated the modern generative AI industry, establishing prompting as the primary software interface for intelligence.

Assumptions

  • Power-law empirical scaling laws hold: larger models with more compute and data exhibit smoothly predictable loss reductions and emergent qualitative abilities

Limitations

  • Extremely high serving latency and compute footprint; prone to sycophancy, bias, and generating plausible falsehoods without human alignment

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: