Skip to main content

> ML_LITERATURE // OUYANG-2022-TRAINING-LANGUAGE-MODELS-TO-FOLLOW-INSTRUCTIONS-WITH-HUMAN-FEEDBACK_v1.0

Training language models to follow instructions with human feedback

Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, Ryan Lowe · Advances in Neural Information Processing Systems (NeurIPS) (2022)

safety-fairness2022foundationalthirdPartyReproduced

Principal Contribution

Pioneered RLHF alignment (SFT + Reward Modeling + PPO), aligning GPT-3 to human intent and giving birth to ChatGPT.

Operational Relevance

The definitive recipe for production conversational assistants, safety guardrails, and instruction following across the entire AI ecosystem.

Assumptions

  • Human labelers can reliably rank comparative completions, and learned scalar rewards generalize across conversational contexts

Limitations

  • Alignment tax: potential slight regression on raw academic knowledge benchmarks; vulnerable to reward hacking and sycophancy

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: