Skip to main content

> ML_LITERATURE // CHRISTIANO-2017-DEEP-REINFORCEMENT-LEARNING-HUMAN-PREFERENCES_v1.0

Deep Reinforcement Learning from Human Preferences

Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei · Advances in Neural Information Processing Systems (NeurIPS) (2017)

foundational2017industry-standardthirdPartyReproduced

Principal Contribution

Formulated RLHF by fitting a learned reward predictor to pairwise human preference comparisons of behavior trajectories, solving complex tasks without formal reward engineering.

Operational Relevance

Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-reinforcement-learning.

Assumptions

  • Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support

Limitations

  • Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: