> ML_LITERATURE // CHRISTIANO-2017-DEEP-REINFORCEMENT-LEARNING-HUMAN-PREFERENCES_v1.0
Deep Reinforcement Learning from Human Preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei · Advances in Neural Information Processing Systems (NeurIPS) (2017)
foundational2017industry-standardthirdPartyReproduced
Principal Contribution
Formulated RLHF by fitting a learned reward predictor to pairwise human preference comparisons of behavior trajectories, solving complex tasks without formal reward engineering.
Operational Relevance
Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-reinforcement-learning.
Assumptions
- Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support
Limitations
- Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology
