> ML_LITERATURE // RAFAILOV-2023-DIRECT-PREFERENCE-OPTIMIZATION-LANGUAGE-MODELS-IMPLICIT-REWARD_v1.0
Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (DPO)
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn · Advances in Neural Information Processing Systems (NeurIPS) (2023)
algorithm2023industry-standardthirdPartyReproduced
Principal Contribution
Derived an exact closed-form substitution of the reward function in the RLHF objective, training aligned language models with simple binary cross-entropy.
Operational Relevance
Widely replaced PPO in open-source LLM post-training (Llama-3, Mistral, Zephyr) due to 3x lower VRAM overhead and extreme training stability.
Assumptions
- The implicit reward can be formulated analytically as r(x, y) = beta * log(pi(y|x) / pi_ref(y|x)), eliminating the need for a separate reward model
Limitations
- Sensitive to preference dataset noise; vulnerable to length-exploitation bias without explicit regularization
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
