Skip to main content

> ML_LITERATURE // RAFAILOV-2023-DIRECT-PREFERENCE-OPTIMIZATION-LANGUAGE-MODELS-IMPLICIT-REWARD_v1.0

Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (DPO)

Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn · Advances in Neural Information Processing Systems (NeurIPS) (2023)

algorithm2023industry-standardthirdPartyReproduced

Principal Contribution

Derived an exact closed-form substitution of the reward function in the RLHF objective, training aligned language models with simple binary cross-entropy.

Operational Relevance

Widely replaced PPO in open-source LLM post-training (Llama-3, Mistral, Zephyr) due to 3x lower VRAM overhead and extreme training stability.

Assumptions

  • The implicit reward can be formulated analytically as r(x, y) = beta * log(pi(y|x) / pi_ref(y|x)), eliminating the need for a separate reward model

Limitations

  • Sensitive to preference dataset noise; vulnerable to length-exploitation bias without explicit regularization

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: