Skip to main content

> ML_ALGORITHM // DIRECT-PREFERENCE-OPTIMIZATION-DPO_v1.0

Direct Preference Optimization (DPO)

Elegant preference alignment algorithm that analytically eliminates the separate reward model in RLHF, optimizing language models directly on paired human preferences.

RL from Human Feedback (RLHF / Alignment)reinforcement-learningblack-boxlarge (>100k)
Back to All Algorithms
Computational Complexity
Training Complexity:O(epochs * batch_size * (policy_pass + reference_pass))
Inference Complexity:O(policy_forward)
Hardware Profile
CPU Friendly:No
Requires GPU:Yes
Memory Footprint:high
Interpretability & Data
Interpretability Tier:black-box
Training Data Needs:large (>100k)

Interpretability Assessment

Directly optimizes the language model policy without requiring training a separate unstable reward model.

Suitable Tasks & Supported Modalities

Suitable Tasks:
llm alignment rlhftext generationpreference tuning
Supported Modalities:
text

Implementing Libraries

TRLHugging Face · v0.10.1
View Spec
TransformersHugging Face · v4.44.2
View Spec
torchtune

Foundational Literature

Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (DPO)Rafael Rafailov, Archit Sharma (2023) · Advances in Neural Information Processing Systems (NeurIPS)
Common Pitfalls & Warnings
  • The model exploits length bias by generating verbose responses because human evaluators often correlate length with quality
  • Overfitting quickly to preference pairs if beta regularization parameter is set too low (<0.05)