> ML_LITERATURE // SCHULMAN-2017-PROXIMAL-POLICY-OPTIMIZATION-ALGORITHMS_v1.0
Proximal Policy Optimization Algorithms (PPO)
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov · arXiv preprint (2017)
algorithm2017industry-standardthirdPartyReproduced
Principal Contribution
Introduced clipped surrogate objective policy optimization, providing TRPO stability with simple first-order stochastic gradient descent.
Operational Relevance
The industry-standard reinforcement learning algorithm powering OpenAI Five, robotics benchmarks, and RLHF for ChatGPT.
Assumptions
- Clipping the policy probability ratio r_t(theta) strictly bounds the incentive for overly large policy updates
Limitations
- On-policy nature requires continuous fresh rollouts; high sample complexity in physical robotics without simulation
