Skip to main content

> ML_LITERATURE // SCHULMAN-2015-TRUST-REGION-POLICY-OPTIMIZATION-TRPO_v1.0

Trust Region Policy Optimization (TRPO)

John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, Philipp Moritz · International Conference on Machine Learning (ICML) (2015)

algorithm2015foundationalthirdPartyReproduced

Principal Contribution

Guaranteed monotonic policy improvement by enforcing a Kullback-Leibler (KL) divergence trust-region constraint between old and new policies solved with conjugate gradient and Fisher information.

Operational Relevance

Serves as qualified theoretical and systems foundation for task-reinforcement-learning.

Assumptions

  • Markovian state dynamics and stationary reward functions hold in target evaluation environments

Limitations

  • Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: