Skip to main content

> ML_LITERATURE // HAARNOJA-2018-SOFT-ACTOR-CRITIC-OFF-POLICY-MAXIMUM-ENTROPY-DEEP-RL_v1.0

Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine · International Conference on Machine Learning (ICML) (2018)

algorithm2018industry-standardthirdPartyReproduced

Principal Contribution

Integrated maximum entropy RL with off-policy actor-critic architectures and twin Q-networks, maximizing reward while maintaining maximum action entropy.

Operational Relevance

The premier benchmark algorithm for continuous control in physical real-world robotics and autonomous vehicle trajectory planning.

Assumptions

  • Augmenting expected return with policy entropy prevents premature policy collapse and encourages robust multi-modal exploration

Limitations

  • Requires automated entropy temperature tuning to prevent either completely random policies or rapid deterministic collapse

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: