> ML_LITERATURE // HAARNOJA-2018-SOFT-ACTOR-CRITIC-OFF-POLICY-MAXIMUM-ENTROPY-DEEP-RL_v1.0
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine · International Conference on Machine Learning (ICML) (2018)
algorithm2018industry-standardthirdPartyReproduced
Principal Contribution
Integrated maximum entropy RL with off-policy actor-critic architectures and twin Q-networks, maximizing reward while maintaining maximum action entropy.
Operational Relevance
The premier benchmark algorithm for continuous control in physical real-world robotics and autonomous vehicle trajectory planning.
Assumptions
- Augmenting expected return with policy entropy prevents premature policy collapse and encourages robust multi-modal exploration
Limitations
- Requires automated entropy temperature tuning to prevent either completely random policies or rapid deterministic collapse
