Skip to main content

> ML_LITERATURE // FUJIMOTO-2018-ADDRESSING-FUNCTION-APPROXIMATION-ERROR-ACTOR-CRITIC-TD3_v1.0

Addressing Function Approximation Error in Actor-Critic Methods (TD3)

Scott Fujimoto, Herke van Hoof, David Meger · International Conference on Machine Learning (ICML) (2018)

algorithm2018foundationalthirdPartyReproduced

Principal Contribution

Introduced Twin Delayed DDPG (TD3), incorporating clipped double Q-learning, delayed policy updates, and target policy smoothing to eliminate actor-critic overestimation.

Operational Relevance

Serves as qualified theoretical and systems foundation for task-reinforcement-learning, task-continuous-control.

Assumptions

  • Markovian state dynamics and stationary reward functions hold in target evaluation environments

Limitations

  • Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: