> ML_LITERATURE // VANHASSELT-2016-DEEP-REINFORCEMENT-LEARNING-WITH-DOUBLE-Q-LEARNING_v1.0
Deep Reinforcement Learning with Double Q-learning (Double DQN)
Hado van Hasselt, Arthur Guez, David Silver · AAAI Conference on Artificial Intelligence (2016)
algorithm2016foundationalthirdPartyReproduced
Principal Contribution
Demonstrated that Q-learning suffers from systematic upward maximization bias, decoupling greedy action selection from target evaluation via Double Q-learning.
Operational Relevance
Serves as qualified theoretical and systems foundation for task-reinforcement-learning.
Assumptions
- Markovian state dynamics and stationary reward functions hold in target evaluation environments
Limitations
- Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning
