> ML_LITERATURE // WATKINS-1992-QLEARNING_v1.0
Q-learning
Christopher J. C. H. Watkins, Peter Dayan · Machine Learning (1992)
foundational1992foundationalthirdPartyReproduced
Principal Contribution
Provided the first mathematical convergence proof of Q-learning to optimal action-values with probability 1 under standard conditions.
Operational Relevance
Foundational theory underlying value-based reinforcement learning and deep Q-networks (DQN).
Assumptions
- Finite state and action spaces; all state-action pairs are continually visited; learning rates satisfy Robbins-Monro conditions
Limitations
- Tabular formulation cannot scale to continuous high-dimensional states without function approximation
