Skip to main content

> ML_LITERATURE // WATKINS-1992-QLEARNING_v1.0

Q-learning

Christopher J. C. H. Watkins, Peter Dayan · Machine Learning (1992)

foundational1992foundationalthirdPartyReproduced

Principal Contribution

Provided the first mathematical convergence proof of Q-learning to optimal action-values with probability 1 under standard conditions.

Operational Relevance

Foundational theory underlying value-based reinforcement learning and deep Q-networks (DQN).

Assumptions

  • Finite state and action spaces; all state-action pairs are continually visited; learning rates satisfy Robbins-Monro conditions

Limitations

  • Tabular formulation cannot scale to continuous high-dimensional states without function approximation

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: