Skip to main content

> ML_LITERATURE // WATKINS-1992-Q-LEARNING_v1.0

Q-learning

Christopher J. C. H. Watkins, Peter Dayan · Machine Learning (1992)

foundational1992foundationalthirdPartyReproduced

Principal Contribution

Proved the mathematical convergence of Q-learning to optimal action-values in discrete Markov Decision Processes without an environment transition model.

Operational Relevance

Serves as qualified theoretical and systems foundation for task-reinforcement-learning.

Assumptions

  • Markovian state dynamics and stationary reward functions hold in target evaluation environments

Limitations

  • Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: