> ML_LITERATURE // WATKINS-1992-Q-LEARNING_v1.0
Q-learning
Christopher J. C. H. Watkins, Peter Dayan · Machine Learning (1992)
foundational1992foundationalthirdPartyReproduced
Principal Contribution
Proved the mathematical convergence of Q-learning to optimal action-values in discrete Markov Decision Processes without an environment transition model.
Operational Relevance
Serves as qualified theoretical and systems foundation for task-reinforcement-learning.
Assumptions
- Markovian state dynamics and stationary reward functions hold in target evaluation environments
Limitations
- Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning
