Skip to main content

> ML_ALGORITHM // Q-LEARNING-TABULAR_v1.0

Tabular Q-Learning

Foundational model-free off-policy reinforcement learning algorithm that iteratively learns optimal action-value functions via temporal difference updates.

Value-Based Model-Free RLreinforcement-learninghigh-intrinsicmedium (1k-100k)
Back to All Algorithms
Computational Complexity
Training Complexity:O(episodes * steps)
Inference Complexity:O(|A|) argmax table lookup
Hardware Profile
CPU Friendly:Yes
Requires GPU:No
Memory Footprint:low
Interpretability & Data
Interpretability Tier:high-intrinsic
Training Data Needs:medium (1k-100k)

Interpretability Assessment

Q-table values directly express expected cumulative discounted future returns per action.

Suitable Tasks & Supported Modalities

Suitable Tasks:
reinforcement learningdiscrete control
Supported Modalities:
tabular

Implementing Libraries

gymnasium
torchrl

Foundational Literature

Q-learningChristopher J. C. H. Watkins, Peter Dayan (1992) · Machine Learning
Common Pitfalls & Warnings
  • Inability to generalize across unseen states; table size scales exponentially with state variables
  • Overestimation of action values due to the max operator