Skip to main content

> ML_ALGORITHM // STATE-ACTION-REWARD-STATE-ACTION-SARSA_v1.0

SARSA (State-Action-Reward-State-Action)

On-policy temporal difference algorithm that updates Q-values using the actual next action taken by the current policy rather than the greedy maximum.

Value-Based Model-Free RLreinforcement-learninghigh-intrinsicmedium (1k-100k)
Back to All Algorithms
Computational Complexity
Training Complexity:O(episodes * steps)
Inference Complexity:O(|A|) lookup
Hardware Profile
CPU Friendly:Yes
Requires GPU:No
Memory Footprint:low
Interpretability & Data
Interpretability Tier:high-intrinsic
Training Data Needs:medium (1k-100k)

Interpretability Assessment

Learned Q-values reflect the safety penalties of the active exploration policy.

Suitable Tasks & Supported Modalities

Suitable Tasks:
reinforcement learningsafe exploration
Supported Modalities:
tabular

Implementing Libraries

gymnasium
torchrl

Foundational Literature

Common Pitfalls & Warnings
  • Converges to sub-optimal conservative paths in dangerous environments compared to Q-learning