> ML_ALGORITHMS_ATLAS_v1.0
Algorithms & Method Families
100 qualified algorithmic method families across 9 disciplines: Classical Supervised, Unsupervised, Time Series, Anomaly Detection, Deep Learning, Recommenders, Causal Inference, Reinforcement Learning, and Specialized Methods.
Gaussian Mixture Models (GMM / Expectation-Maximization)
Outputs probabilistic soft cluster memberships (posteriors) rather than hard assignments.
Isolation Forest (iForest)
Anomaly score is a direct monotonic function of average tree path depth.
Local Outlier Factor (LOF)
LOF ratio directly indicates the factor by which the point is less dense than its local neighborhood.
One-Class Support Vector Machine (OC-SVM)
Decisions are driven by high-dimensional kernel distances from an origin hyperplane.
Independent Component Analysis (FastICA)
Yields an unmixing matrix that recovers physically meaningful underlying independent source signals.
Non-Negative Matrix Factorization (NMF)
Components are non-negative and directly interpretable as additive building blocks of the data.
Masked Language Modeling (MLM / BERT)
Attention maps can be probed post-hoc; internal multi-head representations are highly distributed.
Simple Framework for Contrastive Learning (SimCLR)
Learns a metric embedding space where cosine distance reflects semantic invariant similarity.
Momentum Contrast (MoCo v1 / v2 / v3)
Learns a contrastive metric space uncoupled from mini-batch GPU memory constraints.
Bootstrap Your Own Latent (BYOL)
Generates robust visual representations without relying on negative samples or large contrastive batches.
Self-Distillation with No Labels (DINO & DINOv2)
ViT self-attention heads naturally segment objects and scene semantics without explicit pixel-level supervision.
Masked Autoencoders (MAE)
Reconstruction outputs can be inspected directly to verify semantic visual understanding.
Contrastive Language-Image Pre-training (CLIP)
Enables zero-shot classification via natural language prompt cosine similarity without specialized training heads.
Word2Vec (CBOW & Skip-Gram)
Embedding vector geometry satisfies linear semantic analogies (e.g., King - Man + Woman = Queen).
Global Vectors for Word Representation (GloVe)
Vector dot products directly approximate logarithms of empirical word co-occurrence frequencies.
Causal Autoregressive Next-Token Prediction (GPT)
High capacity enables complex emergent chain-of-thought, but internal mechanistic interpretability remains an active research frontier.
Tabular Q-Learning
Q-table values directly express expected cumulative discounted future returns per action.
SARSA (State-Action-Reward-State-Action)
Learned Q-values reflect the safety penalties of the active exploration policy.
