> ML_ARCHITECTURES_v1.0
Canonical Model Architectures
20 foundational neural architectures defining modern AI: Transformer, ViT, Swin, ResNet, U-Net, DiT, MoE, Mamba, CLIP, and more.
Autoregressive Decoder-Only Transformer (GPT / LLaMA / Mistral)
The dominant foundation model architecture for generative text and code, processing sequences with causal attention masks, RoPE positional embeddings, SwiGLU activations, and RMSNorm.
Sequence-to-Sequence Encoder-Decoder Transformer (T5 / BART / Whisper)
Classic sequence-to-sequence architecture featuring a bidirectional encoder paired with an autoregressive decoder cross-attending to encoder representations, optimal for translation, summarization, and speech.
Bidirectional Encoder-Only Transformer (BERT / RoBERTa / DeBERTa)
Encoder-only architecture conditioning bidirectionally on all tokens simultaneously via masked language modeling, providing the gold standard for embeddings, classification, and token extraction.
Vision Transformer (ViT / DINOv2 / MAE)
Pure transformer architecture operating directly on flattened 16x16 image patches without convolutional inductive biases, unlocking massive scale for visual representation pretraining.
Swin Transformer (Hierarchical Shifted Window Transformer)
Hierarchical vision transformer computing self-attention within local shifted windows, delivering linear computational complexity with respect to image resolution for dense prediction.
Residual Neural Network (ResNet-18 to ResNet-152)
Foundational CNN backbone introducing identity shortcut skip connections (F(x) + x) to eliminate vanishing gradients and enable optimization of networks hundreds of layers deep.
U-Net (Biomedical & Diffusion Convolutional Backbone)
Symmetric contracting and expanding convolutional architecture with cross-level skip connections, foundational in medical segmentation and the backbone of classical latent diffusion.
Diffusion Transformer (DiT / Sora / Flux)
Modern generative architecture replacing convolutional U-Nets with transformer backbones operating on latent spacetime patches, scaling compute predictable in high-fidelity generation.
Sparse Mixture of Experts (MoE / Mixtral / DeepSeekMoE)
Conditional computation architecture replacing dense feed-forward layers with a bank of parallel expert networks gated dynamically per-token, unlocking massive parameter capacity at low inference compute.
Selective State-Space Model (Mamba / Mamba-2)
Sub-quadratic sequence architecture featuring time-varying selection parameters and hardware-aware scan algorithms, delivering constant-memory inference and linear scaling with sequence length.
YOLO (You Only Look Once Real-Time Object Detection Family)
Single-stage anchor-free or anchor-based object detector mapping image pixels directly to spatial bounding box coordinates and class probabilities at ultra-high real-time frame rates.
Contrastive Dual-Encoder (CLIP / OpenCLIP / SigLIP)
Dual-encoder architecture projecting images and text into a shared normalized embedding space trained via symmetric cross-entropy contrastive loss, the backbone of visual search and text-to-image conditioning.
3D Gaussian Splatting (Rasterized Scene Representation)
Explicit 3D radiance field representation modeling scenes as anisotropic 3D Gaussian primitives rasterized using GPU tile sorting, delivering photorealistic real-time interactive 100+ FPS rendering.
Graph Convolutional Network (GCN / GAT / GraphSAGE)
Message-passing neural network iteratively aggregating localized neighbor representations across graph topologies with permutation equivariance, foundational for fraud detection, social networks, and biology.
RWKV (Receptance Weighted Key Value Linear Hybrid)
Hybrid sequence architecture combining the parallelized training advantages of transformers with the constant-time and constant-memory inference efficiency of RNNs via linear attention formulations.
Evoformer & Invariant Point Attention (AlphaFold 2)
Specialized biological transformer exchanging information between multiple sequence alignments (MSAs) and residue pair representations with Invariant Point Attention (IPA) operating in SE(3) space.
Multilayer Perceptron (MLP / Feed-Forward Network)
Foundational universal function approximator composed of stacked fully-connected linear layers interleaved with non-linear activation functions, the foundational building block of deep learning.
Variational Autoencoder (VAE / VQ-VAE)
Probabilistic generative model mapping input data into a parameterized continuous or discrete latent space via the reparameterization trick, regularized using Kullback-Leibler divergence.
Generative Adversarial Network (DCGAN / StyleGAN / HiFi-GAN)
Minimax game-theoretic generative architecture pitting a generator network synthesizing samples against a discriminator network distinguishing real from synthetic distributions, capable of single-pass fast generation.
Gradient Boosted Decision Trees (GBDT / XGBoost / LightGBM / CatBoost)
Ensemble architecture sequentially building shallow decision trees where each successive tree fits the pseudo-residuals (negative gradients) of the loss function, the unrivaled gold standard for tabular data.
