Skip to main content

> ML_ARCHITECTURES_v1.0

Canonical Model Architectures

20 foundational neural architectures defining modern AI: Transformer, ViT, Swin, ResNet, U-Net, DiT, MoE, Mamba, CLIP, and more.

Showing 20 Canonical ArchitecturesFoundational AI Building Blocks
Architecture
text, codeTransformers

Autoregressive Decoder-Only Transformer (GPT / LLaMA / Mistral)

The dominant foundation model architecture for generative text and code, processing sequences with causal attention masks, RoPE positional embeddings, SwiGLU activations, and RMSNorm.

textcode
Architecture
text, audioTransformers

Sequence-to-Sequence Encoder-Decoder Transformer (T5 / BART / Whisper)

Classic sequence-to-sequence architecture featuring a bidirectional encoder paired with an autoregressive decoder cross-attending to encoder representations, optimal for translation, summarization, and speech.

Implementing Tools:
textaudio
Architecture
textTransformers

Bidirectional Encoder-Only Transformer (BERT / RoBERTa / DeBERTa)

Encoder-only architecture conditioning bidirectionally on all tokens simultaneously via masked language modeling, providing the gold standard for embeddings, classification, and token extraction.

Implementing Tools:
text
Architecture
imageVision Transformers

Vision Transformer (ViT / DINOv2 / MAE)

Pure transformer architecture operating directly on flattened 16x16 image patches without convolutional inductive biases, unlocking massive scale for visual representation pretraining.

image
Architecture
imageVision Transformers

Swin Transformer (Hierarchical Shifted Window Transformer)

Hierarchical vision transformer computing self-attention within local shifted windows, delivering linear computational complexity with respect to image resolution for dense prediction.

image
Architecture
imageConvolutional Neural Networks

Residual Neural Network (ResNet-18 to ResNet-152)

Foundational CNN backbone introducing identity shortcut skip connections (F(x) + x) to eliminate vanishing gradients and enable optimization of networks hundreds of layers deep.

image
Architecture
imageConvolutional Neural Networks

U-Net (Biomedical & Diffusion Convolutional Backbone)

Symmetric contracting and expanding convolutional architecture with cross-level skip connections, foundational in medical segmentation and the backbone of classical latent diffusion.

Implementing Tools:
image
Architecture
image, videoGenerative Transformers

Diffusion Transformer (DiT / Sora / Flux)

Modern generative architecture replacing convolutional U-Nets with transformer backbones operating on latent spacetime patches, scaling compute predictable in high-fidelity generation.

Implementing Tools:
imagevideo
Architecture
text, codeSparse Architectures

Sparse Mixture of Experts (MoE / Mixtral / DeepSeekMoE)

Conditional computation architecture replacing dense feed-forward layers with a bank of parallel expert networks gated dynamically per-token, unlocking massive parameter capacity at low inference compute.

Implementing Tools:
textcode
Architecture
text, audio, time-seriesState Space Models

Selective State-Space Model (Mamba / Mamba-2)

Sub-quadratic sequence architecture featuring time-varying selection parameters and hardware-aware scan algorithms, delivering constant-memory inference and linear scaling with sequence length.

Implementing Tools:
textaudiotime series
Architecture
image, videoSingle-Stage Detectors

YOLO (You Only Look Once Real-Time Object Detection Family)

Single-stage anchor-free or anchor-based object detector mapping image pixels directly to spatial bounding box coordinates and class probabilities at ultra-high real-time frame rates.

Implementing Tools:
imagevideo
Architecture
image, textMultimodal Foundation Models

Contrastive Dual-Encoder (CLIP / OpenCLIP / SigLIP)

Dual-encoder architecture projecting images and text into a shared normalized embedding space trained via symmetric cross-entropy contrastive loss, the backbone of visual search and text-to-image conditioning.

imagetext
Architecture
3d-pointclouds, image3D Neural Representations

3D Gaussian Splatting (Rasterized Scene Representation)

Explicit 3D radiance field representation modeling scenes as anisotropic 3D Gaussian primitives rasterized using GPU tile sorting, delivering photorealistic real-time interactive 100+ FPS rendering.

Implementing Tools:
3d pointcloudsimage
Architecture
graphGraph Neural Networks

Graph Convolutional Network (GCN / GAT / GraphSAGE)

Message-passing neural network iteratively aggregating localized neighbor representations across graph topologies with permutation equivariance, foundational for fraud detection, social networks, and biology.

Implementing Tools:
graph
Architecture
text, codeLinear Recurrent Models

RWKV (Receptance Weighted Key Value Linear Hybrid)

Hybrid sequence architecture combining the parallelized training advantages of transformers with the constant-time and constant-memory inference efficiency of RNNs via linear attention formulations.

Implementing Tools:
textcode
Architecture
multimodal, 3d-pointcloudsStructural Biology Transformers

Evoformer & Invariant Point Attention (AlphaFold 2)

Specialized biological transformer exchanging information between multiple sequence alignments (MSAs) and residue pair representations with Invariant Point Attention (IPA) operating in SE(3) space.

Implementing Tools:
multimodal3d pointclouds
Architecture
tabularDense Classical Networks

Multilayer Perceptron (MLP / Feed-Forward Network)

Foundational universal function approximator composed of stacked fully-connected linear layers interleaved with non-linear activation functions, the foundational building block of deep learning.

tabular
Architecture
image, audio, tabularDeep Generative Models

Variational Autoencoder (VAE / VQ-VAE)

Probabilistic generative model mapping input data into a parameterized continuous or discrete latent space via the reparameterization trick, regularized using Kullback-Leibler divergence.

Implementing Tools:
imageaudiotabular
Architecture
image, audioDeep Generative Models

Generative Adversarial Network (DCGAN / StyleGAN / HiFi-GAN)

Minimax game-theoretic generative architecture pitting a generator network synthesizing samples against a discriminator network distinguishing real from synthetic distributions, capable of single-pass fast generation.

Implementing Tools:
imageaudio
Architecture
tabularTree Ensembles

Gradient Boosted Decision Trees (GBDT / XGBoost / LightGBM / CatBoost)

Ensemble architecture sequentially building shallow decision trees where each successive tree fits the pseudo-residuals (negative gradients) of the loss function, the unrivaled gold standard for tabular data.

tabular