Skip to main content

> ML_ARCHITECTURE // SPARSE-MIXTURE-OF-EXPERTS-MOE_v1.0

Sparse Mixture of Experts (MoE / Mixtral / DeepSeekMoE)

Conditional computation architecture replacing dense feed-forward layers with a bank of parallel expert networks gated dynamically per-token, unlocking massive parameter capacity at low inference compute.

Sparse Architecturestextcode
Back to All Architectures

Architecture Overview

Conditional computation architecture replacing dense feed-forward layers with a bank of parallel expert networks gated dynamically per-token, unlocking massive parameter capacity at low inference compute.

Implementing Libraries

TransformersHugging Face · v4.44.2
View Spec
vLLMvLLM Project / UC Berkeley · v0.6.2
View Spec
SGLangLMSYS Org / UC Berkeley · v0.3.1
View Spec

Seminal Papers

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityWilliam Fedus, Barret Zoph (2022) · Journal of Machine Learning Research (JMLR)
Mixtral of ExpertsAlbert Q. Jiang, Alexandre Sablayrolles (2024) · arXiv preprint
Architectural Limitations & Constraints
  • Requires compatible deep learning framework and hardware acceleration for efficient execution.