> ML_ARCHITECTURE // SPARSE-MIXTURE-OF-EXPERTS-MOE_v1.0
Sparse Mixture of Experts (MoE / Mixtral / DeepSeekMoE)
Conditional computation architecture replacing dense feed-forward layers with a bank of parallel expert networks gated dynamically per-token, unlocking massive parameter capacity at low inference compute.
Sparse Architecturestextcode
Back to All ArchitecturesArchitecture Overview
Conditional computation architecture replacing dense feed-forward layers with a bank of parallel expert networks gated dynamically per-token, unlocking massive parameter capacity at low inference compute.
Implementing Libraries
Seminal Papers
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityWilliam Fedus, Barret Zoph (2022) · Journal of Machine Learning Research (JMLR)
Mixtral of ExpertsAlbert Q. Jiang, Alexandre Sablayrolles (2024) · arXiv preprint
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelDeepSeek-AI, Aixin Liu (2024) · arXiv preprint
Architectural Limitations & Constraints
- Requires compatible deep learning framework and hardware acceleration for efficient execution.
