> ML_LITERATURE // JIANG-2024-MIXTRAL-OF-EXPERTS_v1.0
Mixtral of Experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, William El Sayed · arXiv preprint (2024)
Principal Contribution
Introduced 8x7B sparse mixture-of-experts (MoE) routing 2 of 8 experts per token (13B active params out of 47B total), matching or exceeding Llama 2 70B.
Operational Relevance
Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation, task-code-generation.
Assumptions
- Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support
Limitations
- Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology
