> ML_LITERATURE // JIANG-2023-MISTRAL-7B_v1.0
Mistral 7B
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, William El Sayed · arXiv preprint (2023)
seminal-architecture2023industry-standardthirdPartyReproduced
Principal Contribution
Introduced Sliding Window Attention (SWA) and Grouped-query attention (GQA) in a 7B architecture outperforming Llama 2 13B across all benchmarks.
Operational Relevance
Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation, task-code-generation.
Assumptions
- Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support
Limitations
- Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
