> ML_LITERATURE // DUBEY-2024-LLAMA-3-HERD-OF-MODELS_v1.0
The Llama 3 Herd of Models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aravind Srinivas, Arno Candel, Binh Tang, Bryan Catanzaro, Chao Zhou, Chloe Bi, Chris Marra, Dan Bikel, David So, Dhruv Mahajan, Eric Michael Smith, Gautier Izacard, Guillaume Lample, Hakan Inan, Iliyan Zarov, Ishita Dasgupta, Jian Xiang Kuan, Joe Spisak, Jon Carvill, Kalyan Saladi, Kevin Stone, Lukas Blecher, Louis Martin, Mark Tygert, Melanie Kambadur, Naman Goyal, Nikhil Kant, Pushkar Mishra, Rashi Rungta, Ross Taylor, Sainbayar Sukhbaatar, Sergey Edunov, Thomas Scialom, Xavier Martinet, Yuning Mao · arXiv preprint (2024)
Principal Contribution
Trained open foundation models up to 405B dense parameters on 15.6T multilingual tokens with native multimodal support and 128K context window.
Operational Relevance
Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation, task-code-generation, task-multimodal.
Assumptions
- Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support
Limitations
- Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology
