Skip to main content

> ML_LITERATURE // DUBEY-2024-LLAMA-3-HERD-OF-MODELS_v1.0

The Llama 3 Herd of Models

Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aravind Srinivas, Arno Candel, Binh Tang, Bryan Catanzaro, Chao Zhou, Chloe Bi, Chris Marra, Dan Bikel, David So, Dhruv Mahajan, Eric Michael Smith, Gautier Izacard, Guillaume Lample, Hakan Inan, Iliyan Zarov, Ishita Dasgupta, Jian Xiang Kuan, Joe Spisak, Jon Carvill, Kalyan Saladi, Kevin Stone, Lukas Blecher, Louis Martin, Mark Tygert, Melanie Kambadur, Naman Goyal, Nikhil Kant, Pushkar Mishra, Rashi Rungta, Ross Taylor, Sainbayar Sukhbaatar, Sergey Edunov, Thomas Scialom, Xavier Martinet, Yuning Mao · arXiv preprint (2024)

seminal-architecture2024industry-standardthirdPartyReproduced

Principal Contribution

Trained open foundation models up to 405B dense parameters on 15.6T multilingual tokens with native multimodal support and 128K context window.

Operational Relevance

Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation, task-code-generation, task-multimodal.

Assumptions

  • Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support

Limitations

  • Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: