> ML_LITERATURE // TOUVRON-2023-LLAMA-OPEN-EFFICIENT-FOUNDATION-MODELS_v1.0
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, Guillaume Lample · arXiv preprint (2023)
seminal-architecture2023industry-standardthirdPartyReproduced
Principal Contribution
Proved that smaller models (7B-65B) trained on more tokens (1.4T) outperform oversized models at inference, democratizing open weights.
Operational Relevance
The foundational catalyst for the open-source LLM ecosystem: llama.cpp, vLLM, Ollama, fine-tuning, and on-premise AI deployments.
Assumptions
- Hoffmann et al. (Chinchilla) compute-optimal frontier was conservative; training models far beyond compute-optimal tokens yields massive inference efficiency
Limitations
- Initial release contained raw pre-trained models with toxic web text reflections prior to Llama-2 chat alignment
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
