Skip to main content

> ML_LITERATURE // TOUVRON-2023-LLAMA-OPEN-EFFICIENT-FOUNDATION-MODELS_v1.0

LLaMA: Open and Efficient Foundation Language Models

Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, Guillaume Lample · arXiv preprint (2023)

seminal-architecture2023industry-standardthirdPartyReproduced

Principal Contribution

Proved that smaller models (7B-65B) trained on more tokens (1.4T) outperform oversized models at inference, democratizing open weights.

Operational Relevance

The foundational catalyst for the open-source LLM ecosystem: llama.cpp, vLLM, Ollama, fine-tuning, and on-premise AI deployments.

Assumptions

  • Hoffmann et al. (Chinchilla) compute-optimal frontier was conservative; training models far beyond compute-optimal tokens yields massive inference efficiency

Limitations

  • Initial release contained raw pre-trained models with toxic web text reflections prior to Llama-2 chat alignment

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: