> ML_ARCHITECTURE // AUTOREGRESSIVE-DECODER-TRANSFORMER_v1.0
Autoregressive Decoder-Only Transformer (GPT / LLaMA / Mistral)
The dominant foundation model architecture for generative text and code, processing sequences with causal attention masks, RoPE positional embeddings, SwiGLU activations, and RMSNorm.
Transformerstextcode
Back to All ArchitecturesArchitecture Overview
The dominant foundation model architecture for generative text and code, processing sequences with causal attention masks, RoPE positional embeddings, SwiGLU activations, and RMSNorm.
Implementing Libraries
Seminal Papers
Improving Language Understanding by Generative Pre-Training (GPT-1)Alec Radford, Karthik Narasimhan (2018) · OpenAI Technical Report
LLaMA: Open and Efficient Foundation Language ModelsHugo Touvron, Thibaut Lavril (2023) · arXiv preprint
Mistral 7BAlbert Q. Jiang, Alexandre Sablayrolles (2023) · arXiv preprint
Architectural Limitations & Constraints
- Requires compatible deep learning framework and hardware acceleration for efficient execution.
