Skip to main content

> ML_ARCHITECTURE // AUTOREGRESSIVE-DECODER-TRANSFORMER_v1.0

Autoregressive Decoder-Only Transformer (GPT / LLaMA / Mistral)

The dominant foundation model architecture for generative text and code, processing sequences with causal attention masks, RoPE positional embeddings, SwiGLU activations, and RMSNorm.

Transformerstextcode
Back to All Architectures

Architecture Overview

The dominant foundation model architecture for generative text and code, processing sequences with causal attention masks, RoPE positional embeddings, SwiGLU activations, and RMSNorm.

Implementing Libraries

TransformersHugging Face · v4.44.2
View Spec
vLLMvLLM Project / UC Berkeley · v0.6.2
View Spec
SGLangLMSYS Org / UC Berkeley · v0.3.1
View Spec
llama.cppGeorgi Gerganov / Open Source · vb3650
View Spec

Seminal Papers

Improving Language Understanding by Generative Pre-Training (GPT-1)Alec Radford, Karthik Narasimhan (2018) · OpenAI Technical Report
LLaMA: Open and Efficient Foundation Language ModelsHugo Touvron, Thibaut Lavril (2023) · arXiv preprint
Mistral 7BAlbert Q. Jiang, Alexandre Sablayrolles (2023) · arXiv preprint
Architectural Limitations & Constraints
  • Requires compatible deep learning framework and hardware acceleration for efficient execution.