Skip to main content

> ML_ALGORITHM // AUTOREGRESSIVE-SELF-SUPERVISED-GPT_v1.0

Causal Autoregressive Next-Token Prediction (GPT)

The dominant self-supervised paradigm powering modern Large Language Models by training causal Transformer decoders on trillions of tokens via next-token prediction.

Autoregressive Generative Pre-trainingself-supervisedblack-boxmassive (>10M)
Back to All Algorithms
Computational Complexity
Training Complexity:O(tokens * params * 6 FLOPs)
Inference Complexity:O(kv_cache * layers) memory, O(params * 2 FLOPs) per token
Hardware Profile
CPU Friendly:No
Requires GPU:Yes
Memory Footprint:high
Interpretability & Data
Interpretability Tier:black-box
Training Data Needs:massive (>10M)

Interpretability Assessment

High capacity enables complex emergent chain-of-thought, but internal mechanistic interpretability remains an active research frontier.

Suitable Tasks & Supported Modalities

Suitable Tasks:
text generationcode generationin context learningreasoning
Supported Modalities:
textcodemultimodal

Implementing Libraries

vLLMvLLM Project / UC Berkeley · v0.6.2
View Spec
TransformersHugging Face · v4.44.2
View Spec
SGLangLMSYS Org / UC Berkeley · v0.3.1
View Spec
TensorRT-LLMNVIDIA · v0.12.0
View Spec

Foundational Literature

Improving Language Understanding by Generative Pre-Training (GPT-1)Alec Radford, Karthik Narasimhan (2018) · OpenAI Technical Report
Language Models are Few-Shot Learners (GPT-3)Tom B. Brown, Benjamin Mann (2020) · Advances in Neural Information Processing Systems (NeurIPS)
Common Pitfalls & Warnings
  • Assuming factual accuracy without grounding or retrieval; autoregressive generation optimizes token probability rather than veracity
  • Linear KV-cache expansion exhausting GPU VRAM without paging algorithms