Skip to main content

> ML_LITERATURE // DAO-2022-FLASHATTENTION-FAST-MEMORY-EFFICIENT-EXACT-ATTENTION_v1.0

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré · Advances in Neural Information Processing Systems (NeurIPS) (2022)

hardware-compiler2022industry-standardthirdPartyReproduced

Principal Contribution

Tiled exact self-attention computation between fast GPU SRAM and slow HBM using online softmax renormalization, eliminating quadratic O(N^2) memory reads/writes.

Operational Relevance

Integrated into PyTorch (scaled_dot_product_attention), vLLM, Hugging Face, TensorRT-LLM, and every modern LLM training run.

Assumptions

  • GPU compute FLOPs are abundant, but Memory IO (HBM bandwidth) is the fundamental bottleneck in transformer self-attention

Limitations

  • Requires specialized hand-tuned CUDA/Triton kernels targeting specific GPU tensor core microarchitectures (Ampere, Hopper)

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: