> ML_LITERATURE // DAO-2022-FLASHATTENTION-FAST-MEMORY-EFFICIENT-EXACT-ATTENTION_v1.0
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré · Advances in Neural Information Processing Systems (NeurIPS) (2022)
hardware-compiler2022industry-standardthirdPartyReproduced
Principal Contribution
Tiled exact self-attention computation between fast GPU SRAM and slow HBM using online softmax renormalization, eliminating quadratic O(N^2) memory reads/writes.
Operational Relevance
Integrated into PyTorch (scaled_dot_product_attention), vLLM, Hugging Face, TensorRT-LLM, and every modern LLM training run.
Assumptions
- GPU compute FLOPs are abundant, but Memory IO (HBM bandwidth) is the fundamental bottleneck in transformer self-attention
Limitations
- Requires specialized hand-tuned CUDA/Triton kernels targeting specific GPU tensor core microarchitectures (Ampere, Hopper)
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
