> ML_LITERATURE // SHAZEER-2019-FAST-TRANSFORMER-DECODING-MULTI-QUERY-ATTENTION_v1.0
Fast Transformer Decoding: One Write-Head is All You Need
Noam Shazeer · arXiv preprint (2019)
algorithm2019industry-standardthirdPartyReproduced
Principal Contribution
Invented Multi-Query Attention (MQA), sharing a single Key and Value head across all Query heads to dramatically eliminate KV cache memory bandwidth bottlenecks during autoregressive generation.
Operational Relevance
Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation.
Assumptions
- Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support
Limitations
- Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
