Skip to main content

> ML_LITERATURE // SHAZEER-2019-FAST-TRANSFORMER-DECODING-MULTI-QUERY-ATTENTION_v1.0

Fast Transformer Decoding: One Write-Head is All You Need

Noam Shazeer · arXiv preprint (2019)

algorithm2019industry-standardthirdPartyReproduced

Principal Contribution

Invented Multi-Query Attention (MQA), sharing a single Key and Value head across all Query heads to dramatically eliminate KV cache memory bandwidth bottlenecks during autoregressive generation.

Operational Relevance

Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation.

Assumptions

  • Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support

Limitations

  • Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: