> ML_LITERATURE // KWON-2023-EFFICIENT-MEMORY-MANAGEMENT-LARGE-LANGUAGE-MODEL-SERVING-PAGEDATTENTION_v1.0
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica · ACM Symposium on Operating Systems Principles (SOSP) (2023)
systems2023industry-standardthirdPartyReproduced
Principal Contribution
Introduced PagedAttention inspired by OS virtual memory paging, eliminating KV-cache memory fragmentation and enabling 2-4x higher serving throughput.
Operational Relevance
The core innovation powering vLLM, the industry standard high-throughput open-source LLM inference serving engine.
Assumptions
- KV cache memory can be partitioned into non-contiguous physical memory blocks indexed by page tables without impacting attention correctness
Limitations
- Page size tuning balances internal fragmentation vs page table overhead; complex CUDA memory management
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
