Skip to main content

> ML_LITERATURE // KWON-2023-EFFICIENT-MEMORY-MANAGEMENT-LARGE-LANGUAGE-MODEL-SERVING-PAGEDATTENTION_v1.0

Efficient Memory Management for Large Language Model Serving with PagedAttention

Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica · ACM Symposium on Operating Systems Principles (SOSP) (2023)

systems2023industry-standardthirdPartyReproduced

Principal Contribution

Introduced PagedAttention inspired by OS virtual memory paging, eliminating KV-cache memory fragmentation and enabling 2-4x higher serving throughput.

Operational Relevance

The core innovation powering vLLM, the industry standard high-throughput open-source LLM inference serving engine.

Assumptions

  • KV cache memory can be partitioned into non-contiguous physical memory blocks indexed by page tables without impacting attention correctness

Limitations

  • Page size tuning balances internal fragmentation vs page table overhead; complex CUDA memory management

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: