Good Papers

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention applies OS virtual memory and paging to LLM key-value caches, reducing waste and duplication to boost vLLM throughput 2-4x over existing systems.

Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Gabriel Stoica

Published Sep 12, 202350 citations▲ 76 on Hugging FaceCode ★ 86,094arXiv ↗

89%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel16/21reviewers recommend it
lenient 5/5
medium 8/11
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
PagedAttention's OS-inspired paging nearly eliminates KV cache waste and enables flexible sharing for large gains, though its 2-4x throughput claims depend on long sequences and large batches, and independent multi-tenant audits of that near-zero waste remain…

Abstract

High throughput serving of large language models (LLMs) requires batching sufficiently many requests at a time. However, existing systems struggle because the key-value cache (KV cache) memory for each request is huge and grows and shrinks dynamically. When managed inefficiently, this memory can be significantly wasted by fragmentation and redundant duplication, limiting the batch size. To address this problem, we propose PagedAttention, an attention algorithm inspired by the classical virtual memory and paging techniques in operating systems. On top of it, we build vLLM, an LLM serving system that achieves (1) near-zero waste in KV cache memory and (2) flexible sharing of KV cache within and across requests to further reduce memory usage. Our evaluations show that vLLM improves the throughput of popular LLMs by 2-4$\times$ with the same level of latency compared to the state-of-the-art systems, such as FasterTransformer and Orca. The improvement is more pronounced with longer sequences, larger models, and more complex decoding algorithms. vLLM's source code is publicly available at https://github.com/vllm-project/vllm