Historical milestone · 2023
PagedAttention and vLLM
Reviewed through September 18, 2026
2023 · Historical milestone
PagedAttention and vLLM
- Era
- 2020s
- Theme
- Infrastructure, efficiency & open ecosystems
- Evidence form
- Systems paper + serving benchmarks
- School / paradigm
- Inference serving / memory management
- Institution / context
- UC Berkeley
- Researchers
- Woosuk Kwon; collaborators
Understand
Plain-language record, transferred from the reviewed source module.
Theory or experimental setup. Adapted virtual-memory paging to transformer key-value caches and built a high-throughput serving system.
Result / historical claim. Reported two-to-four-times throughput at comparable latency in tested serving workloads.
Apply
Professional implication, only where the reviewed record states one.
The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: Infrastructure, efficiency, and open ecosystems.
Verify
Evidence status, stated limitations, and the external sources this record actually carries.
Evidence form. Systems paper + serving benchmarks
Limitation / debate. Gains depend on request mix, hardware, model, batching, and software version and do not alter model quality.
Source status. The source link below is the verified link our reviewed topic research already carries for this milestone.
Reproduce
A reproduction tutorial is linked only when one exists for this exact record.
A reproduction tutorial is not yet available for this entry. The closest reviewed material is Infrastructure, efficiency, and open ecosystems.
Cite or share
APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.
BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.
