Historical milestone · 2022
FlashAttention
Reviewed through September 18, 2026
2022 · Historical milestone
FlashAttention
- Era
- 2020s
- Theme
- Infrastructure, efficiency & open ecosystems
- Evidence form
- Algorithm + hardware benchmarks
- School / paradigm
- IO-aware algorithms / accelerator kernels
- Institution / context
- Stanford University; University at Buffalo
- Researchers
- Tri Dao; collaborators
Understand
Plain-language record, transferred from the reviewed source module.
Theory or experimental setup. Tiled exact attention to reduce reads and writes between accelerator memory levels.
Result / historical claim. Improved speed and memory use on tested GPUs without approximating the attention result.
Apply
Professional implication, only where the reviewed record states one.
The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: Infrastructure, efficiency, and open ecosystems.
Verify
Evidence status, stated limitations, and the external sources this record actually carries.
Evidence form. Algorithm + hardware benchmarks
Limitation / debate. Left quadratic arithmetic and hardware- and kernel-specific constraints in place.
Source status. The source link below is the verified link our reviewed topic research already carries for this milestone.
Reproduce
A reproduction tutorial is linked only when one exists for this exact record.
A reproduction tutorial is not yet available for this entry. The closest reviewed material is Infrastructure, efficiency, and open ecosystems.
Cite or share
APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.
BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.
