AI Inference & Serving
FlashAttention
An attention algorithm that reorganizes computation into tiles to reduce transfers between different levels of device memory. It computes exact attention in the algorithmic sense rather than replacing it with a sparse approximation. Floating-point rounding, supported masks, hardware and implementation details still matter; exact does not mean universally bit-for-bit identical.
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.