2.2b: FlashAttention — Online Softmax
Blog post from Hugging Face
FlashAttention’s online softmax computes exact, numerically stable attention incrementally, avoiding the need to materialize a full attention matrix or hold all scores in fast memory. Standard softmax subtracts the maximum score, denoted m, before exponentiation to prevent overflow, while l denotes the resulting sum of exponentials used for normalization. When scores are processed in blocks, the running maximum and normalization sum can be updated exactly; if a later block introduces a larger maximum, prior contributions must be rescaled by exp(m_old − m_new) so they share the new reference point. For attention, the same rescaling applies both to l and to the accumulated unnormalized weighted-value output O, ensuring that final normalization O/l matches standard softmax attention. Each query maintains independent m, l, and O state, allowing parallel blockwise processing while preserving mathematical equivalence to conventional attention.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.