Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs
Blog post from Google Cloud
Video diffusion inference is slowed by many denoising steps and especially by quadratic self-attention costs at long video sequence lengths, making attention an increasingly dominant bottleneck as resolution rises. The Sparse VideoGen approach exploits the structured behavior of attention heads, dynamically routing them to spatial or temporal sparse masks that preserve important interactions, including first-frame context, while eliminating many low-value query-key pairs. On TPU hardware, however, logical sparsity alone initially performed worse than dense attention because elementwise masking within visited compute tiles imposed substantial overhead; performance improved by separating fully valid tiles from boundary tiles, tuning tile sizes, and slightly aligning mask boundaries to hardware tiles. These kernel changes reduced single-device sparse attention latency from 78.70 ms for dense attention to 32.76 ms, a 2.40× speedup. Efficient integration also required static tensor shapes for dynamic head routing, temporal token permutation to make temporal access contiguous, and applying layout changes after distributed head exchanges to avoid excess communication. In end-to-end generation on eight TPU v6e chips, aggressive sparsity reduced 720p denoising time by 22% to 119.86 seconds, while larger resolutions yielded greater gains: 1.49× at 1080p and 1.69× at 1440p, saving more than 16 minutes per video while maintaining comparable output similarity.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| TPUs | 14 | 4 | 2 | 1 | -92% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.