Presentation
SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
DescriptionVideo diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end-to-end gains. We present \textsc{SPADE}, a training-free sparse-attention engine of three parts: (i) \textsc{vDiT-SSR}, specification defining 3D blocking candidates and formalizing dynamic masks via \emph{Summarizer}/\emph{Estimator} expressions; (ii) runtime \textsc{scheme generation} using SICS and a head-wise policy; and (iii) an executor with low-overhead index search, flash block-sparse attention, and kernel grouping. Across Hunyuan-Video, Wan~2.1/2.2, \textsc{SPADE} raises sparsity and speed and preserves quality, accelerating attention $2.26{\times}$--$3.40{\times}$ and end-to-end $1.49{\times}$--$1.80{\times}$.
Event Type
Research Manuscript
TimeMonday, July 275:18pm - 5:30pm PDT
LocationMtg Room 101B
Similar Presentations
