MEGA Hub

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

Authors

Do you know Shanghao Liu?You can claim authorship or link another user.Do you know Renze Chen?You can claim authorship or link another user.Do you know Size Zheng?You can claim authorship or link another user.Do you know Yuanqiang Liu?You can claim authorship or link another user.Do you know Yun?You can claim authorship or link another user.Do you know Liang?You can claim authorship or link another user.Do you know Hailong Yang?You can claim authorship or link another user.

Abstract

Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end-to-end gains. We present SPADE, a training-free sparse-attention engine of three parts: (i) vDiT-SSR, a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions; (ii) runtime scheme generation using SICS and a head-wise policy; and (iii) an executor with low-overhead index search, flash block-sparse attention, and kernel grouping. Across Hunyuan-Video and Wan 2.1/2.2 for text-to-video and image-to-video generation, SPADE raises sparsity and speed while preserving quality, accelerating attention by 2.26x-3.40x and end-to-end inference by 1.49x-1.80x. Our code is open-sourced at https://github.com/6somehow/DAC-SPADE.

Community

00

Publication notes

Author note
Published in the 63rd ACM/IEEE Design Automation Conference (DAC '26). 7 pages, 6 figures, 3 tables
Journal
63rd ACM/IEEE Design Automation Conference (DAC '26), 2026, 7 pages
DOI
10.1145/3770743.3804015