MEGA Hub

MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

Authors

Do you know Jianlin Yu?You can claim authorship or link another user.Do you know Jing Lin?You can claim authorship or link another user.Do you know Linghui Kong?You can claim authorship or link another user.Do you know Aiyue Chen?You can claim authorship or link another user.Do you know Weiyi Sun?You can claim authorship or link another user.Do you know Chenyu Zeng?You can claim authorship or link another user.Do you know Wangli Lan?You can claim authorship or link another user.Do you know Jinxi Li?You can claim authorship or link another user.Do you know Zhuo Zheng?You can claim authorship or link another user.Do you know Ziyang Yue?You can claim authorship or link another user.Do you know Danning Ke?You can claim authorship or link another user.Do you know Fei Yi?You can claim authorship or link another user.Do you know Tianchi Hu?You can claim authorship or link another user.Do you know Yuan Ding?You can claim authorship or link another user.Do you know Yiwu Yao?You can claim authorship or link another user.Do you know Junsong Wang?You can claim authorship or link another user.

Abstract

The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quantization (PNQ), which quantizes unnormalized softmax exponentials before row-wise summation to preserve normalization by construction. Experiments on Wan2.2 and HunyuanVideo show that MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, substantially improves frame-level similarity, and preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics. MXAttention also achieves performance competitive with strong NVFP4-based baselines with negligible overhead when fused into the attention pipeline. The implementation is publicly available in MindIE-SD.

Community

00