MEGA Hub

FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

Authors

Do you know Jiatong Li?You can claim authorship or link another user.Do you know Leo Liang?You can claim authorship or link another user.Do you know Linghe Kong?You can claim authorship or link another user.Do you know Yulun Zhang?You can claim authorship or link another user.

Abstract

Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.

Community

00

Publication notes

Author note
Code is available at: https://github.com/jiatongli2024/FreqForcing