MEGA Hub

SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense

Authors

Do you know Siyuan Liang?You can claim authorship or link another user.Do you know Yupeng Qiu?You can claim authorship or link another user.Do you know Junfeng Fang?You can claim authorship or link another user.Do you know Rong-Cheng Tu?You can claim authorship or link another user.Do you know Jiaxing Huang?You can claim authorship or link another user.Do you know Dacheng Tao?You can claim authorship or link another user.

Abstract

Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense approaches mainly rely on input filtering or reconstruction, which not only incur high computational latency but also tend to distort semantics. To address these issues, we experimentally and systematically analyze the differences between clean and jailbreak samples in the cross-attention feature space, revealing for the first time a cumulative separation effect and a progressively increasing trend of linear separability between the two during the diffusion process. Based on this insight, we propose SafeCA, a feature-level defense mechanism for safe cross-attention localization and regularization. Firstly, we identify key defensive regions and values through attention stability analysis using cross-attention features collected from clean prompts within a single inference. Secondly, SafeCA mitigates anomalous activations via attention masking with energy normalization and introduces a lightweight semantic-space adapter to redirect abnormal semantic flows. Furthermore, we detect and suppress potentially malicious tokens by back-propagating feature anomaly signals to the input cue words, thereby enhancing the deployability of the defense in commercial models. Experimental results show that SafeCA reduces the jailbreak success rate by about 20% on mainstream T2V models, adds almost no inference overhead (+0.1s), and maintains good text-video semantic consistency. Overall, SafeCA provides an architecture-level, deployable protection paradigm for T2V generation models.

Community

00

Publication notes

Author note
10 pages, 4 figures