MEGA Hub

On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

Authors

Do you know Yuchen Ren?You can claim authorship or link another user.Do you know Zhengyu Zhao?You can claim authorship or link another user.Do you know Chenhao Lin?You can claim authorship or link another user.Do you know Bo Yang?You can claim authorship or link another user.Do you know Chao Shen?You can claim authorship or link another user.

Abstract

Vision-Language Pre-training Models (VLPMs) are known to be vulnerable to adversarial attacks. Recent transferable attacks on VLPMs have followed a common pipeline with complicated loss functions or multi-stage text/image attacks. However, in this paper, we demonstrate that such a sophisticated attack pipeline can be simpler yet more successful. Specifically, we identify three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations. To address them, we propose the Simple Vision-Language Attack (SimVLA) pipeline, which observably improves transferability and efficiency. Experiments on four datasets and three downstream tasks validate the superiority of our pipeline. For instance, on Flickr30k text-image retrieval dataset, our SimVLA outperforms the SOTA baseline in R@1 transferability by 8.01\%-14.71\%, while consuming only about 35.73\% of the time and 46.26\% of the max VRAM. Overall, the superiority of our SimVLA highlights the importance of leveraging domain knowledge (e.g., our proposed cross-modal word identification), while blindly pursuing intricate operations (e.g, complex loss functions and redundant multi-stage designs) may even be harmful. We hope our SimVLA can serve as a simple yet effective backbone for future extensions. Code is available at https://github.com/RYC-98/SimVLA.

Community

00

Publication notes

Author note
Accepted for publication in IEEE Transactions on Information Forensics and Security (TIFS)
DOI
10.1109/TIFS.2026.3714129