MEGA Hub

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

Authors

Do you know Xuanhui Lin?You can claim authorship or link another user.Do you know Junhao Dong?You can claim authorship or link another user.Do you know Mingrong Gong?You can claim authorship or link another user.Do you know Yucheng Chen?You can claim authorship or link another user.Do you know Xinghua Qu?You can claim authorship or link another user.Do you know Yew-Soon Ong?You can claim authorship or link another user.

Abstract

Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attacks typically follow single-trajectory gradient optimization or task-specific objectives, limiting search-space exploration and cross-task transferability. We propose an evolutionary-computation-guided cross-modal attack framework for unified VLMs. The framework adaptively searches both textual and visual spaces. On the textual side, it evolves hard negative semantic embeddings around the source-category representation to provide diverse cross-modal repulsion. On the visual side, it maintains a population of object-region perturbations and combines momentum-based gradient updates with evolutionary selection, mutation, and crossover to more reliably explore multiple feasible trajectories. Jointly optimizing semantic negative guidance and localized perturbations generates adversarial examples that consistently shift source-object semantics toward target categories across vision-language tasks. Theoretical analyses show that the co-evolutionary search preserves perturbation feasibility, prevents degradation of the best observed fitness, and increases the probability of reaching high-margin adversarial regions compared with single-trajectory optimization. Experiments on Florence-2, OFA, and UnifiedIO-2 demonstrate strong overall attack performance across image captioning, object detection, region categorization, and object localization. Ablation studies further verify the complementary effectiveness of text-side semantic evolution and image-side perturbation evolution, as well as the framework's efficiency and cross-task transferability.

Community

00

Publication notes

Author note
15 pages, 7 figures, and 8 tables; includes supplementary material