MEGA Hub

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

Authors

Do you know Jiahui Han?You can claim authorship or link another user.Do you know Yuhui Yao?You can claim authorship or link another user.Do you know Xin Wang?You can claim authorship or link another user.Do you know Jiafei Cao?You can claim authorship or link another user.Do you know Mingxuan Zhang?You can claim authorship or link another user.Do you know Danfeng Shan?You can claim authorship or link another user.Do you know Huiqi Deng?You can claim authorship or link another user.Do you know Guanchu Wang?You can claim authorship or link another user.Do you know Xia Hu?You can claim authorship or link another user.

Abstract

Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for VLA models. DURA supports both white-box and black-box attack settings, where the black-box setting requires only the predicted actions of the victim model. By optimizing along the latent trajectory of a pretrained diffusion model, DURA generates visually natural patches while steering the robot toward attacker-specified target actions. Extensive experiments in both simulation and the real physical world show that DURA consistently outperforms existing methods. Our findings expose a safety risk for physically deployed VLA models and call for stronger defenses.

Community

00