MEGA Hub

Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

Authors

Do you know Yi Tang?You can claim authorship or link another user.Do you know Xinyi Shang?You can claim authorship or link another user.Do you know Jiacheng Cui?You can claim authorship or link another user.Do you know Sondos Mahmoud Bsharat?You can claim authorship or link another user.Do you know Jiacheng Liu?You can claim authorship or link another user.Do you know Xiaohan Zhao?You can claim authorship or link another user.Do you know Tran Dinh Tien?You can claim authorship or link another user.Do you know Ahmed Elhagry?You can claim authorship or link another user.Do you know Salwa K. Al Khatib?You can claim authorship or link another user.Do you know Tianjun Yao?You can claim authorship or link another user.Do you know Yonina C. Eldar?You can claim authorship or link another user.Do you know Jing-Hao Xue?You can claim authorship or link another user.Do you know Hao Li?You can claim authorship or link another user.Do you know Salman Khan?You can claim authorship or link another user.Do you know Zhiqiang Shen?You can claim authorship or link another user.

Abstract

Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions. We propose a simple yet effective domain-generalized training framework built on two practical strategies. First, we introduce a balanced minibatch sampling scheme that strategically samples tampered and real images in each minibatch, preventing biased optimization toward either manipulated artifacts or clean-image priors and avoiding training collapse, ensuring that each optimization step receives proper sampled gradient signals. Second, we adopt a simple late-injection strategy, where the detector is first trained on large-scale base data until stable convergence, and then exposed to a small amount of newly selected supporting data from emerging VLM distributions, improving adaptability without overfitting to limited new domains. Together, these components provide a simple yet strong recipe for improving pixel-level tampering localization and OOD robustness across modern VLMs. Despite the conceptual simplicity, our framework outperforms the prior state-of-the-art PIXAR by a large margin of 26.1% and 26.8% relative improvement in average gIoU and cIoU, respectively, across OOD VLMs of GPT-Images-2.0, Gemini-3.1, FLUX.2, and Seedream 4.5. Our code is available at https://github.com/VILA-Lab/PIXAR-DG

Community

00

Publication notes

Author note
Our code is available at https://github.com/VILA-Lab/PIXAR-DG