MEGA Hub

ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

Authors

Do you know Yuanhao Sun?You can claim authorship or link another user.Do you know Huawei Ji?You can claim authorship or link another user.Do you know Jiaxin Ding?You can claim authorship or link another user.Do you know Luoyi Fu?You can claim authorship or link another user.Do you know Xinbing Wang?You can claim authorship or link another user.

Abstract

Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in lightweight VLMs. Existing methods only focus on the visual modality and fail to dynamically preserve the integrity of prompt-relevant regions, limiting performance. In this work, we observe that the early-layer image-text entropy of cross-modal attention strongly correlates with answer grounding quality and task accuracy. Building on this finding, we propose \textbf{ENCORE}, an entropy-guided framework with two components: At inference, an \textbf{Entropy-based Cropping Strategy} (ECS) evaluates a small set of candidate crops and selects the one with minimal entropy, preserving contiguous regions relevant to the prompt. At training, \textbf{Entropy Regularization Training} (ERT) augments next-token prediction with an entropy term that sharpens attention on key visual tokens while down-weighting irrelevant ones. Experiments on ten VQA benchmarks show that ENCORE, fine-tuning only 0.14\% of parameters, achieves an average 1.43\% accuracy gain and state-of-the-art performance among recent 2B-parameter VLMs. Our code is released in https://github.com/baokou-fw2/ENCORE.

Community

00

Publication notes

Journal
ICASSP 2026
DOI
10.1109/ICASSP55912.2026.11463726