MEGA Hub

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

Authors

Do you know Junyi Hu?You can claim authorship or link another user.Do you know Tian Bai?You can claim authorship or link another user.Do you know Fengyi Wu?You can claim authorship or link another user.Do you know Yian Huang?You can claim authorship or link another user.Do you know Wei Wen?You can claim authorship or link another user.Do you know Zaoli Li?You can claim authorship or link another user.Do you know Junli Lin?You can claim authorship or link another user.Do you know Xingchen Li?You can claim authorship or link another user.Do you know Zhenming Peng?You can claim authorship or link another user.Do you know Yi Zhang?You can claim authorship or link another user.

Abstract

Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective of unified open-vocabulary grounding and identify representation degeneration as a key obstacle to scaling a single generalist model. To preserve representation diversity, we propose a holistic data-model co-design framework. Architecturally, we introduce the Modulated Attention-Contrastive Head (mACH) for efficient token-level vision-language alignment and a text-conditioned JEPA auxiliary stream that provides complementary gradient support to preserve alignment-active representations without inference overhead. On the data side, we introduce Objects365-Caption, enriching Objects365 with context-aware referring expressions for large-scale language supervision. We further provide a theoretical analysis showing that complementary gradient subspaces preserve alignment capacity and thereby scale representation diversity. Extensive experiments demonstrate that our single-checkpoint framework achieves highly competitive performance on standard REC benchmarks while exhibiting strong generalization across heterogeneous grounding datasets without benchmark-specific adaptation.

Community

00

Publication notes

Author note
21 pages, 10 figures, 6 tables