MEGA Hub

DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering

Authors

Do you know Yue Zhang?You can claim authorship or link another user.Do you know Xiangyu Li?You can claim authorship or link another user.Do you know Wanshu Fan?You can claim authorship or link another user.Do you know Xin Yang?You can claim authorship or link another user.Do you know Dongsheng Zhou?You can claim authorship or link another user.

Abstract

Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propose a question-guided dynamic regional-level module (QGDR) that combines complementary image context through ROI Align and text content, enabling precise localization of text-related visual content. Moreover, we present a cross-modal multi-scale aggregation module (CMA) that enhances feature fusion between pixel-level and region-level features, facilitating the effective localization of visual content associated with grounded answers. Furthermore, we fuse the located visual content with text features to locate the region and provide answers to questions posed about the image. Experimental results demonstrate that our DDVT outperforms state-of-the-art methods on several widely-used benchmarks.

Community

00

Publication notes

Author note
Accepted by CGI