MEGA Hub

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Authors

Do you know Huanyao Zhang?You can claim authorship or link another user.Do you know Jiepeng Zhou?You can claim authorship or link another user.Do you know Runhao Zhao?You can claim authorship or link another user.Do you know Yanzhe Shan?You can claim authorship or link another user.Do you know Jiaoyang Chen?You can claim authorship or link another user.Do you know Bowen Zhou?You can claim authorship or link another user.Do you know Bo Li?You can claim authorship or link another user.Do you know Fang Wang?You can claim authorship or link another user.Do you know Jialong Wu?You can claim authorship or link another user.Do you know Zhengwei Tao?You can claim authorship or link another user.Do you know Lang Mei?You can claim authorship or link another user.Do you know Xiaohan Yu?You can claim authorship or link another user.Do you know Liyan Liu?You can claim authorship or link another user.Do you know Chong Chen?You can claim authorship or link another user.Do you know Wentao Zhang?You can claim authorship or link another user.

Abstract

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.

Community

00