MEGA Hub

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Authors

Do you know Zhen Fang?You can claim authorship or link another user.Do you know Yu Zeng?You can claim authorship or link another user.Do you know Wenxuan Huang?You can claim authorship or link another user.Do you know Yiming Zhao?You can claim authorship or link another user.Do you know Shiting Huang?You can claim authorship or link another user.Do you know Tianfei Ren?You can claim authorship or link another user.Do you know Qi Lu?You can claim authorship or link another user.Do you know Qingnan Ren?You can claim authorship or link another user.Do you know Qisheng Su?You can claim authorship or link another user.Do you know Lionel Z. Wang?You can claim authorship or link another user.Do you know Qingyu Yin?You can claim authorship or link another user.Do you know Shuang Chen?You can claim authorship or link another user.Do you know Zehui Chen?You can claim authorship or link another user.Do you know Lin Chen?You can claim authorship or link another user.Do you know Zhenfei Yin?You can claim authorship or link another user.Do you know Yao Hu?You can claim authorship or link another user.Do you know Shaohui Lin?You can claim authorship or link another user.Do you know Wanli Ouyang?You can claim authorship or link another user.Do you know Shaosheng Cao?You can claim authorship or link another user.Do you know Feng Zhao?You can claim authorship or link another user.

Abstract

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.

Community

00