MEGA Hub

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Authors

Do you know Wenqi Liu?You can claim authorship or link another user.Do you know Shijie Ma?You can claim authorship or link another user.Do you know Yunxiao Wang?You can claim authorship or link another user.Do you know Meng Liu?You can claim authorship or link another user.Do you know Qile Su?You can claim authorship or link another user.Do you know Han Liu?You can claim authorship or link another user.Do you know Bohan Hou?You can claim authorship or link another user.Do you know Xuanyu Zheng?You can claim authorship or link another user.Do you know Changyi Liu?You can claim authorship or link another user.Do you know Tianke Zhang?You can claim authorship or link another user.Do you know Haonan Fan?You can claim authorship or link another user.Do you know Kaiyu Jiang?You can claim authorship or link another user.Do you know Yingxin Li?You can claim authorship or link another user.Do you know Jiankang Chen?You can claim authorship or link another user.Do you know Xu Wang?You can claim authorship or link another user.Do you know Bin Wen?You can claim authorship or link another user.Do you know Tingting Gao?You can claim authorship or link another user.Do you know Han Li?You can claim authorship or link another user.Do you know Jianhua Yin?You can claim authorship or link another user.Do you know Yinwei Wei?You can claim authorship or link another user.Do you know Xuemeng Song?You can claim authorship or link another user.

Abstract

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are typically developed in isolation. We introduce VideoRover, a unified Video Deep Research framework that iteratively coordinates video cropping, multimodal search, and webpage browsing. Given a video-question pair, VideoRover uses each tool result to select the next action, so localized video clips guide external retrieval and retrieved evidence triggers further video inspection and verification. To develop this capability, we construct an automated data curation pipeline, producing 26K verified SFT trajectories and 3K challenging RL instances. We also introduce VideoRover-Bench, a benchmark stratified by video duration and research difficulty. Experiments on VideoDR and VideoRover-Bench show that our VideoRover-8B-RL achieves performance comparable to proprietary models in the direct-answer setting without tool use while outperforming larger open-source models equipped with the same tool suite. Ablation studies and training dynamics further validate the complementary roles of active video grounding, external retrieval, and long-horizon reinforcement learning.

Community

00