MEGA Hub

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Authors

Do you know Ke Li?You can claim authorship or link another user.Do you know Jiayu Chen?You can claim authorship or link another user.Do you know Maoliang Li?You can claim authorship or link another user.Do you know Zihao Zheng?You can claim authorship or link another user.Do you know Hailong Zou?You can claim authorship or link another user.Do you know Hengyi Zhang?You can claim authorship or link another user.Do you know Xuanzhe Liu?You can claim authorship or link another user.Do you know Xiang Chen?You can claim authorship or link another user.

Abstract

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to search for additional candidate evidence. Specifically, entropy-gated budget scheduling uses output uncertainty to stop early when the current evidence is sufficient or progressively expand the frame budget otherwise. Meanwhile, attention-guided candidate proposal converts frame-level attention into a temporal prior, enabling dense local search in informative regions while preserving global coverage when attention is diffuse. Experiments on Video-MME, LongVideoBench, and MLVU demonstrate that EcoFrame achieves a better accuracy--efficiency trade-off across multiple VLM backbones. On Qwen2.5-VL, EcoFrame achieves an average accuracy of 64.4, surpassing BOLT at 63.5, while providing a $1.85\times$ speedup over AKS and BOLT. Compared with the agent-based A.I.R., EcoFrame maintains comparable accuracy with up to a $13.5\times$ inference speedup. Code will be available at https://github.com/AK-DREAM/EcoFrame.

Community

00