MEGA Hub

Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering

Authors

Do you know Hsiang-Wei Huang?You can claim authorship or link another user.Do you know Fu-Chen Chen?You can claim authorship or link another user.Do you know Li-Wu Tsao?You can claim authorship or link another user.Do you know Cheng-Han Lee?You can claim authorship or link another user.Do you know Che-Chun Su?You can claim authorship or link another user.Do you know Lu Xia?You can claim authorship or link another user.Do you know Ronghui Peng?You can claim authorship or link another user.Do you know Jenq-Neng Hwang?You can claim authorship or link another user.Do you know Min Sun?You can claim authorship or link another user.Do you know Cheng-Hao Kuo?You can claim authorship or link another user.

Abstract

Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input. However, visual search methods are inefficient because they require visual search among thousands of video frames for each individual user query. In this work, we propose a memory tree guided key frame selection paradigm for efficient 3D question answering in embodied scenarios. Our method leverages a compact and reusable 3D scene representation, termed MemTree3D, which supports real-time online construction leveraging camera 6-DoF poses. MemTree3D captures multi-level 3D scene information, enabling a Large Language Model to efficiently query and retrieve question-relevant key frames through our scoring-based frame selection without reprocessing the entire video stream. On OpenEQA, our method improves the LLM-Match of GPT-4o by 17.4%, LLaVA-OneVision-7B by 5.8%, outperforms existing visual search methods. Our code is available at https://github.com/hsiangwei0903/MemTree3D

Community

00

Publication notes

Author note
ECCV 2026