MEGA Hub

Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering

Authors

Do you know Zikui Cai?You can claim authorship or link another user.Do you know Kaushal Janga?You can claim authorship or link another user.Do you know Tan Dat Dao?You can claim authorship or link another user.Do you know Seungjae Lee?You can claim authorship or link another user.Do you know Shivin Dass?You can claim authorship or link another user.Do you know Mingyo Seo?You can claim authorship or link another user.Do you know Kaiyu Yue?You can claim authorship or link another user.Do you know Mintong Kang?You can claim authorship or link another user.Do you know Nandhu Pillai?You can claim authorship or link another user.Do you know Monte Hoover?You can claim authorship or link another user.Do you know Aadi Palnitkar?You can claim authorship or link another user.Do you know Ruchit Rawal?You can claim authorship or link another user.Do you know Ruijie Zheng?You can claim authorship or link another user.Do you know Bo Li?You can claim authorship or link another user.Do you know Yuke Zhu?You can claim authorship or link another user.Do you know Roberto Martín-Martín?You can claim authorship or link another user.Do you know Tom Goldstein?You can claim authorship or link another user.Do you know Furong Huang?You can claim authorship or link another user.

Abstract

Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. However, real-world robots operate continuously and must accumulate, retain, and selectively reuse information acquired from prior interactions. Despite this practical requirement, the architectural mechanisms needed to support sequential memory in EQA remain underexplored. In this work, we investigate how different memory architectures behave when EQA agents are evaluated sequentially, with multiple questions answered in the same scene while memory is carried forward across queries. We find that simply preserving existing memory is often insufficient. Agents that retain only traversability information, such as 2D occupancy maps, remember where the robot has explored but not the visual-semantic evidence needed for later questions. Agents trained on short-horizon episodic data face a different challenge: when exposed to continuous, multi-query histories, their inherited context suffers from severe temporal mismatch, rather than forming a reusable scene representation. To overcome this architectural bottleneck, we highlight the necessity of structured, spatially grounded memory: architectures that map persistent visual observations onto metric 3D geometry preserve visual-semantic evidence in a coherent scene representation. Extensive experiments in simulated environments reveal that this form of memory breaks the accuracy-efficiency tradeoff in sequential settings, simultaneously achieving higher answer accuracy and lower navigation costs. We further validate these findings on a real-world mobile robot, demonstrating that spatially grounded visual memory is critical for enabling continuous, intelligent operation in physical environments.

Community

00

Publication notes

Author note
Accepted to IROS 2026