MEGA Hub

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

Authors

Do you know Weitao Chen?You can claim authorship or link another user.Do you know Hu Jiaxin?You can claim authorship or link another user.Do you know Xie Tianyidan?You can claim authorship or link another user.Do you know Yang Li?You can claim authorship or link another user.Do you know Yuyi Qian?You can claim authorship or link another user.Do you know Banghao Xu?You can claim authorship or link another user.Do you know Ziheng Tang?You can claim authorship or link another user.Do you know Shenyi Wang?You can claim authorship or link another user.Do you know Mingyue Yu?You can claim authorship or link another user.Do you know Duo Li?You can claim authorship or link another user.Do you know Jiacheng Shi?You can claim authorship or link another user.Do you know Gao Wang?You can claim authorship or link another user.Do you know Zhan Xu?You can claim authorship or link another user.Do you know Zhicheng Qiu?You can claim authorship or link another user.Do you know Xuanfu Li?You can claim authorship or link another user.Do you know Jian Yang?You can claim authorship or link another user.Do you know Lanjun Wang?You can claim authorship or link another user.Do you know Zili Yi?You can claim authorship or link another user.

Abstract

Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.

Community

00

Publication notes

Author note
21 pages, 4 figures, 6 tables, including appendices