MEGA Hub

Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation

Authors

Do you know Qiaohui Chu?You can claim authorship or link another user.Do you know Haoyu Zhang?You can claim authorship or link another user.Do you know Meng Liu?You can claim authorship or link another user.Do you know Yisen Feng?You can claim authorship or link another user.Do you know Haoxiang Shi?You can claim authorship or link another user.Do you know Liqiang Nie?You can claim authorship or link another user.

Abstract

Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enables proactive and context-aware AI assistance. However, existing approaches suffer from three key limitations: 1) underutilization of fine-grained visual cues from hand-object interactions, 2) neglect of semantic dependencies between verbs and nouns, and 3) lack of explicit cognitive reasoning, limiting generalization and long-term forecasting ability. To overcome these challenges, we propose INSIGHT, a unified two-stage framework for egocentric action anticipation. In the first stage, INSIGHT focuses on extracting semantically rich features from hand-object interaction regions and enhances action representations using a verb-noun co-occurrence matrix. In the second stage, it introduces a reinforcement learning-based module that simulates explicit cognitive reasoning through a structured process: visual perception (think) -> intention inference (reason) -> action anticipation (answer). Extensive experiments on Ego4D, EPIC-Kitchens-55, and EGTEA Gaze+ benchmarks show that INSIGHT achieves state-of-the-art performance, demonstrating its effectiveness and strong generalization capability.

Community

00

Publication notes

Author note
Accepted by AAAI-2026