MEGA Hub

Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos

Authors

Do you know Bingyi Xia?You can claim authorship or link another user.Do you know Han Bao?You can claim authorship or link another user.Do you know Zhewei Chen?You can claim authorship or link another user.Do you know Hanjing Ye?You can claim authorship or link another user.Do you know Jingwen Yu?You can claim authorship or link another user.Do you know Yuhan Pang?You can claim authorship or link another user.Do you know Wenjun Xu?You can claim authorship or link another user.Do you know Jiankun Wang?You can claim authorship or link another user.

Abstract

Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific data collection and the limited coverage of rare yet safety-critical scenarios. To address these challenges, we present a scalable framework for learning point-goal urban navigation from web-scale in-the-wild egocentric videos while systematically exposing its long tail. The framework automatically annotates uncurated web videos with metric trajectories and structured navigation semantics, which are then used to train a vision-language-action policy for interpretable navigation planning. We characterize the long tail based on model performance and the distribution of perception-motion patterns, and employ reflection-based analysis to diagnose recurring failure modes. Experiments on web-video data and real-world urban navigation tasks demonstrate effective knowledge transfer from unconstrained videos and reveal coherent long-tail structures beyond aggregate navigation performance.

Community

00