MEGA Hub

VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

Authors

Do you know Jiabin Lou?You can claim authorship or link another user.Do you know Haopeng Wang?You can claim authorship or link another user.Do you know Yuanshuai Wang?You can claim authorship or link another user.Do you know Xinyu Liu?You can claim authorship or link another user.Do you know Xuxin Lv?You can claim authorship or link another user.Do you know Yuxin Guo?You can claim authorship or link another user.Do you know Lei Huang?You can claim authorship or link another user.Do you know Rongye Shi?You can claim authorship or link another user.Do you know Wenjun Wu?You can claim authorship or link another user.

Abstract

Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encode spatial priors, such as orientation, distance, and layout, that are not explicitly available from onboard sensing at deployment in open, GPS-denied environments. Benchmark performance under such interfaces therefore jointly reflects visual navigation ability and the use of route structure explicitly supplied by the task description. As a complementary formulation, we propose Vision-Only Long-Horizon Navigation (VoLN), which shifts route-relevant information from externally supplied instructions and global guidance to locally observable in-scene cues. In VoLN, goal views specify the destination, while route-relevant information is available only through locally observable in-scene cues that the agent must detect, interpret, and select online. We instantiate VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection. We further provide VoLN-MLLM as an initial reference baseline. It aligns self-supervised visual features with a structured semantic space and predicts short-horizon waypoint segments from observation history, goal views, retrieved visual--semantic tokens, and proprioception. On the five-environment Test-Unseen split, it obtains success rates of 7.4%, 4.5%, and 1.8% on Easy, Normal, and Hard episodes, respectively. These results provide an initial evaluation of VoLN and reveal substantial remaining challenges in long-horizon evidence integration, cross-view goal matching, and closed-loop stability. Project page: https://admire-ljb.github.io/VoLN-UAV/

Community

00

Publication notes

Author note
10 pages, 7 figures, 2 tables. Project page: https://admire-ljb.github.io/VoLN-UAV/