MEGA Hub

SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation

Authors

Do you know Yi Wu?You can claim authorship or link another user.Do you know Junjie An?You can claim authorship or link another user.Do you know Xiao Liu?You can claim authorship or link another user.Do you know Yiqun Zhou?You can claim authorship or link another user.Do you know Yuechen Wu?You can claim authorship or link another user.Do you know Xiaoqing Guan?You can claim authorship or link another user.Do you know Shuyang Yu?You can claim authorship or link another user.Do you know You Wang?You can claim authorship or link another user.Do you know Guang Li?You can claim authorship or link another user.

Abstract

In goal-directed embodied navigation, where an agent must locate a specified target in an unseen environment, 3D scene understanding and navigation reasoning must work in concert. Current approaches transmit 3D scene information to vision-language models (VLMs) through text, suggesting a representation gap in our tested configurations; a controlled ablation confirms that direct embedding-level transfer significantly outperforms the evaluated text serialization formats. We introduce SoftNav, which injects entity-level 3D continuous representations -- one token per detected object or frontier -- into a VLM's hidden space as soft tokens through a lightweight projector. With the 3D encoder and VLM frozen, only ~1,200 samples and ~17M trainable parameters are needed. On HM3D-OVON, SoftNav achieves 74.2%/68.3%/66.7% SR across three splits, surpassing all prior methods in both SR and SPL; the same navigation policy transfers zero-shot to GOAT-Bench (67.2% SR), SG3D (47.2% s-SR), and real-world robot deployment without retraining or architectural modification. Injecting 3D scene tokens directly into VLMs bridges the representation gap, enabling transferable navigation with minimal training.

Community

00

Publication notes

Author note
8 pages, 6 figures. Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026