MEGA Hub

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies

Authors

Do you know Jiarui Yang?You can claim authorship or link another user.Do you know Yehao Lu?You can claim authorship or link another user.Do you know Yuning Su?You can claim authorship or link another user.Do you know Yufeng Xie?You can claim authorship or link another user.Do you know Yu Zhong?You can claim authorship or link another user.Do you know Haiyu Lan?You can claim authorship or link another user.Do you know Tianjing Hao?You can claim authorship or link another user.Do you know Kaixiang Lu?You can claim authorship or link another user.Do you know Peiwen Lin?You can claim authorship or link another user.Do you know Chuang Wang?You can claim authorship or link another user.Do you know Enyu Li?You can claim authorship or link another user.Do you know Junwei Liang?You can claim authorship or link another user.

Abstract

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that jointly capture the end effector, objects, and targets under occlusion. Existing multi-camera VLAs usually concatenate view tokens, leaving action representations weak in metric depth and inconsistent across cameras. We introduce Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views. A coordinate-query depth objective makes metric depth recoverable, while a preprocessing-aware correspondence objective aligns tokens observing the same physical point from different cameras. Both directly shape the hidden states consumed by the action module. After geometry injection, depth, camera calibration, and auxiliary heads are removed, so deployment uses the original RGB-only graph with no extra inference FLOPs. Held-out probes confirm stronger depth recovery and cross-view matching. Under matched GR00T-N1.6 settings, MVUCF reaches 98.9% on LIBERO, improves LIBERO-Plus by 22.4 points, and raises success by 23.3 points across six RoboTwin tasks spanning three action families: touch, move-and-place, and contact interaction. Real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.

Community

00