MEGA Hub

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Authors

Do you know Yuhao Pan?You can claim authorship or link another user.Do you know Haosong Peng?You can claim authorship or link another user.Do you know Zhengshen Zhang?You can claim authorship or link another user.Do you know Zhengyang Yan?You can claim authorship or link another user.Do you know Yalun Dai?You can claim authorship or link another user.Do you know Fushuo Huo?You can claim authorship or link another user.Do you know Chujie Wang?You can claim authorship or link another user.Do you know Tianyu Qi?You can claim authorship or link another user.Do you know Xiucheng Wang?You can claim authorship or link another user.Do you know Nan Cheng?You can claim authorship or link another user.Do you know Wenchao Xu?You can claim authorship or link another user.

Abstract

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.

Community

00