MEGA Hub

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

Authors

Do you know Siyuan Ma?You can claim authorship or link another user.Do you know Boshi Zhang?You can claim authorship or link another user.Do you know Yutian Zhang?You can claim authorship or link another user.Do you know Qinglian Wu?You can claim authorship or link another user.Do you know Jiaqi Zhai?You can claim authorship or link another user.Do you know Dong Wei?You can claim authorship or link another user.Do you know Qiaojun Yu?You can claim authorship or link another user.

Abstract

Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM backbone and trains residual adapters, an action-equivalent future bottleneck distilled from privileged observations, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. We further introduce ARMDOG, a real-robot dataset that synchronizes video, whole-body state and action, and language. On a fixed replay protocol, DECOWAM improved both future-video and action prediction over FastWAM, reducing action MSE by 21.7% with 25.95M trainable adaptation parameters. Across 79 closed-loop trials per method, it achieved the highest observed whole-body coordination and base-displacement robustness among the compared systems, while task completion remained comparable to the strongest baseline. These results show that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.

Community

00

Publication notes

Author note
8 pages, 5 figures. Introduces DECOWAM, a decoupled whole-body world-action model for legged mobile manipulation, and the ARMDOG real-robot dataset