MEGA Hub

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Authors

Do you know Zehua Fan?You can claim authorship or link another user.Do you know Junjie He?You can claim authorship or link another user.Do you know Wenxuan Song?You can claim authorship or link another user.Do you know Xi Wang?You can claim authorship or link another user.Do you know Wenqi Lyu?You can claim authorship or link another user.Do you know Linge Zhao?You can claim authorship or link another user.Do you know Fuhao Li?You can claim authorship or link another user.Do you know Zihan You?You can claim authorship or link another user.Do you know Yifei Yang?You can claim authorship or link another user.Do you know Kaiming Xu?You can claim authorship or link another user.Do you know Qi Jiang?You can claim authorship or link another user.Do you know Yue Jiang?You can claim authorship or link another user.Do you know Haoang Li?You can claim authorship or link another user.Do you know Cheng Chi?You can claim authorship or link another user.Do you know Bailin Li?You can claim authorship or link another user.Do you know Yan Wang?You can claim authorship or link another user.

Abstract

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation amid scene-scale dynamics, yet is still dominated by dynamics-blind visual encoders with hand-crafted coordination. We bridge this gap with MobileWAM, a mixture-of-transformers architecture that fuses a pretrained video diffusion transformer with a lightweight action expert through layerwise joint attention, translating internet-scale motion priors into whole-body control. To reconcile the heterogeneous dynamics of moving and manipulating, each feed-forward layer of the action expert becomes a three-expert mixture of shared, locomotion, and manipulation experts, softly routed by the motion intent in the action tokens. To densify supervision, we further propose Chain-of-Foresight (CoF): intermediate representations sequentially predict a chain of future latent chunks, each step conditioned on its predecessor. CoF pairs naturally with our decoupled video--action denoising scheme. At deployment, the WAM serves as a pure current-frame encoder; foresight acts only through gradients, so at inference the foresight chain and video generation are discarded, leaving only policy-level cost. MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization. Code will be released soon.

Community

00