MEGA Hub

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

Authors

Do you know Fan Yang?You can claim authorship or link another user.Do you know Yuting Su?You can claim authorship or link another user.Do you know Xiaobo Wang?You can claim authorship or link another user.Do you know Yuncheng You?You can claim authorship or link another user.Do you know Fugui Fan?You can claim authorship or link another user.Do you know Yuting Wu?You can claim authorship or link another user.Do you know Minghui Wu?You can claim authorship or link another user.Do you know Chenxu Zhao?You can claim authorship or link another user.Do you know JiaHong Ning?You can claim authorship or link another user.Do you know Peiguang Jing?You can claim authorship or link another user.

Abstract

World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. Pixel-space methods often allocate substantial capacity to visual details that may not be directly relevant to control, while some latent-space methods require multi-stage training to construct the reasoning space. The resulting training cost can make such methods difficult to train under modest computational budgets. In this work, we propose LiLa-WAM, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU. Its core design is a compact latent reasoning space jointly shaped by future-state prediction and action generation, which keeps the model lightweight while remaining well aligned with control. For task specification, we further propose the Visual Transition Token(VTT), a language-free task representation that encodes each task as a direction in visual feature space. Experiments on RoboTwin~2.0, LIBERO, and real-robot tasks demonstrate LiLa-WAM's effectiveness, achieving 90.48\% success across 50 RoboTwin tasks with single-GPU training.

Community

00