MEGA Hub

JEPA-WAM: Stage-Level Joint-Embedding Prediction for World-Action Models in Robot Manipulation

Authors

Do you know Xiao Liu?You can claim authorship or link another user.Do you know Yuguang Yang?You can claim authorship or link another user.Do you know Xi Wang?You can claim authorship or link another user.Do you know Kai Jiang?You can claim authorship or link another user.Do you know Cheng Chi?You can claim authorship or link another user.Do you know Yong Xu?You can claim authorship or link another user.Do you know Wenchao Ding?You can claim authorship or link another user.Do you know Yilun Chen?You can claim authorship or link another user.Do you know Yan Wang?You can claim authorship or link another user.

Abstract

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, but it does not explicitly describe the stage-level future that specifies how a task should progress from its current stage to the next. We therefore distinguish two complementary futures for robot manipulation: a short-term physical future to capture local scene evolution and a stage-level semantic future to represent task progress. We introduce JEPA-WAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Given the current observation and task instruction, Stage-JEPA uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, JEPA-WAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.

Community

00