MEGA Hub

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Authors

Do you know Yixiang Chen?You can claim authorship or link another user.Do you know Jiabing Yang?You can claim authorship or link another user.Do you know Yuan Xu?You can claim authorship or link another user.Do you know Qisen Ma?You can claim authorship or link another user.Do you know Keji He?You can claim authorship or link another user.Do you know Peiyan Li?You can claim authorship or link another user.Do you know Kai Wang?You can claim authorship or link another user.Do you know Ziheng He?You can claim authorship or link another user.Do you know Xiangnan Wu?You can claim authorship or link another user.Do you know Jing Liu?You can claim authorship or link another user.Do you know Nianfeng Liu?You can claim authorship or link another user.Do you know Yan Huang?You can claim authorship or link another user.Do you know Liang Wang?You can claim authorship or link another user.

Abstract

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.

Community

00