MEGA Hub

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

Authors

Do you know Guiyu Zhao?You can claim authorship or link another user.Do you know Longteng Guo?You can claim authorship or link another user.Do you know Yanghong Mei?You can claim authorship or link another user.Do you know Zilin Zhu?You can claim authorship or link another user.Do you know Yu Zhang?You can claim authorship or link another user.Do you know Bin Cao?You can claim authorship or link another user.Do you know Mingming Yu?You can claim authorship or link another user.Do you know Xingjian He?You can claim authorship or link another user.Do you know Jie Jiang?You can claim authorship or link another user.Do you know Jing Liu?You can claim authorship or link another user.

Abstract

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.

Community

00