MEGA Hub

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

Authors

Do you know AlayaWorld Team?You can claim authorship or link another user.Do you know Kaipeng Zhang?You can claim authorship or link another user.Do you know Chuanhao Li?You can claim authorship or link another user.Do you know Yifan Zhan?You can claim authorship or link another user.Do you know Yongtao Ge?You can claim authorship or link another user.Do you know Yuanyang Yin?You can claim authorship or link another user.Do you know Jiaming Tan?You can claim authorship or link another user.Do you know Kang He?You can claim authorship or link another user.Do you know Liaoyuan Fan?You can claim authorship or link another user.Do you know Mingliang Zhai?You can claim authorship or link another user.Do you know Ruicong Liu?You can claim authorship or link another user.Do you know Xiaojie Xu?You can claim authorship or link another user.Do you know Xuangeng Chu?You can claim authorship or link another user.Do you know Zhen Li?You can claim authorship or link another user.Do you know Zhengyuan Lin?You can claim authorship or link another user.Do you know Zhixiang Wang?You can claim authorship or link another user.Do you know Zian Meng?You can claim authorship or link another user.Do you know Zihui Gao?You can claim authorship or link another user.

Abstract

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two major changes. First, we replace the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer. Second, we redesign the conditioning pipeline so that visual conditions are encoded in the same causal-VAE latent space, with temporal statistics consistent with those of the generated video. Concretely, the new version introduces six modifications: (1) replacing static-frame image conditioning with motion-aware latent conditioning; (2) causally encoding re-rendered spatial memory as a continuous sequence; (3) aligning the temporal-memory window in pixel space; (4) adopting hard memory dropout that removes memory tokens rather than zeroing them; (5) unifying the VAE encoding and decoding protocol across training and inference; and (6) removing the camera AdaLN branch, such that viewpoint control is provided entirely through the re-rendered spatial condition.

Community

00

Publication notes

Author note
Authors are listed alphabetically by the first name and their role. See the contribution section for details