MEGA Hub

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

Authors

Do you know AlayaWorld Team?You can claim authorship or link another user.Do you know Kaipeng Zhang?You can claim authorship or link another user.Do you know Chuanhao Li?You can claim authorship or link another user.Do you know Yifan Zhan?You can claim authorship or link another user.Do you know Yongtao Ge?You can claim authorship or link another user.Do you know Yuanyang Yin?You can claim authorship or link another user.Do you know Jiaming Tan?You can claim authorship or link another user.Do you know Kang He?You can claim authorship or link another user.Do you know Liaoyuan Fan?You can claim authorship or link another user.Do you know Mingliang Zhai?You can claim authorship or link another user.Do you know Ruicong Liu?You can claim authorship or link another user.Do you know Xiaojie Xu?You can claim authorship or link another user.Do you know Xuangeng Chu?You can claim authorship or link another user.Do you know Zhen Li?You can claim authorship or link another user.Do you know Zhengyuan Lin?You can claim authorship or link another user.Do you know Zhixiang Wang?You can claim authorship or link another user.Do you know Zian Meng?You can claim authorship or link another user.Do you know Zihui Gao?You can claim authorship or link another user.

Abstract

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.

Community

00