MEGA Hub

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Authors

Do you know Nan Duan?You can claim authorship or link another user.Do you know Haoyang Huang?You can claim authorship or link another user.Do you know Weiyang Jin?You can claim authorship or link another user.Do you know Haoran Li?You can claim authorship or link another user.Do you know Yaowei Li?You can claim authorship or link another user.Do you know Yuming Li?You can claim authorship or link another user.Do you know Yijun Liu?You can claim authorship or link another user.Do you know Xin Lu?You can claim authorship or link another user.Do you know Xiaoxiao Ma?You can claim authorship or link another user.Do you know Yanwen Ma?You can claim authorship or link another user.Do you know Yaofeng Su?You can claim authorship or link another user.Do you know Yilang Sun?You can claim authorship or link another user.Do you know Haoyu Wang?You can claim authorship or link another user.Do you know Zeyue Xue?You can claim authorship or link another user.Do you know Songchun Zhang?You can claim authorship or link another user.Do you know Junhao Zhuang?You can claim authorship or link another user.

Abstract

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.

Community

00

Publication notes

Author note
Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/