MEGA Hub

Video = World + Event Stream

Authors

Do you know Lianghua Huang?You can claim authorship or link another user.Do you know Zhi-Fan Wu?You can claim authorship or link another user.Do you know Yupeng Shi?You can claim authorship or link another user.Do you know Wei Wang?You can claim authorship or link another user.Do you know Mengyang Feng?You can claim authorship or link another user.Do you know Cheng Yu?You can claim authorship or link another user.Do you know Chen Liang?You can claim authorship or link another user.Do you know Junjie He?You can claim authorship or link another user.Do you know Chen-Wei Xie?You can claim authorship or link another user.Do you know Yu Liu?You can claim authorship or link another user.Do you know Jingren Zhou?You can claim authorship or link another user.Do you know Ang Wang?You can claim authorship or link another user.Do you know Bang Zhang?You can claim authorship or link another user.Do you know Baole Ai?You can claim authorship or link another user.Do you know Chongyang Zhong?You can claim authorship or link another user.Do you know Jinwei Qi?You can claim authorship or link another user.Do you know Kai Zhu?You can claim authorship or link another user.Do you know Pandeng Li?You can claim authorship or link another user.Do you know Peng Zhang?You can claim authorship or link another user.Do you know Wenyuan Zhang?You can claim authorship or link another user.Do you know Xinhua Cheng?You can claim authorship or link another user.Do you know Yitong Huang?You can claim authorship or link another user.Do you know Yun Zheng?You can claim authorship or link another user.Do you know Yuxiang Bao?You can claim authorship or link another user.Do you know Yuzheng Wang?You can claim authorship or link another user.Do you know Zhiwei Lin?You can claim authorship or link another user.Do you know Zoubin Bi?You can claim authorship or link another user.

Abstract

We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.

Community

00

Publication notes

Author note
website; https://wan-streamer.com/