MEGA Hub

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Authors

Do you know Vorch Team?You can claim authorship or link another user.Do you know Xiaoyu Chen?You can claim authorship or link another user.Do you know Yang Ding?You can claim authorship or link another user.Do you know Cong Han?You can claim authorship or link another user.Do you know Menglin Han?You can claim authorship or link another user.Do you know Yuxin Hong?You can claim authorship or link another user.Do you know Jiebo Hou?You can claim authorship or link another user.Do you know Zequn Jie?You can claim authorship or link another user.Do you know Xiang Li?You can claim authorship or link another user.Do you know Jing Liu?You can claim authorship or link another user.Do you know Qi Liu?You can claim authorship or link another user.Do you know Yulei Lu?You can claim authorship or link another user.Do you know Siyuan Luo?You can claim authorship or link another user.Do you know Lin Ma?You can claim authorship or link another user.Do you know Xin Ma?You can claim authorship or link another user.Do you know Yinlong Qian?You can claim authorship or link another user.Do you know Peng Shi?You can claim authorship or link another user.Do you know Fang Wan?You can claim authorship or link another user.Do you know Siqi Wang?You can claim authorship or link another user.Do you know Yaohui Wang?You can claim authorship or link another user.Do you know Yaole Wang?You can claim authorship or link another user.Do you know Yidi Wu?You can claim authorship or link another user.Do you know Siqian Yang?You can claim authorship or link another user.Do you know Mingyu Yin?You can claim authorship or link another user.Do you know Haoran Yu?You can claim authorship or link another user.Do you know Gang Yue?You can claim authorship or link another user.Do you know Lisai Zhang?You can claim authorship or link another user.Do you know Yuting Zhang?You can claim authorship or link another user.

Abstract

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.

Community

00

Publication notes

Author note
Project Page: https://vorch-project.github.io/Vorch-Omni-project/