MEGA Hub

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Authors

Do you know Jun Zhan?You can claim authorship or link another user.Do you know Chen Yang?You can claim authorship or link another user.Do you know Yitian Gong?You can claim authorship or link another user.Do you know Donghua Yu?You can claim authorship or link another user.Do you know Kuangwei Chen?You can claim authorship or link another user.Do you know Wenbo Zhang?You can claim authorship or link another user.Do you know Kexin Huang?You can claim authorship or link another user.Do you know Qi Luo?You can claim authorship or link another user.Do you know Zhe Xu?You can claim authorship or link another user.Do you know Ying Zhu?You can claim authorship or link another user.Do you know Jin Wang?You can claim authorship or link another user.Do you know Tengyue Zhang?You can claim authorship or link another user.Do you know Qi Chen?You can claim authorship or link another user.Do you know Cheng Chang?You can claim authorship or link another user.Do you know Songlin Wang?You can claim authorship or link another user.Do you know Junqi Dai?You can claim authorship or link another user.Do you know Jiasheng Ye?You can claim authorship or link another user.Do you know Xiaogui Yang?You can claim authorship or link another user.Do you know Tianyi Liang?You can claim authorship or link another user.Do you know Xiangyu Peng?You can claim authorship or link another user.Do you know Zhaoye Fei?You can claim authorship or link another user.Do you know Shimin Li?You can claim authorship or link another user.Do you know Qinyuan Cheng?You can claim authorship or link another user.Do you know Xie Chen?You can claim authorship or link another user.Do you know Xinchi Chen?You can claim authorship or link another user.Do you know Xipeng Qiu?You can claim authorship or link another user.

Abstract

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1

Community

00