MEGA Hub

Qwen-Audio-3.0-Gen-Preview Technical Report

Authors

Do you know Junyu Dai?You can claim authorship or link another user.Do you know Xiaoyue Duan?You can claim authorship or link another user.Do you know Xinyue Fan?You can claim authorship or link another user.Do you know Yihan Feng?You can claim authorship or link another user.Do you know Xiangang Li?You can claim authorship or link another user.Do you know Yunjia Li?You can claim authorship or link another user.Do you know Lejun Min?You can claim authorship or link another user.Do you know Yufei Shi?You can claim authorship or link another user.Do you know Xingchen Song?You can claim authorship or link another user.Do you know Yiran Wang?You can claim authorship or link another user.Do you know Cheng Wen?You can claim authorship or link another user.Do you know Menglin Wu?You can claim authorship or link another user.Do you know Bajian Xiang?You can claim authorship or link another user.Do you know Huaicheng Zhang?You can claim authorship or link another user.Do you know Han Zhao?You can claim authorship or link another user.Do you know Ruichen Zheng?You can claim authorship or link another user.

Abstract

Existing single-domain and multi-task audio systems remain limited in directly organizing speech, music, sound effects, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion Transformer (DiT) and a shared variational autoencoder (VAE) to generate the complete mixed waveform. Prompt enhancement converts free-form requests into structured temporal records that are rendered as textual conditions, while a two-stage data curriculum and semantic conditional views train the proposed model to use these conditions across domains. A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences and incorporates semantic supervision, providing one representation for speech, music, sound effects, and their mixtures. On Seed-TTS-Eval, speaker similarity is the proposed model's clearest strength across all three subsets, and on the multi-speaker benchmark, the proposed model shows higher cross-turn consistency than Seed-Audio-1.0 in both languages. On AudioCaps, its advantages are concentrated in evaluations using large audio-language models and AudioBox. Relative to Seed-Audio-1.0, it achieves stronger temporal localization. Using approximately 10% music data of a dedicated in-house model, the proposed model remains close across all seven SongBench components and leads in three while retaining speech and general-audio capabilities. These results demonstrate the potential of unified generation for temporally structured, multi-domain audio.

Community

00