MEGA Hub

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

Authors

Do you know Bajian Xiang?You can claim authorship or link another user.Do you know Cheng Wen?You can claim authorship or link another user.Do you know Han Zhao?You can claim authorship or link another user.Do you know Hao Wang?You can claim authorship or link another user.Do you know Haoxu Wang?You can claim authorship or link another user.Do you know Jiawei Jin?You can claim authorship or link another user.Do you know Jiayan Cui?You can claim authorship or link another user.Do you know Jie Chen?You can claim authorship or link another user.Do you know Mengxi Nie?You can claim authorship or link another user.Do you know Tianyu Zhao?You can claim authorship or link another user.Do you know Weiqin Li?You can claim authorship or link another user.Do you know Xiang Lv?You can claim authorship or link another user.Do you know Xiangang Li?You can claim authorship or link another user.Do you know Yang Xiang?You can claim authorship or link another user.Do you know Yang Zhou?You can claim authorship or link another user.

Abstract

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated language model (LM) and flow-matching model (FM) optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.

Community

00

Publication notes

Author note
19 pages