MEGA Hub

Speaker-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement

Authors

Do you know Hanlin Zhang?You can claim authorship or link another user.Do you know Daxin Tan?You can claim authorship or link another user.Do you know Dehua Tao?You can claim authorship or link another user.Do you know Chengxi Deng?You can claim authorship or link another user.Do you know Xiao Chen?You can claim authorship or link another user.Do you know Linqi Song?You can claim authorship or link another user.

Abstract

Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic Token Purification (ISTP), an alternating speech-to-unit (S2U) and text-to-unit (T2U) training procedure guided by text predictability. Starting from an initial S2U tokenizer, each iteration trains a T2U model on its deduplicated token sequences. The decoded T2U predictions then serve as connectionist temporal classification targets for a newly initialized S2U model, whose outputs supervise the next T2U model. This cycle progressively aligns the two token generators and biases the token space toward information recoverable from text. Experiments on Mandarin and English show substantially improved S2U--T2U agreement. Independently trained de-tokenizers further show that the refined S2U and T2U tokens retain sufficient content for high-intelligibility voice conversion and text-to-speech synthesis. In voice conversion, the generated speaking rate follows the reference more closely. The refined tokens also exhibit substantially improved cross-speaker consistency and reduced probe-recoverable speaker information.

Community

00