MEGA Hub

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Authors

Do you know Oluwanifemi Bamgbose?You can claim authorship or link another user.Do you know Simon Rosen?You can claim authorship or link another user.Do you know Jash Shah?You can claim authorship or link another user.Do you know Lindsay Devon Brin?You can claim authorship or link another user.Do you know Hoang H Nguyen?You can claim authorship or link another user.Do you know Anke Koelzer?You can claim authorship or link another user.Do you know Rachel Hansen?You can claim authorship or link another user.Do you know Tara Bogavelli?You can claim authorship or link another user.Do you know Fanny Riols?You can claim authorship or link another user.

Abstract

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.

Community

00

Publication notes

Author note
Work in progress