MEGA Hub

The Role of Disfluencies in Speech Translation

Authors

Do you know Maike Züfle?You can claim authorship or link another user.Do you know Maria Teleki?You can claim authorship or link another user.Do you know Fabian Retkowski?You can claim authorship or link another user.Do you know Vilém Zouhar?You can claim authorship or link another user.Do you know Oliver Grabner?You can claim authorship or link another user.Do you know Alexander Waibel?You can claim authorship or link another user.Do you know James Caverlee?You can claim authorship or link another user.Do you know Jan Niehues?You can claim authorship or link another user.

Abstract

Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them. We show this comes at a cost: disfluencies carry meaning that gets lost when speech is cleaned up. To study this systematically, we introduce Uh-Mazing, a benchmark of human-translated, disfluency-annotated Switchboard speech covering English into eight target languages. Across these languages and several architectures, we find that false starts and self-repairs, not filled pauses or discourse markers, drive most of the translation-quality loss, and that models which fail to preserve a disfluency tend to omit it rather than mistranslate it. We show inference-time decoding can mitigate this without retraining, and release the benchmark and code.

Community

00