MEGA Hub

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

Authors

Do you know David Ayllon?You can claim authorship or link another user.Do you know Alice Baird?You can claim authorship or link another user.Do you know Jeffrey Brooks?You can claim authorship or link another user.Do you know Franc Camps-Febrer?You can claim authorship or link another user.Do you know Jakub Piotr Cłapa?You can claim authorship or link another user.Do you know Theo Lebryk?You can claim authorship or link another user.Do you know Jens Madsen?You can claim authorship or link another user.Do you know Olya Ossipova?You can claim authorship or link another user.Do you know Sharath Rao?You can claim authorship or link another user.Do you know Hoon Shin?You can claim authorship or link another user.Do you know Tigran Soghbatyan?You can claim authorship or link another user.Do you know Georg Streich?You can claim authorship or link another user.Do you know Rashish Tandon?You can claim authorship or link another user.Do you know Panagiotis Tzirakis?You can claim authorship or link another user.

Abstract

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

Community

00

Publication notes

Author note
Benchmark and leaderboard: https://huggingface.co/spaces/HumeAI/rw-voice-eq