MEGA Hub

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions

Authors

Do you know Fernando López?You can claim authorship or link another user.Do you know Ana Ayala?You can claim authorship or link another user.Do you know Guillermo Segovia?You can claim authorship or link another user.Do you know Fernando Ibáñez?You can claim authorship or link another user.Do you know Ana Martínez?You can claim authorship or link another user.Do you know Pablo Gómez?You can claim authorship or link another user.Do you know Jordi Luque?You can claim authorship or link another user.

Abstract

As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LALMs across heterogeneous acoustic conditions and reasoning abilities. ESCUCHA comprises 1,000 human-curated questions paired with audio, totaling 162.9 hours sourced directly ``from the wild'' rather than drawn from existing datasets, with durations ranging from a few seconds to over 80 minutes. The benchmark emphasizes reasoning, spanning 9 perceptual and 10 reasoning categories, and it captures linguistic diversity through multiple Spanish accents and non-normative speech. ESCUCHA further includes multi-audio questions, spoken questions, and audio instructions, and it flags which questions support open-ended evaluation. Benchmarking several state-of-the-art multimodal and speech models reveals substantial performance gaps relative to trained humans.

Community

00

Publication notes

Author note
Under review