MEGA Hub

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

Authors

Do you know Nhan Phan?You can claim authorship or link another user.Do you know Ilona Lähteenmäki?You can claim authorship or link another user.Do you know Anna von Zansen?You can claim authorship or link another user.Do you know Olli-Pekka Pauna?You can claim authorship or link another user.Do you know Yaroslav Getman?You can claim authorship or link another user.Do you know Tamás Grósz?You can claim authorship or link another user.Do you know Mikko Kurimo?You can claim authorship or link another user.

Abstract

Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.

Community

00

Publication notes

Author note
To be submitted to ICASSP 2027. Code is available at https://github.com/aalto-speech/casa