MEGA Hub

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

Authors

Do you know Junyu Wang?You can claim authorship or link another user.Do you know Siyuan Zhang?You can claim authorship or link another user.Do you know Peiyuan Jiang?You can claim authorship or link another user.Do you know Jian Zong?You can claim authorship or link another user.Do you know Jingyu Zhang?You can claim authorship or link another user.Do you know Tianrui Wang?You can claim authorship or link another user.Do you know Yuqin Lin?You can claim authorship or link another user.Do you know Zhenghui Chen?You can claim authorship or link another user.Do you know Shuqing Xie?You can claim authorship or link another user.Do you know Ziyang Ma?You can claim authorship or link another user.Do you know Meng Ge?You can claim authorship or link another user.Do you know Xiaobao Wang?You can claim authorship or link another user.Do you know Longbiao Wang?You can claim authorship or link another user.Do you know Jianwu Dang?You can claim authorship or link another user.

Abstract

Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.

Community

00

Publication notes

Author note
Accepted at ACM Multimedia 2026 (MM '26)