MEGA Hub

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

Authors

Do you know Yuqi Li?You can claim authorship or link another user.Do you know Yi-Cheng Lin?You can claim authorship or link another user.Do you know Xianglong Wang?You can claim authorship or link another user.Do you know Kuo Yang?You can claim authorship or link another user.Do you know Xiaoqin Feng?You can claim authorship or link another user.Do you know Yixuan Wang?You can claim authorship or link another user.Do you know Huiran Duan?You can claim authorship or link another user.Do you know Yingli Tian?You can claim authorship or link another user.

Abstract

On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.

Community

00