MEGA Hub

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

Authors

Do you know Rong Chao?You can claim authorship or link another user.Do you know Sung-Feng Huang?You can claim authorship or link another user.Do you know Moreno La Quatra?You can claim authorship or link another user.Do you know Sabato Marco Siniscalchi?You can claim authorship or link another user.Do you know Wen-Huang Cheng?You can claim authorship or link another user.Do you know Szu-Wei Fu?You can claim authorship or link another user.Do you know Yu Tsao?You can claim authorship or link another user.

Abstract

We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidth-efficient long-form inference. We further introduce a progressive knowledge distillation (KD) strategy that compresses an 8-layer teacher into a shallow 1-layer student by jointly distilling complex spectral outputs and intermediate representations. On Voicebank-DEMAND, the 8-layer RT-SEMamba achieves 3.32 PESQ with a 25 ms algorithmic latency constraint, and the distilled 1-layer student improves over a naive 1-layer baseline from 3.06 to 3.18 PESQ while preserving the same steady-state RTF, delivering a 2.75x speedup over the teacher. These results demonstrate that state-space models with progressive KD provide a competitive quality-latency trade-off for real-time SE.

Community

00

Publication notes

Author note
Accepted to INTERSPEECH 2026