MEGA Hub

Test-Time Training for Speech Enhancement

Authors

Do you know Avishkar Behera?You can claim authorship or link another user.Do you know Riya Ann Easow?You can claim authorship or link another user.Do you know Venkatesh Parvathala?You can claim authorship or link another user.Do you know K. Sri Rama Murty?You can claim authorship or link another user.

Abstract

This paper introduces a novel application of Test-Time Training (TTT) for Speech Enhancement, addressing the challenges posed by unpredictable noise conditions and domain shifts. This method combines a main speech enhancement task with a self-supervised auxiliary task in a Y-shaped architecture. The model dynamically adapts to new domains during inference time by optimizing the proposed self-supervised tasks like noise-augmented signal reconstruction or masked spectrogram prediction, bypassing the need for labeled data. We further introduce various TTT strategies offering a trade-off between adaptation and efficiency. Evaluations across synthetic and real-world datasets show consistent improvements across speech quality metrics, outperforming the baseline model. This work highlights the effectiveness of TTT in speech enhancement, providing insights for future research in adaptive and robust speech processing.

Community

00

Publication notes

Author note
Published in the Proceedings of Interspeech 2025
Journal
Proceedings of Interspeech 2025, pp. 2375-2379
DOI
10.21437/Interspeech.2025-2725