MEGA Hub

DNN-Based Frequency-Dependent Estimation of Speech, Music, and Noise Power in Acoustic Mixtures for Hearing-Aid Scene Analysis

Authors

Do you know Mats Lang?You can claim authorship or link another user.Do you know Thomas Haubner?You can claim authorship or link another user.Do you know Nina Kiessling?You can claim authorship or link another user.Do you know Christoph Hoog Antink?You can claim authorship or link another user.Do you know Henning Puder?You can claim authorship or link another user.

Abstract

Acoustic scene analysis is essential for adapting hearing-aid signal processing algorithms to the current listening environment. However, state-of-the-art (SOTA) systems typically rely on multiple independent estimators for tasks such as scene classification, Voice Activity Detection (VAD), or Signal-to-Noise Ratio estimation, which increases computational complexity and fails to exploit dependencies between related tasks. To address this problem, we propose a unified and interpretable acoustic scene representation by decomposing the observed mixture spectrum into speech, music, and noise power components. This is motivated by the typical listening targets of hearing-aid users. In particular, we estimate time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing. In this work, we validate the proposed representation using VAD as a representative downstream task and show performance comparable to a SOTA estimator while providing a substantially richer scene description.

Community

00

Publication notes

Author note
Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2026