MEGA Hub

A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones

Authors

Do you know Wiebke Middelberg?You can claim authorship or link another user.Do you know Svantje Voit?You can claim authorship or link another user.Do you know Simon Doclo?You can claim authorship or link another user.Do you know Ryan Corey?You can claim authorship or link another user.

Abstract

Mask-based beamforming is a popular geometry-agnostic approach for speech enhancement, typically applying a single mask across all microphones to estimate the required covariance matrices. While effective for compact arrays, this strategy may be suboptimal for spatially distributed microphones, where signal characteristics may vary strongly across microphones. To effectively capture the spatial diversity across microphones, we extend the mask-based beamformer to a multi-channel formulation, where each microphone is pre-filtered by a separate mask before covariance estimation. To address time-varying acoustic scenes, caused by spectro-temporal nonstationarity, we adopt a frame-causal online implementation with a sliding window. Experiments with simulated compact arrays and distributed microphones show that multi-channel masking yields a benefit over using a single mask when microphone signals differ substantially, while retaining similar performance in compact arrays. We further demonstrate the robustness of the multi-channel masking approach by comparing oracle ideal ratio masks to blind DNN-based mask estimation.

Community

00

Publication notes

Author note
Accepted for publication at IWAENC 2026