MEGA Hub

Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

Authors

Do you know Zhenglong Liu?You can claim authorship or link another user.Do you know Wangyou Zhang?You can claim authorship or link another user.Do you know Chenda Li?You can claim authorship or link another user.Do you know Yanmin Qian?You can claim authorship or link another user.

Abstract

Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they largely fail to exploit explicit array geometry priors when available, missing a crucial cue for optimal spatial filtering. A Geometry-Aware Dynamic Convolution (Geo-DConv) framework is proposed, which explicitly leverages microphone coordinates to transform standard fixed-array SE models into robust array-invariant systems. Experiments are conducted on the recent real-recorded RealMAN multi-channel speech dataset. Results demonstrate that the proposed architecture enables two widely used fixed-array models to adapt to array-invariant settings, with consistent performance improvements across diverse array topologies.

Community

00

Publication notes

Author note
Comments: 5 pages, 1 figure, 2 tables. Accepted at Interspeech 2026