MEGA Hub

Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation

Authors

Do you know Yuzhu Wang?You can claim authorship or link another user.Do you know Archontis Politis?You can claim authorship or link another user.Do you know Konstantinos Drossos?You can claim authorship or link another user.Do you know Tuomas Virtanen?You can claim authorship or link another user.

Abstract

Long speech separation typically employs a segment-separation-stitch paradigm where recordings are divided into short segments, processed independently, and stitched together. Its challenge lies in predicting cross-segment permutations. This paper proposes a training-free dynamic clustering approach for cross-segment permutation alignment using speaker embedding reference pools. The method predicts the permutation using the cosine similarity between current segment embeddings and the reference pools. The approach updates reference pools by retaining the most representative speaker embeddings based on their overall cosine similarity with existing references. As a plug-and-play post-processing module compatible with existing separation models, the proposed method demonstrates superior performance compared to existing methods on dense and sparse long speech scenarios, particularly in challenging sparse scenarios with extended utterance gaps, and further shows robustness to speaker count estimation errors in unknown speaker count scenarios.

Community

00

Publication notes

Author note
6 pages, accepted by MLSP 2026