MEGA Hub

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Authors

Do you know Shuai Wang?You can claim authorship or link another user.Do you know Zihan Qian?You can claim authorship or link another user.Do you know Ke Zhang?You can claim authorship or link another user.Do you know Jiangyu Han?You can claim authorship or link another user.Do you know Zikai Liu?You can claim authorship or link another user.Do you know Xiaoyang Yu?You can claim authorship or link another user.Do you know Haoyu Li?You can claim authorship or link another user.Do you know Marc Delcroix?You can claim authorship or link another user.Do you know Kai Yu?You can claim authorship or link another user.Do you know Lei Xie?You can claim authorship or link another user.Do you know Ming Li?You can claim authorship or link another user.Do you know Haizhou Li?You can claim authorship or link another user.

Abstract

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.

Community

00

Publication notes

Author note
Overview paper of Real-TSE Challenge