MEGA Hub

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

Authors

Do you know Dehui Gao?You can claim authorship or link another user.Do you know Zhixian Zhao?You can claim authorship or link another user.Do you know Zhennan Lin?You can claim authorship or link another user.Do you know Yujie Liao?You can claim authorship or link another user.Do you know Yuhang Dai?You can claim authorship or link another user.Do you know Yike Zhu?You can claim authorship or link another user.Do you know Longshuai Xiao?You can claim authorship or link another user.Do you know Hui Bu?You can claim authorship or link another user.Do you know Xin Xu?You can claim authorship or link another user.Do you know Xie Chen?You can claim authorship or link another user.Do you know Shuai Wang?You can claim authorship or link another user.Do you know Liumeng Xue?You can claim authorship or link another user.Do you know Zhonghua Fu?You can claim authorship or link another user.Do you know Jun Du?You can claim authorship or link another user.Do you know Eng-Siong Chng?You can claim authorship or link another user.Do you know Jun Zhou?You can claim authorship or link another user.Do you know Lei Xie?You can claim authorship or link another user.

Abstract

Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentric platform for continuous audio sensing and assistance. However, speech recognition and understanding in this setting remain challenging because of dynamic acoustic conditions, speaker overlap, and the spatial ambiguity introduced by wearer-centered recording geometry. To support systematic evaluation in this setting, we introduce the IEEE SLT 2026 SmartGlasses Challenge for egocentric multi-speaker speech processing. The challenge consists of two tracks, Dyadic Dialogue Understanding and Multi-party Meeting Understanding, and jointly evaluates Time-Stamped Speaker-Attributed Automatic Speech Recognition (TSA-ASR) and Spoken Language Understanding (SLU). It is built on a 106-hour four-channel egocentric speech dataset containing 714 sessions collected in real-world scenarios. This paper describes challenge tasks, dataset construction, submissions, and summarizes the main findings from the shared evaluation. The results show that heavy speaker overlap remains a major factor affecting TSA-ASR performance, while paralinguistic acoustic understanding continues to be difficult for current audio-language models in complex SLU settings. Further details can be found on the official challenge website.

Community

00

Publication notes

Author note
7 pages, 7 figures