MEGA Hub

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation

Authors

Do you know Leiye Liu?You can claim authorship or link another user.Do you know Miao Zhang?You can claim authorship or link another user.Do you know Jiahong Jiang?You can claim authorship or link another user.Do you know Jingjing Li?You can claim authorship or link another user.Do you know Jialong Zhong?You can claim authorship or link another user.Do you know Kai Peng?You can claim authorship or link another user.Do you know Tingwei Liu?You can claim authorship or link another user.Do you know Wei Ji?You can claim authorship or link another user.Do you know Yongri Piao?You can claim authorship or link another user.Do you know Huchuan Lu?You can claim authorship or link another user.

Abstract

Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing methods struggle to match overlapping acoustic events with visual instances and handle asynchronous audio-visual dynamics. Therefore, two critical questions arise: how can a model establish precise correspondence between overlapping sound sources and visual instances, and how can a model maintain robust tracking when audio and visual signals are temporally misaligned?This paper proposes Hear to See (H2S), addressing these challenges through two mechanisms. The Acoustic-Semantic Projector (ASP) disentangles mixed audio and establishes hierarchical correspondence from semantic to spatial domains. The Asynchronous Dynamics Modulator (ADM) adaptively adjusts state transitions via audio-modulated Mamba, prioritizing current information during dynamic variations and maintaining continuity in stable periods.Experiments on AVISeg show H2S achieves SOTA performance, attaining 48.54 mAP with a COCO pretrained ResNet50 and surpassing the previous by 7.8\%. The code will be open-sourced once the paper is accepted. The source code will be publicly available at https://github.com/leiyeliu/H2S.

Community

00

Publication notes

Author note
Accepted by ACM MM 2026