MEGA Hub

Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge

Authors

Do you know Yiwen Ren?You can claim authorship or link another user.Do you know Jianing Liu?You can claim authorship or link another user.Do you know Yingxin Wang?You can claim authorship or link another user.Do you know Kexin Zhang?You can claim authorship or link another user.Do you know Licheng Jiao?You can claim authorship or link another user.Do you know Lingling Li?You can claim authorship or link another user.Do you know Xu Liu?You can claim authorship or link another user.

Abstract

The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the described target is absent. We present a simple staged solution. Qwen3-ASR first converts speech into text. Several video mask tracks are then produced with complementary grounding and segmentation models. Instead of trusting a single prediction, we select the track that has the highest average mask agreement with the other candidates. A small set of explicit direction, count, and plural rules corrects queries that require more than ordinary single-object tracking. Finally, a video-level classifier combines visual, audio-visual, and within-video query scores to decide whether any target is present. The submitted system obtains 0.5952 J &F, 0.7931 no-target accuracy, 0.9205 target accuracy, and a final score of 0.769589. The challenge organizers notified our team that this result ranked first in the track.

Community

00