MEGA Hub

MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

Authors

Do you know Liangtao Shi?You can claim authorship or link another user.Do you know Jinxia Xie?You can claim authorship or link another user.Do you know Xiantao Hu?You can claim authorship or link another user.Do you know Ting Liu?You can claim authorship or link another user.

Abstract

In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the task into several stages and identify suitable foundation models for each stage. Without introducing additional model training or task-specific fine-tuning, our approach leverages the strong multimodal reasoning capabilities of MLLMs to model text-visual correspondence and employs SAM-based models for accurate object mask generation. The proposed framework demonstrates the effectiveness of leveraging foundation models for audio-guided video segmentation and achieves competitive performance in the MeViS-Audio Track of the 8th LSVOS Challenge.

Community

00

Publication notes

Author note
5 pages