MEGA Hub

VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)

Authors

Do you know Canyang Wu?You can claim authorship or link another user.Do you know Jinrong Zhang?You can claim authorship or link another user.Do you know Xusheng He?You can claim authorship or link another user.Do you know Ce Bian?You can claim authorship or link another user.Do you know Xianjing Han?You can claim authorship or link another user.Do you know Jianlong Wu?You can claim authorship or link another user.

Abstract

Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides strong promptable mask propagation, a uniform inference path remains unreliable for tiny targets with insufficient visual evidence and semantic-dominated targets whose identities depend on explicit attributes. To this end, we present VOS-Agent, a collaborative framework that retains SAM3 as the shared dense segmentation module and conditionally activates specialized agents according to target characteristics. A Target Perception and Routing Agent assigns each sequence to a regular, tiny, or semantic-dominated route. Tiny targets are supported by a Visual Tracking Agent through confidence-aware box prompts, while semantic-dominated targets are handled by an MLLM-based Semantic Agent through description-guided localization and candidate verification. On the MOSEv2 test set, VOS-Agent achieves 69.82% on the official $\mathcal{J}\&\dot{\mathcal{F}}$ metric and ranks first in the MOSEv2 Track of the 8th LSVOS Challenge at ECCV 2026.

Community

00

Publication notes

Author note
1st Place Solution for the 8th LSVOS MOSEv2 Challenge (ECCV 2026 Workshop)