MEGA Hub

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

Authors

Do you know Jihyun Lee?You can claim authorship or link another user.Do you know Cheol-Ho Cho?You can claim authorship or link another user.Do you know Woojin Jun?You can claim authorship or link another user.Do you know Woojin Jeong?You can claim authorship or link another user.Do you know Jae-Pil Heo?You can claim authorship or link another user.

Abstract

Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. Despite advances in pre-trained vision-language models and multimodal large language models, existing ZMR methods still heavily depend on query-to-video content similarity, making them vulnerable to modality and language-style gaps. These gaps lead to unreliable span proposals and unstable moment retrieval results. To address this issue, we propose Self-Similarity-based Moment Proposal and Scoring that instead exploits intrinsic relationships within videos, enabling robust span generation and scoring. By deriving self-similarity only from the video content, we circumvent the noisy and mismatched patterns of query-frame or query-caption similarities, thereby mitigating both modality and language-style gaps. Furthermore, we introduce a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video. Extensive experiments demonstrate that Self-SiMS achieves state-of-the-art performance across ZMR benchmarks.

Community

00

Publication notes

Author note
ECCV 2026 (* These authors contributed equally.)