MEGA Hub

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

Authors

Do you know Kyeongmo Chae?You can claim authorship or link another user.Do you know Jihoon Lee?You can claim authorship or link another user.Do you know Sangtae Ahn?You can claim authorship or link another user.

Abstract

This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional deterministic point matching. By modeling both text and video embeddings as Gaussian distributions defined by mean and variance, DAB explicitly accounts for modality-specific uncertainty. We employ a deterministic, diffusion-inspired bridge to iteratively refine text distributions toward their target video distributions through a truncated refinement process. This approach unifies probabilistic embedding and distributional transformation into a cohesive, end-to-end trainable system. To optimize cross-modal similarity, we introduce a distribution-aware contrastive loss based on Kullback-Leibler divergence. Extensive evaluations on MSR-VTT, MSVD, and VATEX benchmarks confirm that DAB significantly outperforms existing probabilistic and diffusion-based baselines, while providing calibrated uncertainty-aware ranking through bridge-induced distributional margins.

Community

00

Publication notes

Author note
ECCV 2026