MEGA Hub

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Authors

Do you know Tengfei Liu?You can claim authorship or link another user.Do you know Yang Shi?You can claim authorship or link another user.Do you know Yuran Wang?You can claim authorship or link another user.Do you know Xiaohan Zhang?You can claim authorship or link another user.Do you know Yuqing Wen?You can claim authorship or link another user.Do you know Yuqi Tang?You can claim authorship or link another user.Do you know Qixun Wang?You can claim authorship or link another user.Do you know Zhuoran Zhang?You can claim authorship or link another user.Do you know Xuanyu Zhu?You can claim authorship or link another user.Do you know Weihong Lin?You can claim authorship or link another user.Do you know Xinlei Yu?You can claim authorship or link another user.Do you know Yujie Wei?You can claim authorship or link another user.Do you know Xinwei Long?You can claim authorship or link another user.Do you know Fengxiang Wang?You can claim authorship or link another user.Do you know Xinlong Chen?You can claim authorship or link another user.Do you know Yue Ding?You can claim authorship or link another user.Do you know Jialu Chen?You can claim authorship or link another user.Do you know Haotian Wang?You can claim authorship or link another user.Do you know Yuanxing Zhang?You can claim authorship or link another user.

Abstract

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing $20,000$ videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.

Community

00

Publication notes

Author note
https://github.com/pkucs-Ltf/RefCaptioner