MEGA Hub

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Authors

Do you know Xinyu Liu?You can claim authorship or link another user.Do you know Shihao Li?You can claim authorship or link another user.Do you know Weihong Lin?You can claim authorship or link another user.Do you know Xinlong Chen?You can claim authorship or link another user.Do you know Yang Shi?You can claim authorship or link another user.Do you know Yujin Han?You can claim authorship or link another user.Do you know Yiyang Cai?You can claim authorship or link another user.Do you know Yanghao Wang?You can claim authorship or link another user.Do you know Ruibin Yuan?You can claim authorship or link another user.Do you know Yuanxing Zhang?You can claim authorship or link another user.Do you know Pengfei Wan?You can claim authorship or link another user.Do you know Wenhan Luo?You can claim authorship or link another user.Do you know Yike Guo?You can claim authorship or link another user.

Abstract

Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordinate information from multiple visual sources accurately. We identify a critical deficiency in existing approaches. Existing editing instructions lack explicit reference relationships, and most multimodal large language models (MLLMs) cannot generate them reliably. To address this problem, we propose ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing. Our key insight is embedding reference tokens at semantic positions to eliminate ambiguity and establish precise bindings between visual attributes and their sources. We develop ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attributes and their reference sources through a two-stage progressive scheme for precise reference relationships. We further develop ReBind-Edit, which enables lightweight adaptation of text-to-video models to coordinate multiple references by binding visual attributes to their designated sources. Extensive experiments demonstrate that ReBind substantially outperforms general-purpose MLLMs in instruction quality and achieves state-of-the-art performance among open-source methods on reference image conditioned video editing. Our project webpage: https://rebind-mrv2v.github.io/.

Community

00

Publication notes

Author note
Project Page: https://rebind-mrv2v.github.io/