MEGA Hub

VicEdit: Learning to Edit Videos from Visual In-Context Examples

Authors

Do you know Yuji Wang?You can claim authorship or link another user.Do you know Teng Hu?You can claim authorship or link another user.Do you know Yuheng Chen?You can claim authorship or link another user.Do you know Ran Yi?You can claim authorship or link another user.Do you know Han Feng?You can claim authorship or link another user.Do you know Weijian Cao?You can claim authorship or link another user.Do you know Chengjie Wang?You can claim authorship or link another user.Do you know Lizhuang Ma?You can claim authorship or link another user.Do you know Jiangning Zhang?You can claim authorship or link another user.

Abstract

Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEditBench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.

Community

00