MEGA Hub

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

Authors

Do you know Fuchen Long?You can claim authorship or link another user.Do you know Cong Wang?You can claim authorship or link another user.Do you know Zitao Gao?You can claim authorship or link another user.Do you know Wenhao Zhong?You can claim authorship or link another user.Do you know Yu Cheng?You can claim authorship or link another user.Do you know Xiaolu Hou?You can claim authorship or link another user.Do you know Yan Li?You can claim authorship or link another user.Do you know Xiao Cao?You can claim authorship or link another user.Do you know Xinlong Sun?You can claim authorship or link another user.Do you know Xi Chen?You can claim authorship or link another user.Do you know Yu Liu?You can claim authorship or link another user.

Abstract

The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.

Community

00