MEGA Hub

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Authors

Do you know Yicheng Xiao?You can claim authorship or link another user.Do you know Wenxun Dai?You can claim authorship or link another user.Do you know Xinran Qin?You can claim authorship or link another user.Do you know Lin Song?You can claim authorship or link another user.Do you know Maoquan Zhang?You can claim authorship or link another user.Do you know Hang Xu?You can claim authorship or link another user.Do you know Yukang Chen?You can claim authorship or link another user.Do you know Yitong Li?You can claim authorship or link another user.Do you know Guohui Zhang?You can claim authorship or link another user.Do you know Yuan Zhang?You can claim authorship or link another user.Do you know Xuying Zhang?You can claim authorship or link another user.Do you know Tommy Zhang?You can claim authorship or link another user.Do you know Jianlong Yuan?You can claim authorship or link another user.Do you know Peihao Li?You can claim authorship or link another user.Do you know Shuai Lu?You can claim authorship or link another user.Do you know Siming Fu?You can claim authorship or link another user.Do you know Chuyang Zhao?You can claim authorship or link another user.Do you know Xin Han?You can claim authorship or link another user.Do you know Jie Huang?You can claim authorship or link another user.Do you know Wenbo Li?You can claim authorship or link another user.Do you know Guoqing Ma?You can claim authorship or link another user.Do you know Wei Huang?You can claim authorship or link another user.Do you know Xiaojuan Qi?You can claim authorship or link another user.Do you know Haoyang Huang?You can claim authorship or link another user.Do you know Nan Duan?You can claim authorship or link another user.

Abstract

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.

Community

00

Publication notes

Author note
Code: https://github.com/jd-opensource/JoyAI-Video-Edit