MEGA Hub

RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

Authors

Do you know Yaowei Guo?You can claim authorship or link another user.Do you know Zeng Tao?You can claim authorship or link another user.Do you know Yuxin Jiang?You can claim authorship or link another user.Do you know Yunuo Chen?You can claim authorship or link another user.Do you know Zhiyang Dou?You can claim authorship or link another user.Do you know Yuxiang Ma?You can claim authorship or link another user.Do you know Yin Yang?You can claim authorship or link another user.Do you know Demetri Terzopoulos?You can claim authorship or link another user.Do you know Ying Jiang?You can claim authorship or link another user.Do you know Chenfanfu Jiang?You can claim authorship or link another user.

Abstract

Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning.

Community

00

Publication notes

Author note
14 pages, 13 figures. Supplementary material included