MEGA Hub

Robots Acquire Manipulation Skills in Seconds from a Single Human Video

Authors

Do you know Guangyan Chen?You can claim authorship or link another user.Do you know Meiling Wang?You can claim authorship or link another user.Do you know Te Cui?You can claim authorship or link another user.Do you know Zichen Zhou?You can claim authorship or link another user.Do you know Qi Shao?You can claim authorship or link another user.Do you know Shalfun Li?You can claim authorship or link another user.Do you know Hang Su?You can claim authorship or link another user.Do you know Roy Gan?You can claim authorship or link another user.Do you know Hao Wang?You can claim authorship or link another user.Do you know Mengyin Fu?You can claim authorship or link another user.Do you know Yi Yang?You can claim authorship or link another user.Do you know Yufeng Yue?You can claim authorship or link another user.

Abstract

The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. In this paper, we introduce HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills. HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot's progress within the demonstrated task, then translates the upcoming progression into the robot's own future observations, and finally derives actions from these predicted observations. This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's embodiment. HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate. It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster. Additional information about HOST is available on the project website.

Community

00