MEGA Hub

SurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos

Authors

Do you know Xinhao Chen?You can claim authorship or link another user.Do you know JuoTung Chen?You can claim authorship or link another user.Do you know Nigel Nelson?You can claim authorship or link another user.Do you know Antony Goldenberg?You can claim authorship or link another user.Do you know Jesse Haworth?You can claim authorship or link another user.Do you know Sean D. Huver?You can claim authorship or link another user.Do you know Axel Krieger?You can claim authorship or link another user.

Abstract

Learning-based surgical robot autonomy requires large-scale demonstrations with synchronized videos and robot actions, but such data are exceedingly rare in clinical or realistic tissue settings because robot kinematics are typically inaccessible outside controlled research systems. In contrast, phantom data collected on research platforms provide accurate action labels but lack the visual diversity of real tissue. We propose SurgVIL, a framework for scaling surgical robot imitation learning using open-source surgical videos. SurgVIL combines kinematically labeled phantom robot demonstrations with surgical videos from open-source datasets and online sources for policy learning. Since these videos lack robot motion labels, we estimate approximate kinematics as weak supervision. We evaluate SurgVIL on two da Vinci robot tasks: needle pick-up and cholecystectomy cutting. Across ACT, $π_0$, and GR00T-H backbones, adding surgical videos substantially improves generalization to real-tissue and out-of-distribution settings, suggesting a scalable path from phantom training toward generalizable surgical robot policies.

Community

00