MEGA Hub

ContactFlow: A video action conditioning that transfers across embodiments

Authors

Do you know Sami Azirar?You can claim authorship or link another user.Do you know Enrico Pallotta?You can claim authorship or link another user.Do you know Jan Nogga?You can claim authorship or link another user.Do you know Jürgen Gall?You can claim authorship or link another user.Do you know Sven Behnke?You can claim authorship or link another user.Do you know Hermann Blum?You can claim authorship or link another user.

Abstract

World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact. Further, their action conditioning is often constrained to specific embodiments such as parallel grippers. We propose \emph{Contact Flow}, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object. By discarding actor-specific appearance and kinematics, Contact Flow provides a shared conditioning signal for both human demonstrations and robotic execution. Therefore, we can train a large-scale video generative model on both human and robotic object interaction videos conditioned on Contact Flow, yielding a world model that predicts physically plausible manipulation outcomes. We integrate this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model before execution. Experiments on the DROID dataset and real-world tabletop manipulation tasks demonstrate that Contact Flow enables transfer between human demonstrations and different robotic embodiments.

Community

00