MEGA Hub

Hydra-0: Action Flow for Generalist World Modeling and Control

Authors

Do you know Hongyu Li?You can claim authorship or link another user.Do you know Bowen Wen?You can claim authorship or link another user.Do you know Xinghao Zhu?You can claim authorship or link another user.Do you know Yixuan Wang?You can claim authorship or link another user.Do you know Yilun Du?You can claim authorship or link another user.Do you know Yunzhu Li?You can claim authorship or link another user.Do you know George Konidaris?You can claim authorship or link another user.Do you know Stan Birchfield?You can claim authorship or link another user.Do you know Soha Pouya?You can claim authorship or link another user.Do you know Chenran Li?You can claim authorship or link another user.Do you know Yan Chang?You can claim authorship or link another user.

Abstract

We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.

Community

00