MEGA Hub

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

Authors

Do you know Dongchi Huang?You can claim authorship or link another user.Do you know Hongyin Zhang?You can claim authorship or link another user.Do you know Bohan Hou?You can claim authorship or link another user.Do you know Siteng Huang?You can claim authorship or link another user.Do you know Zhian Su?You can claim authorship or link another user.Do you know Hang Guo?You can claim authorship or link another user.Do you know Tong Lu?You can claim authorship or link another user.Do you know Zhaofeng Xu?You can claim authorship or link another user.Do you know Jiahao Tang?You can claim authorship or link another user.Do you know Jianfei Yang?You can claim authorship or link another user.Do you know Donglin Wang?You can claim authorship or link another user.Do you know Peixi Peng?You can claim authorship or link another user.Do you know Mingxiu Chen?You can claim authorship or link another user.Do you know Deli Zhao?You can claim authorship or link another user.Do you know Xin Li?You can claim authorship or link another user.

Abstract

General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's tau_a of 0.675 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.

Community

00

Publication notes

Author note
23 pages, 5 figures