MEGA Hub

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Authors

Do you know Yuyang Liu?You can claim authorship or link another user.Do you know Yanqing Shen?You can claim authorship or link another user.Do you know Ruike Chen?You can claim authorship or link another user.Do you know Jifan Zhao?You can claim authorship or link another user.Do you know Yuxuan Tian?You can claim authorship or link another user.Do you know Yichi Zhang?You can claim authorship or link another user.Do you know Tianfeng Long?You can claim authorship or link another user.Do you know Zixuan Yin?You can claim authorship or link another user.Do you know Yipu Wang?You can claim authorship or link another user.Do you know Ziheng Qin?You can claim authorship or link another user.Do you know Wenxing Tan?You can claim authorship or link another user.Do you know Yang Shi?You can claim authorship or link another user.Do you know Mingyu Cao?You can claim authorship or link another user.Do you know Runze Xiao?You can claim authorship or link another user.Do you know Ziqi Wang?You can claim authorship or link another user.Do you know Zhixin Yin?You can claim authorship or link another user.Do you know Shiwei Chu?You can claim authorship or link another user.Do you know Yi-Fan Zhang?You can claim authorship or link another user.Do you know Yao Mu?You can claim authorship or link another user.Do you know Yuheng Ji?You can claim authorship or link another user.Do you know Yihao Wang?You can claim authorship or link another user.Do you know Jun Yan?You can claim authorship or link another user.Do you know Zhongyuan Wang?You can claim authorship or link another user.Do you know Pengwei Wang?You can claim authorship or link another user.Do you know Xiaolong Zheng?You can claim authorship or link another user.

Abstract

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.

Community

00