MEGA Hub

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Authors

Do you know Weiliang Chen?You can claim authorship or link another user.Do you know Haowen Sun?You can claim authorship or link another user.Do you know Jun Gao?You can claim authorship or link another user.Do you know Jiawei Chi?You can claim authorship or link another user.Do you know Hanyang Wang?You can claim authorship or link another user.Do you know Qiyu Dai?You can claim authorship or link another user.Do you know Yihao Li?You can claim authorship or link another user.Do you know Hao Li?You can claim authorship or link another user.Do you know Jingnan Gao?You can claim authorship or link another user.Do you know Yi-Hsin Hung?You can claim authorship or link another user.Do you know Xingzhuo Guo?You can claim authorship or link another user.Do you know Shangchen Miao?You can claim authorship or link another user.Do you know Zhiyuan Shi?You can claim authorship or link another user.Do you know Xiang Li?You can claim authorship or link another user.Do you know Fengrui Tian?You can claim authorship or link another user.Do you know Weihua Du?You can claim authorship or link another user.Do you know Ziqi Huang?You can claim authorship or link another user.Do you know Shenyuan Gao?You can claim authorship or link another user.Do you know Siqiao Huang?You can claim authorship or link another user.Do you know Mingyu Liu?You can claim authorship or link another user.Do you know Yifei Li?You can claim authorship or link another user.Do you know Shizun Wang?You can claim authorship or link another user.Do you know Xi Wang?You can claim authorship or link another user.Do you know Tianqi Zhang?You can claim authorship or link another user.Do you know Xue Luo?You can claim authorship or link another user.Do you know Xiyin Ren?You can claim authorship or link another user.Do you know Jinshan Ren?You can claim authorship or link another user.Do you know Xiaoyang Shen?You can claim authorship or link another user.Do you know Xiaobo Hu?You can claim authorship or link another user.Do you know Zhiyang Dou?You can claim authorship or link another user.Do you know Mingyu Ding?You can claim authorship or link another user.Do you know Yichao Yan?You can claim authorship or link another user.Do you know Xinchao Wang?You can claim authorship or link another user.Do you know Yizhou Wang?You can claim authorship or link another user.Do you know Shilong Liu?You can claim authorship or link another user.Do you know Wenzhao Zheng?You can claim authorship or link another user.Do you know Yueqi Duan?You can claim authorship or link another user.Do you know Yuan Gong?You can claim authorship or link another user.Do you know Ziwei Liu?You can claim authorship or link another user.Do you know Ming-Yu Liu?You can claim authorship or link another user.Do you know Jialong Wu?You can claim authorship or link another user.Do you know Jiangran Lyu?You can claim authorship or link another user.Do you know Fangfu Liu?You can claim authorship or link another user.

Abstract

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.

Community

00

Publication notes

Author note
Project Page: https://mirros-lab.github.io/HarnessEval-W