MEGA Hub

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

Authors

Do you know Jun Zhang?You can claim authorship or link another user.Do you know Qiao Zhao?You can claim authorship or link another user.Do you know Cheng Cui?You can claim authorship or link another user.Do you know Jianying Qu?You can claim authorship or link another user.Do you know Zhongkai Sun?You can claim authorship or link another user.Do you know Jianwen Yang?You can claim authorship or link another user.Do you know Changda Zhou?You can claim authorship or link another user.Do you know ZhuoXin Liu?You can claim authorship or link another user.Do you know Shubin Han?You can claim authorship or link another user.

Abstract

While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.

Community

00