MEGA Hub

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

Authors

Do you know Chenxu Du?You can claim authorship or link another user.Do you know Kang An?You can claim authorship or link another user.Do you know Tengyue Wang?You can claim authorship or link another user.Do you know Zhongyu Yang?You can claim authorship or link another user.Do you know Xinqi Yang?You can claim authorship or link another user.Do you know Yuanchi Zhu?You can claim authorship or link another user.Do you know Hebao Zhu?You can claim authorship or link another user.Do you know Ziliang Wang?You can claim authorship or link another user.Do you know Faqiang Qian?You can claim authorship or link another user.Do you know Yunli Yang?You can claim authorship or link another user.Do you know Qibing Ren?You can claim authorship or link another user.

Abstract

Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers. Its $1{,}212$ short-answer items are produced by a decoupled planner--writer pipeline and validated through automated screening, a blind adversarial audit, and expert review, so that answering requires perceiving the relevant evidence, identifying the governing principle, and applying it, not exploiting textual or single-figure shortcuts. Evaluating $18$ open-weight and proprietary MLLMs against a domain-expert panel, we find a wide gap: the strongest open-source model attains about $30\%$ and the best proprietary system $52\%$, while human experts reach $95\%$, more than forty points ahead. Our error analysis shows that failures concentrate in applying principles and combining evidence across figures rather than in locating it, pointing to substantial headroom for future research. Code and data are available at https://dcx-swjtu.github.io/MMArch/.

Community

00