MEGA Hub

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Authors

Do you know Zichao Lin?You can claim authorship or link another user.Do you know Yifeng Xie?You can claim authorship or link another user.Do you know Bowen Qu?You can claim authorship or link another user.Do you know Haiming Wang?You can claim authorship or link another user.Do you know Jia Li?You can claim authorship or link another user.Do you know Haoning Wu?You can claim authorship or link another user.Do you know Yuhao Dong?You can claim authorship or link another user.Do you know Zuhao Yang?You can claim authorship or link another user.Do you know Jinguo Zhu?You can claim authorship or link another user.Do you know Haoyu Lu?You can claim authorship or link another user.Do you know Zijia Zhao?You can claim authorship or link another user.Do you know Tongtian Yue?You can claim authorship or link another user.Do you know Zhangyang Qi?You can claim authorship or link another user.Do you know Junwei Yang?You can claim authorship or link another user.Do you know Mengfan Dong?You can claim authorship or link another user.Do you know Peizhou Cao?You can claim authorship or link another user.Do you know Chenzhuang Du?You can claim authorship or link another user.Do you know Zaida Zhou?You can claim authorship or link another user.Do you know Haotian Yao?You can claim authorship or link another user.Do you know Hao Yang?You can claim authorship or link another user.Do you know Hongcheng Gao?You can claim authorship or link another user.Do you know Lin Sui?You can claim authorship or link another user.Do you know Weihong Li?You can claim authorship or link another user.Do you know Xinxing Zu?You can claim authorship or link another user.Do you know Jia Chen?You can claim authorship or link another user.Do you know Yao Wang?You can claim authorship or link another user.Do you know Xiaoxue Wu?You can claim authorship or link another user.Do you know Yalin Wang?You can claim authorship or link another user.Do you know Y. Charles?You can claim authorship or link another user.Do you know Yiping Bao?You can claim authorship or link another user.Do you know Yangyang Liu?You can claim authorship or link another user.Do you know Zhiqi Huang?You can claim authorship or link another user.Do you know Xinyu Zhou?You can claim authorship or link another user.

Abstract

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.

Community

00