MEGA Hub

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Authors

Do you know Siyu Yan?You can claim authorship or link another user.Do you know Zhuoran Yan?You can claim authorship or link another user.Do you know Haiying Xu?You can claim authorship or link another user.Do you know Panhao Zhou?You can claim authorship or link another user.Do you know Jingyu Chen?You can claim authorship or link another user.Do you know Chenhao Ji?You can claim authorship or link another user.Do you know Shuo Cao?You can claim authorship or link another user.Do you know Yongheng Zhang?You can claim authorship or link another user.Do you know Haoze Liu?You can claim authorship or link another user.Do you know Siyu Zhang?You can claim authorship or link another user.Do you know Xiwen Gu?You can claim authorship or link another user.Do you know Yihao Liu?You can claim authorship or link another user.Do you know Alex Jinpeng Wang?You can claim authorship or link another user.

Abstract

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.

Community

00