MEGA Hub

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

Authors

Do you know Xiwei Liu?You can claim authorship or link another user.Do you know Yulong Li?You can claim authorship or link another user.Do you know Xinlin Zhuang?You can claim authorship or link another user.Do you know Xuhui Li?You can claim authorship or link another user.Do you know Zhixiang Lu?You can claim authorship or link another user.Do you know Haolin Yang?You can claim authorship or link another user.Do you know Imran Razzak?You can claim authorship or link another user.Do you know Yutong Xie?You can claim authorship or link another user.

Abstract

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires precise objects localization, or can bypass explicit localization through global layout cues. In this work, we investigate two representative model families, LLaVA-1.5 and Qwen2.5-VL, using a suite of mechanistic interpretability tools, including token ablation, layer-wise probing, attention knockout, and causal mediation analysis. We find that spatial relation prediction follows a staged grounding-to-reasoning process in which object-aligned tokens establish coarse target-reference anchors, while precise bounding-box boundaries are not required. Positional information becomes decodable before relation decisions emerge, and a small set of attention heads mediates the causal effects of both localization and spatial reasoning. The two tasks share early grounding-related processing but ultimately rely on partially distinct specialized pathways. Through rigorous experiments, we provide a token-, layer-, and head-level account of how VLMs transform object grounding into spatial relations, showing that knowing where objects are is not equivalent to knowing how they relate.

Community

00