MEGA Hub

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

Authors

Do you know Rui Tang?You can claim authorship or link another user.Do you know Wentao Yang?You can claim authorship or link another user.Do you know Peirong Zhang?You can claim authorship or link another user.Do you know Yongxin Shi?You can claim authorship or link another user.Do you know Shun Zhang?You can claim authorship or link another user.Do you know Huiguo He?You can claim authorship or link another user.Do you know Lianwen Jin?You can claim authorship or link another user.

Abstract

Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement. To accurately identify this anchor without oracle labels, we introduce Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards. To further improve representation robustness and localization precision, we introduce Directional Embedding Alignment (DEA) to suppress unstable norm bias by decoupling feature magnitude and direction, and Patch-Enhanced Decoding (PED) to fuse the routed anchor with language semantics and cross-attend over the full-image feature map for geometry-aware boundary regression beyond coordinate-space surrogates. Extensive experiments demonstrate that SPaTS consistently and significantly outperforms both frontier closed-source MLLMs and OCR MLLMs. Code will be released soon.

Community

00

Publication notes

Author note
15 pages, 11 figures. Accepted to ACM Multimedia 2026
Journal
Proceedings of the 34th ACM International Conference on Multimedia (MM '26), 2026
DOI
10.1145/3767308.3835174