MEGA Hub

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

Authors

Do you know Xiuyuan Zhu?You can claim authorship or link another user.Do you know Ke Lu?You can claim authorship or link another user.Do you know Kun Dong?You can claim authorship or link another user.Do you know Siwen Jiao?You can claim authorship or link another user.Do you know Hao Wu?You can claim authorship or link another user.Do you know Zijin Du?You can claim authorship or link another user.Do you know Shun Mao?You can claim authorship or link another user.Do you know Dongming Zhang?You can claim authorship or link another user.Do you know Jian Xue?You can claim authorship or link another user.

Abstract

Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify this representation as an important source of error in visual grounding. Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM architecture. Hi-GAR complements this representation with a geometry-based reward for Group Relative Policy Optimization (GRPO), using box overlap and coordinate accuracy at multiple scales. Controlled comparisons under matched training conditions show that Hi-Token improves localization throughout the evaluated IoU range. Hi-GAR further reduces low-overlap predictions and is used only during training. Experiments on three VLM backbones and the RefCOCO family show consistent gains across models and benchmarks. Hi-R1 achieves higher values than strong specialist baselines on most reported metrics. Analyses of token frequency, digit boundaries, object scale, and IoU distributions explain the effects of coordinate representation and reward training. The results show that structured coordinate generation provides an effective approach to generative visual grounding.

Community

00

Publication notes

Author note
15 pages, 7 figures, 15 tables