MEGA Hub

From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

Authors

Do you know Shiyong Meng?You can claim authorship or link another user.Do you know Bolei Chen?You can claim authorship or link another user.Do you know Ping Zhong?You can claim authorship or link another user.Do you know Yang Wan?You can claim authorship or link another user.Do you know Rongzhi Wang?You can claim authorship or link another user.Do you know Jiazhi Xia?You can claim authorship or link another user.Do you know Jianxin Wang?You can claim authorship or link another user.

Abstract

Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.

Community

00