MEGA Hub

Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

Authors

Do you know Zhuo Song?You can claim authorship or link another user.Do you know Lian Xu?You can claim authorship or link another user.Do you know Runqing Jiang?You can claim authorship or link another user.Do you know Yongjian Zhang?You can claim authorship or link another user.Do you know Kunhong Li?You can claim authorship or link another user.Do you know Ye Zhang?You can claim authorship or link another user.Do you know Yulan Guo?You can claim authorship or link another user.

Abstract

Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.

Community

00