MEGA Hub

Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

Authors

Do you know Qi Ming?You can claim authorship or link another user.Do you know Yuyang Wang?You can claim authorship or link another user.Do you know Mingjing Zhao?You can claim authorship or link another user.Do you know Yifan Xiao?You can claim authorship or link another user.Do you know Zhixin Guo?You can claim authorship or link another user.Do you know Zhiqiang Zhou?You can claim authorship or link another user.Do you know Peng Sun?You can claim authorship or link another user.Do you know Juan Fang?You can claim authorship or link another user.Do you know Fuqiang Yang?You can claim authorship or link another user.Do you know Xudong Zhao?You can claim authorship or link another user.

Abstract

Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently addressed. In this paper, we propose a Joint Feature-domain Registration and Detection network (JFRDet), an end-to-end visible-infrared oriented object detector tailored for severely cross-modal geometric discrepancies. JFRDet introduces a Cross-Modal Affine Alignment (CMAA) module to estimate an image-level affine transformation for explicit multi-level feature alignment. Note that illumination changes directly affect the reliability of RGB cues, an Illumination-Guided Complementary Fusion (IGCF) module adaptively exploits modality reliability under varying illumination conditions for cross-modal fusion. Then, an Alignment Quality-Consistency Gating (AQCG) strategy stabilizes joint optimization by modulating detection supervision according to alignment reliability and gradient consistency. We further construct DroneVehicle Misaligned (DVMA), a benchmark for evaluating visible-infrared oriented object detection under severe cross-modal geometric misalignment. The proposed JFRDet achieves 69.7\% $\mathrm{mAP}_{50}$ on DVMA, which represents state-of-the-art (SOTA) performance. The code and dataset will be available on GitHub.

Community

00