MEGA Hub

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Authors

Do you know Andong Lu?You can claim authorship or link another user.Do you know Ziyi Zha?You can claim authorship or link another user.Do you know Jiandong Jin?You can claim authorship or link another user.Do you know Shihao Li?You can claim authorship or link another user.Do you know Chenglong Li?You can claim authorship or link another user.Do you know Jin Tang?You can claim authorship or link another user.Do you know Bin Luo?You can claim authorship or link another user.

Abstract

Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition, current approaches exhibit limited flexibility in processing both missing and complete data. To overcome these limitations, we propose a Spatio-temporal Conditional Denoising Transformer (SCDT), which integrates the spatial cues and the temporal context to adaptively perform information reconstruction of missing modalities and feature enhancement of weak modalities in a unified framework, for robust modality-missing RGBT tracking. In particular, SCDT leverages the short-term temporal cues from recent historical frames to capture the fine-grained temporal correlations and the long-term temporal cues encoding modality evolution to capture the global context. By jointly exploiting long short-term temporal contexts as the conditions, SCDT progressively guides noisy features of available modalities to learn reliable and temporally consistent multimodal representations. Furthermore, SCDT introduces a noisemodulated adaptation mechanism that dynamically adjusts its behavior according to the modal availability, enabling a single framework to unify feature learning under both modality-missing and complete scenarios without changing the architecture or parameters. Extensive experiments on three public benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods. The code is available here.

Community

00

Publication notes

Author note
Accepted by CVPR2026