MEGA Hub

TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

Authors

Do you know Fang Li?You can claim authorship or link another user.Do you know Shihao Zou?You can claim authorship or link another user.Do you know Weixin Si?You can claim authorship or link another user.Do you know Yang Gao?You can claim authorship or link another user.Do you know Shuai Li?You can claim authorship or link another user.Do you know Aimin Hao?You can claim authorship or link another user.

Abstract

Understanding complex surgical scenes requires recognizing multiple interdependent entities, such as instruments, actions, and targets, while maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and inter-frame temporal semantics in a unified manner. To address these limitations, we propose a unified framework that integrates spatial, relational, and temporal cues for robust surgical triplet recognition. Specifically, class-specific spatial priors are first extracted through a multi-scale encoder. These priors are then refined by a Label Correlation Modeling module with multi-scale class activation map-guided relational extraction (MS-CAMRE), enabling the model to capture both static co-occurrence patterns and dynamic contextual dependencies among triplet components. Furthermore, a Bidirectional Temporal-Relational Fusion Attention (BTRFA) module harmonizes temporal and relational representations to achieve coherent temporal reasoning. We also introduce a new evaluation metric, the Triplet Consistency Error Rate (TCER), which quantitatively measures the model's ability to preserve causal and semantic consistency across triplets. Extensive experiments on the CholecT45 and ProstaTD datasets show that our method achieves state-of-the-art performance, improving AP_IVT by 5.1 percent and 7.8 percent, respectively. Moreover, according to TCER, our approach achieves relative reductions of more than 36 percent and 25 percent on the two datasets, respectively, demonstrating the effectiveness of our framework in temporal-relational co-reasoning.

Community

00

Publication notes

Author note
code: https://github.com/Neesky/TRCoRSurg