MEGA Hub

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

Authors

Do you know Wanhao Liu?You can claim authorship or link another user.Do you know Jinsong Lin?You can claim authorship or link another user.Do you know Rulin Zhou?You can claim authorship or link another user.Do you know Chi Kit Ng?You can claim authorship or link another user.Do you know Wenbin Pan?You can claim authorship or link another user.Do you know Zhiqing Tang?You can claim authorship or link another user.Do you know Dongyue Li?You can claim authorship or link another user.Do you know Liwei Luo?You can claim authorship or link another user.Do you know Yanshen Wu?You can claim authorship or link another user.Do you know Panshuo Li?You can claim authorship or link another user.Do you know Zhiyong Xiong?You can claim authorship or link another user.Do you know Huxin Gao?You can claim authorship or link another user.Do you know Tamas Haidegger?You can claim authorship or link another user.Do you know Hongliang Ren?You can claim authorship or link another user.

Abstract

Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother--Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship. Unlike conventional multi-view fusion that assumes symmetric information exchange, we formulate \textbf{role-asymmetric dual-scope future prediction}, where cross-view evidence is selectively transferred according to the prediction target and its underlying spatial requirements. We propose \textbf{CrossScope}, a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions. CrossScope learns two complementary communication directions: geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. This design allows each scope to contribute task-relevant evidence without compromising its view-specific representation. To evaluate this problem, we establish a paired dual-scope benchmark comprising synchronized phantom and real-world ERCP episodes, with evaluations assessing visual fidelity, structural preservation, target localization, and motion consistency. Experiments demonstrate that CrossScope consistently outperforms strong surgical video generation baselines, validating the importance of role-aware evidence routing for multi-observer visual world modeling.

Community

00