MEGA Hub

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

Authors

Do you know Jiajun Cheng?You can claim authorship or link another user.Do you know Subarna Tripathi?You can claim authorship or link another user.Do you know Sainan Liu?You can claim authorship or link another user.Do you know Xiaofan Yu?You can claim authorship or link another user.Do you know Shan Lin?You can claim authorship or link another user.

Abstract

Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.

Community

00