MEGA Hub

GuidedAttention: Interpretable and Correctable Visual Attention for OOD-Robust Robot Manipulation via Imitation Learning

Authors

Do you know Masaki Murooka?You can claim authorship or link another user.Do you know Ryoichi Nakajo?You can claim authorship or link another user.Do you know Keisuke Shirai?You can claim authorship or link another user.Do you know Tomohiro Motoda?You can claim authorship or link another user.Do you know Hanbit Oh?You can claim authorship or link another user.Do you know Ryo Hanai?You can claim authorship or link another user.Do you know Yukiyasu Domae?You can claim authorship or link another user.

Abstract

End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions.

Community

00