MEGA Hub

S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

Authors

Do you know Kaneyoshi Hiratsuka?You can claim authorship or link another user.Do you know Benjamin Yen?You can claim authorship or link another user.Do you know Ryosuke Kojima?You can claim authorship or link another user.

Abstract

Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and $π_0$, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.

Community

00

Publication notes

Author note
Project page: https://azuma413.github.io/projects/s2a2