MEGA Hub

Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots

Authors

Do you know Hung Nguyen?You can claim authorship or link another user.Do you know Kim Nhat Minh Nguyen?You can claim authorship or link another user.Do you know Van Duc Vu?You can claim authorship or link another user.Do you know Van-Danh Le?You can claim authorship or link another user.Do you know Hoang Huy Le?You can claim authorship or link another user.Do you know Dinh Tuan Nguyen?You can claim authorship or link another user.Do you know Pham Tuyen Le?You can claim authorship or link another user.Do you know Van-Truong Nguyen?You can claim authorship or link another user.Do you know Quan Nguyen?You can claim authorship or link another user.

Abstract

Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inputs. In this paper, we investigate whether a well-established text-conditioned model can be transferred to speech in a data-efficient manner. Using ALBEF as a case study, we conduct diagnostic analyses showing that a lightweight MLP-based projector effectively adapts it to speech, while preserving semantic discrimination and robustness. Motivated by these findings, we introduce Speech2Grasp, a framework for data-efficient transfer of text-conditioned grasp detection to speech. Real-world humanoid robot experiments show that Speech2Grasp outperforms cascaded ASR-based pipeline, while reducing inference latency. Our findings suggest a practical paradigm for extending established text-conditioned systems to speech.

Community

00