MEGA Hub

FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation

Authors

Do you know Ruicheng Li?You can claim authorship or link another user.Do you know Qixiu Li?You can claim authorship or link another user.Do you know Ruichun Ma?You can claim authorship or link another user.Do you know Yu Deng?You can claim authorship or link another user.Do you know Lin Luo?You can claim authorship or link another user.Do you know Zhiying Du?You can claim authorship or link another user.Do you know Jianfeng Xiang?You can claim authorship or link another user.Do you know Huizhi Liang?You can claim authorship or link another user.Do you know Ruicheng Wang?You can claim authorship or link another user.Do you know Jiaolong Yang?You can claim authorship or link another user.Do you know Baining Guo?You can claim authorship or link another user.

Abstract

Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal events are visually ambiguous, e.g., pushing a button multiple times with small movements. We propose FM-VLA, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation. We encode force histories into compact force memory tokens with a variational autoencoder (VAE) pretrained with force time series reconstruction. By projecting force latent representations and short state history as additional conditioning tokens to the action expert module, we enable VLAs to leverage accumulated contact event history to guide manipulation. We evaluate FM-VLA on three memory-dependent tasks, including finding a hidden block, pressing a button, and wiping a dish for a specific number of times. Our lightweight force memory achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches. Project page: https://qft-333.github.io/FM-VLA-Page/

Community

00