MEGA Hub

What Matters for Latent Actions in Robot Learning

Authors

Do you know Xizhou Bu?You can claim authorship or link another user.Do you know Qingda Hu?You can claim authorship or link another user.Do you know Lei Zhou?You can claim authorship or link another user.Do you know Lingfeng Zhang?You can claim authorship or link another user.Do you know Yingbo Tang?You can claim authorship or link another user.Do you know Zihao Liu?You can claim authorship or link another user.Do you know Xinyi Tao?You can claim authorship or link another user.Do you know Zhiqiang Ma?You can claim authorship or link another user.Do you know Qingqiu Huang?You can claim authorship or link another user.Do you know Chufeng Tang?You can claim authorship or link another user.Do you know Hongbo Wang?You can claim authorship or link another user.Do you know Jing Zhang?You can claim authorship or link another user.Do you know Jiayi Ma?You can claim authorship or link another user.Do you know Hangjun Ye?You can claim authorship or link another user.Do you know Wei Li?You can claim authorship or link another user.Do you know Xiaoshuai Hao?You can claim authorship or link another user.

Abstract

Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite rapid progress, research on LAM remains highly fragmented, with existing methods evaluating different design choices in isolation under inconsistent experimental settings, making it difficult to identify the factors that truly determine downstream robotic manipulation performance. In this work, we present the first comprehensive empirical study of latent action learning for robotic manipulation. We unify representative LAM methods within a common autoencoding framework and systematically investigate 41 LAM design choices across three dimensions, including latent action modeling paradigms, learning objectives and regularization methods, and latent action integration strategies. We further examine four proxy metrics for evaluating latent action quality and assess their ability to reliably predict downstream robotic manipulation performance. Extensive experiments on three widely used benchmarks provide strong empirical evidence that fine-tuning vision-language model (VLM) backbones with latent actions provides a stronger initialization for downstream policy learning, with further validation on real-world robot manipulation tasks.

Community

00

Publication notes

Author note
Project page: https://carldegio.github.io/latent_action.github.io