MEGA Hub

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Authors

Do you know Xiaomi Robotics Team?You can claim authorship or link another user.Do you know Jun Guo?You can claim authorship or link another user.Do you know Piaopiao Jin?You can claim authorship or link another user.Do you know Jason Li?You can claim authorship or link another user.Do you know Peiyan Li?You can claim authorship or link another user.Do you know Yingyan Li?You can claim authorship or link another user.Do you know Futeng Liu?You can claim authorship or link another user.Do you know Wanli Peng?You can claim authorship or link another user.Do you know Optimus Qin?You can claim authorship or link another user.Do you know Yifei Su?You can claim authorship or link another user.Do you know Nan Sun?You can claim authorship or link another user.Do you know Qiao Sun?You can claim authorship or link another user.Do you know Runze Suo?You can claim authorship or link another user.Do you know Heyun Wang?You can claim authorship or link another user.Do you know Yunhong Wang?You can claim authorship or link another user.Do you know Rujie Wu?You can claim authorship or link another user.Do you know Caoyu Xia?You can claim authorship or link another user.Do you know Lina Zhang?You can claim authorship or link another user.Do you know Jack Zhao?You can claim authorship or link another user.Do you know Guoliang Chen?You can claim authorship or link another user.Do you know Wenlong Chen?You can claim authorship or link another user.Do you know Xinze He?You can claim authorship or link another user.Do you know Bin Li?You can claim authorship or link another user.Do you know Qing Li?You can claim authorship or link another user.Do you know Zhuorong Li?You can claim authorship or link another user.Do you know Heng Qu?You can claim authorship or link another user.Do you know Wenxuan Song?You can claim authorship or link another user.Do you know Diyun Xiang?You can claim authorship or link another user.Do you know Yifan Xie?You can claim authorship or link another user.Do you know Peiran Xu?You can claim authorship or link another user.Do you know Hangjun Ye?You can claim authorship or link another user.Do you know Wen Ye?You can claim authorship or link another user.Do you know Han Zhao?You can claim authorship or link another user.Do you know Quanyun Zhou?You can claim authorship or link another user.

Abstract

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html

Community

00