MEGA Hub

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Authors

Do you know Hengyi Xie?You can claim authorship or link another user.Do you know Chenfei Yao?You can claim authorship or link another user.Do you know Xianjin Wu?You can claim authorship or link another user.Do you know Xuanyang Xi?You can claim authorship or link another user.Do you know Yiping Tang?You can claim authorship or link another user.Do you know Di Xu?You can claim authorship or link another user.Do you know Yingying Zhu?You can claim authorship or link another user.Do you know Dingkang Liang?You can claim authorship or link another user.Do you know Xiang Bai?You can claim authorship or link another user.Do you know Han Ding?You can claim authorship or link another user.

Abstract

Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional $V \to L \to A$ pathway as a direct $V + L \to A$ mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.

Community

00

Publication notes

Author note
Code is available at https://github.com/H-EmbodVis/TurboVLA