MEGA Hub

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

Authors

Do you know Zhuoran Song?You can claim authorship or link another user.Do you know Haozhe Jiang?You can claim authorship or link another user.Do you know Chunyu Qi?You can claim authorship or link another user.Do you know Minnan Pei?You can claim authorship or link another user.Do you know Gang Li?You can claim authorship or link another user.Do you know Xiaoyao Liang?You can claim authorship or link another user.Do you know Haibing Guan?You can claim authorship or link another user.

Abstract

Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot's execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation.

Community

00