MEGA Hub

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference

Authors

Do you know Zheng Liu?You can claim authorship or link another user.Do you know Zeyu Guo?You can claim authorship or link another user.Do you know Zihan Liu?You can claim authorship or link another user.Do you know Anbang Wu?You can claim authorship or link another user.Do you know Han Zhao?You can claim authorship or link another user.Do you know Fangxin Liu?You can claim authorship or link another user.Do you know Zhezhi He?You can claim authorship or link another user.Do you know Yinhe Han?You can claim authorship or link another user.Do you know Jingwen Leng?You can claim authorship or link another user.Do you know Minyi Guo?You can claim authorship or link another user.Do you know Yiming Gan?You can claim authorship or link another user.Do you know Yu Feng?You can claim authorship or link another user.

Abstract

Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices. In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference. First, we exploit the temporal similarity of consecutive inputs and propose a \textit{temporal-aware bit-sparsity} algorithm that computes only the differences between consecutive inputs, eliminating redundant bit-level operations. To further address the extra off-chip traffic introduced by our algorithm, we propose a \textit{speculative inference} technique, which amortizes data loading across multiple control steps. Lastly, to support these techniques, we co-design a dedicated accelerator with customized 1D systolic bit-serial PE arrays that eliminate PE workload imbalance. Our evaluation shows that Deltoris achieves up to 34.2$\times$ speedup over mobile GPUs and 6.1$\times$ over prior accelerators, while maintaining comparable accuracy.

Community

00