MEGA Hub

VibeVoice-ASR-BitNet Technical Report

Authors

Do you know Songchen Xu?You can claim authorship or link another user.Do you know Ting Song?You can claim authorship or link another user.Do you know Shaohan Huang?You can claim authorship or link another user.Do you know Zhiliang Peng?You can claim authorship or link another user.Do you know Yan Xia?You can claim authorship or link another user.Do you know Yujie Tu?You can claim authorship or link another user.Do you know Xin Huang?You can claim authorship or link another user.Do you know Jianwei Yu?You can claim authorship or link another user.Do you know Li Dong?You can claim authorship or link another user.Do you know Furu Wei?You can claim authorship or link another user.

Abstract

We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization (I8_S) with kernel fusion and SIMD optimization, while the autoregressive language model adopts BitNet-style ternary weights (I2_S). To preserve accuracy under aggressive compression, we employ a progressive quantization-aware training strategy. For inference, we implement custom SIMD kernels and fused operators within the ggml framework targeting both ARM and x86 platforms, achieving real-time recognition with RTF < 1 using as few as 3 CPU threads. VibeVoice-ASR-BitNet is 1.6-2.3x faster than Whisper.cpp at comparable model sizes (~1.6 GB), with only modest accuracy degradation compared to the FP16 baseline.

Community

00

Publication notes

Author note
Technical Report