MEGA Hub

Unified Lookup-Table Inference with Signed-Digit K/V Caches for Ternary LLMs

Authors

Do you know Ziang Duan?You can claim authorship or link another user.Do you know Jiajun Wu?You can claim authorship or link another user.Do you know Zetian Chen?You can claim authorship or link another user.Do you know Hao Song?You can claim authorship or link another user.Do you know Yanwen Deng?You can claim authorship or link another user.Do you know Zixuan Shen?You can claim authorship or link another user.Do you know Nuobei Xie?You can claim authorship or link another user.Do you know Simo Wu?You can claim authorship or link another user.Do you know Bolun Wang?You can claim authorship or link another user.Do you know Peng Zhou?You can claim authorship or link another user.Do you know Chao Wang?You can claim authorship or link another user.

Abstract

Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separate higher-precision engine. Compressing this cache alone does not resolve the mismatch. To execute attention with the same lookup-table machinery as ternary projections, values accumulated in one reduction must retain a compatible representation and scale. This requirement also differs for keys and values during causal decoding, because newly generated values may belong to an unfinished cache block. This work develops a unified lookup-table inference approach for ternary LLMs. It stores runtime K/V states as scaled multi-plane signed digits organized around the reduction structure of attention. The resulting digit planes are consumed directly by activation-derived tables, avoiding dense K/V materialization between cache storage and attention computation. The design combines online K/V formation, bounded handling of incomplete value blocks, and a shared multi-stream datapath for Linear projections and attention. A constraint-guided search selects the representation and execution policy for a target quality--efficiency trade-off. Experiments on native and post-training ternary models validate the approach across cache capacity, model quality, and hardware efficiency.

Community

00