MEGA Hub

MoNe: Modular Neural Memory for Efficient Long Context Inference

Authors

Do you know Wonguk Cho?You can claim authorship or link another user.Do you know Kyubyung Chae?You can claim authorship or link another user.Do you know Tribhuvanesh Orekondy?You can claim authorship or link another user.Do you know Sunghyun Park?You can claim authorship or link another user.Do you know Hyoungwoo Park?You can claim authorship or link another user.Do you know Jeongho Kim?You can claim authorship or link another user.Do you know Arash Behboodi?You can claim authorship or link another user.Do you know Kyuwoong Hwang?You can claim authorship or link another user.Do you know Sungrack Yun?You can claim authorship or link another user.

Abstract

We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.

Community

00