MEGA Hub

Learning What to Remember: Test-Time Training via Context Distillation

Authors

Do you know Zixuan Wang?You can claim authorship or link another user.Do you know Xingyu Dang?You can claim authorship or link another user.Do you know Rui-Jie Zhu?You can claim authorship or link another user.Do you know Zixin Wen?You can claim authorship or link another user.Do you know Hengyu Fu?You can claim authorship or link another user.Do you know Wenhao Chai?You can claim authorship or link another user.Do you know Jason D. Lee?You can claim authorship or link another user.

Abstract

Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-context modeling, yet existing TTT methods only optimize either reconstruction or online adaptation objectives without considering the future utility of retained information. In this work, we propose \textbf{T}est-\textbf{T}ime \textbf{C}ontext \textbf{D}istillation (TTCD), a TTT framework that introduces a self-supervised objective for allocating limited memory capacity for future use. Specifically, TTCD uses a long-window teacher to supervise the fast weights of a short-window student, where the hidden-state discrepancy between them offers a dense, self-supervised signal guiding the model to memorize the contextual information crucial for future token predictions. We focus on an in-place variant: In-Place TTCD (IP-TTCD), which uses the existing MLP parameters as the fast weights. Experiments on long-context language modeling tasks show IP-TTCD consistently outperforms DeltaNet, Gated DeltaNet, sliding-window attention, and TTT when pre-trained from scratch. Furthermore, IP-TTCD allows pre-trained transformer models to adapt their parameters during inference through continual pre-training, gaining long-context capabilities with only a lightweight architectural augmentation. Our results position TTCD as a step toward architectural continual learning.

Community

00