MEGA Hub

Progressive Cramming: Reliable Token Compression and What It Reveals

Authors

Do you know Dmitrii Tarasov?You can claim authorship or link another user.Do you know Timofei Lashukov?You can claim authorship or link another user.Do you know Elizaveta Goncharova?You can claim authorship or link another user.Do you know Andrey Kuznetsov?You can claim authorship or link another user.

Abstract

Token cramming compresses sequences into learned embeddings with near-perfect reconstruction, but fixed token budgets and 99\% accuracy thresholds leave it unclear whether residual errors reflect optimization failures or fundamental limits. We introduce progressive cramming, which grows the target prefix token-by-token, stopping only when reconstruction is no longer achievable within a fixed optimization budget. Progressive trajectories occupy low-dimensional structure in embedding space. Prepending a crammed embedding causes a moderate but consistent accuracy drop on multiple-choice benchmarks even with the original prefix in context, and collapses capability almost entirely under generative evaluation. Causal attention-knockout interventions trace this degradation to the embedding's interactions in the model's early layers. These results position progressive cramming as a tool for studying compression limits and show that perfect reconstruction - achievable through brittle steering rather than transferable semantics - is insufficient for meaningful compression.

Community

00