MEGA Hub

LoopMTP: A looped transformer guided by latent multi-token prediction

Authors

Do you know Behzad Shomali?You can claim authorship or link another user.Do you know Markus Frey?You can claim authorship or link another user.Do you know David Berghaus?You can claim authorship or link another user.Do you know Joachim Koehler?You can claim authorship or link another user.Do you know Mehdi Ali?You can claim authorship or link another user.

Abstract

Looped transformers have emerged as a parameter-efficient alternative to scaling depth for strong reasoning. By reusing one stack of layers across $T$ iterations, they attain the effective depth and reasoning capabilities of larger models at a fixed parameter count. Yet existing approaches suffer from latent overthinking and undifferentiated computation, largely because intermediate representations receive no guidance across loops. Multi-token prediction (MTP) supplies exactly the dense, forward-looking supervision the loop is missing. We propose \textsc{LoopMTP}, which links the two through a structural correspondence in latent space: a model that loops $T$ times can anticipate $T$ future tokens. \textsc{LoopMTP} realizes this by softly aligning the hidden state of loop $t$ with the embedding of the token $t$ steps ahead, while a lightweight gate preserves useful information across iterations. \textsc{LoopMTP} improves average accuracy by up to 8.1\% (relative) over the non-looped baseline, with training remaining stable for up to 15 loops.

Community

00