MEGA Hub

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Authors

Do you know Yutong Chen?You can claim authorship or link another user.Do you know Shouqian Shi?You can claim authorship or link another user.Do you know Xinran Liu?You can claim authorship or link another user.Do you know Haochen Wang?You can claim authorship or link another user.Do you know Jiaying Wang?You can claim authorship or link another user.Do you know Tianxing Xu?You can claim authorship or link another user.Do you know Yuanxi Wang?You can claim authorship or link another user.Do you know Zirui Ding?You can claim authorship or link another user.

Abstract

Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while reducing measured inference latency. These results show that latent refinement can be localized to a narrow decoder interval, reducing repeated full-decoder execution without generating a long visible reasoning trace and providing a practical accuracy-efficiency tradeoff for decoder-only Transformer models.

Community

00

Publication notes

Author note
8 pages, 2 figures