MEGA Hub

Optimal Stopping of Self-Refining Foundation Models

Authors

Do you know Kim Hammar?You can claim authorship or link another user.Do you know Tansu Alpcan?You can claim authorship or link another user.Do you know Emil C. Lupu?You can claim authorship or link another user.

Abstract

Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in an iterative loop where it generates outputs, receives feedback from verifiers, and refines its responses through in-context learning. Following a novel approach, we formalize this process as an optimal stopping problem where the number of refinement iterations is decided based on expected improvement relative to cost. We derive optimal stopping policies and show that they can be efficiently computed through stochastic approximation. To evaluate our approach experimentally, we apply it to a coding benchmark for foundation models. The empirical results show that our stopping policies are significantly more cost-efficient than stopping policies proposed in prior work.

Community

00

Publication notes

Author note
Accepted at 65th IEEE Conference on Decision and Control (CDC 2026)