Optimal stopping of self-refining foundation models
File(s) cdc_camera_ready_kim_tansu_emil.pdf (486.76 KB)
Accepted version
Author(s)
Hammar, Kim
Alpcan, Tansu
Lupu, Emil
Type
Conference Paper
Abstract
Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in an iterative loop where it generates outputs, receives feedback from verifiers, and refines its responses through in-context learning. Following a novel approach, we formalize this process as an optimal stopping problem where the number of refinement iterations is decided based on expected improvement relative to cost. We derive optimal stopping policies and show that they can be efficiently computed through stochastic approximation. To evaluate our approach experimentally, we apply it to a coding benchmark for foundation models. The empirical results show that our stopping policies are significantly more cost-efficient than stopping policies proposed in prior work.
Date Acceptance
2026-07-15
Citation
IEEE Conference on Decision and Control (CDC)
Publisher
IEEE
Journal / Book Title
IEEE Conference on Decision and Control (CDC)
Copyright Statement
Subject to copyright. This paper is embargoed until publication. Once published the author’s accepted manuscript will be made available under a CC-BY License in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy).
License URL
Source
IEEE Conference on Decision and Control (CDC) 2026
Publication Status
Accepted
Start Date
2026-12-15
Finish Date
2026-12-18
Coverage Spatial
Honolulu, Hawaii
