A new paper asks a blunt question: when an AI model updates itself mid-search, is it actually getting smarter, or just riding a bigger pile of search data?
The paper looks at automatic heuristic design, where a large language model proposes and refines heuristics - rules of thumb paired with runnable code - and a task-specific evaluator scores each attempt. Most of these systems freeze the model and just keep searching. A few, including EvoTune and CALM, instead retrain the model using reinforcement learning with verifiable rewards, feeding it signals built from evaluated candidates. The authors test several ways of turning program validity, performance, and the surrounding prompt and search context into those training signals, using shared rollouts and matched training budgets so the comparisons are apples to apples.
The real test is whether online updating beats simply running more search with a frozen generator, once compute cost is held equal. That matters because "self-improving" claims in this field often rest on end-to-end search results alone, without separating whether the model itself got better from whether the search process just accumulated more useful history to lean on.
It is a useful dose of methodological hygiene for a subfield that has been quick to credit the model when the search loop may be doing most of the work.