A finished reinforcement learning run may already contain a better model than the one you deployed.
Researchers built SURGE (Scaling Up RL Gradient-free via Eigenspace fusion), a method that fuses two checkpoints from a single, already-completed RL run: a high-accuracy "anchor" and a "donor" that produces shorter answers. It treats each checkpoint as a set of weight changes from their shared starting point, then uses spectral decomposition to keep the anchor's strongest component while folding in the donor's. No extra training, no extra inference cost per query. Tested on two 1.5B math-reasoning histories (DeepSeek, Nemotron) and one 7B coding history (OLMo), the fused model beat every checkpoint the original run ever produced - scoring 54.17% on DeepSeek's AIME24 benchmark against a 50.83% ceiling from the actual training curve, while using fewer reasoning tokens.
That's the interesting part. Most efforts to improve AI reasoning assume you need to spend more: longer RL runs, bigger models, more inference compute per answer. SURGE got a gain by excavating value from a run that had already finished and been filed away.
It's a small-model, narrow-benchmark demonstration, not proof this scales to frontier systems - but if mining old training runs for a free upgrade holds up, it is a lot cheaper than training a new one.