Researchers built a smarter traffic controller for AI training, and it squeezes a bit more accuracy out of the same compute budget.
FSPO is a feedback-state controller for reinforcement-learning post-training of large language models, the kind of training that adjusts knobs like rollout temperature, group size, and verifier allocation while a model trains under a fixed compute budget. It replaces ad hoc tuning with a risk model trained to match the exact controller making future decisions, a calibration method called DCTC that recalibrates risk estimates as the controller's own choices shift the data it sees, and a feasibility checker called PRCC that blocks any action that would leave the remaining budget unable to finish the run. Tested under a matched compute budget against PB2, the strongest adaptive baseline in the paper, FSPO reached 66.11% held-out accuracy and 59.43% out-of-distribution accuracy, versus 64.47% and 57.03% for PB2. The same components also cut calibration error (ECE) from 0.108 to 0.053 and, on an 18-action catalog, eliminated false-feasible admissions that previously happened 19.7% of the time.
RL post-training runs are expensive, and most resource controllers either tune knobs with simple heuristics or risk promising more compute than they can actually deliver mid-run. FSPO's real contribution is structural: it targets specific failure modes, like a risk model that doesn't match the controller it's paired with, or a budget plan that quietly becomes infeasible, rather than just chasing a bigger accuracy number.
The accuracy edge over PB2 is real but modest, 1.64 points held-out and 2.40 points out-of-distribution, so the more convincing result here is operational: trajectory failures dropped from 18.1% to 8.3% once all three components were switched on.