AI/ reinforcement-learning · llm-post-training · ai-research · compute-budget

A new controller for AI post-training beats its best rival

A new budget-aware RL controller, FSPO, edges past the strongest prior baseline on accuracy while cutting miscalibration and blocking infeasible moves.

Researchers built a smarter traffic controller for AI training, and it squeezes a bit more accuracy out of the same compute budget.

FSPO is a feedback-state controller for reinforcement-learning post-training of large language models, the kind of training that adjusts knobs like rollout temperature, group size, and verifier allocation while a model trains under a fixed compute budget. It replaces ad hoc tuning with a risk model trained to match the exact controller making future decisions, a calibration method called DCTC that recalibrates risk estimates as the controller's own choices shift the data it sees, and a feasibility checker called PRCC that blocks any action that would leave the remaining budget unable to finish the run. Tested under a matched compute budget against PB2, the strongest adaptive baseline in the paper, FSPO reached 66.11% held-out accuracy and 59.43% out-of-distribution accuracy, versus 64.47% and 57.03% for PB2. The same components also cut calibration error (ECE) from 0.108 to 0.053 and, on an 18-action catalog, eliminated false-feasible admissions that previously happened 19.7% of the time.

RL post-training runs are expensive, and most resource controllers either tune knobs with simple heuristics or risk promising more compute than they can actually deliver mid-run. FSPO's real contribution is structural: it targets specific failure modes, like a risk model that doesn't match the controller it's paired with, or a budget plan that quietly becomes infeasible, rather than just chasing a bigger accuracy number.

The accuracy edge over PB2 is real but modest, 1.64 points held-out and 2.40 points out-of-distribution, so the more convincing result here is operational: trajectory failures dropped from 18.1% to 8.3% once all three components were switched on.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →