AI/ ai · reinforcement-learning · llm-training · research

A Frozen Critic Can Replace Reward Labels in LLM Training

A new arXiv paper (2609.37119) shows a frozen critic can match PPO's results without reward labels, cutting compute for long reasoning tasks.

A new training method throws out reward labels entirely and still matches a labeled baseline.

Researchers posted "Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training" to arXiv on September 30, 2026, as arXiv:2609.37119, cross-listed in cs.AI. The paper pushes back on a recent trend: stripping the critic model out of reinforcement learning post-training to cut memory use and training instability. The authors argue a well-trained critic - one that predicts whether an unfinished answer will eventually succeed - is more valuable kept than thrown away once training ends. Their method, Reward-Free Policy Optimization (RFPO), freezes a single calibrated critic and reuses it three ways: as the reward for a finished rollout, as a baseline for advantage estimates, and as a forecaster that scores incomplete prefixes before generation is done. Binarizing the critic's score, the authors say, stops the policy from gaming the critic's bias toward longer outputs, and the resulting system matches supervised PPO - the standard reinforcement-learning baseline that depends on labeled reward data - without a single label in the training loop.

For long chain-of-thought tasks, where a model can generate thousands of tokens before anyone knows if the answer is right, waiting for a full rollout to score it is expensive. RFPO scores trajectories before they finish, so training no longer pays to wait on every generation - which the authors say cuts both compute and memory overhead. That is a direct challenge to two years of RL post-training work that has treated critics as the unstable, memory-hungry piece worth removing.

If independent replication holds up, it flips a design assumption baked into a lot of today's RLHF tooling: that a critic is dead weight once training ends, rather than a forecasting tool worth keeping.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →