AI/ reinforcement-learning · llm-agents · ai-research · fine-tuning

A Cheaper Way to Fine-Tune AI Agents Without a Critic Model

A new paper swaps GRPO's repeated rollouts and critic models for ordinal filtering on replay data, matching performance on agent benchmarks.

A new training method lets AI agents learn from trial and error without the usual critic model or repeated rollouts.

A team describes Follow the Winners, a critic-free reinforcement fine-tuning algorithm for agentic large language models. Instead of GRPO's approach of running several rollouts per task and averaging them into a baseline, FTW filters samples from a replay buffer using an ordinal cutoff borrowed from the cross-entropy method. The authors frame the method through a control-as-inference lens that also explains GRPO and DPO as special cases, with GRPO staying risk-neutral while DPO and FTW share a bounded risk-seeking offset. Tested on the Sokoban and Search-R1 benchmarks, FTW matched GRPO and PPO in performance while trading a value model or group rollouts for plain CPU memory.

That trade matters for agents operating in environments where repeated rollouts are expensive or impossible, like live production services or security sandboxes, since you cannot rerun a real customer session five times just to compute a baseline. It also suggests critic-free post-training does not have to mean accepting more variance or noisier updates on long, sparsely rewarded tasks.

The usual caveat applies: two toy benchmarks are a long way from the messy, expensive environments this method is pitched for, and matching a baseline is not the same as beating it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →