AI/ reinforcement-learning · ai-training · entropy-collapse · llm-reasoning

Researchers Find Fewer Training Rollouts Improve AI Reasoning

A new method called GRPODropout drops some high-probability rollouts before each update, boosting accuracy and exploration with less data.

A tweak to a popular AI training algorithm improves reasoning models by training on fewer examples, not more.

Researchers propose GRPODropout, a modification to GRPO (Group Relative Policy Optimization), the reinforcement learning method widely used to sharpen large language model reasoning. GRPO training often suffers from "entropy collapse" - the model's outputs get less diverse over time, which chokes off exploration and stalls further gains. Instead of changing the reward function or adding regularization like prior fixes, GRPODropout targets which generated rollouts get used in each update: it discards a small set of high-probability, positive-advantage rollouts and recenters the remaining ones before the policy update. The authors back the design with a rollout-level theoretical analysis that guides how many rollouts to drop.

In experiments, GRPODropout beat standard GRPO on accuracy while keeping actor entropy higher - meaning the model retained more diverse reasoning paths - despite using fewer rollouts per update. That matters because most entropy-collapse fixes add complexity, like extra regularization terms or reward shaping, rather than simplifying what's already there. This result suggests rollout selection, not just reward design, is an underused lever for making reasoning-model training more efficient.

The code is on GitHub, but the gains are reported on the authors' own benchmarks - the real test is whether this holds up once bigger labs with messier training pipelines try it on their own models.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →