AI/ reinforcement-learning · grpo · ai-training · research

FastRL Prunes Weak Rollouts to Speed Up Vision Model Training

A new arXiv paper details FastRL, which prunes weak training rollouts to roughly double GRPO-style RL speed and modestly lift accuracy.

FastRL trims reinforcement learning training time nearly in half by learning which practice attempts are worth keeping.

According to a new arXiv paper (arXiv:2609.36932), the framework targets Group Relative Policy Optimization (GRPO) and its variants, which train models by sampling many attempted solutions per question and scoring them against each other, an approach that's accurate but computationally expensive. FastRL adds two tricks: an advantage-aware pruning step that keeps only the most informative attempts while preserving diversity between them, and an adaptive sampling mechanism that adjusts how many attempts to generate based on how much pruning happened in earlier rounds. The authors report it plugs into GRPO, DAPO, and GSPO without modification. On two visual reasoning benchmarks, Geometry3K and GeoQA8K-R1V, they measured an average 2.07x training speedup and a roughly 1.64% accuracy gain, per the paper.

That combination, faster and more accurate at once, is the notable part, since most RL efficiency tricks trade one for the other. If the pruning genuinely discards redundant signal rather than useful signal, it's a rare free lunch for labs burning GPU hours on reasoning-focused post-training, a cost that scales quickly as models handle longer, multimodal outputs.

Those numbers come from the paper's own experiments on two geometry-focused benchmarks, not independent replication, and the promised code release has not landed yet. Reinforcement learning gains have a habit of shrinking once other labs try to reproduce them on different models and tasks.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →