A new study maps exactly where Group Relative Policy Optimization training helps small AI models, and where it stalls out.
Researchers ran a systematic test of GRPO, a memory-efficient reinforcement fine-tuning method for reasoning tasks, on language models between 1.5 and 7 billion parameters. The whole study ran on a single machine with eight A100 GPUs, a deliberately modest setup next to the GPU clusters usually associated with reinforcement learning research. They trained and evaluated multiple model families across three domains: math problems, coding tasks, and multiple-choice science questions, while tracking how group size affects training stability and how updates propagate through the model's layers. They also tested which LoRA (low-rank adaptation) modules respond best to this kind of training.
The first round of GRPO-tuned models beat their untrained base versions on roughly 80% of math benchmarks, but barely improved coding and multiple-choice science scores. That split matters for anyone outside a frontier lab trying to cheaply add reasoning skills to a small model: a technique billed as broadly useful for reasoning turns out to be domain-picky, and naive use of it can waste a GPU budget on tasks it was not going to fix.
The researchers used their own tensor-level diagnostics to retune the LoRA setup and reward shaping, closing some of that gap - proof that GRPO's reputation as a plug-and-play reasoning booster needs an asterisk.