AI/ ai · reinforcement-learning · small-language-models · open-source

GRPO Fine-Tuning Boosts Small AI Models at Math, Falters Elsewhere

A new study found a popular reinforcement fine-tuning method sharpens small AI models at math but barely helps with coding or science quizzes.

A new study maps exactly where Group Relative Policy Optimization training helps small AI models, and where it stalls out.

Researchers ran a systematic test of GRPO, a memory-efficient reinforcement fine-tuning method for reasoning tasks, on language models between 1.5 and 7 billion parameters. The whole study ran on a single machine with eight A100 GPUs, a deliberately modest setup next to the GPU clusters usually associated with reinforcement learning research. They trained and evaluated multiple model families across three domains: math problems, coding tasks, and multiple-choice science questions, while tracking how group size affects training stability and how updates propagate through the model's layers. They also tested which LoRA (low-rank adaptation) modules respond best to this kind of training.

The first round of GRPO-tuned models beat their untrained base versions on roughly 80% of math benchmarks, but barely improved coding and multiple-choice science scores. That split matters for anyone outside a frontier lab trying to cheaply add reasoning skills to a small model: a technique billed as broadly useful for reasoning turns out to be domain-picky, and naive use of it can waste a GPU budget on tasks it was not going to fix.

The researchers used their own tensor-level diagnostics to retune the LoRA setup and reward shaping, closing some of that gap - proof that GRPO's reputation as a plug-and-play reasoning booster needs an asterisk.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →