New research puts a price tag on what reinforcement learning does to language models after post-training: more consistency, fewer ideas.
Researchers studied 14 pairs of base and post-trained large language models across four model families and three agentic benchmarks, 42 cases total. Post-training raised single-attempt accuracy, but it narrowed the range of tasks a model could solve when given many attempts. Base models paired with a lightweight prompting setup often solved more total problems than their post-trained counterparts once allowed a large sampling budget, despite much lower first-try accuracy. The team calls this gap the "sharpening tax" and argues post-training pushes tasks toward two extremes: always solved or never solved.
That split matters because production AI agents rarely get one shot. They take multiple turns, call tools, and retry. If post-training trades away a model's capacity to explore different solution paths, teams optimizing only for benchmark accuracy may be capping how much their systems improve with extra compute at run time. The researchers' proposed fix, a sampling method that adjusts temperature per prompt based on estimated difficulty, recovered some lost coverage in two agentic environments without giving up the accuracy gains.
Fine-tuning is not a free upgrade. It reshapes what a model can do, and this study is a tally of what gets reshaped away.