A new study pokes a hole in one of AI reasoning's cleverer tricks: grafting good steps into bad chains to rescue stalled logic does basically nothing.
Researchers tested PRM-Pruned Fragment Grafting, a technique where a process reward model scores parallel chains of reasoning, then copies the best-performing fragment into a chain that looks stuck. Running it on Qwen2.5-7B-Instruct with the Math-Shepherd reward model across the full MATH500 benchmark (500 problems, three seeds), the grafting method came back statistically indistinguishable from a plain parallel chain-of-thought baseline with no grafting at all. Digging into why, the team found only 14% of grafting events actually hit a chain that was genuinely struggling - the rest landed on chains that had already succeeded, were nearly done, or were stuck on a flat reward plateau that a graft can't fix. A random-injection control matched the same results while firing 2.4 times more often, and the pattern held across three base models, six benchmarks, and a second reward model.
The finding matters because parallel chain-of-thought sampling is a go-to fix for AI reasoning collapse, and fragment grafting has been treated as a near-free upgrade to it. This paper shows the targeting problem, not the grafting mechanism, is the bottleneck - a reward model usually can't find the chain that needs rescuing before it's too late to matter. A hindsight oracle, with perfect foresight, could only squeeze out a 0.13 percentage point gain over doing nothing special.
It's a useful reminder that an intervention sounding mechanistically sensible is not the same as it working - sometimes the fancy fix and the coin flip land in the same place.