Rewarding an AI for the right answer does not tell it how to produce that answer again from different inputs, and a new study pins down exactly what has to happen before reward-based training actually works.
Researchers tested the idea on pretrained Qwen2.5 checkpoints across tasks requiring step-by-step state tracking and memory retrieval. When models received correct source information and first-operation supervision, they hit 82.61% success on a set of eight test environments, compared to 44.15% for a control trained on randomized, task-irrelevant source data. A second, independent set of eight environments produced a similar gap: 75.32% success when memory replay preserved retrieval ability during reward adaptation, versus 49.86% when models were trained on mismatched retrieval examples instead. The team also ran the comparison on GSM8K math problems and HotpotQA reading-comprehension questions, separately measuring how much a model already knew before reward training started, how much it gained, and where it ended up.
The takeaway is that reward signals are thin instructions. They can confirm an answer is correct, but they cannot, by themselves, teach a model the underlying computation or retrieval steps needed to get there on a new problem. That groundwork has to come from pretraining, or a later stage the paper calls midtraining, before reward-based fine-tuning is applied. For anyone fine-tuning a model with rewards, the finding suggests throwing more reward data at a model will not fix gaps that pretraining never filled.
Fine-tuning with rewards, in other words, edits a model. It does not teach it to think.