A new paper shows AI models can learn from tasks they almost never solve, just by writing themselves a short note after each failed attempt.
Reinforcement learning with verifiable rewards, the standard way labs fine-tune models today, works by having a model attempt a task many times and rewarding the attempts that succeed. That approach falls apart when a task is so hard the model succeeds zero or nearly zero times out of 128 tries, since there's nothing to reward. Researchers built a method called RLTL;DR that instead shows the model its own failed attempt plus the verifier's feedback, and asks it to write a one-line takeaway. Each new attempt is conditioned on every takeaway written so far, and the model keeps trying until it solves the problem or runs out of chances.
On coding and tool-calling tasks picked specifically because a baseline model solved zero out of 128 attempts, standard training on a Qwen 3.5 9B Thinking model stayed stuck at 0-1% success. Adding RLTL;DR's self-written notes pushed that to 14-31% success during training, and the model still solved 12-13% of problems even when its own notes weren't shown to it at test time, meaning the lesson had been baked into the weights rather than propped up by a crutch. A stripped-down version, trained on just 4,000 task-and-note pairs with no attempt at the actual problem, recovered almost all of that gain.
It's a neat trick for turning total failure into partial success, but 31% is still a coin flip at best, and it's one arXiv preprint on one model family, not yet a fix for the broader ceiling on what today's models can reason through.