AI labs training large language models are sitting on a pile of discarded work, and a new paper argues they should be reusing it.
Researchers behind a preprint called ROSS (Relearning from Self-Generated Rollouts through Selective Supervision) found that the practice attempts a model produces during reinforcement learning and on-policy distillation don't have to be thrown away once the model improves. Instead of retraining from scratch, ROSS keeps an old attempt's full trajectory as context but applies its training signal only to the specific continuations worth keeping, filtering out the mistakes, abandoned tries, and redundant moves baked into raw rollouts. The results come from the paper's own authors and haven't been independently verified, but across math, code generation, instruction following, and software engineering tasks, the method consistently improved already-trained upstream checkpoints over standard baselines. On the Qwen3.6-35B-A3B model, it pushed a six-benchmark average from 58.40% to 62.20% and the SWE-bench Verified coding benchmark from 64.20% to 68.40%.
The appeal is efficiency: this is offline supervised fine-tuning on data the model already generated, not another expensive round of reinforcement learning. That matters because RL fine-tuning for coding and agentic models is one of the more compute-hungry parts of building modern LLMs, and ROSS suggests some labs are throwing away usable training signal every time they move on from an old policy.
It's not a new model or a new architecture, just a cheaper way to wring more performance out of ones that already exist - the kind of unglamorous efficiency trick that tends to get quietly folded into every lab's training pipeline once it holds up outside the paper that proposed it.