[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-find-old-ai-training-runs-still-have-more-to-give":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},8568,"researchers-find-old-ai-training-runs-still-have-more-to-give","Researchers Find Old AI Training Runs Still Have More To Give","ROSS recycles discarded AI training attempts as fresh lessons, lifting one model's SWE-bench score from 64.2% to 68.4% with no new rollouts.","AI labs training large language models are sitting on a pile of discarded work, and a new paper argues they should be reusing it.\n\nResearchers behind a preprint called ROSS (Relearning from Self-Generated Rollouts through Selective Supervision) found that the practice attempts a model produces during reinforcement learning and on-policy distillation don't have to be thrown away once the model improves. Instead of retraining from scratch, ROSS keeps an old attempt's full trajectory as context but applies its training signal only to the specific continuations worth keeping, filtering out the mistakes, abandoned tries, and redundant moves baked into raw rollouts. The results come from the paper's own authors and haven't been independently verified, but across math, code generation, instruction following, and software engineering tasks, the method consistently improved already-trained upstream checkpoints over standard baselines. On the Qwen3.6-35B-A3B model, it pushed a six-benchmark average from 58.40% to 62.20% and the SWE-bench Verified coding benchmark from 64.20% to 68.40%.\n\nThe appeal is efficiency: this is offline supervised fine-tuning on data the model already generated, not another expensive round of reinforcement learning. That matters because RL fine-tuning for coding and agentic models is one of the more compute-hungry parts of building modern LLMs, and ROSS suggests some labs are throwing away usable training signal every time they move on from an old policy.\n\nIt's not a new model or a new architecture, just a cheaper way to wring more performance out of ones that already exist - the kind of unglamorous efficiency trick that tends to get quietly folded into every lab's training pipeline once it holds up outside the paper that proposed it.","[\"ai training\",\"reinforcement learning\",\"llm benchmarks\",\"arxiv research\"]","2026-09-30T04:00:00.000Z","2026-09-30T11:54:49.607Z","2026-09-30T11:54:53.921Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add explicit attribution for the benchmark figures — name the arXiv paper (ID\u002Fdate, e.g. arXiv:2609.35954) and note these are the authors' own self-reported results, since the draft currently cites specific numbers to unnamed 'researchers' with no citable source for verification.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The attribution issue is fixed, but the closing paragraph is caveat-only and ends abruptly — fold the preprint\u002Fself-reported caveat earlier (e.g., into the benchmark paragraph) and close with a substantive final thought instead of leaving the piece on a bare caveat.","ai",[36,37,38,39],"ai training","reinforcement learning","llm benchmarks","arxiv research",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.35954",0,{"sections":46},[47,50,54,58,63,68,73,78,83,87,92,97,102,107],{"name":48,"slug":34,"count":49,"latest_published_at":18},"AI",5105,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",785,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Policy","policy",417,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]