[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-agents-learn-terminal-tasks-from-recycled-lab-software":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},10007,"ai-agents-learn-terminal-tasks-from-recycled-lab-software","AI Agents Learn Terminal Tasks From Recycled Lab Software","Researchers reused existing scientific software as training fodder for terminal AI agents, lifting a smaller model's benchmark score by nearly six points.","A new paper shows how to turn old scientific software into free training data for AI terminal agents, and the trick actually moves the needle.\n\nThe method, called software-in-the-loop reconstruction, builds training environments from 500 existing software workflows spanning 46 families across six scientific domains. For each workflow, researchers ran several input configurations and split the results into public examples and hidden tests. An AI model, Qwen3.8-Max, was given the task instructions, the input schema, and the public input-output pairs, then asked to rebuild a working program without ever seeing the original source code. A multi-part verifier checked the rebuilt program's output against the hidden test cases for semantic correctness, structural validity, and attempts to game the test. Across three attempts per task, the model solved 838 of these task instances, producing 1,422 verified solution runs that were oversampled into 3,000 training examples.\n\nThat matters because building agent training environments usually means hand-writing a reference answer and a custom checker for every single task, which is why most agent benchmarks stay confined to software engineering. Harvesting supervision from software that already exists sidesteps that bottleneck and lets it reach into chemistry, biology, and other lab domains. After fine-tuning a smaller model, Qwen3.8-27B, on the resulting data, its average score on Terminal-Bench 2 climbed from 47.94% to 53.56% across three seeds, beating three other training sets of equal token count on every evaluation reported.\n\nA six-point jump from recycling someone else's old code is a real result, not hype, but it is one model family tested on one new benchmark. Whether borrowed software as a verifier generalizes to messier, less-documented lab tools is the open question this paper doesn't answer yet.","[\"llm training\",\"terminal agents\",\"benchmarks\",\"self-supervised learning\"]","2026-10-05T04:00:00.000Z","2026-10-05T17:53:55.867Z","2026-10-05T17:54:02.115Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The body states the model solved 838 of the tasks after saying only 500 workflows\u002Ftasks were tested, an internally inconsistent number.","resolved","ai",[32,33,34,35],"llm training","terminal agents","benchmarks","self-supervised learning",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.02710",0,{"sections":42},[43,46,50,55,60,65,69,74,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",6233,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",868,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",177,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":18},"Dev Tools","dev-tools",97,{"name":79,"slug":80,"count":77,"latest_published_at":81},"Software","software","2026-10-04T10:00:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]