AI/ llm training · terminal agents · benchmarks · self-supervised learning

AI Agents Learn Terminal Tasks From Recycled Lab Software

Researchers reused existing scientific software as training fodder for terminal AI agents, lifting a smaller model's benchmark score by nearly six points.

A new paper shows how to turn old scientific software into free training data for AI terminal agents, and the trick actually moves the needle.

The method, called software-in-the-loop reconstruction, builds training environments from 500 existing software workflows spanning 46 families across six scientific domains. For each workflow, researchers ran several input configurations and split the results into public examples and hidden tests. An AI model, Qwen3.8-Max, was given the task instructions, the input schema, and the public input-output pairs, then asked to rebuild a working program without ever seeing the original source code. A multi-part verifier checked the rebuilt program's output against the hidden test cases for semantic correctness, structural validity, and attempts to game the test. Across three attempts per task, the model solved 838 of these task instances, producing 1,422 verified solution runs that were oversampled into 3,000 training examples.

That matters because building agent training environments usually means hand-writing a reference answer and a custom checker for every single task, which is why most agent benchmarks stay confined to software engineering. Harvesting supervision from software that already exists sidesteps that bottleneck and lets it reach into chemistry, biology, and other lab domains. After fine-tuning a smaller model, Qwen3.8-27B, on the resulting data, its average score on Terminal-Bench 2 climbed from 47.94% to 53.56% across three seeds, beating three other training sets of equal token count on every evaluation reported.

A six-point jump from recycling someone else's old code is a real result, not hype, but it is one model family tested on one new benchmark. Whether borrowed software as a verifier generalizes to messier, less-documented lab tools is the open question this paper doesn't answer yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →