NeutronGym grades AI agents on building real scientific instruments - and most of them flunk.
Researchers built an executable environment where language-model agents design neutron scattering instruments, with McStas simulation software ray-tracing their designs and a tiered grading system checking syntax, runtime, structure, and physics - no human or LLM judge involved. On a curated set of 16 tasks drawn from published instruments, seven tested models solved at most 7, none could retrieve a reference design from memory, and none hit the researchers' improvement target. Reinforcement learning did better. Training the smaller Qwen3-8B model on the environment's reward signal took it from 11% to 77% success on held-out versions of one instrument family, beating an untrained Qwen3-32B, and similar gains held across three more families.
The partial-credit grading turns out to matter more than the headline number. Strip it out and performance drops 60 points, meaning the model is learning incremental physics reasoning rather than gaming a pass-fail signal. The comparison that puts the 77% in perspective: that score only ties a classical optimizer's 81% when the optimizer is also handed the closed-form physics equations for the problem, a gap the researchers say is not statistically significant at this sample size. Hand the same tasks to frontier models and they solve 98-99% of them.
So the trained model has learned to approximate equations it was never given, which is real progress - just not the kind that beats a textbook, let alone a frontier lab's best model.