[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-neutrongym-tests-ai-agents-on-real-instrument-design":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9957,"neutrongym-tests-ai-agents-on-real-instrument-design","NeutronGym Tests AI Agents on Real Instrument Design","A benchmark grades AI agents on neutron instrument design, and RL training nearly matches a classical optimizer only when given the same equations.","NeutronGym grades AI agents on building real scientific instruments - and most of them flunk.\n\nResearchers built an executable environment where language-model agents design neutron scattering instruments, with McStas simulation software ray-tracing their designs and a tiered grading system checking syntax, runtime, structure, and physics - no human or LLM judge involved. On a curated set of 16 tasks drawn from published instruments, seven tested models solved at most 7, none could retrieve a reference design from memory, and none hit the researchers' improvement target. Reinforcement learning did better. Training the smaller Qwen3-8B model on the environment's reward signal took it from 11% to 77% success on held-out versions of one instrument family, beating an untrained Qwen3-32B, and similar gains held across three more families.\n\nThe partial-credit grading turns out to matter more than the headline number. Strip it out and performance drops 60 points, meaning the model is learning incremental physics reasoning rather than gaming a pass-fail signal. The comparison that puts the 77% in perspective: that score only ties a classical optimizer's 81% when the optimizer is also handed the closed-form physics equations for the problem, a gap the researchers say is not statistically significant at this sample size. Hand the same tasks to frontier models and they solve 98-99% of them.\n\nSo the trained model has learned to approximate equations it was never given, which is real progress - just not the kind that beats a textbook, let alone a frontier lab's best model.","[\"ai\",\"reinforcement-learning\",\"benchmarks\",\"physics\"]","2026-10-05T04:00:00.000Z","2026-10-05T15:19:18.330Z","2026-10-05T15:19:23.272Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The physics-equation sentence inverts the source's finding — the source says the trained model only ties the classical optimizer when it IS handed the closed-form physics equations, not when it isn't; fix this reversed conditional before publication.","resolved","ai",[30,32,33,34],"reinforcement-learning","benchmarks","physics",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.03631",0,{"sections":41},[42,45,49,54,59,64,68,73,77,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",6170,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",860,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",177,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Dev Tools","dev-tools",97,{"name":78,"slug":79,"count":76,"latest_published_at":80},"Software","software","2026-10-04T10:00:00.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]