[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-dataset-grades-ai-science-agents-on-how-they-think":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},6281,"new-dataset-grades-ai-science-agents-on-how-they-think","New Dataset Grades AI Science Agents on How They Think","A public dataset of AI research agent trajectories shows top models succeed at similar rates but fail in very different ways.","Researchers have released a dataset that grades AI science agents on their reasoning, not just their final answers.\n\nCalled OpenDiscoveryTrace, the dataset logs complete step-by-step trajectories from AI agents tackling 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. Each step records nine fields: thoughts, tool calls, observations, errors, revision triggers, and confidence scores. The breakdown covers three frontier models (GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro) at 124 trajectories each, four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, Qwen2.5-1.5B) at 30 each, plus 60 live-retrieval variants, for 552 trajectories total. It's released under CC BY 4.0 with a trace schema, agent harness, and five benchmark tasks with baseline models included.\n\nMost AI benchmarks only check whether the output is right, which rewards lucky guesses as much as sound method. This one exposes the difference. In a pilot analysis of 363 trajectories, the three frontier models all succeeded 84-89% of the time, but Claude Opus 4.6 made 30 times more errors per trajectory than GPT-5.4 (2.5 vs. 0.08). Their mistakes looked nothing alike either: Claude's errors were mostly tool misuse, GPT-5.4's were mostly reasoning errors.\n\nThat's the real finding here: success rate alone would have told you these models are basically interchangeable for scientific work. It took the process trace to show they fail in opposite ways, which matters a lot if you're the one auditing an AI-generated hypothesis before it reaches a lab bench.","[\"ai-agents\",\"benchmarks\",\"ai-research\",\"open-source\"]","2026-09-11T04:00:00.000Z","2026-09-11T04:13:56.422Z","2026-09-11T04:14:08.271Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek and opening state 558 total trajectories, but the body's own breakdown (3 frontier models × 124 tasks + 4 open-weight models × 30 tasks = 492) doesn't add up to that total — include the source's 60 live-retrieval variant trajectories or otherwise reconcile the stated total with the listed components before publishing.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The body now includes the 60 live-retrieval trajectories, but the components still don't sum to 558 (124x3 + 30x4 + 60 = 552, not 558) — reconcile the stated total with the listed components or drop the precise total claim.","ai",[36,37,38,39],"ai-agents","benchmarks","ai-research","open-source",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.09203",0,{"sections":46},[47,50,54,58,63,68,73,76,81,85,90,95,100,105],{"name":48,"slug":34,"count":49,"latest_published_at":18},"AI",3507,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",636,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Policy","policy",338,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",153,"2026-09-09T15:12:32.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":74,"slug":75,"count":71,"latest_published_at":18},"Science","science",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":18},"Dev Tools","dev-tools",70,{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]