[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-ai-benchmark-exposes-cracks-in-automated-science-curation":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8091,"new-ai-benchmark-exposes-cracks-in-automated-science-curation","New AI Benchmark Exposes Cracks in Automated Science Curation","FlyAOC, a new Drosophila gene curation benchmark, finds AI agents' performance swings with harness design and tool use, not just model choice.","A new benchmark called FlyAOC checks whether AI agents can do the unglamorous work of curating scientific databases, and the results depend more on how the agent is built than on which model powers it.\n\nFlyAOC evaluates agents on the full curation workflow instead of isolated subtasks like named entity recognition. Given a gene symbol, a short FlyBase description, a 16,898-paper corpus, and ontology resources, an agent has to search the literature and reconstruct curator-grade annotations, including standardized function terms, expression patterns, and historical name synonyms. The benchmark grades that output against 7,397 expert-curated annotations covering 100 genes drawn from FlyBase, the Drosophila research community's knowledge base. The researchers tested four agent designs, a memorization baseline, a fixed pipeline, a single agent, and a multi-agent setup, and found that performance shifted with harness design, model family, and how reliably each agent could use its own tools.\n\nDatabases like FlyBase are quiet infrastructure for biology research, and increasingly for the AI systems being built on top of that research. FlyAOC's core finding, that results swing with scaffolding as much as with the underlying model, matters because most agent benchmarks report a single score per model, burying the system-level breakdown that would actually trip up a real deployment.\n\nIt's a reminder that bolting a stronger model onto a shaky agent scaffold won't fix it, a lesson coding and web-browsing agent benchmarks have already been teaching for a while.","[\"ai-agents\",\"benchmarks\",\"bioinformatics\",\"drosophila\"]","2026-09-28T04:00:00.000Z","2026-09-28T09:05:13.778Z","2026-09-28T09:05:20.173Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline\u002Fdek\u002Flede claim agents 'flunked' and 'still struggle to replicate' curators' work, but the body never cites any actual accuracy, score, or pass-rate figures from the paper — only that performance was 'sensitive' to harness\u002Fmodel\u002Ftool-use design; either pull in the paper's actual performance numbers or soften the framing to match what the source substantiates.","resolved","ai",[32,33,34,35],"ai-agents","benchmarks","bioinformatics","drosophila",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.09163",0,{"sections":42},[43,46,50,55,60,65,69,74,79,84,89,94,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",4791,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",762,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",151,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":95,"slug":96,"count":92,"latest_published_at":97},"General","general","2026-09-26T17:02:42.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]