A new benchmark called FlyAOC checks whether AI agents can do the unglamorous work of curating scientific databases, and the results depend more on how the agent is built than on which model powers it.
FlyAOC evaluates agents on the full curation workflow instead of isolated subtasks like named entity recognition. Given a gene symbol, a short FlyBase description, a 16,898-paper corpus, and ontology resources, an agent has to search the literature and reconstruct curator-grade annotations, including standardized function terms, expression patterns, and historical name synonyms. The benchmark grades that output against 7,397 expert-curated annotations covering 100 genes drawn from FlyBase, the Drosophila research community's knowledge base. The researchers tested four agent designs, a memorization baseline, a fixed pipeline, a single agent, and a multi-agent setup, and found that performance shifted with harness design, model family, and how reliably each agent could use its own tools.
Databases like FlyBase are quiet infrastructure for biology research, and increasingly for the AI systems being built on top of that research. FlyAOC's core finding, that results swing with scaffolding as much as with the underlying model, matters because most agent benchmarks report a single score per model, burying the system-level breakdown that would actually trip up a real deployment.
It's a reminder that bolting a stronger model onto a shaky agent scaffold won't fix it, a lesson coding and web-browsing agent benchmarks have already been teaching for a while.