Researchers built a system that automatically redesigns the software scaffolding around AI agents - and it beats every human-tuned version they tested.
A new paper describes MILO (Meta-evolutionary Island Orchestration), a framework that evolves both an AI agent's "harness" - the code controlling how a model executes tasks and interacts with its environment - and the search strategy used to find better harnesses. It works through island-based lineage trees that treat rejected mutations as useful negative evidence, mutator agents that rewrite entire harnesses using global search history, and an orchestrator that reassigns mutators and revises the curriculum as search progresses. Tested on three benchmarks (Terminal-Bench 2.1, PaperBench, and DeepSWE) using both Opus 4.8 and the open-weight gpt-oss-120b, MILO-discovered harnesses beat eight existing harnesses and six prior search methods. On Terminal-Bench 2.1 it scored 86.1%, topping the official leaderboard's prior best of 83.8% while using 26% fewer tokens than the harness it started from.
Harness design - prompts, tool definitions, retry logic, memory strategies - has become as big a performance lever as the underlying model, but tuning it by hand is slow and has to be redone every time a new model ships. Automating that search, and improving the search process itself rather than just the harness, points toward agent systems that retune themselves as models change instead of requiring a team to rerun benchmarks after every release.
Earlier automated approaches mostly tweaked prompts or skills in isolation and plateaued fast; MILO's actual trick is treating the search strategy as something worth evolving too, not just the thing being searched.