A new study says the model behind your AI agent might matter less than how you've set it up.
Researchers built a benchmark of four scientific tasks where a coding agent has to find and correctly run a published specialist model. They tested five configuration choices: the task information given to the agent, whether it reasons through steps, whether it self-verifies, how much time it gets, and which backbone model runs it. The team ran over 18,000 agent trajectories and found something uncomfortable: about 54% of the variation in outcomes came from just rerunning the exact same setup, not from changing anything. Of the five levers, the information given to the agent mattered most, more than time budget or model size, and it also cut costs and improved how well the agent judged its own accuracy.
This matters because most agent evaluations compare models while treating everything else as fixed, which this study suggests is the wrong frame. The same model can look great or mediocre depending on prompt detail and tooling, which means a lot of published agent benchmarks may be measuring configuration luck as much as model capability. The researchers also found that telling an agent to verify its own work barely changed its behavior, but giving it an actual verification tool did.
That last point is the real takeaway: if you want an agent to double-check itself, build the checking in rather than asking nicely. Prompting is not a substitute for system design.