AI/ ai · dev-tools · code-generation · software-reliability

AI Coding Agents Nail the Code but Botch the Dependencies

A new study finds coding agents agree on correct dependencies for the same task as little as 7 percent of the time, even as their code runs fine.

AI coding agents can write code that runs - but good luck getting it to run anywhere else.

A new study tested three AI coding agents across four programming languages and fifty tasks, checking not just whether their generated code worked but whether the agents correctly specified what it needed to run. Researchers built a three-layer framework comparing declared dependencies, the ones agents listed, against what actually got installed at runtime and what was truly necessary. The gaps were stark: for the same task, different agents agreed on the right dependency set as little as 7 percent of the time. Newer agents did no better than older ones, and the biggest mismatch showed up between what agents claimed they needed and what they actually pulled in while running.

Functional tests only check if code works today, on the machine that built it. They miss whether someone else's environment, or next year's dependency versions, will let that code run at all. That is a real cost for anyone pasting agent-generated code into a production repo without pinning versions first - it works in the demo, then breaks in deployment.

The paper pins the blame on loose "environment priors" agents memorized from training data rather than real reasoning about requirements - which sounds less like an engineering oversight and more like models guessing at a problem they were never graded on.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →