AI coding agents write code that's harder for other AI agents to work with later.
A paper posted to arXiv, "Is Agent Code Less Maintainable Than Human Code?" (arXiv:2606.21804), introduces a framework called CodeThread to test this directly. Researchers ran four frontier coding agents against four existing coding benchmarks, having each agent build on top of code written either by humans or by other agents. When agents built on agent-written code instead of human code, their task resolve rate dropped by as much as 13.1%. The paper's authors found that standard software maintainability metrics - things like complexity scores - didn't explain most of the gap.
The clearer predictors were subtler: differences in how agents handle input validation and error handling, plus gaps in downstream code size and task difficulty. That matters because more dev teams are chaining agents together to handle multi-step work, on the assumption that code quality stays flat across handoffs.
If one agent can't reliably build on another agent's work, "autonomous engineering" pipelines may need a maintainability check between every handoff, not just a pass-fail grade at the end.