AI/ google · program-repair · ai-agents · developer-tools

Google's Bug Fixing AI Works, But Developers Barely Use It

Google's FlowAgent fixed 67% of test failures it attempted, but most of its 295,508 suggestions went unopened by developers.

Google built an AI that proposes fixes for failing tests while a developer is still staring at the red X, and most developers still don't click apply.

The tool is called FlowAgent, and it runs inside Critique and Cider, Google's internal code review and development tools. It targets the pre-submit phase, catching test failures in continuous integration before a developer loses focus and switches tasks. FlowAgent uses a ReAct-style generate-and-validate loop, plus abstention filters before and after it proposes a fix, to keep suggestions accurate under tight latency limits. In a manual evaluation of 195 real test failures, it proposed a correct fix 67.18% of the time. After a company-wide rollout, it suggested fixes on 295,508 changes; developers previewed 65,069 of those and applied only 28,554.

That gap between a roughly two-in-three accuracy rate and a one-in-ten adoption rate is the real finding here. Most automated program repair research measures whether a model can produce correct code. Google's numbers show correctness isn't the bottleneck once a tool ships internally; habit and trust are. Developers have to notice the suggestion, believe it's worth a look, and then decide it's safe enough to merge.

Google says internal interviews describe the reception as positive, which is a low bar when the alternative was no suggestion at all; the harder number to watch is whether that preview rate climbs as engineers learn when to trust the agent.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →