AI/ ai-benchmarks · coding-assistants · llm-agents · software-development

New Benchmark Shows AI Coding Assistants Struggle With Non-Coders

A new benchmark testing coding assistants on long multi-turn tasks finds they pass over 75% of tests for architects but under 25% for non-coders.

A new benchmark says AI coding assistants are great at talking to software architects and bad at everyone else.

Researchers built a benchmark called SWE-Journey to test coding assistants like Claude Code and Codex on the kind of work real developers actually do: long chains of tasks across an evolving codebase, with back-and-forth clarification instead of one clean prompt. The team automated the creation of long-horizon coding tasks and built a simulated user that mimics four real user personas drawn from actual interaction data, including software architects and non-coders. Models were graded on how many requested features passed their tests after the full exchange. The gap was stark: assistants passed over 75% of tests when paired with simulated software architects, but fewer than 25% when paired with simulated non-coders.

That split exposes the limit of today's demos, which tend to showcase a skilled engineer steering an already-capable model. The moment the person on the other end can't specify requirements precisely, read the resulting code, or catch a wrong turn early, these tools lose most of their reliability. The researchers trace the failure to three specific skills - asking the right clarifying questions, finding the right code to change, and fixing it correctly - that assistants manage fine with expert guidance and poorly without it.

In other words, the "AI replaces programmers" pitch still needs a programmer in the loop.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →