A new benchmark says AI coding assistants are great at talking to software architects and bad at everyone else.
Researchers built a benchmark called SWE-Journey to test coding assistants like Claude Code and Codex on the kind of work real developers actually do: long chains of tasks across an evolving codebase, with back-and-forth clarification instead of one clean prompt. The team automated the creation of long-horizon coding tasks and built a simulated user that mimics four real user personas drawn from actual interaction data, including software architects and non-coders. Models were graded on how many requested features passed their tests after the full exchange. The gap was stark: assistants passed over 75% of tests when paired with simulated software architects, but fewer than 25% when paired with simulated non-coders.
That split exposes the limit of today's demos, which tend to showcase a skilled engineer steering an already-capable model. The moment the person on the other end can't specify requirements precisely, read the resulting code, or catch a wrong turn early, these tools lose most of their reliability. The researchers trace the failure to three specific skills - asking the right clarifying questions, finding the right code to change, and fixing it correctly - that assistants manage fine with expert guidance and poorly without it.
In other words, the "AI replaces programmers" pitch still needs a programmer in the loop.