AI agents keep getting benchmarked on conversations that don't happen in real life.
Researchers built Drift-Bench++, a benchmark that tests how well AI agents handle users who miscommunicate, change their minds, or simply lose patience mid-task. Instead of assuming a user states one clear, fixed goal and sticks to it, the benchmark simulates diverse users with finite patience and intent that can silently shift partway through a task. To score agents, the team built an evaluation protocol called GRIP, short for Grounding, Realism, Inquiry, and Pivoting: how well an agent's actions stay grounded in the actual task, how realistic the simulated users behave, how effectively the agent asks clarifying questions, and how well it pivots when the user's intent changes. Across multiple models and environments, better interaction habits helped, but no model came close to matching how it performed with a perfectly clear, unchanging user.
Most agent benchmarks still assume what the researchers call oracle communication - a user who states exactly what they want and never wavers. That is not how people actually talk to chatbots: they hedge, backtrack, and get impatient. The researchers checked their findings against real sessions from a deployed system called ProdAgent and found the same failure patterns show up often enough to matter in production, not just in a lab.
That last detail is the real finding here: a lot of agentic AI demos assume users show up with tidy, stable requests. This benchmark is a reminder that they don't, and that the gap between demo performance and deployed performance is still wide.