The simulated "users" powering AI agent benchmarks are quietly breaking the scripts they're supposed to follow.
A preprint posted to arXiv this month (arXiv:2609.38043), which has not been peer reviewed, introduces UserProxyBench, a new evaluation layer for the popular tau-bench family of agent benchmarks. The researchers built a User Fidelity Score (UFS) that checks whether the AI model playing "the user" in these tests actually follows its private instructions, instead of just scoring whether the agent completed the task. Holding the agent constant at GPT-5.5 and swapping only the simulated user across 375 enterprise tasks shifted the mean task reward by 15.2 points. In 24.4% of episodes where the agent succeeded, the user proxy had violated its own instructions, most often by blurting out information before the agent asked for it.
That matters because agent benchmarks assume the simulated user is a neutral, consistent stand-in for a real person. If the user model itself is unreliable, some of what looks like agent performance, or agent failure, is really benchmark noise. The premature-disclosure bug barely changed success rates but cut the number of tool calls agents needed by about one per task, meaning the test was measuring a different interaction than intended even when the final grade looked the same.
The paper also maps a cost-fidelity frontier across seven proxy models, useful if you're building your own agent tests. But the bigger lesson is familiar: before trusting a leaderboard, check who is grading it and who is playing the customer, and remember this framework itself is fresh, unpublished research that hasn't yet faced outside scrutiny.