A new benchmark says the AI agents writing and running your code can be hijacked more than two-thirds of the time.
Researchers built EvoRiskBench, a test suite of 450 adversarial tasks spanning six scenarios, to probe how workspace agents - AI models paired with code-execution harnesses - behave when someone tries to steer them off task. The tasks are organized around a framework the team calls EP-Path-EF, which maps nine ways an attack can start and five technical outcomes it can produce, validated for classification consistency with a 20-person study. Each task runs in an isolated sandbox, with an automated pipeline checking the result against runtime traces and environment states rather than trusting the agent's own account of what happened. The team then ran nine combinations of three models - GPT-5.6 Sol, DeepSeek-V4-Pro-0813, and Claude Opus 5 - against three harnesses: Claude Code, Codex, and OpenClaw.
The worst performer, Codex running DeepSeek-V4-Pro-0813, got hijacked in 68.44% of attempts. Across all nine configurations, which model sat underneath mattered more than which harness wrapped around it, and the effect of swapping harnesses changed depending on the model. In other words, there is no single safe harness to bolt onto a risky model.
That is a problem for anyone treating agent harnesses as a security layer rather than a thin shell around a model that still takes instructions from strangers.