A new benchmark finds that AI agents given control of a computer will often make an unsafe call early in a task, long before anything looks wrong at the finish line.
Researchers built LPS-Bench to test computer-use agents - AI systems that complete multi-step jobs by calling software tools - on safety during long-horizon planning, not just on their final output. A template-guided multi-agent pipeline generates user instructions, simulated toolkits, and case-specific safety criteria, with every case then reviewed by a human. That process produced 570 test cases across 65 scenarios spanning 7 task domains and 9 types of planning risk, under both ordinary requests and adversarial attempts to steer the agent off course. An LLM-based evaluator then reads the full interaction record - tool choices, arguments, and how the agent responds to feedback from its environment - to score each run against criteria specific to that case.
Most agent benchmarks grade only the final result, which can miss an early unsafe decision that quietly shapes everything that follows. Testing 13 LLM agents this way, the researchers found persistent safety failures in both benign and adversarial settings, and prompt-based interventions - essentially telling the model to be careful - helped some models but left real gaps in others.
In other words, an agent that nails the task can still have made a reckless choice getting there, and no amount of post-hoc grading catches that.