AI agents are bad at managing their own clock, according to a new benchmark.
Researchers built AgentTime, a benchmark of 222 tasks pulled from 18 sources spanning coding, computer use, agentic work, and automated research. The tests give an agent a single instruction specifying how long to work, with requests ranging from about a minute to multiple days, then check how closely the agent matches that duration. Results vary widely: Fable 5.1 running in Claude Code missed requested runtimes by a typical factor of 2.9x, while GPT-6 Astra in Codex missed by only 1.2x. In a review of 158 Astra transcripts, 14 runs showed the agent appearing to finish early and then sleeping instead of stopping.
Hitting a time target isn't the same as working the whole time. Some agents coast or idle instead of producing output, a failure mode normal pass or fail grading would miss entirely. The benchmark also found agents tend to overestimate how long a task will take before starting, and get noticeably worse at guessing elapsed time once timestamps are stripped from their context, which suggests they lean on external cues rather than any real internal clock.
Running an agent for days unsupervised, the pitch behind most autonomous-agent products, means trusting a sense of time it doesn't reliably have.