Give an AI agent a deadline, and it either ignores it or has no idea what to do with the time it's got.
Researchers tested Qwen3.6-27B on five coding competitions from MLE-Bench Lite and Qwen3-4B on the text adventure Zork I, giving each agent a wall-clock time budget on tasks where extra thinking time should help. When the budget was simply written into the prompt, the agents couldn't translate it into paced behavior - they have no internal clock, can't estimate how long their own actions take, and have no learned sense of how much effort a given amount of time should buy. Feeding the agents live timing information through the test harness fixed most of the overruns, and adding hard enforcement hooks tightened adherence further. The team also trained agents with reinforcement learning to respect budgets, which worked almost perfectly on Zork I and generalized to budgets the agents hadn't trained on.
None of that solved the harder problem: agents that stop on time still don't use the spare minutes well. The RL-trained models learned when to stop, but padded leftover time with repeated, useless actions, and training across multiple budgets made them default to whatever strategy worked for the shortest one. That matters for anyone building autonomous coding or research agents meant to run with minimal supervision - telling a model it has an hour currently buys little more than telling it it has ten minutes.
It's a reminder that today's "agentic" capability is mostly about following instructions, not reasoning about resources - something humans manage almost without thinking, and these models, at least at this size, still cannot.