AI agents can hold it together for a single task, but a new study finds they slowly lose the plot the longer they run.
Researchers ran 84,540 trajectories across 8 model families through a 20-step task modeled on delayed-gratification experiments: at each step, an agent could keep waiting for a bigger reward or grab a smaller one now and end the episode early. They varied whether the agent's choice was visible to others, added persona-based pressure, and changed how much deliberation the agent was allowed, then used survival analysis - the statistical method typically used to study time until failure - to see how these factors shifted the odds of an agent caving early. To explain why agents gave in, the team built a seven-category taxonomy from 13,780 deliberation transcripts, with human reviewers checking the AI-assisted labels and reaching a Cohen's kappa of 0.83, a standard chance-corrected measure of inter-rater agreement, not a raw percentage score.
The rationales shifted with time and context: early failures read as impulsive, later ones as fatigue and cost-benefit framing, and public settings pushed agents toward justifying choices in terms of social norms. More strikingly, agents that deliberated longer before failing were also more likely to contradict themselves mid-reasoning, undercutting the assumption that longer chains of thought mean more reliable output.
For anyone deploying agents on long-running tasks, that's a bigger red flag than any single wrong answer: the failure isn't random, it has a shape, and right now every model's shape is different.