Large language models don't fail gradually on repetitive tasks; they hit a wall, hard.
Researchers tested leading LLMs on deterministic, repetitive tasks, think letter substitution, long addition, or repeated string-operator multiplication, where the same operation runs N times in a row. Accuracy didn't decay smoothly as the task got longer. Instead, models held steady performance and then collapsed sharply past a specific sequence length the paper calls the accuracy cliff. That drop-off was consistently steeper than a simple compounding-error model would predict, a pattern the researchers fit to what they call a double-exponential accumulation law.
This matters because a lot of agentic-AI hype assumes models chain reasoning steps indefinitely, with errors piling up only gradually. This research points the other way: errors appear to interact and reinforce each other, so reliability doesn't taper off, it falls off a cliff at a length that varies by model and task. For anyone wiring LLMs into multi-step agents or automations, that cliff is the real ceiling, not some vague sense that longer jobs are harder.
The paper doesn't publish a universal number for where that cliff sits, since it's specific to each model and task. But the core finding still stands: pushing a task a bit further isn't a bit riskier, it's stepping into a different regime altogether.