AI/ llm-reliability · ai-research · agentic-ai · benchmarks

Study Finds LLMs Hit a Sudden Accuracy Cliff on Repetitive Tasks

A new study finds large language models stay accurate on repeated simple tasks, then collapse sharply once output length crosses a predictable threshold.

Large language models don't fail gradually on repetitive tasks; they hit a wall, hard.

Researchers tested leading LLMs on deterministic, repetitive tasks, think letter substitution, long addition, or repeated string-operator multiplication, where the same operation runs N times in a row. Accuracy didn't decay smoothly as the task got longer. Instead, models held steady performance and then collapsed sharply past a specific sequence length the paper calls the accuracy cliff. That drop-off was consistently steeper than a simple compounding-error model would predict, a pattern the researchers fit to what they call a double-exponential accumulation law.

This matters because a lot of agentic-AI hype assumes models chain reasoning steps indefinitely, with errors piling up only gradually. This research points the other way: errors appear to interact and reinforce each other, so reliability doesn't taper off, it falls off a cliff at a length that varies by model and task. For anyone wiring LLMs into multi-step agents or automations, that cliff is the real ceiling, not some vague sense that longer jobs are harder.

The paper doesn't publish a universal number for where that cliff sits, since it's specific to each model and task. But the core finding still stands: pushing a task a bit further isn't a bit riskier, it's stepping into a different regime altogether.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →