AI/ ai agents · llm · machine learning · research

Why Feeding AI Agents Their Own Mistakes Backfires

Sentry lets AI agents access failure lessons only when relevant, beating prior fixes by double-digit margins in testing.

AI agents that learn from their mistakes often learn the wrong lesson at the wrong time, and a new system called Sentry tries to fix that by being pickier about what it remembers.

Sentry runs alongside an AI agent as a separate layer rather than feeding every past failure into the agent's working memory. When it detects a failure, it searches an external playbook for a matching lesson, applies it, and checks whether the agent actually recovered. It only writes a new lesson to the playbook if the recovery worked, and the full playbook never enters the agent's context. Across several agent benchmarks, this beat the best runtime-only recovery method by 37% on average and the best context-evolving method by 39% on benchmarks where both were tested, with extra gains when the two were combined.

The real finding here is about memory, not the agent itself. Failure lessons are conditional: they help when the matching failure recurs, but dumping the whole playbook into the agent's context measurably hurt performance, even when the relevant lesson was still available on demand. That's a pointed counterpoint to the common approach of just growing an agent's context with every past mistake and assuming more history equals more reliability.

Call it a spell-checker for agents: it stays silent until something breaks, and it only writes the fix down after confirming it worked.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →