AI/ ai agents · agent memory · llm · benchmarks

New Method Lets AI Agents Learn From Their Own Failures

DAEDALUS pairs an explorer and solver agent to turn self-generated practice failures into reusable memory, without labeled training tasks or human guides.

A new research method lets AI agents teach themselves how to use unfamiliar tools, instead of relying on humans to write the manual.

Researchers built DAEDALUS, a system that pairs two AI agents: an explorer that invents tasks for a new environment, and a solver that attempts them. When the solver fails, the system extracts a heuristic, a short rule of thumb, from that failure. A heuristic only gets saved if the solver can use it to succeed on similar tasks repeatedly. Those saved heuristics accumulate into a memory bank the agent can draw on later, and the explorer uses solver performance to calibrate how hard the next batch of tasks should be.

Most agent memory systems need either hand-written playbooks or a stack of labeled training tasks checked by a verifier, both expensive to produce for a new environment. DAEDALUS needs neither. Across three benchmarks, AppWorld, tau2-bench, and AutomationBench, it raised success rates by up to 15.9 percentage points over an agent with no memory, and the heuristics it generated also worked for agents built on other model families.

That portability is the interesting part: the real product here isn't a smarter agent, it's a transferable cheat sheet an agent writes for itself.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →