AI/ ai · ai-agents · prompt-engineering · research

Researchers Find Failed AI Prompts Are Worth Keeping

Mara Chain reuses rejected AI prompt and harness edits instead of discarding them, cutting rollouts needed to improve performance.

A new paper argues that AI teams tuning prompts and agent harnesses have been throwing away their most useful data: the failed attempts.

Researchers introduce Mara Chain, a refinement procedure for optimizing deployed AI systems, meaning the prompts, skills, harnesses, and code that increasingly substitute for retraining model weights. Standard propose-evaluate-select loops score candidate edits and discard anything that doesn't meet an acceptance bar, but Mara Chain instead keeps rejected candidates and iteratively refines them using evidence from earlier attempts, capping each refinement chain at a fixed depth and trimming the pool with Pareto-filtered Top-N selection. Tested on AppWorld skill optimization, TerminalBench 2.1 harness optimization, and MuSiQue retrieval-pipeline optimization, it beat existing optimizers GEPA, ACE, and SkillOpt-Lite by up to 20.5% in relative performance on AppWorld, reaching the target score with 65.5% fewer rollouts than GEPA. On TerminalBench 2.1 it raised pass rates by 20.2 and 22.5 percentage points over two rival methods, and on MuSiQue it beat a hand-written retrieval pipeline by 0.104 on nDCG@10 and 0.131 on Recall@10.

The interesting claim isn't the benchmark wins, it is the premise behind them: a rejected configuration still carries information about which failure modes to avoid, and tossing it forces the next round of proposals to rediscover the same dead ends. For teams running automated prompt or harness tuning at scale, that is a direct cut in wasted rollouts, and rollouts cost real compute. It also reflects a wider shift toward treating prompt and harness design as a formal search problem rather than manual trial and error.

The gains come from benchmarks chosen by the paper's own authors, so the harder test, as always, is whether other teams see the same numbers when they try it on their own systems.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →