[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-hunts-longer-for-why-ai-agents-fail":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8681,"new-method-hunts-longer-for-why-ai-agents-fail","New Method Hunts Longer for Why AI Agents Fail","A paper on arXiv, now in its second revision, finds that nudging AI judges to keep searching logs uncovers failure causes single-pass reviews miss.","AI agents are now running tasks long enough to generate execution logs no human wants to read, and figuring out why they failed is starting to look like a needle-in-a-haystack problem.\n\nA paper posted to arXiv, currently on its second revision, argues that automated root-cause attribution has been asking large language models to make a single judgment call on failure logs, and that approach falls apart as traces grow longer. The researchers' fix, called Continual Search, pushes the same judge model to take multiple passes over a trace instead of stopping at the first plausible-looking cause. To test it at scale, they built a new benchmark, MegaRCA-Mix, with 50 human-annotated failure trials drawn from long-horizon, execution-heavy agent tasks - filling a gap the paper says existing RCA benchmarks don't cover. On that benchmark, Continual Search lifted Opus-4.8's F1 score from 0.471 to 0.608, a 29 percent jump, and the gains held across multiple model families and existing benchmark suites too.\n\nThe more interesting finding buried in the numbers: within the same model family, a lower-tier model using Continual Search could outperform its higher-tier sibling running the old one-shot approach. That undercuts the assumption that better diagnosis just means a bigger model. It's a search problem, not a raw-capability problem - which matters for anyone building agent monitoring tools, since it suggests budget goes further spent on iteration than on model upgrades.\n\nNone of this makes debugging agents easy. It's still a paper-stage method tested on researcher-curated benchmarks, not something bolted onto a production observability stack yet. But as agents take on longer-horizon jobs, reading the raw logs was never going to scale, and this is one of the more concrete attempts at fixing that.","[\"ai agents\",\"root-cause analysis\",\"llm research\",\"arxiv\"]","2026-09-30T04:00:00.000Z","2026-09-30T19:05:56.476Z","2026-09-30T19:06:01.397Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings explicitly to the source — name arXiv and the paper (e.g., 'a paper posted to arXiv, currently on its second revision') since the draft cites specific figures and claims without ever naming the publication or that it's a preprint replacement.","resolved","ai",[32,33,34,35],"ai agents","root-cause analysis","llm research","arxiv",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.13463",0,{"sections":42},[43,46,50,54,59,64,68,73,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5184,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",791,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",155,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]