[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-when-to-evolve-an-ai-agents-harness-vs-train-its-weights":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},10887,"when-to-evolve-an-ai-agents-harness-vs-train-its-weights","When to Evolve an AI Agent's Harness vs Train Its Weights","A new study argues some AI agent failures are fixed by tweaking the harness, while others require retraining the model's weights.","A new study gives AI agent builders a way to decide whether to patch the software wrapped around a model or retrain the model's weights.\n\nResearchers tested long-horizon planning agents and found their failures split into two types: process failures, where the agent gets stuck in loops, blocked by tool calls, or runs out of steps, and content failures, where the agent delivers a plan but the plan itself is bad. Letting a harness - the code, prompts, and tool-calling logic around a frozen model - evolve to fix process failures lifted the held-out score of Qwen3.5-4B from 0.16 to 0.30 on the DeepPlanning benchmark, and Qwen3.5-9B's from 0.32 to 0.44. For the 4B model, the rate of successfully delivered plans jumped from 55% to 90%. Training LoRA adapters on trajectories produced by that improved harness then baked the gains into the weights: run under the original, unevolved harness, the adapters alone added 0.13 points, and for the 9B model matched the full harness-evolution gain while cutting content failures from a quarter of trajectories to one in twenty.\n\nThe practical upshot is a cheap sanity check before reaching for expensive fine-tuning. If an agent keeps hitting step limits or looping, that's scaffolding to fix, not a reason to retrain a model. If it finishes the task but hands back a weak plan, that's what the extra training cost should buy.\n\nThe effect also carried over to 117 unseen WebArena-Lite tasks, a modest 0.09-point gain, and a control adapter trained on shuffled answers performed worse than the untrained model - a reminder that not every fine-tuning run is worth the compute it burns.","[\"ai-agents\",\"fine-tuning\",\"benchmarks\",\"llm-research\"]","2026-10-09T04:00:00.000Z","2026-10-09T20:06:15.764Z","2026-10-09T20:06:19.515Z","published",null,[],"ai",[26,27,28,29],"ai-agents","fine-tuning","benchmarks","llm-research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.11655",0,{"sections":36},[37,40,44,49,54,59,63,68,73,78,83,88,93,98],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",6619,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",926,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",486,"2026-10-08T22:40:11.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",474,"2026-10-08T22:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",229,"2026-10-08T20:47:10.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",192,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",181,"2026-10-08T23:26:35.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Startups","startups",117,"2026-10-08T16:45:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"General","general",66,"2026-10-09T04:46:11.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",58,"2026-10-08T20:08:45.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]