AI/ ai-agents · fine-tuning · benchmarks · llm-research

When to Evolve an AI Agent's Harness vs Train Its Weights

A new study argues some AI agent failures are fixed by tweaking the harness, while others require retraining the model's weights.

A new study gives AI agent builders a way to decide whether to patch the software wrapped around a model or retrain the model's weights.

Researchers tested long-horizon planning agents and found their failures split into two types: process failures, where the agent gets stuck in loops, blocked by tool calls, or runs out of steps, and content failures, where the agent delivers a plan but the plan itself is bad. Letting a harness - the code, prompts, and tool-calling logic around a frozen model - evolve to fix process failures lifted the held-out score of Qwen3.5-4B from 0.16 to 0.30 on the DeepPlanning benchmark, and Qwen3.5-9B's from 0.32 to 0.44. For the 4B model, the rate of successfully delivered plans jumped from 55% to 90%. Training LoRA adapters on trajectories produced by that improved harness then baked the gains into the weights: run under the original, unevolved harness, the adapters alone added 0.13 points, and for the 9B model matched the full harness-evolution gain while cutting content failures from a quarter of trajectories to one in twenty.

The practical upshot is a cheap sanity check before reaching for expensive fine-tuning. If an agent keeps hitting step limits or looping, that's scaffolding to fix, not a reason to retrain a model. If it finishes the task but hands back a weak plan, that's what the extra training cost should buy.

The effect also carried over to 117 unseen WebArena-Lite tasks, a modest 0.09-point gain, and a control adapter trained on shuffled answers performed worse than the untrained model - a reminder that not every fine-tuning run is worth the compute it burns.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →