AI/ ai agents · harness optimization · arxiv research · self-improving ai

Researchers Teach AI Agents to Tune Their Own Setup

A new paper shows AI agents can rewrite their own setup per task instead of using one fixed version for everything.

A new research framework lets AI agents rewrite their own operating setup for each task, instead of running one fixed configuration for every job.

The paper, posted to arXiv as Turbo Harness: Instance-Adaptive Harness Optimization, reuses leftover data from a prior harness-optimization run to help agents adapt on a per-task basis. Harness optimization searches for the best scaffolding, meaning the prompts, tools, and rules wrapped around a model, but most methods settle on one harness applied to every task. Turbo Harness compiles the original run's artifacts into what the authors call a structured playbook, then trains a separate harness editor to read that playbook and patch the global setup for each new instance at inference time. The authors tested this across seven benchmarks covering interactive agent tasks, software engineering, and long-horizon terminal work, reporting consistent gains over existing harness-optimization baselines.

Harness design has quietly become as consequential as model choice for agent performance, yet most teams still tune it once and ship the same setup everywhere. An approach that customizes the harness per task, without a fresh optimization run for each one, is a modest but concrete step toward agents that adjust their own scaffolding, the kind of recursive self-improvement researchers have been chasing.

The comparisons here are against other harness-optimization methods, not against simply hand-tuning a setup for each task, so how much of this gain holds up outside the authors' seven benchmarks remains to be seen.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →