One AI system just got better at its job by learning to design the workspace for another AI, not by getting smarter itself.
A paper posted to arXiv on September 30, 2026 (arXiv:2609.38143) describes an experiment in what the authors call test-time AI4AI: a "Builder" model constructs the execution environment, or harness, that a "Target" model uses to complete tasks, while both models' underlying weights stay frozen. Instead of hard-coding rules, the Builder extracts what the researchers term Meta-Skills, reusable principles about when a task needs extra support and what resources to hand over, by watching the Target's feedback on a set of development tasks. That skill bank is then frozen and reused to build harnesses for tasks the Builder has never seen. Across two benchmarks, Harness-Bench and NewtonBench, the full meta-skill bank lifted macro-average performance by 8.95 percentage points over giving the Target no scaffolding, and by 12.02 points over simply handing the Target the same skill bank directly rather than using it to construct a harness.
This is a quieter, more useful framing than the usual "agent gets smarter" story: it treats the environment around a model, not the model's weights, as the thing worth optimizing, which matches what practitioners building agent frameworks have long suspected: the same model performs wildly differently depending on scaffolding. The detail that stands out is the 12-point gap between building a harness and just forwarding the skill bank as instructions; a model apparently benefits more from someone else doing the environment design than from being told the design rules outright. The paper also notes gains persist when a single model plays both Builder and Target, which is the closest thing here to a self-improvement result.
That is still self-improvement in a narrow sense, better environment design, not better reasoning, and it is demonstrated on two purpose-built benchmarks, not the messy tasks agents actually get deployed on.