[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-train-one-ai-to-build-better-workspaces-for-another-ai":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8547,"researchers-train-one-ai-to-build-better-workspaces-for-another-ai","Researchers Train One AI to Build Better Workspaces for Another AI","A new arXiv paper finds a frozen AI model learns reusable rules for building test environments that beat direct instructions to another model.","One AI system just got better at its job by learning to design the workspace for another AI, not by getting smarter itself.\n\nA paper posted to arXiv on September 30, 2026 (arXiv:2609.38143) describes an experiment in what the authors call test-time AI4AI: a \"Builder\" model constructs the execution environment, or harness, that a \"Target\" model uses to complete tasks, while both models' underlying weights stay frozen. Instead of hard-coding rules, the Builder extracts what the researchers term Meta-Skills, reusable principles about when a task needs extra support and what resources to hand over, by watching the Target's feedback on a set of development tasks. That skill bank is then frozen and reused to build harnesses for tasks the Builder has never seen. Across two benchmarks, Harness-Bench and NewtonBench, the full meta-skill bank lifted macro-average performance by 8.95 percentage points over giving the Target no scaffolding, and by 12.02 points over simply handing the Target the same skill bank directly rather than using it to construct a harness.\n\nThis is a quieter, more useful framing than the usual \"agent gets smarter\" story: it treats the environment around a model, not the model's weights, as the thing worth optimizing, which matches what practitioners building agent frameworks have long suspected: the same model performs wildly differently depending on scaffolding. The detail that stands out is the 12-point gap between building a harness and just forwarding the skill bank as instructions; a model apparently benefits more from someone else doing the environment design than from being told the design rules outright. The paper also notes gains persist when a single model plays both Builder and Target, which is the closest thing here to a self-improvement result.\n\nThat is still self-improvement in a narrow sense, better environment design, not better reasoning, and it is demonstrated on two purpose-built benchmarks, not the messy tasks agents actually get deployed on.","[\"ai-agents\",\"arxiv\",\"self-improving-ai\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T10:26:53.936Z","2026-09-30T10:26:58.452Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings explicitly to the source — name arXiv (with the paper's identifier, e.g. arXiv:2609.38143) and note the posting date, since the draft currently presents all figures and quotes as if from an anonymous 'new study' with no named publication or date.","resolved","ai",[32,33,34,35],"ai-agents","arxiv","self-improving-ai","research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38143",0,{"sections":42},[43,46,50,54,59,64,69,74,79,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5105,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",785,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]