[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-teach-ai-agents-to-manage-their-own-reasoning":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},8548,"researchers-teach-ai-agents-to-manage-their-own-reasoning","Researchers Teach AI Agents to Manage Their Own Reasoning","A new arXiv paper adds a controller layer that plans an AI agent's own execution, beating direct control across several benchmarks.","A new technique lets AI agents pause and think about how they're thinking, not just what they're doing next.\n\nIn a paper titled \"Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning,\" posted to arXiv on September 30, 2026 (arXiv:2609.38147), researchers introduce agentic meta-reasoning: an inference-time harness that splits the workers doing task-level computation from a controller that decides what to build on, when to start fresh, and when to stop. The controller tracks only a compact summary of the run instead of replaying its full history, then dispatches work using context pulled from persistent memory. The team benchmarked it against production coding agents including Codex and Claude Code, plus a Direct Control Agent running on the same workers and compute budget.\n\nOn ProgramBench, a long-horizon program-reconstruction test, meta-reasoning scored 71.5% with GPT-5.5 versus 58.0% for Codex, and 67.2% with Opus 4.8 versus 65.5% for Claude Code. Across other benchmarks covering abstract reasoning, multi-domain reasoning, and proof generation, it beat direct control by 3.6 to 4.2 points on average across three frontier models, and kept improving as compute budgets grew in ranges where direct control plateaued.\n\nThat last part is the real finding: as agent runs get longer, managing attention matters almost as much as the model doing the work. The next arms race in agent tooling may not be about bigger models but better bureaucracy - and even a coding agent, it turns out, still needs a boss.","[\"ai\",\"ai-agents\",\"agentic-ai\",\"arxiv-research\"]","2026-09-30T04:00:00.000Z","2026-09-30T10:31:06.452Z","2026-09-30T10:31:12.824Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Explain what ProgramBench actually measures (per the source: long-horizon agentic capability via program reconstruction) before citing the 71.5%\u002F58.0% and 67.2%\u002F65.5% figures, since the draft gives the comparison numbers but never states what the benchmark evaluates.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"Add attribution — name the research team\u002Finstitution, the paper title (\"Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning\"), and the arXiv source\u002Fdate, since all figures are currently presented without any named source.","ai",[34,36,37,38],"ai-agents","agentic-ai","arxiv-research",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38147",0,{"sections":45},[46,49,53,57,62,67,72,77,82,86,91,96,101,106],{"name":47,"slug":34,"count":48,"latest_published_at":18},"AI",5105,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Security","security",785,{"name":54,"slug":55,"count":56,"latest_published_at":18},"Policy","policy",417,{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]