[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-finds-ai-harness-diversity-adds-little-over-repetition":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},8454,"study-finds-ai-harness-diversity-adds-little-over-repetition","Study Finds AI Harness Diversity Adds Little Over Repetition","A controlled study of AI-generated problem-solving harnesses finds their apparent gains mostly reflect repeated execution, not real task specialization.","Running the same AI program on a math problem three times can boost your score almost as much as building eight different specialized programs to solve it, according to a new controlled study.\n\nResearchers pitted eight AI-generated problem-solving harnesses against a baseline of nine identical copies of one program, each executed three times on 386 problems from the MATH-500 benchmark. Simply repeating the same program three times already recovered 2.16 percentage points of extra correct answers, just from run-to-run variation. The generated harnesses did produce more consistent right-or-wrong patterns than the baseline copies, but that consistency mostly meant consistently wrong: 100 tasks failed on every one of the three attempts across all generated harnesses, while only a single task was solved reliably, and even that result depended on how the system extracted the final answer. A method for picking the best harness in advance, before running anything, gained exactly 0.00 percentage points.\n\nHere is the number that needed unpacking: both groups eventually hit the same 98.70% oracle coverage ceiling at 27 total executions - but that 27 is the baseline's own tally, nine identical copies run three times each, not the generated harnesses' natural total of 24 runs across eight programs. Matched against that same budget, the fancier harnesses do not pull ahead; they just catch up to what brute repetition already delivers. That is a real problem for any product pitching harness generation as smarter reasoning rather than expensive dice-rolling.\n\nIf your AI tool's improvement mostly shows up after you let it try three times, it might not be reasoning better - it might just be rolling more dice.","[\"ai\",\"llm-evaluation\",\"benchmarks\",\"ai-research\"]","2026-09-30T04:00:00.000Z","2026-09-30T04:43:31.245Z","2026-09-30T04:43:36.913Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"Numeric inconsistency: the article states eight generated harnesses (which at three runs each is 24 executions) yet claims both the generated harnesses and the nine-copy baseline converge after the same '27 executions,' which doesn't reconcile with the stated harness counts.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"editor-r2","editor",2,"publisher-r1 is still unresolved: the piece says both groups plateau at '27 total runs, the point at which the experiment's full execution budget was spent,' but eight harnesses at three runs each is only 24 executions, so clarify explicitly what the 27-execution mark refers to (e.g., that it's the baseline's execution count, not the generated harnesses') rather than glossing over the mismatch with vague phrasing.","ai",[35,37,38,39],"llm-evaluation","benchmarks","ai-research",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.35873",0,{"sections":46},[47,50,54,58,63,68,73,78,83,88,93,98,103,108],{"name":48,"slug":35,"count":49,"latest_published_at":18},"AI",5029,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",780,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Policy","policy",417,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":109,"slug":110,"count":111,"latest_published_at":112},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]