[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-build-a-better-filter-for-self-taught-ai-agents":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5204,"researchers-build-a-better-filter-for-self-taught-ai-agents","Researchers Build a Better Filter for Self-Taught AI Agents","A new admission gate for AI systems that build their own skill libraries improves accuracy, but the study also shows it breaks down on real-world data.","AI agents that solve optimization problems can get better over time by saving skills that worked - but only if they can tell a good skill from a broken one, and nobody hands them an answer key.\n\nA new system called AdmitOR tries to solve that by checking behavioral evidence instead of known answers. It runs candidate solutions from three different model families, prompting strategies, and solver stacks against resampled versions of a problem, then looks for agreement across those runs before deciding whether to accept, abstain, or escalate a skill for human review. Tested against logs from an existing skill-learning system, AdmitOR hit 92.7% admission precision, versus 87.1% for simple majority vote and 72.6% for just checking whether code runs without erroring. Do the error-rate math and that is roughly 1.8 times fewer poisoned admissions than majority vote and about 3.8 times fewer than execution-only checks - a real improvement, though smaller than the paper's own headline multiplier implies. The resulting skill library was also the smallest of the bunch, yet scored highest on macro accuracy across five public benchmarks: 58.4, beating majority vote's 54.8 and even a library built from ground-truth labels at 53.9.\n\nThat last number is the interesting part. A filter with no access to correct answers outperformed one that had them, at least on this metric. It suggests that quality control on what an AI learns from itself may matter more than how the training data was labeled in the first place.\n\nThe researchers also admit their method's core promise - a calibrated false-discovery guarantee - held up on their own test data but broke down on a fresh, unlabeled stream, largely because some benchmark problems do not actually match their own stated answers. Worth remembering next time a benchmark score gets cited as fact.","[\"ai agents\",\"llm\",\"optimization modeling\",\"ai research\"]","2026-08-18T04:00:00.000Z","2026-08-18T09:30:50.311Z","2026-08-18T09:31:02.099Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The claimed '3 to 8 times fewer bad skills' does not follow from the stated precision figures (0.927 vs 0.871 vs 0.726), which imply error-rate reductions of roughly 1.8x to 3.75x, not 3x to 8x.","resolved","ai",[32,33,34,35],"ai agents","llm","optimization modeling","ai research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15565",0,{"sections":42},[43,47,51,56,61,66,71,76,81,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":46},"Security","security",435,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]