[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-agent-rsi-master-beats-human-tuned-qwen3-model-on-coding-test":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},9250,"ai-agent-rsi-master-beats-human-tuned-qwen3-model-on-coding-test","AI Agent RSI-Master Beats Human-Tuned Qwen3 Model on Coding Test","RSI-Master, an AI agent that runs its own training experiments, beat Qwen3's human-tuned 35B Instruct model on two benchmarks.","An AI agent that designs its own training experiments just outperformed a model humans fine-tuned by hand.\n\nRSI-Master is built to let an agent run open-ended post-training experiments on a base language model without a human steering every step - and without quietly cutting corners to inflate its own scores. To stop that kind of gaming, the researchers added an \"Experiment OS\" that logs every action in a traceable record, plus a reviewer system that compares results across a growing tree of experiments so the agent doesn't lock onto one approach too early. On PostTrainBench, using a 4-billion-parameter Qwen3 base model, RSI-Master averaged a score of 54.49 against 46.53 for the best rival agent, with a reported 0.0% hacking rate. Scaled up to a 35-billion-parameter model, RSI-Master's resulting model beat Qwen3's own human-tuned Instruct version on the LiveCodeBench-v6 coding benchmark (41.21 vs. 37.36) and on SciCode, and posted a nonzero score on HorizonMath, a test of unsolved research problems where most frontier models score close to zero.\n\nThe real finding isn't the benchmark scores - it's the zero hacking rate. Letting an AI system redesign its own training process is only useful if it isn't also learning to cheat the metrics that measure success, and that's the failure mode this architecture is explicitly built to catch. If the approach holds up at larger scale, it suggests autonomous post-training can be made auditable, not just fast.\n\nA nonzero score on HorizonMath sounds unimpressive until you remember most frontier models can't clear that bar at all.","[\"ai\",\"ai agents\",\"benchmarks\",\"model training\"]","2026-10-01T04:00:00.000Z","2026-10-02T04:44:06.298Z","2026-10-02T04:44:09.901Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The sentence about the science-coding benchmark result is cut off mid-thought — it says the model beat the Instruct version 'on a separate science-coding benchmark' but never states the actual scores or outcome.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"editor-r2","editor",2,"The dek says RSI-Master 'beats human-tuned models' (plural) but the body only shows it beating one model (the 35B Instruct version, across two benchmarks) — fix the dek to singular\u002Faccurate wording.","ai",[35,37,38,39],"ai agents","benchmarks","model training",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.35561",0,{"sections":46},[47,50,54,58,63,68,72,77,82,86,91,96,101,106],{"name":48,"slug":35,"count":49,"latest_published_at":18},"AI",5659,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",818,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Policy","policy",430,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":18},"Science","science",163,{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":83,"slug":84,"count":80,"latest_published_at":85},"Software","software","2026-09-30T21:41:11.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]