[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-test-splits-ai-agent-gains-into-arriving-and-solving":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},6771,"new-test-splits-ai-agent-gains-into-arriving-and-solving","New Test Splits AI Agent Gains Into Arriving and Solving","A new method clones an AI agent's mid-task state for another, showing reinforcement learning mainly improves solving, not just positioning.","A new testing method separates two things AI agent benchmarks usually blur together: whether an agent reaches a good position, and whether it knows what to do once it gets there.\n\nResearchers built checkpoint handoff, a technique that clones the exact state one trained AI agent reached partway through a task and hands it to a different agent, with no retraining involved. That split lets them measure REACH, how often an agent arrives at a state confirmed to be a fixed number of actions from success, separately from SOLVE, how often it finishes the job from that identical state. Testing across multiple benchmarks and training pipelines, they found REACH and SOLVE consistently reinforced each other for reinforcement-learning-trained agents. On the ALFWorld benchmark specifically, the RL-trained agent both reached better positions and executed better once there, and it never failed from a spot where a supervised-fine-tuned agent succeeded.\n\nMost agentic RL papers report a single combined success rate, which hides whether a model actually got smarter or just got lucky landing somewhere easier to finish from. The researchers also found that restricting comparisons to only the states both agents reach, a common shortcut, does not fix this problem and can even flip the sign of the measured effect.\n\nIt is a plumbing fix, not a flashier model or a bigger benchmark score, but it is exactly the kind of plumbing that decides whether the next wave of agent leaderboard claims deserves to be believed.","[\"ai\",\"reinforcement-learning\",\"evaluation-methods\",\"benchmarks\"]","2026-09-18T04:00:00.000Z","2026-09-18T15:58:34.123Z","2026-09-18T15:58:46.025Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The claim that RL checkpoints beat SFT checkpoints on both REACH and SOLVE in all five conditions overstates the source, which only confirms both metrics improved together on ALFWorld — elsewhere the source reports only that the reacher-by-solver interaction effect was positive, a distinct and more limited statistical finding; rewrite to accurately distinguish the ALFWorld-specific result from the broader interaction-effect finding.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The body states the finding held in 'all five tested conditions' but earlier says the study covered only two benchmarks and two training pipelines, which implies four conditions, not five — an unresolved numerical inconsistency.","ai",[35,37,38,39],"reinforcement-learning","evaluation-methods","benchmarks",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.19636",0,{"sections":46},[47,50,54,59,64,68,72,77,82,87,92,97,102,107],{"name":48,"slug":35,"count":49,"latest_published_at":18},"AI",3959,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",652,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Hardware","hardware",155,{"name":69,"slug":70,"count":71,"latest_published_at":18},"Science","science",116,{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",76,"2026-09-18T01:04:54.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]