[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-cloud-scpo-trims-labeled-data-needed-for-llm-reasoning-training":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},5091,"cloud-scpo-trims-labeled-data-needed-for-llm-reasoning-training","Cloud-ScPO Trims Labeled Data Needed for LLM Reasoning Training","A new technique mines preference data for math reasoning from a small labeled set and hidden-state geometry, cutting reliance on human annotation.","A new preference-optimization method for math-reasoning models leans on the shape of a model's internal hidden states, not just human-verified answers, to decide which reasoning chains to reward.\n\nResearchers behind Cloud-ScPO show that when large language models generate many reasoning trajectories across different math problems, the correct and incorrect ones organize into distinct geometric clusters, or \"clouds,\" in the model's hidden-state space. Cloud-ScPO uses a small labeled set to build reference clouds of correct and incorrect trajectories, then scores each new trajectory by how well it connects to those clusters using a soft k-nearest-neighbor measure. That score is combined with self-consistency, the standard trick of picking the answer most models agree on, to filter chosen-versus-rejected training pairs by confidence margin. On GSM8K and MATH-Numeric, across four different model setups, the method beat the earlier ScPO baseline by up to 4.49 and 4.19 percentage points, respectively.\n\nPreference optimization for reasoning models usually depends on either verified ground-truth answers, human annotators, or a separate reward model to judge which output is better - all expensive to produce at scale. Cloud-ScPO's bet is that a model's own hidden states already encode a rough map of correctness, and that map can substitute for a chunk of that supervision once a small seed of labels exists.\n\nIt is not a label-free method - it still needs that seed set to anchor its reference clouds - and the gains so far are confined to two math benchmarks. Whether the same geometric signal holds up on messier, non-math reasoning tasks is the open question.","[\"ai\",\"llm-reasoning\",\"preference-optimization\",\"semi-supervised-learning\"]","2026-08-17T04:00:00.000Z","2026-08-17T10:04:55.275Z","2026-08-17T10:05:07.106Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims the method works 'without extra labels,' but the body states Cloud-ScPO builds its reference clouds from 'a small labeled set' — reconcile the dek with that detail so it doesn't overstate the label-free claim.","resolved","ai",[30,32,33,34],"llm-reasoning","preference-optimization","semi-supervised-learning",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.01014",0,{"sections":41},[42,46,50,55,60,65,70,75,80,85,90,95,100,105],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":45},"Security","security",435,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]