[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-algorithm-closes-efficiency-gap-in-preference-learning":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8937,"new-algorithm-closes-efficiency-gap-in-preference-learning","New Algorithm Closes Efficiency Gap in Preference Learning","A new algorithm called GINOP shows that learning from pairwise preferences can be just as statistically efficient as learning from direct reward signals.","A new algorithm claims that learning from preferences alone is no less data-efficient than learning from direct reward scores.\n\nThe paper studies preference-based bandits: a system repeatedly picks two options and learns only which one a rater preferred, with preferences modeled using the Bradley-Terry framework common in tournament ranking. That setup covers recommender systems, ranking tournaments, and training AI on human feedback, where people reliably say which answer is better but struggle to score it numerically. Earlier approaches to this problem were mostly stuck with simple linear reward models and were slowed by a troublesome constant in the math linking preferences to rewards, one that can grow large and inflate the data needed to learn. The researchers introduce a new complexity measure, the locally sensitive eluder dimension, and an algorithm called GINOP (Generic INformative OPtimism) that builds confidence sets from log-loss and picks arm pairs to balance curiosity about uncertain options against confidence in good ones.\n\nThe practical hook is reinforcement learning from human feedback, the method behind tuning chatbots like ChatGPT and Claude. If preference data is genuinely as efficient as direct reward data, as this paper's regret bound suggests, that undercuts the assumption that RLHF is a statistically wasteful workaround for not having real reward signals.\n\nThe claims rest on theoretical bounds and benchmark comparisons against baselines, not a live run on an actual language model, so the real test is whether the math holds up once human raters start disagreeing with each other.","[\"ai\",\"reinforcement-learning\",\"bandit-algorithms\",\"rlhf\"]","2026-10-01T04:00:00.000Z","2026-10-01T11:43:33.220Z","2026-10-01T11:43:39.718Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Define what the GINOP acronym stands for (Generic INformative OPtimism) the first time it's introduced in the body, since it's currently used undefined.","resolved","ai",[30,32,33,34],"reinforcement-learning","bandit-algorithms","rlhf",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39351",0,{"sections":41},[42,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5350,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]