[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-correct-answers-arent-always-the-best-ai-training-data":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},9623,"correct-answers-arent-always-the-best-ai-training-data","Correct Answers Aren't Always the Best AI Training Data","A new framework shows that AI training data which looks most correct is not always the data that most improves a model.","Picking AI training data by how correct it looks turns out to be the wrong instinct, according to new research.\n\nA paper called FAER tackles a problem in how language models get further trained after their initial build: when old model outputs are cached and replayed for more training, the usual selection methods rank them by format compliance, confidence scores, or freshness, not by whether they actually help the model learn. Researchers tested this on GSM8K math problems using Qwen2.5-1.5B-Instruct, a 1.5-billion-parameter model, over 128 training updates. A selector built around format feedback picked responses that were correct 69.53% of the time, far more often than the 35.94% correctness rate of FAER's own baseline selector. Despite picking worse-looking data, FAER's baseline still trained a better model, scoring 0.6329 on the paper's quality measure versus 0.6037 for the format-feedback selector and 0.5482 for random selection.\n\nThat gap matters because most post-training pipelines assume correct-looking outputs make better training examples. FAER's results say that assumption does not reliably hold; what makes a training example useful is a different property than how accurate its answer is. A metadata-only version of FAER's calibrated selector pushed quality to 0.6476 on average across eight test runs, and a fuller version of the same approach reached 0.6624.\n\nNone of it is free. The metadata-only version needed 189,642 tokens of total compute to produce that result, even though only 63,276 of those tokens went toward the actual target task; the rest quietly covered tuning the selector itself, a cost easy to leave out of a headline number.","[\"ai\",\"llm-training\",\"benchmarks\",\"research\"]","2026-10-02T04:00:00.000Z","2026-10-03T03:34:08.947Z","2026-10-03T03:34:14.324Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The 3.48 GPU-hours figure is misattributed to the fixed selector — the source assigns that cost (and its 189,642 tokens) to the separate 'metadata-only cross-fitted calibration' variant (which scored 0.6476, not 0.6329), so fix the GPU-hour comparison to match the correct entity before republishing.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The r1 GPU-hour misattribution is now fixed, but the closing line's claim that accuracy gains cost 'roughly triple the compute of the baseline selector' is unsupported — the source gives no GPU-hour or token cost for the training-free fixed selector, so either drop this comparison or re-ground it in the actual figures (e.g., the metadata-only variant's 189,642-token full cost versus its own 63,276-token training cost).","ai",[34,36,37,38],"llm-training","benchmarks","research",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00385",0,{"sections":45},[46,49,53,57,62,66,70,75,80,85,90,95,100,105],{"name":47,"slug":34,"count":48,"latest_published_at":18},"AI",5976,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Security","security",842,{"name":54,"slug":55,"count":56,"latest_published_at":18},"Policy","policy",438,{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":18},"Hardware","hardware",199,{"name":67,"slug":68,"count":69,"latest_published_at":18},"Science","science",173,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]