[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-compares-17-fixes-for-ai-models-that-memorize-training-data":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},9255,"study-compares-17-fixes-for-ai-models-that-memorize-training-data","Study Compares 17 Fixes for AI Models That Memorize Training Data","Researchers tested 17 fixes for AI memorization and found an unlearning method called BalancedSubnet works best.","Language models sometimes memorize their training data well enough to repeat it back verbatim, and a new paper tests 17 ways to make them stop.\n\nResearchers evaluated three methods based on regularizers (training penalties that discourage memorization), three based on fine-tuning (retraining a model after the fact), and eleven based on machine unlearning (surgically deleting specific information from a model's weights). Five of those eleven unlearning methods are new, introduced for this project. The team also built TinyMem, a suite of small, cheap-to-run language models meant for quickly testing memorization fixes before trying them on production-scale models.\n\nThe regularizer methods were slow and barely worked. Fine-tuning worked better but cost too much, especially if you wanted the model to stay accurate on its normal tasks. Unlearning-based methods were both faster and more effective, letting researchers pinpoint and remove memorized data from a model's weights before it ever answers a query. One of the new techniques, BalancedSubnet, beat every other method at stripping out memorized data while preserving performance on the target task.\n\nThis matters because memorization is a real liability, not a hypothetical one: a model that spits out verbatim chunks of its training data can leak private records, copyrighted text, or anything else it happened to see. Most fixes to date have been too crude or too expensive to deploy at scale, which has kept memorization in the unsolved column even as models get bigger and swallow more data.\n\nCall it progress, not a cure: the paper shows removal can be surgical, but it still depends on knowing what to cut in the first place.","[\"ai\",\"machine-learning\",\"data-privacy\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-02T05:01:10.480Z","2026-10-02T05:01:14.878Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Remove or attribute the unsupported claim that 'memorization research has a long history of fixes that work in the lab and then get bypassed by a slightly different prompt' — it's not in the source — and stop equating fine-tuning with 'retraining from scratch,' since the source distinguishes fine-tuning from full retraining and never uses that phrase.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The headline and dek claim 14 methods, but the body enumerates three regularizer-based, three fine-tuning-based, and eleven unlearning-based methods, which sums to 17 — fix the count to match the source.","ai",[34,36,37,38],"machine-learning","data-privacy","research",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2410.02159",0,{"sections":45},[46,49,53,57,62,67,71,76,81,85,90,95,100,105],{"name":47,"slug":34,"count":48,"latest_published_at":18},"AI",5659,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Security","security",818,{"name":54,"slug":55,"count":56,"latest_published_at":18},"Policy","policy",430,{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":18},"Science","science",163,{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":82,"slug":83,"count":79,"latest_published_at":84},"Software","software","2026-09-30T21:41:11.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]