[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-framework-flags-hidden-leaks-in-ai-test-data":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9033,"new-framework-flags-hidden-leaks-in-ai-test-data","New Framework Flags Hidden Leaks in AI Test Data","ReLaG is an open-source framework that spots hidden links between related samples to stop inflated AI benchmark results.","A new open-source tool wants to stop AI benchmarks from quietly lying to you.\n\nResearchers released ReLaG, a framework for splitting datasets into training and test sets when samples are secretly related - think molecules or proteins that share a common ancestor or scaffold. Standard random splitting assumes every sample is independent, but in biochemistry and similar fields, related samples often end up on both sides of the split, inflating how well a model appears to generalize. ReLaG detects these hidden relationships using proximity graphs and community detection, then groups related samples so they land entirely in training or entirely in testing. On molecular and protein datasets, it matches the accuracy of existing relation-aware splitting methods while scaling to dataset sizes those methods can't handle. It ships as a pip-installable package.\n\nThis matters because data leakage has been a known embarrassment in drug-discovery and protein-modeling benchmarks for years, producing headline accuracy numbers that don't survive contact with real deployment. ReLaG also adds a label-free way to tune split granularity to match production data, narrowing the gap between benchmark and reality, and its grouping can double as a cheap estimate of how much genuinely diverse data a dataset actually contains.\n\nLeakage has inflated machine-learning results quietly for a decade; a scalable fix only matters if researchers actually bother to use it instead of the benchmark that makes their model look best.","[\"machine-learning\",\"benchmarking\",\"open-source\",\"bioinformatics\"]","2026-10-01T04:00:00.000Z","2026-10-01T16:43:13.682Z","2026-10-01T16:43:16.982Z","published",null,[],"ai",[26,27,28,29],"machine-learning","benchmarking","open-source","bioinformatics",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38538",0,{"sections":36},[37,40,44,49,54,59,63,68,73,77,82,87,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",5487,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",809,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",162,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":74,"slug":75,"count":71,"latest_published_at":76},"Software","software","2026-09-30T21:41:11.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]