[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-llms-outdo-purpose-built-redaction-tools-on-hospital-records":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},5691,"llms-outdo-purpose-built-redaction-tools-on-hospital-records","LLMs Outdo Purpose-Built Redaction Tools on Hospital Records","A benchmark on pediatric oncology notes finds LLMs catch institution-specific PHI that de-identification tools and gold-standard labels both miss.","A new study finds that AI language models catch protected health information that dedicated redaction software, and the human-made gold standard used to grade it, both overlook.\n\nResearchers tested eight large language models against two purpose-built de-identification systems, Stanford's TiDE and OpenMed PII, plus two pattern-based baselines, on 100 pediatric oncology notes from Texas Children's Hospital containing 5,322 labeled PHI spans. Each LLM ran three prompts of increasing specificity: a standard HIPAA-based instruction, that instruction plus a list of institution-specific categories the model had missed, and a final version telling the model not to over-redact. The best LLM configuration reached an F1 score of 0.918, well ahead of TiDE's 0.779, with most of the gap coming from local, contextual details like hospital building names, internal codes, and abbreviations. Naming the missed categories in the prompt recovered 79 percent of them, and the anti-over-redaction instruction restored precision without costing recall.\n\nThe more interesting finding is what the LLMs caught that even the human-annotated reference dataset missed: 414 candidate gaps, of which re-annotation confirmed 227 real PHI spans the original gold standard had left unmarked. That means the benchmark researchers use to grade de-identification tools was itself under-labeled, and a well-prompted general-purpose LLM ended up doing double duty as redactor and auditor. Fourteen multi-agent and ensemble configurations, despite the extra compute, never beat a single well-calibrated prompt.\n\nIt's a reminder that institution-specific is doing a lot of the work here, a prompt tuned to one hospital's abbreviations and building names won't generalize to the next one for free.","[\"llms\",\"healthcare data\",\"privacy\",\"de-identification\"]","2026-08-19T04:00:00.000Z","2026-08-19T11:39:55.134Z","2026-08-19T11:40:07.052Z","published",null,[],"ai",[26,27,28,29],"llms","healthcare data","privacy","de-identification",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.17051",0,{"sections":36},[37,41,45,50,55,60,65,70,75,80,85,90,95,100],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":40},"Security","security",435,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]