[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-document-ai-extraction-needs-real-error-guarantees-study-finds":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},5296,"document-ai-extraction-needs-real-error-guarantees-study-finds","Document AI Extraction Needs Real Error Guarantees, Study Finds","A new study of an AI document-extraction pipeline finds the standard error-check method misses its own risk budget in nearly half of repeated trials.","A study of an AI receipt-reading system finds its standard error-check quietly fails almost half the time.\n\nResearchers tested Claude Sonnet 5 on 13,859 fields pulled from 800 receipts in the CORD dataset, where the model got 49% of fields right. They identified three ways the standard accept\u002Freject check breaks: treating fields from the same receipt as independent when they are not, reusing the same data to both calibrate and test the system, and duplicate confidence scores that collapse the decision threshold. The researchers organized fixes into a tiered framework, but even the most common fix, splitting data into a calibration set and a separate validation set, misses its own 10% error budget in 47.5% of repeated trials, per the paper's own numbers. A stricter method, using per-group statistical bounds with exact tail guarantees, does deliver a genuine certificate, but only by being so cautious it approves a small fraction of fields.\n\nDocument extraction quietly underpins invoicing, expense reports, and compliance pipelines, so a claimed error guarantee that fails nearly half the time is riskier than admitting there is no guarantee at all. It is the same gap that has tripped up confidence scoring elsewhere in machine learning: a bound that holds on average across many runs is not the same as a bound that holds for the one run in front of you, and most teams have no way to tell which they got.\n\nA human-verified spot check on this particular run measured actual risk at just 1.3%, well inside the 10% cap - a good outcome, but by the paper's own math, more lucky than proven, since the same check misses its budget in roughly half of comparable trials.","[\"ai\",\"document-extraction\",\"risk-control\",\"research\"]","2026-08-18T04:00:00.000Z","2026-08-18T13:56:11.803Z","2026-08-18T13:56:23.561Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims 'a stricter fix mostly works,' but the body says the strictest (document-level) guarantee only covers 6% of fields and calls it 'near-vacuous' — revise the dek so it doesn't overstate the strict tier and instead reflects that only the looser 'practical' tier actually held up in the audit.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The article contradicts itself on whether the practical split-based risk check is reliable: mid-body it states real-world risk exceeds the 10% budget in nearly half of resampled trials (\"not a real guarantee\"), yet the dek and closing audit paragraph claim this same check \"holds up\" and is confirmed safe.","ai",[35,37,38,39],"document-extraction","risk-control","research",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.14639",0,{"sections":46},[47,51,55,60,65,70,75,80,85,89,94,99,104,109],{"name":48,"slug":35,"count":49,"latest_published_at":50},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":50},"Security","security",435,{"name":56,"slug":57,"count":58,"latest_published_at":59},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":90,"slug":91,"count":92,"latest_published_at":93},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":110,"slug":111,"count":112,"latest_published_at":113},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]