[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-pdf-datasets-count-documents-not-the-text-inside-them":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},5454,"pdf-datasets-count-documents-not-the-text-inside-them","PDF Datasets Count Documents, Not the Text Inside Them","An arXiv analysis finds that PDF training corpora report stats by document count, hiding a truncation problem that erases most of a document's text.","Most PDF datasets brag about their size in tokens, but a new analysis shows their internal stats are measured in something else entirely: documents. That mismatch hides a lot.\n\nA study of CC-MAIN-2021-31-PDF-UNTRUNCATED, a 7.9 million-document, 32.6-billion-token web PDF corpus, found the text is wildly concentrated: 3.02% of documents hold half the tokens, and documents over 50 pages make up just 5.00% of the corpus but 53.53% of its text. PDFs built with a TeX toolchain are only 1.66% of documents yet 4.05% of the text. The bigger issue is Common Crawl's truncation cap, which affected 23.06% of documents but 63.08% of the text -- and when researchers tried to recover the missing content, two widely used extraction libraries pulled back only 11.4% and 1.4% of it, with 72% to 97% of affected documents yielding nothing at all. Even under the larger 5 MiB cap Common Crawl adopted in March 2025, 30.19% of tokens are still cut off, and recovery only improves from 3.3% to 13.2%.\n\nAny dataset statistic reported per document -- coverage, OCR routing, language mix -- is quietly describing a small slice of the actual text, since a handful of long documents carry most of the tokens. For anyone building or auditing an LLM's training data, that means headline corpus sizes can overstate what's usable, and truncation isn't a rounding error -- it's more than half the corpus's text, gone.\n\nThe paper's fix is unglamorous: report stats in both units, documents and tokens. Not flashy, just accurate -- which turns out to be rarer than it should be.","[\"training-data\",\"pdf-corpora\",\"common-crawl\",\"data-quality\"]","2026-08-18T04:00:00.000Z","2026-08-18T20:54:51.855Z","2026-08-18T20:55:03.753Z","published",null,[],"ai",[26,27,28,29],"training-data","pdf-corpora","common-crawl","data-quality",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.16390",0,{"sections":36},[37,41,45,50,55,60,65,70,75,79,84,89,94,99],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":40},"Security","security",435,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]