[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-technique-prunes-audio-tokens-before-ai-model-runs":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9078,"new-technique-prunes-audio-tokens-before-ai-model-runs","New Technique Prunes Audio Tokens Before AI Model Runs","A new method called Triage predicts how much attention audio tokens will get before the language model even starts, cutting compute without hurting accuracy.","A new technique called Triage predicts, before a language model runs a single layer, which audio tokens it will actually pay attention to.\n\nAudio language models typically convert one minute of speech into 750 to 1,500 tokens and feed every one into the model, which is costly. Image models already prune unneeded tokens early, right after the first couple of layers, because image tokens draw little attention there. Audio tokens behave differently: they draw heavy attention from the start, and their final importance ranking is still unsettled, so there is no safe point to prune them after the fact. The researchers found that a simple linear model, fit without labeled training data, can predict an audio token's eventual attention ranking directly from the audio encoder's output, achieving a correlation of .69 or better on 11 of 13 models tested.\n\nBecause the cuts happen before the expensive part of the model runs, the savings are real, not just theoretical: at a conservative setting, word error rate and accuracy stay within .04 of using full, uncompressed audio, and at a more aggressive setting Triage beats every baseline across all twelve transcription tests. On multiple-choice tasks, it also outperforms DART, an existing audio-compression method that was previously the best-performing baseline, by .043 in average accuracy at compression ratios of 2.2x to 5x. That same compression roughly triples usable context, pushing Qwen2.5-Omni-3B's audio limit from 21.8 minutes to about 62 minutes, and lets one GPU serve four times as many concurrent five-minute streams.\n\nIt's unglamorous plumbing work, not a flashy new model launch, but figuring out what to ignore before you pay to process it remains one of the more durable ways to make AI inference cheaper.","[\"ai\",\"audio ai\",\"token pruning\",\"inference efficiency\"]","2026-10-01T04:00:00.000Z","2026-10-01T18:50:03.535Z","2026-10-01T18:50:06.049Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Introduce DART with enough context (what kind of method\u002Fbaseline it is) before citing the .043 accuracy comparison — right now it's used as an unexplained acronym\u002Fname.","resolved","ai",[30,32,33,34],"audio ai","token pruning","inference efficiency",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38878",0,{"sections":41},[42,45,49,54,59,64,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5488,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",809,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",162,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":79,"slug":80,"count":76,"latest_published_at":81},"Software","software","2026-09-30T21:41:11.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]