[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-selector-trims-vision-model-tokens-94-percent":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},6403,"new-selector-trims-vision-model-tokens-94-percent","New Selector Trims Vision Model Tokens 94 Percent","A training-free method called StackTok keeps 95 percent of a vision model's accuracy using just 5.6 percent of its image tokens, per a new arXiv paper.","A new token-selection method lets vision-language models skip most of the image data they used to process, with barely a dent in accuracy.\n\nResearchers describe StackTok, a training-free technique for trimming the visual token sequences that vision-language models (VLMs) chew through during inference, in a paper posted to arXiv (arXiv:2609.16841v1, https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.16841). Instead of picking tokens by one fixed rule, StackTok switches on the fly between prioritizing tokens relevant to the specific question and tokens that cover more of the image overall, based on how narrow or spread-out the query seems. Tested across five VLMs and ten image-understanding benchmarks, it beat every other training-free selector in the paper's comparisons. On the high-resolution LLaVA-NeXT-7B model, StackTok kept 95.26% of full-token accuracy while using only 160 of 2,880 visual tokens - a 94.4% cut.\n\nThat matters because image resolution keeps climbing, and with it the token counts driving up VLM inference costs. A fix that needs no retraining is the cheap kind: it can slot into models already in production instead of forcing a costly do-over, unlike prior methods that bake a fixed relevance-versus-coverage balance in from the start.\n\nTraining-free efficiency tricks have a habit of looking great in benchmark tables and less great once they meet real production traffic and hardware quirks. Worth watching for independent replication before anyone bets a deployment on it.","[\"ai\",\"vision-language-models\",\"inference\",\"research\"]","2026-09-16T04:00:00.000Z","2026-09-17T14:15:28.885Z","2026-09-17T14:15:40.812Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the headline's '95 Percent' figure — it conflates the 95.26% performance-retention stat with the actual token reduction, which is 94.4% (160 of 2,880 tokens kept, i.e. 5.6%), so the headline states a different number than the body supports.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The headline's 94 percent figure now matches the body's 94.4 percent token-reduction stat, but the body still attributes the work to unnamed 'researchers' without naming the paper or including the arXiv ID\u002Flink (arXiv:2609.16841v1) — add that citation for attribution.","ai",[34,36,37,38],"vision-language-models","inference","research",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.16841",0,{"sections":45},[46,50,55,60,65,69,73,78,83,87,92,97,102,107],{"name":47,"slug":34,"count":48,"latest_published_at":49},"AI",3853,"2026-09-17T08:27:09.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Security","security",648,"2026-09-17T04:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":54},"Hardware","hardware",154,{"name":70,"slug":71,"count":72,"latest_published_at":54},"Science","science",114,{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":54},"Dev Tools","dev-tools",73,{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]