[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-predicts-llm-output-length-using-attention-weights":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5207,"new-method-predicts-llm-output-length-using-attention-weights","New Method Predicts LLM Output Length Using Attention Weights","ESTP blends token entropy with attention based importance scores to predict how long an LLM response will run before it finishes generating.","A new technique called ESTP wants LLM servers to stop guessing how long a response will be and start predicting it.\n\nESTP, short for Entropy and Semantic Token Pooling, comes from a paper posted to arXiv on August 18, 2026. The problem it targets is mundane but expensive: LLM serving systems often pad every sequence to a fixed maximum length, which wastes compute and slows throughput. Earlier fixes leaned on entropy guided token pooling, using token by token uncertainty as the main signal for predicting output length. ESTP adds a second signal, pulling attention based importance scores straight from the self attention weights computed during the prefill phase, so it can weigh semantically important tokens instead of just noisy ones.\n\nThe reuse of prefill activations is the clever part. It means ESTP adds almost no extra memory overhead and only minimal latency, which matters because a length predictor nobody can afford to run is not a length predictor anyone will ship. On the ForeLen benchmark, ESTP beat baseline methods on prediction accuracy and error rate in most scenarios, and paired with a length aware scheduler it improved throughput and cut the padding ratio in end to end tests.\n\nThis is an infrastructure paper, not a model paper, and that is exactly why it matters. Every efficiency gain in serving translates directly into lower inference costs, which is the line item every LLM provider is fighting to shrink right now. The catch is the same one every entropy-plus-something paper faces: \"in most scenarios\" is doing real work in that sentence, and the paper does not spell out where ESTP loses. Worth watching whether independent benchmarks confirm the gains outside ForeLen before serving teams bet infrastructure on it.","[\"llm-serving\",\"inference-optimization\",\"arxiv\",\"attention-mechanisms\"]","2026-08-18T04:00:00.000Z","2026-08-18T09:38:03.767Z","2026-08-18T09:38:15.612Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The body states ESTP 'beat baseline methods on prediction accuracy and error rate' as an unqualified win, but the source specifies this held only 'in most scenarios' — restore that qualifier so the claim matches the paper's actual reported results.","resolved","ai",[32,33,34,35],"llm-serving","inference-optimization","arxiv","attention-mechanisms",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15592",0,{"sections":42},[43,47,51,56,61,66,71,76,81,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":46},"Security","security",435,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]