[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-technique-cuts-ai-reasoning-inference-costs-288-percent":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},5158,"new-technique-cuts-ai-reasoning-inference-costs-288-percent","New Technique Cuts AI Reasoning Inference Costs 28.8 Percent","Funnel of Thoughts prunes unproductive AI reasoning trajectories early, cutting total inference cost by 28.8 percent while matching majority-vote accuracy.","AI reasoning models waste a lot of compute rambling before they give up on a bad answer. A new method fixes that by cutting off the bad answers early.\n\nResearchers built Funnel of Thoughts (FoT), an inference-time technique that prunes unpromising reasoning trajectories before they finish generating. After analyzing 115,000 reasoning traces from six large reasoning models, the team found that trajectories destined to be wrong tend to repeat hesitation words like \"Wait,\" \"Actually,\" and \"perhaps,\" sometimes spiraling into loops that never reach an answer. FoT uses that lexical pattern, with no extra model calls or retraining required, to kill those trajectories early. Applied to the standard practice of generating 32 samples and taking a majority vote, it matches that method's accuracy while cutting attention FLOPs during generation by 56.1%, wall time by 37.6%, and total inference cost by 28.8%.\n\nThat last number matters more than the flashier ones. Multi-sample voting is becoming the default way to make reasoning models reliable, and reliability at scale is expensive. A training-free trick that shaves nearly 30% off the real inference bill, and reportedly transfers to other model architectures and tasks without retuning, is a meaningful dent in that cost curve rather than a benchmark curiosity.\n\nStill, a heuristic built on catching AI models saying \"actually\" and \"wait\" feels more like a clever patch than a fundamental fix, and the big headline savings apply to attention FLOPs during generation, not the total bill, so temper the excitement accordingly.","[\"ai\",\"llm inference\",\"reasoning models\",\"efficiency\"]","2026-08-18T04:00:00.000Z","2026-08-18T07:33:18.473Z","2026-08-18T07:33:30.251Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims the technique cuts 'inference costs by half,' but the body reports only a 28.8% reduction in total inference cost — the ~56% figure applies solely to attention FLOPs during generation, not overall cost — so rewrite the dek to state the 28.8% total inference cost reduction accurately.","resolved","ai",[30,32,33,34],"llm inference","reasoning models","efficiency",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15065",0,{"sections":41},[42,46,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":45},"Security","security",435,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]