[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-new-attack-makes-ai-inference-quietly-slower-and-pricier":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},10966,"a-new-attack-makes-ai-inference-quietly-slower-and-pricier","A New Attack Makes AI Inference Quietly Slower and Pricier","Researchers show adversarial text suffixes can hijack speculative decoding, the trick behind fast LLM inference, to inflate a victim's compute costs.","A new attack can quietly make AI chatbots slower and more expensive to run, without changing their answers.\n\nResearchers describe Speculative Rejection Attacks, or SRAs, two methods called Speedbump-P and Speedbump-D that target speculative decoding, the technique most large LLM providers use to speed up responses by having a small draft model guess ahead while a larger target model checks its work in one pass. The attacks append an adversarial suffix to content the attacker controls, tuned so the draft and target models disagree more often. That forces the target model to redo more work token by token, and in some tests pushed speculative decoding to run slower than ordinary autoregressive decoding. The suffixes keep working under random sampling, and they transfer: Speedbump-P suffixes fool other draft models, while Speedbump-D suffixes work across target models that share the same drafter.\n\nSpeculative decoding exists specifically to cut inference cost and latency, so an attack that reverses those gains turns a performance feature into a billing liability. The tradeoff the researchers found is the sharper detail: pushed to maximum slowdown, the attack visibly degrades output quality, and reining it in to stay stealthy gives back most of the slowdown too.\n\nThink of it as a sponge attack for the agentic era: instead of poisoning what a model says, it poisons what a model costs, and it only needs attacker-controlled text sitting somewhere an LLM will read it, like a webpage or an uploaded document.","[\"speculative-decoding\",\"llm-security\",\"adversarial-attacks\",\"ai-inference\"]","2026-10-09T04:00:00.000Z","2026-10-09T23:47:51.264Z","2026-10-09T23:47:57.383Z","published",null,[],"security",[26,27,28,29],"speculative-decoding","llm-security","adversarial-attacks","ai-inference",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.10929",0,{"sections":36},[37,41,44,49,54,58,62,67,72,77,82,87,92,97],{"name":38,"slug":39,"count":40,"latest_published_at":18},"AI","ai",6708,{"name":42,"slug":24,"count":43,"latest_published_at":18},"Security",931,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",486,"2026-10-08T22:40:11.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",474,"2026-10-08T22:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",231,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",192,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",181,"2026-10-08T23:26:35.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Startups","startups",117,"2026-10-08T16:45:00.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"General","general",66,"2026-10-09T04:46:11.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",58,"2026-10-08T20:08:45.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]