[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-prompt-injection-rankings-flip-depending-on-where-attacks-land":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9315,"prompt-injection-rankings-flip-depending-on-where-attacks-land","Prompt Injection Rankings Flip Depending on Where Attacks Land","Moving an identical injection payload from a tool's output to its description flips which AI model looks safer, a new study finds.","Where you hide an attack on an AI agent matters as much as the attack itself.\n\nA new study tested 13 large language models using AgentDojo, a benchmark that measures how well AI agents resist prompt-injection attacks - malicious instructions smuggled into text an agent reads. Researchers took one byte-identical malicious payload and planted it in two spots: inside a tool's output, and inside a tool's description. That single change flipped the relative security ranking for 44.9% of all model pairs, across all four task suites in the benchmark. Models that looked resistant to injection via tool outputs often looked far more exposed when the identical payload arrived through tool descriptions, and sometimes the reverse.\n\nThis matters because most prompt-injection benchmarks test one surface and publish a single attack-success-rate number as if it described the model itself. The researchers found real predictive structure behind the instability - using three task suites to guess a model's riskier surface predicted the vulnerable surface on an unseen suite with 76.9% accuracy - meaning this isn't noise, it's a systematic blind spot in how agent security gets measured. Defenses that blocked tool-output attacks sometimes left tool-description attacks wide open.\n\nA security score that only checks one door isn't a security score. It's a guess with a confidence interval attached.","[\"prompt-injection\",\"ai-security\",\"llm-agents\",\"benchmarking\"]","2026-10-01T04:00:00.000Z","2026-10-02T08:58:55.943Z","2026-10-02T08:58:57.613Z","published",null,[],"security",[26,27,28,29],"prompt-injection","ai-security","llm-agents","benchmarking",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.30454",0,{"sections":36},[37,41,44,48,53,58,62,67,72,76,81,86,91,96],{"name":38,"slug":39,"count":40,"latest_published_at":18},"AI","ai",5690,{"name":42,"slug":24,"count":43,"latest_published_at":18},"Security",820,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",430,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",308,"2026-10-01T12:30:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",165,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",150,"2026-10-01T11:59:27.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":73,"slug":74,"count":70,"latest_published_at":75},"Software","software","2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]