[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-finds-llm-agents-leak-data-through-side-doors":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6641,"new-benchmark-finds-llm-agents-leak-data-through-side-doors","New Benchmark Finds LLM Agents Leak Data Through Side Doors","A new evaluation framework shows that checking only an AI agent's final answer or attacker report misses nearly half of the private data it actually exposes.","AI agents that use tools can leak more private data than the tests meant to catch them realize.\n\nA new paper introduces ASLEval, a benchmark for measuring what its authors call \"privacy exposure displacement\" - the gap between what a privacy test actually checks and what an AI agent exposes across an entire multi-step session. Instead of grading only a single designated action or a chatbot's final answer, ASLEval pre-registers a hidden set of sensitive targets, then tracks every declared \"visible exit\" - the outlets, reports, and tool paths an agent could use to leak them - across enterprise-style test environments and multiple independently built runtimes. The researchers found that checking only the expected outlet misses 46.9% of the exposure that shows up when you count every visible exit. Attacker self-reports, a common shortcut for grading these tests, mix real omissions with a high rate of false alarms. Internal, schema-aligned evidence of a leak usually shows up before it becomes visible at the request or probe level, and clamping down on what the model is allowed to return does reduce leakage - but can also break the agent's ability to do its actual job.\n\nThis matters because most privacy audits of agentic AI still grade a single output, not the full session, while enterprises are already routing agents through email, databases, and internal tools where a single narrow check offers false comfort. A benchmark that assumes one correct exit point can certify an agent as safe while missing leaks in tool calls or secondary responses nobody thought to inspect.\n\nThe catch, buried in the paper's own results, is that the fix and the feature are in tension: the more tightly you restrict an agent's visible output to stop leaks, the more likely you are to stop it from finishing the task it was built for.","[\"ai agents\",\"privacy\",\"security\",\"benchmarks\"]","2026-09-17T04:00:00.000Z","2026-09-18T03:58:11.078Z","2026-09-18T03:58:22.985Z","published",null,[],"ai",[26,27,28,29],"ai agents","privacy","security","benchmarks",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.18864",0,{"sections":36},[37,41,44,49,54,58,62,67,72,76,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3852,"2026-09-17T08:27:09.000Z",{"name":42,"slug":28,"count":43,"latest_published_at":18},"Security",648,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",154,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",114,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":18},"Dev Tools","dev-tools",73,{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]