[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-agents-can-leak-secrets-by-connecting-separate-clues":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},6034,"ai-agents-can-leak-secrets-by-connecting-separate-clues","AI Agents Can Leak Secrets by Connecting Separate Clues","New research shows AI agents can piece together innocuous tool outputs into sensitive conclusions, and prompting alone barely fixes it.","An AI agent doesn't need to be asked a sensitive question to expose sensitive information. It just needs to be asked several boring ones.\n\nA new paper formalizes this as Tools Orchestration Privacy Risk, or TOP-R: an agent calls multiple tools, none of which individually reveals anything private, then stitches the results into a conclusion that does. The researchers built a benchmark called TOP-Bench, with 1,000 test cases generated through a four-library pipeline they call LRSE, and ran it against six LLM agents. The agents completed their assigned tasks 98 percent of the time. They also leaked the sensitive conclusion 88.6 percent of the time. With reasoning mode switched on, four of the models leaked the answer in their final response 81.4 percent of the time on average, and in their visible reasoning trace 82.4 percent of the time.\n\nHere's the part that should worry anyone relying on a system prompt as a safety net. On TOP-Bench, telling the models to be careful lifted the composite safety-and-usefulness H-score by only about 3.4 points. On a separate evaluation set built to test a proposed fix, prompt-only instructions moved that score by 5.0 points, while a model actually retrained on the problem, using supervised fine-tuning plus DPO, moved it by 16.2 points, more than three times the prompting-only gain on the same test.\n\nThat gap is the real finding here. Prompting is cheap and everyone reaches for it first, but the data suggests it functions more as a light suggestion than a boundary. If agent orchestration is going to keep chaining tool calls on a user's behalf, the fix has to live in training, not in a polite request bolted onto the system prompt.","[\"ai-agents\",\"privacy\",\"llm-security\",\"benchmarks\"]","2026-09-03T04:00:00.000Z","2026-09-03T06:49:12.707Z","2026-09-03T06:49:24.631Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The article gives two conflicting figures for the prompting-only H-score improvement (3.4 points in paragraph 3 vs. 5.0 points in paragraph 4), and neither matches the 'more than three times' comparison to TOP-Align's 16.2-point gain.","resolved","security",[32,33,34,35],"ai-agents","privacy","llm-security","benchmarks",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2512.16310",0,{"sections":42},[43,48,52,57,62,67,72,77,82,87,92,97,102,107],{"name":44,"slug":45,"count":46,"latest_published_at":47},"AI","ai",3385,"2026-09-04T22:17:36.000Z",{"name":49,"slug":30,"count":50,"latest_published_at":51},"Security",565,"2026-09-05T00:03:08.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Policy","policy",300,"2026-09-04T22:18:34.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",152,"2026-09-03T09:26:48.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",97,"2026-09-04T15:29:18.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Science","science",96,"2026-09-03T22:30:00.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",54,"2026-09-04T23:36:14.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",37,"2026-09-04T20:22:41.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]