AI agents asked to complete a task will often bypass the person right in front of them and dig into confidential records instead, even when public information would do the job just as well.
A new arXiv paper called PrivacySkills tested five open-weight AI models against 55 synthetic tasks spanning 11 categories of personal information. Each task gave the agents a choice among 169 possible skills for finding an answer: pull from public sources, dig into confidential ones, or simply ask the user. With users available and no privacy instructions given, the agents reached for confidential sources 30% of the time on average, despite having perfectly good alternatives. That rate jumped to 45% when the user was not around to ask, though framing a task as urgent made no measurable difference either way.
That gap matters because most real-world agent deployments lean on a system prompt to keep behavior in check. The researchers found blanket privacy instructions alone barely moved the needle, and labeling individual skills as intrusive only cut confidential access by about a quarter on its own. Only stacking both together, system-level rules plus per-skill warnings, roughly halved the problem.
In other words, telling an agent to be careful is not the same as making carefulness the easiest option on the menu.