A research team has built a "digital twin" that decides, on the fly, what an AI coding agent is allowed to touch.
The paper describes Pincer, a defense that sits at the resource layer rather than the tool-call layer most existing safeguards operate on. It uses an isolated-context model, trained on a new user-centric dataset of multi-day user-agent transcripts, to learn an individual's preferences over time. That digital twin then acts as a stand-in for the user, approving or denying the permission requests a long-running coding agent makes as it works through a task. The researchers say this is needed because today's deployed agents, including Claude and Codex, mostly lean on user-maintained permission lists that decay or on automatic tool-call classifiers that were never built to catch adversarial setups and learn nothing user-specific.
The useful admission here is that permission fatigue is itself a security hole, not just an annoyance. Every "always allow" click from a tired developer becomes a permanent policy decision, and Pincer's bet is that a model quietly watching real behavior over days can approximate what a user would actually refuse, better than a static allowlist or a context-blind classifier can.
The paper benchmarks Pincer against several LLM-judge setups and against adaptations of Conseca, a resource-authorization system for agents proposed at HotOS 2025, and reports Pincer winning by wide margins on some attack types while still outperforming on the rest. That is a strong claim from one paper's own tests, and it will mean more once independent teams try to reproduce it.