Security/ ai · security · ai-agents · coding-agents

Researchers build a digital twin to police AI coding agents

A new framework called Pincer uses a per-user digital twin to approve or block an AI coding agent's resource requests, and it beats rival defenses in testing.

A research team has built a "digital twin" that decides, on the fly, what an AI coding agent is allowed to touch.

The paper describes Pincer, a defense that sits at the resource layer rather than the tool-call layer most existing safeguards operate on. It uses an isolated-context model, trained on a new user-centric dataset of multi-day user-agent transcripts, to learn an individual's preferences over time. That digital twin then acts as a stand-in for the user, approving or denying the permission requests a long-running coding agent makes as it works through a task. The researchers say this is needed because today's deployed agents, including Claude and Codex, mostly lean on user-maintained permission lists that decay or on automatic tool-call classifiers that were never built to catch adversarial setups and learn nothing user-specific.

The useful admission here is that permission fatigue is itself a security hole, not just an annoyance. Every "always allow" click from a tired developer becomes a permanent policy decision, and Pincer's bet is that a model quietly watching real behavior over days can approximate what a user would actually refuse, better than a static allowlist or a context-blind classifier can.

The paper benchmarks Pincer against several LLM-judge setups and against adaptations of Conseca, a resource-authorization system for agents proposed at HotOS 2025, and reports Pincer winning by wide margins on some attack types while still outperforming on the rest. That is a strong claim from one paper's own tests, and it will mean more once independent teams try to reproduce it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →