Security/ ai-agents · ai-safety · llm-security · academic-research

Researchers Build a Honeypot That Predicts AI Agent Attacks

A new framework has small AI models predict an agent's next moves to catch multi-turn attacks before they fully play out, not just after the fact.

Researchers have built a honeypot that guesses what an AI agent will do next, then checks if it was right.

The paper describes a system called Speculative Safety Honeypot, or SSH. It borrows its core idea from speculative decoding, a technique normally used to make large language models run faster. Instead of predicting the next word, SSH uses a cluster of small language models to simulate and predict an AI agent's future actions, building out a tree of likely next moves before the agent actually takes them. A second verification stage then compares the agent's real behavior against that predicted tree, pruning branches that do not match and using the mismatches to catch attacks that only reveal themselves across several turns.

Most agent-security tools work backward, scanning conversation history for red flags after the fact. That approach misses attacks deliberately split across turns so no single message looks dangerous on its own. By scoring risk against a whole predicted trajectory rather than one exchange, SSH is designed to flag trouble earlier and lean less on any single detector being perfect.

It is a clever reuse of an inference-speed trick for security, but it is still a paper, not a product - and like any prediction system, its usefulness depends on how often its guesses about bad behavior turn out to be right.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →