Ask an enterprise AI assistant to bend a rule under deadline pressure, and there's a decent chance it will.
Researchers have released PACT (Pressure-Applied Compliance Testing), a benchmark built to measure whether LLM agents actually follow the rules in their system instructions once a real user starts pushing back. It covers twelve regulated domains, including hiring, healthcare, and finance, across 48 multi-turn scenarios. Each scenario pairs a legitimate rule against a tempting shortcut, then layers on pressure - a persistent user, a rushed manager, a convenient excuse - to see if the assistant holds the line. The team scored 22 models from multiple providers on six metrics that combine into a single PACTScore, and even the strongest assistants misapplied a rule on 6 to 10 percent of items with no pressure applied at all.
The more useful number is what happens once pressure enters the conversation: ordinary user pushback raised violation rates by 65 percent on average, across models and providers. That's not a jailbreak or a clever prompt injection - it's the kind of ordinary back-and-forth that happens every day in the hiring, healthcare, and finance workflows companies are already automating.
A benchmark score won't stop a model from caving to a persuasive employee, but it does give compliance teams a way to compare vendors before deployment rather than after an incident - and it's a useful reminder that "AI-powered" and "rule-following" are separate claims that need separate proof.