AI/ ai · security · adversarial-ai · benchmarks

Benchmark Shows AI Decision Models Bend to One Planted Opinion

A new arXiv benchmark called JevAdvBench shows that a single unverified opinion slipped into an AI decision model's input can flip over one in ten verdicts.

A new benchmark says AI models built to hand back a single trustworthy verdict, a probability, a choice, or a score, can be knocked off course by one planted opinion.

The paper, "JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models" (arXiv:2609.31142), posted September 28, 2026, tests models trained with reinforcement learning for calibrated decisions, or RLCD, such as Jev. These models answer a typed question about an input, the state, and software acts on that answer without a person checking it. The authors built JevAdvBench: 812 typed questions across 66 scenarios, plus 9,744 single-edit attack variants that each change one part of a request. On jev-1.13.0, simple rewording stayed within 1.2 percentage points of a clean re-run baseline, but appending one unverified opinion to the state flipped 12.1% of decisions, statistically tied with the strongest injected command at 10.1%, and pushed 38% of confident answers below the 0.8 confidence threshold meant to trigger human review.

That last number is the one that should worry anyone shipping these systems. The whole appeal of a calibrated-decision model is that high-confidence answers skip human review, that is the automation payoff. If a single sentence of opinion can knock over a third of those answers below the review line, the automation shortcut becomes the attack surface.

It is the same prompt-injection problem that has dogged chatbots for years, just wearing a suit and now making decisions nobody double-checks.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →