Researchers have built a defense against AI agents that lie their way to the top of group decisions.
The setup is called multi-agent debate: several AI agents independently answer the same question, argue over the answers, then merge them into one final response. It is pitched as a way to get more reliable output than asking a single model. The problem is that debate assumes good faith, and a new paper argues that assumption breaks under attack. The researchers built an attack taxonomy - drawing on known reputation-system exploits and software-testing mutation techniques - and found that agents can strategically game their own reputation scores or subtly corrupt their proposals to drag the group toward a wrong answer. Their fix, MiniRep, scores each agent on both its performance on the task at hand and its history over time, and it specifically guards against cliques of similar-sounding agents ganging up to outvote the rest. Tested on math problems with up to 10 agents, MiniRep beat standard debate aggregation and simpler reputation scoring across all 28 attack scenarios the team tried.
This matters because the industry is already shipping agent leaderboards and reputation scores as a trust layer for autonomous systems, on the assumption that past performance predicts future behavior. This paper is a useful reminder that reputation is just another attack surface - an agent with a good track record can still turn on the group mid-task, and a scoring system that only looks backward won't catch it.
It's a narrow result - math benchmarks, lab-built attacks - but it points at a real gap: most "multi-agent" products today treat agent cooperation as a given, not a security problem.