AI models built to fact-check claims can be tricked into reversing a verdict simply by relabeling a source's trustworthiness, even when the evidence itself never changes.
That finding comes from an arXiv preprint, 2609.36611, posted September 30 and not yet peer-reviewed. Its authors built a test called TrustSwap that keeps every piece of evidence text identical while swapping, lowering, or removing the HIGH or LOW trust label attached to a source. Confidence scores and the decision to keep searching for more evidence moved the way they should in 49 of 50 comparisons, but the verdict itself flipped in 4 to 23 percent of confident cases for Qwen3 models and in up to 50 percent of cases for one existing RL-trained fact-checker, based on nothing but the label swap. Standard GRPO reinforcement fine-tuning made that shortcut worse in all six settings tested at 8 billion parameters, and the authors' proposed fix, trust-swap augmentation, cut the flip rate by 7 to 35 percent at 4 billion parameters but stopped working reliably once models scaled up to 8 billion.
That's a familiar shape of failure: reinforcement learning tends to reward whatever cheap signal correlates with the right answer, and a trust label is a much easier tell to grab than actually weighing evidence. For anyone hoping to hand fact-checking off to an AI system, that undoes the whole point: a bot that changes its mind based on who supposedly said something, rather than what was said, is reproducing the exact bias it was built to catch.
It's a reminder that RL fine-tuning teaches a model to get reward, not to read, and accuracy scores alone cannot tell the difference between the two.