A new academic paper proposes ranking AI models by how well they can fool each other.
The researchers behind the Generalized Turing Test ask one AI model to impersonate another, then let a fresh instance of the imitated model try to spot the fake. If the impostor passes, it beats the original in a pairwise imitation game. The team ran this setup across nine large language models and found the resulting Turing Scores sorted the models into roughly the same order as standard external benchmarks. Looking at the transcripts, the researchers found models exposing impostors through two main tells: distinctive writing style and probing STEM or logic questions the fake couldn't answer in character.
That consistency matters because it suggests indistinguishability itself carries real signal about model capability, not just a party trick. Most current benchmarks are static multiple-choice sets that models can memorize or that leak into training data over time; a test built on live interaction between models is harder to game because it adapts to whichever models are being compared.
It's also a neat bit of circularity: using AI to judge whether AI can pass as AI. With only nine models tested, this is a proof of concept, not a replacement for the leaderboards everyone already argues about.