AI/ ai · research · human-ai collaboration · llm accuracy

Half of AI's Accuracy Advantage Never Reaches the User

A 535-person study found only about half of an AI model's raw accuracy edge actually reaches the human using it, and that share varies by model.

Pairing with a chatbot on a reasoning test doesn't get you the chatbot's full brainpower - you get about half of it.

A study posted to arXiv this week put 535 people through a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies. Some worked alone; others were required to consult one of four models, GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3, on every assisted item. Each model also took the same battery solo, answering every question 100 times, giving researchers a clean baseline for how good each model actually was. When they compared a model's own accuracy gains to what its human partners achieved, only about half of the model's edge carried through to the team, and the exact share depended on which model was doing the advising.

That gap matters because this is the realistic case: a person weighing a chatbot's answer, not a model running unsupervised. The study also found that people's confidence after getting advice was a worse predictor of being right than their confidence before asking, so gut feeling does not reliably flag when to override the model. Deference to the AI rose as the model got better at a given task type, but that pattern varied enough across tasks that no simple rule of thumb protected participants from bad advice.

Vendors love to publish solo benchmark scores. This study is a reminder that the number worth watching is how much of that score survives contact with an actual, skeptical user.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →