AI/ ai · ai-bias · benchmarks · llms

New Benchmark Finds AI Models Attribute Quotes Unevenly by Race

A new benchmark shows 11 top language models attribute and omit quotes unevenly across race, gender, and intersectional groups.

Ask a chatbot who said something, and the answer can depend on who said it.

Researchers built AttriBench, the first quote-attribution benchmark balanced by both fame and demographics, then tested 11 widely used LLMs on it across multiple prompt setups. Even frontier models struggled to correctly credit quotes to their original authors. The team found large, systematic gaps in attribution accuracy across race, gender, and intersectional groups, not just random noise. They also identified a separate failure mode they call suppression: models omitting attribution entirely, even when they clearly had access to authorship information.

As search tools and research assistants lean harder on LLMs to summarize and cite sources, who gets credited stops being a purely technical question and becomes a fairness one. Suppression is the sneakier half of the problem: a model that quietly drops attribution for certain demographic groups won't show up in a simple accuracy score, which is likely why standard benchmarks have missed it until now.

If chatbots are becoming the front door to information, uneven crediting isn't a rounding error - it's a quieter rerun of the bias problems search engines have been fighting for years.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →