AI models that explain their face matches in plain English often can't be trusted to explain them accurately, a new benchmark shows.
Researchers built a testing framework for vision-language models, AI systems that pair image understanding with text generation, used in face recognition. Some forensic and identification tools favor these models because they don't just output a similarity score, they generate a written justification, something like "the eyes and nose shape match." The new benchmark checks two things: whether that justification actually points to identity-stable features, called relevance, and whether it accurately describes what's visible in the photos rather than making things up, called faithfulness. The team tested several families of open-weight models, found consistent shortcomings on both counts, and released the evaluation harness as open-source.
Forensic face-matching decisions are supposed to be auditable, and a text explanation only helps if it holds up under scrutiny. This work is a useful check on a growing assumption in the field: that giving a model the ability to write a sentence about its own reasoning means that sentence is true. It usually doesn't, and the gap matters far more in a courtroom than in a chatbot.
The tools are getting more articulate long before they're getting more honest.
