AI/ ai · explainability · multi-agent-systems · machine-learning-research

Researchers Build Multi-Agent System for Honest AI Explanations

MEA trains an explanation-writing agent against a perturbation-based faithfulness score, beating standard AI explainers across tabular, text, and vision data.

A new research system tries to stop AI explanations from quietly lying about how a model actually works.

The system, called MEA, splits the job between two agents. A Proposer agent looks at the question being asked and the type of data involved - tabular, text, or image - and picks and configures the right explanation tool for the job. An Actor agent then turns that tool's raw output into a plain-language explanation. The Actor is trained end-to-end against a reward tied to faithfulness, where faithfulness itself is scored by perturbing a model's inputs and checking whether the explanation's claims hold up to those changes. Across six datasets and three task types - feature attribution, counterfactual reasoning, and spurious-feature detection - the researchers found that frontier LLMs left to explain models unsupervised routinely produce explanations that sound convincing but do not match real model behavior.

That perturbation-based training produced measurable gains over the untrained version of the same system: faithfulness improved 28% on tabular data, 21% on text, and 34% on vision tasks. That matters because today's explainability tools are mostly rigid, single-purpose utilities that require an ML background to run and interpret correctly - this points toward natural-language explanations a non-specialist could actually use.

Worth remembering: the faithfulness score here comes from a perturbation test in a research setting, not from a doctor, loan officer, or regulator checking whether the explanation holds up in the field.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →