A new research system tries to stop AI explanations from quietly lying about how a model actually works.
The system, called MEA, splits the job between two agents. A Proposer agent looks at the question being asked and the type of data involved - tabular, text, or image - and picks and configures the right explanation tool for the job. An Actor agent then turns that tool's raw output into a plain-language explanation. The Actor is trained end-to-end against a reward tied to faithfulness, where faithfulness itself is scored by perturbing a model's inputs and checking whether the explanation's claims hold up to those changes. Across six datasets and three task types - feature attribution, counterfactual reasoning, and spurious-feature detection - the researchers found that frontier LLMs left to explain models unsupervised routinely produce explanations that sound convincing but do not match real model behavior.
That perturbation-based training produced measurable gains over the untrained version of the same system: faithfulness improved 28% on tabular data, 21% on text, and 34% on vision tasks. That matters because today's explainability tools are mostly rigid, single-purpose utilities that require an ML background to run and interpret correctly - this points toward natural-language explanations a non-specialist could actually use.
Worth remembering: the faithfulness score here comes from a perturbation test in a research setting, not from a doctor, loan officer, or regulator checking whether the explanation holds up in the field.