A new benchmark shows most tools built to explain graph neural networks aren't actually reading the model's real logic.
The team found that popular GNN explainability benchmarks assume a model trained to spot a planted pattern, or motif, actually relies on that motif - but that assumption often fails, since something as simple as node degree can solve the task instead. To get a benchmark they could trust, they skipped training altogether. They built Gracr, a compiler that turns graded modal logic formulas directly into GNN weights, so a model's behavior is fixed by the formula instead of learned from data. With the correct answer now known exactly, they built GracrBench and tested eleven explainer methods across six tasks.
That matters because explainer tools are usually graded on plausibility - how closely their output matches a guessed ground truth - rather than whether they're reading the model correctly. A GNN that's quietly keying off degree statistics will still make a weak explainer look good, since the explainer and the benchmark share the same blind spot. Most of the eleven explainers tested broke down once a formula had indirect influences or was implemented a different way, even when the underlying logic was identical.
It's a tidy illustration of grading your own homework: an explanation tool checked against its own assumptions will usually pass, until someone builds a model it can't talk its way around.