A new black-box test exposes machine unlearning tools that claim to erase a feature's influence but quietly leave it behind.
Researchers built a testing method called CAFE that probes a deployed model without touching its parameters or training code. Instead of just checking whether a feature was used directly, CAFE intervenes on that feature, propagates the change through any downstream features it feeds into, and watches whether the model's predictions actually shift. Tested on two causal-network benchmarks across four different unlearning methods, CAFE ranked residual influence with 0.92-0.93 pairwise accuracy, compared with at most 0.71 for existing direct-input checks. On real census data, it caught influence that survived supposed unlearning and slipped past the older tests entirely.
Machine unlearning is the mechanism companies lean on when a user invokes a right-to-be-forgotten request or a flawed data source needs scrubbing from a live model. The problem CAFE highlights is that most audits only look at whether a feature is used directly, so a model can pass a compliance check while still leaning on that feature's ghost through correlated, downstream signals.
Removing a column from a dataset is not the same as removing its fingerprints, and this research is a pointed reminder that unlearning claims deserve more scrutiny than a quick direct-input scan.