A new data poisoning technique can hide malicious training data so well that no known filter catches it.
Researchers describe an attack called Phantom Transfer that adapts a technique known as subliminal learning so it works in realistic training setups, not just lab conditions. The poisoned data looks ordinary to a human or a model checking it, yet it still pushes a trained model toward attacker-chosen behavior. The attack doesn't care which model generated the poisoned samples, which model later trains on them, or what the attacker is trying to achieve. In tests against 11 separate data-level defenses, including one that has a second model paraphrase every sample before training, Phantom Transfer got through all of them.
That paraphrasing defense was supposed to be a strong catch-all: rewrite everything and any hidden signal should wash out. It didn't. The researchers also showed the attack can plant a password-triggered behavior, code that only activates when a specific trigger phrase appears, while still slipping past every defense tested.
Their fix isn't a better filter. It's auditing the trained model itself after training, which says something about how far behind detection currently sits.