Six AI models trained to read brainwaves ace clean lab tests but stumble once the data gets messy, according to a new evaluation.
Researchers tested six EEG foundation models, plus a standard supervised baseline, across ten datasets, looking past raw accuracy to robustness, interpretability, and representational capacity. They hit the models with added noise, random channel dropout, and region-specific signal loss, and found no single model survived every failure type: the model most resistant to noise was among the most fragile when electrodes dropped out entirely. Using attribution methods, the team found the models generally focus on the brain regions neuroscience says should matter for a given task. They also found that early layers already hold task-relevant information, and that weak performance previously blamed on bad pretrained representations was actually caused by how token embeddings get pooled before classification.
That matters because foundation models get sold on the promise that scale and pretraining produce broadly reliable representations, a pitch borrowed from the language-model world. For EEG, where electrode noise and signal dropout are routine in real headsets and clinical settings, a model that only looks robust on one axis is not actually deployment-ready.
An accuracy leaderboard score, in other words, says almost nothing about whether a model survives someone's EEG cap slipping half an inch.