A new study pries open AI speech emotion detectors and finds they are mostly reading the transcript, not the voice.
Researchers adapted concept bottleneck models, a technique borrowed from image classification, to speech emotion recognition. Instead of feeding raw audio straight into a model, they extracted explicit concepts first: transcripts, descriptions of acoustic qualities, and speaker attributes. Testing three large language models on three benchmark datasets (CREMA-D, IEMOCAP, and MELD), they found that in a zero-shot setting the models leaned so hard on the transcript that performance on CREMA-D collapsed, with the Macro-F1 score dropping from 27.8 to 5.8. Fine-tuning reversed that effect: once fine-tuned, adding the transcript back in actually helped, lifting Macro-F1 from 41.8 to 45.1.
The more interesting finding is buried past the headline numbers. Stripping out a single concept, like speech rate, flipped 48 percent of "Neutral" labels to "Disgust" on CREMA-D. Separately, removing intensity level on MELD reshuffled many individual predictions even though the overall Macro-F1 score barely changed.
That is a useful reminder for anyone shopping an off-the-shelf model for sentiment or emotion detection: the leaderboard number is not the whole story, and sometimes not even most of it.