A fine-tuned Llama model now reads emotional tone in speech and text better than GPT does, at a fraction of the size.
Researchers benchmarked six language models from the LLaMA, GPT, and Qwen families on IEMOCAP, a dataset of recorded conversations rated for valence, arousal, and dominance, or VAD, the three dimensions used to map emotional intensity and polarity. They compared zero-shot prompting, few-shot prompting, and low-rank adaptation, or LoRA, a method that fine-tunes a small slice of a model's parameters rather than retraining the whole thing. The LoRA-tuned LLaMA models beat prompt-engineered GPT models at both classifying emotions and scoring them along the VAD scale, despite GPT's larger size, and the best configuration hit a valence score of 0.7822, a new high for the benchmark. Written descriptions of vocal tone also boosted accuracy for the smaller models, though they barely moved the needle for the largest one.
The result is a data point for an argument that's been building across AI research: targeted fine-tuning on the right data can beat raw scale for a specific skill. That matters for anyone building voice assistants or support-ticket triage that needs to gauge how a customer feels, because it means a cheaper, smaller model tuned on the right dataset can outperform a giant general-purpose one. It's a real cost lever, not a marketing claim.
Worth noting: the dataset's own human annotators didn't fully agree on these emotion ratings either, and the models' errors tracked that disagreement almost dimension for dimension. Call it emotion recognition if you like, but it's closer to modeling consensus noise than reading feelings.