AI still can't reliably tell a performed emotion from a felt one.
Researchers behind a new paper introduce TRACE, a framework that treats an emotional moment as three linked stages: the Condition that triggers it, the internal Affect a person experiences, and the external Effect, the regulated display and its social consequences. Built on that model, TRACE-Bench tests multimodal AI systems on 3,746 structured questions drawn from 646 real-world videos, covering five tasks: spotting grounded affect, decoding how someone is regulating their expression, reasoning about what caused a feeling, reasoning about what follows from it, and reconstructing the whole chain. Compared against human annotators, every tested model showed a sizable performance gap. Oddly, models built specifically for emotion recognition generally did worse than general-purpose multimodal models not designed for the task.
The failures are specific and telling. Models routinely treat a put-on smile as genuine happiness, and when asked to trace a long emotional chain, they invent events that never happened in the video. That matters for anyone building on these systems for things like customer-service sentiment analysis or wellness apps, where mistaking performance for feeling produces bad advice. The researchers' own fix, TRACER, forces the model to justify each inference against explicit observations and prior conclusions, and it beat every baseline model on all five tasks.
It is a useful gut check for anyone marketing emotionally intelligent AI. Recognizing a face is easy. Understanding why it is making that expression is still mostly guesswork.