A new spoken-dialogue AI called LoopSLM figures out how you're feeling without narrating its thought process first.
Researchers built LoopSLM on looped Transformers, reusing a single decoder block across multiple passes to refine its read on acoustic cues like tone, pace, and emotion, instead of writing out a chain-of-thought explanation before replying. A two-stage training process separates learning to reason about those cues from learning to generate a response, so the model reasons internally and answers directly. On the EchoMind benchmark, it beat Qwen2.5-Omni-7B on paralinguistic understanding, reasoning, and reply quality. Against a chain-of-thought baseline, it gained more than 20 points in reasoning accuracy while producing 64.5% fewer tokens at half the latency, and it topped Qwen3-Omni-Thinking on most empathetic-reply metrics with 34 times lower latency.
Chain-of-thought has become the default trick for making AI sound thoughtful, but writing it out takes time - a real problem for a live phone call or voice assistant, where a pause reads as awkward rather than deliberate. LoopSLM's results suggest that reasoning habit, borrowed from text models, doesn't transfer cleanly to speech, where latency is part of the experience itself. It's also worth noting the model trained only on dialogue data but still improved on general audio benchmarks it never saw during training.
This is a single academic paper, not a shipping product, so treat "34x faster" as a lab result until someone runs it on a real customer-service line.