Spotify's new conversational recommendation agent learned to chat about music by practicing on conversations nobody actually had.
Researchers at Spotify built a pipeline that generates synthetic multi-turn conversations from single-turn prompts, like "recommend Italian indie artists I haven't heard before," to test how a recommendation agent plans and sequences its tool calls before any real users touch it. They paired that with a self-improvement loop that uses variance-based contrastive optimization and a coding agent to automatically spot and fix planning errors. The approach lifted output quality 8% over an already-tuned manual prompt. Spotify says the system is now in production and sped up the agent's path to launch.
Cold-start is the perennial problem for any AI feature that needs real conversations to get good, and most companies just eat the awkward early version while users complain. Spotify's answer - generate the awkward conversations synthetically and let a coding agent fix the failures before launch - is a workable template other recommendation and support bots could copy. The payoff shows up in the numbers: a 14% jump in listening time, 5% more weekly active users, and a 5% drop in skip rate versus the older session-only refinement flow.
Whether that self-improvement loop generalizes beyond song recommendations, or just quietly overfits to Spotify's own catalog quirks, is the open question the paper doesn't answer.