A new benchmark says today's AI shopping assistants are good talkers but bad listeners.
Researchers introduced RealWorldShop, a benchmark built on 3.28 million real, grounded products, complete with structured shopping episodes and a simulated user who reveals and revises preferences over a conversation rather than stating them all upfront. Testing existing systems against it exposed a consistent flaw: they generate responses that sound reasonable in isolation but fail to track what a shopper has already said, update constraints as the conversation evolves, or anchor suggestions firmly to the catalog. The problems got worse in messier, more realistic scenarios - ambiguous requests, bundled purchases, multiple competing goals in one chat. The same team also built RealShopAgent, a framework that adds explicit state tracking, shopping-flow control, catalog-grounded retrieval, and runtime guards, and reports it beats baseline systems on their own benchmark.
Every retailer chasing a shopping copilot is implicitly promising customers a system that remembers what they said five messages ago. This research suggests that promise is still mostly aspirational - the gap between a slick demo and a conversation that holds up under real, shifting intent is wider than the marketing suggests.
A benchmark built in-house by the same team proposing the fix is worth reading with one eyebrow raised, but the underlying critique - that chatbots forget what you told them - matches what anyone who has used one for more than two exchanges has probably already suspected.