AI/ ai-agents · benchmarks · llm-evaluation · reinforcement-learning

AI Agents Still Lose to Scripts at Running an Online Store

StoreBench pits seven top AI models against a simulated apparel store, and the best one still can't beat a basic scripted strategy.

Researchers built a live online store to see if AI agents can actually run a business, not just chat about one.

StoreBench puts an AI agent in charge of a mid-size apparel store running on a real commerce backend, complete with round-the-clock orders, suppliers that reprice or go dark, and market shocks with little or no warning. The agent manages all of it through the same 29 tools a human merchant would use, and a budget system ties simulated time to actions taken so slower models can't stall the clock. Researchers tested seven frontier models across 11 scenarios ranging from 30-day sprints to a full simulated year, measuring results against scripted benchmark policies. The best performer, DeepSeek-V4-Pro, cleared only 49% of task-seed cells, compared with 97% for a simple scripted smart-triage policy, and human experts working the same tools still beat every model on average.

The gap shows a familiar pattern: these models can narrate a plan but struggle to execute one under sustained uncertainty, where a consistent heuristic still wins. There's a hint of promise in the training results, though: a small amount of targeted reinforcement learning took one model, Qwen3.5-27B, from a composite score of 0.136 to 0.373 after training on just five tasks, and most models improved markedly over a simulated year of practice.

The researchers are keeping the full environment under wraps to stop future models from simply memorizing the test, which says plenty about how fast benchmarks get gamed once they leak.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →