A new benchmark says most AI commentators go quiet exactly when a game gets interesting.
Researchers built GameCommBench, a benchmark that spans board games, traditional sports, and esports, with commentary tagged by type rather than lumped into one generic category. Alongside it they built Type-Aware Commentary Evaluation, or TACE, a scoring framework meant to judge each type of commentary on its own terms instead of grading everything with the same yardstick. The team checked that TACE's scores line up with how human judges rate commentary, then used it to test a set of existing AI commentary generators. The results were uneven: some commentary types came out strong, others barely held up.
The weak spots are the ones that matter most to a viewer. Live observation, describing what is happening as it happens, and strategic analysis, explaining why a move matters, were the two biggest bottlenecks across the systems tested. Scripted-sounding recap and color commentary is apparently easier to fake than real-time insight.
That tracks. Reciting stats is easy. Calling the game as it unfolds, and explaining why the crowd should care, is the part human broadcasters still get paid for.