AI/ llms · ai storytelling · benchmarks · generative ai

Benchmark Shows LLMs Can't Keep AI Stories Consistent and Rich

A new benchmark reveals that large language models trade off consistency and narrative richness in evolving story worlds, with no model excelling at both.

A new benchmark finds that even top language models can't juggle a story's memory and its imagination at the same time.

Researchers built WSE-bench, a benchmark that scores AI storytelling not just on the finished text but on the process of generating it. It separately measures three things: how much of a planned story a model actually completes, how well it keeps facts and character states consistent as the plot unfolds, and how meaningfully the story branches based on a reader's or player's choices. Testing frontier models, the researchers found that consistency and narrative richness don't trade off in a simple, predictable way: some middling configurations beat both more consistent and more elaborate ones, and no single weighting of the two scores could pick a clear winner. Adding more narrative structure could make stories richer, but it just as often broke continuity or cut them short, and bigger models mainly generated more content without getting more coherent or creative.

That's a real problem for the growing pile of AI dungeon masters, choice-driven games, and living-world chatbot products that lean on LLMs to run persistent narratives. It suggests that simply using a bigger model, the default fix for most AI shortcomings, won't solve continuity errors or shallow branching on its own. Builders of these systems likely need dedicated memory and state-tracking machinery, not just a beefier base model.

It's a useful reality check for any pitch promising an AI game master that remembers everything and never repeats itself, since right now no model manages both reliably.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →