A hobbyist blogger is using the 1989 platformer Prince of Persia to size up how far AI models have actually come.
The post, published Sept. 25 on blog.priyan.in, is titled "Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia." It treats the decades-old game as an informal benchmark for tracking capability gains across frontier AI models. The piece picked up 14 points and 7 comments on Hacker News in its first day, a modest but real signal of developer interest. It's the kind of personal, low-stakes test that circulates in technical circles far more than in mainstream AI coverage.
Standard AI benchmarks get gamed, memorized, or saturated fast, which is part of why quirky, homemade tests like this keep showing up. A single favorite game won't replace a formal eval suite, but it captures something leaderboards often miss: whether a model can reason through a concrete, familiar task rather than just score well on a multiple-choice set. That kind of grassroots scrutiny is hard for official benchmark makers to replicate.
Whether frontier models actually dodge spike traps and outduel palace guards better than they did a year ago is a fun question. Whether it tells us anything a rigorous benchmark doesn't is a separate one.