Two AI training methods just fought to a draw in a chess variant most people have never heard of.
Researchers re-implemented the engine for Dragonchess, a three-dimensional chess game, rewriting it from Python's PyGame into C++. The switch made the simulation fast enough to run 10,000 games with confidence intervals and significance tests, instead of the single small tournament that this kind of research usually settles for. They then pitted two adaptive approaches, evolutionary transfer learning and TD(lambda), a learning algorithm that traces back to 1990s backgammon research, against a field of other game-playing agents in a round-robin. Both adaptive methods beat every other agent in the tournament.
The more interesting result is what didn't happen: evolution and learning produced statistically indistinguishable performance. That implies the hard part of building AI for complex, non-standard games is adapting to the environment at all, not picking the right adaptation mechanism. For anyone designing evaluation functions for messy, computationally heavy game domains, that is a useful, if modest, data point.
Dragonchess will never get the research budget chess or Go enjoy. The real achievement here might just be the plumbing: proving that rigorous, large-sample testing is possible in an obscure game, not the specific numbers it produced.