A team of researchers has built a world model that teaches groups of AI agents to coordinate by predicting abstract representations of what will happen next, not pixels.
The system, called MA-JEPA, is designed for multi-agent reinforcement learning - the kind of training used when several AI agents need to act together, like a group of game units or robots. Instead of reconstructing raw observations, it uses a self-supervised technique called joint-embedding prediction to guess the next useful representation of the world. A categorical latent state and a causal Transformer handle the prediction, and a training-only module that sees every agent's state and actions feeds those predictions into each agent's local model, even though each agent still acts independently once training ends. The researchers tested it on SMAC, a standard benchmark for multi-agent StarCraft II micromanagement, where it matched or beat the best previously reported results on four of eight maps.
That four-out-of-eight split matters more than the headline number suggests. It signals that representation-based world models, which have already reshaped single-agent reinforcement learning, are starting to migrate into coordination problems where agents must also guess what their teammates are doing - a much harder modeling target than predicting your own next move.
Half a benchmark is not a breakthrough, but it is a serviceable proof that skipping pixel reconstruction can hold up even when the thing being predicted is a moving target shaped by other learning agents.