Nvidia says it can generate photorealistic driving footage in real time, frame by frame, based on what a virtual car does next.
The company built OmniDreams by mid- and post-training its Cosmos diffusion model on 21,000 hours of driving data. The system autoregressively generates action-conditioned video: it looks at past frames, the current simulator state, and the car's next move, then renders what the sensors would see. Nvidia plugged it into a closed-loop setup with its Alpamayo 1 driving policy and an orchestrator called AlpaSim, so the simulated car's actions directly shape what it "sees" next. The pitch is that OmniDreams can conjure scenarios - extreme weather, erratic pedestrians - that reconstruction-based simulators never captured because they were never filmed.
That matters because the hardest part of testing self-driving software isn't the common cases, it's the long tail: the one-in-a-million scenario that never showed up in recorded data. A generative world model that can synthesize those situations, rather than replay them, is a real shift from stitching together captured sensor logs. Nvidia also reports a world-action model built on OmniDreams beat its own Alpamayo 1.5 research policy on a benchmark dataset while using a fifth of the parameters, suggesting the world model itself might double as a policy backbone.
Still, this is a preprint from Nvidia about Nvidia's own stack, tested against Nvidia's own benchmark and prior model - independent verification is the part missing from the abstract.