A new preprint shows exactly when AI video generators start learning to track objects across frames - and it turns out only a tiny sliver of the network does the work.
In an arXiv preprint (arXiv:2609.31654v2, posted September 30, 2026), researchers ran a checkpoint-by-checkpoint census of every temporal-attention head in nine training runs of Open-Sora's STDiT video diffusion model, spanning three sizes from 306 million to 1.03 billion parameters. Rather than averaging attention patterns across the whole network - a method that can hide small but important signals - they scored each individual head with a new entropy-based metric for cross-frame attention concentration. The result: only about 4 to 13 percent of heads develop strong frame-tracking behavior, and those heads consistently show up in the first temporal block of the network rather than scattered randomly throughout. Across different training runs, these heads converge on the same small set of tricks: attending to a single frame, or scanning a narrow band of neighboring frames.
That's useful for anyone building or debugging video-generation models: it suggests the mechanism that keeps a video's frames consistent over time is not spread evenly through the network, but concentrated in a handful of predictable spots. The catch, which the researchers flag themselves, is that they only found a correlation - their ablation tests did not prove these heads actually cause better video quality.
It is a reminder that most interpretability work on AI models happens after training is done, on finished checkpoints. Watching the wiring specialize while training is still running is rarer, and rarer still to find the effect this concentrated in so few components.