AI video generation just got a speed boost from a smarter way of deciding which pixels actually need to talk to each other.
Researchers built a technique called Parameterized Stripe Attention, or PSA, that speeds up Diffusion Transformers, the models behind tools like HunyuanVideo and Wan 2.1. The slowdown in video generation comes from full spatio-temporal attention: every frame's pixels comparing themselves against every other frame's pixels, an expensive process. The team found this attention naturally forms repeating diagonal stripe patterns across time and space, then built a single CUDA kernel that exploits that structure instead of guessing at sparse patterns on the fly. A training-free search algorithm tunes how aggressively to trim each attention head without exceeding a set error tolerance.
This matters because earlier sparse-attention methods forced a trade-off. Fixed masks ran fast but couldn't adapt; flexible masks adapted but ran slow. PSA claims to dodge that trade-off, delivering up to 1.57x speedup over FlashAttention-3 on HunyuanVideo and 1.37x on Wan 2.1, with what the researchers call "acceptable" quality loss.
"Acceptable" is doing a lot of work in that sentence, and this year every efficiency paper benchmarks against FlashAttention-3. The real test is whether video model builders actually bolt this kernel on, not whether it beats a baseline in a paper.