A new technique lets AI models watching live video streams predict what's about to happen, then adjust how hard they work in real time - without any retraining.
The system, called FORESIGHT, pairs two identical large language models that share weights, input encoders, and a memory cache. One model processes the incoming video feed continuously. The other runs ahead of it, anticipating future events and deciding when the main model should reason more carefully, what it should check for, and how densely it should sample frames. A lightweight reconfiguration protocol applies those plans to the live model without interrupting the stream, and the whole thing runs on a frozen Qwen3-VL-8B backbone with zero extra training.
That training-free framing matters because it beat the strongest trained baseline on the OmniPro Online benchmark by 9.5%, and posted its largest gain - 18.7 points - specifically when key evidence showed up late in a video. That's the exact situation where today's fixed-computation streaming models fall apart, since they can't scale up attention once something actually becomes important.
It's a smart patch for a real architectural limit, but it's still one backbone, one set of benchmarks, and a code release that hasn't happened yet.