An AI can now decide, question by question, whether to answer immediately or rewind to a moment it already watched.
Researchers built a system called Watch-Think-Interact (WTI) that processes streaming video and, for each question, checks whether its current view and compressed notes are enough to answer. If not, it either keeps watching or jumps back to a specific stretch of earlier footage tagged with a timestamp, then answers. The team trained this behavior with a new 82,335-question dataset spanning 4,812 video sequences and a reinforcement-learning method called Stream-GDPO that rewards good timing, accurate recall, and useful memory updates across a full multi-question session. On two standard benchmarks, WTI scored 83.3% on StreamingBench and 73.6% weighted accuracy on OVO-Bench, topping other open-source streaming systems tested.
The core problem with streaming video AI has always been foresight: a model has to decide what's worth remembering before it knows what will be asked later, so compressed summaries quietly drop details that turn out to matter. WTI's answer is to stop treating memory as a one-shot compression problem and instead let the model go back and re-watch the original footage when its notes fall short. That's a meaningfully different design than systems that just try to write a better running summary.
Still, 'state of the art' here means beating other open-source streaming systems on benchmarks built by the same research community, not necessarily outperforming proprietary systems that never get compared openly.