AI researchers built a way for models to skim long recordings instead of swallowing them whole.
A new framework called LEAP tackles a real bottleneck in audio-visual AI: feeding an hour of video into a model either blows past its context limit or forces heavy compression that erases small but crucial details. LEAP splits a recording into fixed-length blocks, scans each one cheaply to flag the most promising short windows, then re-encodes only those windows for the final answer. The localization pass works off pre-computed transcripts rather than raw video frames, keeping it fast, while the answer pass pulls from the actual audio-visual stream so fine details survive. The researchers trained the system in two stages (one to pick better windows, one to answer better from them) and tested it across two different omni-modal AI backbones, Qwen3-Omni-30B-A3B and MiniCPM-o 4.5.
The result is a model whose context usage stays flat no matter how long a recording runs, which matters once you are dealing with hour-scale audio-visual content rather than short clips. LEAP outperformed the Qwen baseline by 4.5 to 16.8 percent and beat MiniCPM-o 4.5's published numbers by 3.1 to 13 percent on audio-visual question-answering benchmarks, without needing retraining for streaming use.
It's less a smarter model than a better search engine bolted onto one: the gains come from knowing where to look, not from reasoning harder once it gets there.