[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-squeeze-long-video-ai-memory-into-256-kib":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},7351,"researchers-squeeze-long-video-ai-memory-into-256-kib","Researchers Squeeze Long-Video AI Memory Into 256 KiB","A new method called PREM compresses long video into a compact memory once, letting AI models answer many questions later without extra prompt tokens.","A new memory scheme lets AI models watch long videos without blowing through their token budget.\n\nResearchers introduce Prefix-Steered Recurrent Memory, or PREM, a framework built for frozen vision-language models. It splits video processing into two stages: a recurrent writer compresses the video stream into a compact 256 KiB multi-slot memory, then a question-conditioned readout steers the model's existing key-value cache during prefill instead of appending new tokens. That write-once, query-many design was tested across six long-video benchmarks in both offline and streaming settings. Under a tight 16-frame budget, it lifted macro-average accuracy by 3.06% on Qwen2.5-VL-3B, with an 11.0% jump on action antonym identification and 9.9% on localized needle retrieval, while tuning only 0.24% of the model's parameters and adding just 0.03 GiB of peak GPU memory.\n\nMost long-video fixes so far have meant compressing frames, bolting on memory tokens, or rewriting the model's internal KV cache directly - approaches that either throw away detail or make every new question expensive. PREM's split between ingesting the video once and querying it repeatedly is the more interesting move here: the same compact memory state can field many different questions without re-reading the video or lengthening the prompt each time, which matters for anything latency-sensitive, like video search or streaming review tools.\n\nThe catch is that all of this is demonstrated on one 3-billion-parameter model. Whether the gains survive at larger scale, or against the closed video-understanding systems built by bigger labs, is still an open question.","[\"ai\",\"video-understanding\",\"vision-language-models\",\"research\"]","2026-09-23T04:00:00.000Z","2026-09-23T08:58:21.362Z","2026-09-23T08:58:27.698Z","published",null,[],"ai",[24,26,27,28],"video-understanding","vision-language-models","research",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.23601",0,{"sections":35},[36,39,43,48,53,57,61,66,71,76,81,86,91,96],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",4297,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",710,{"name":44,"slug":45,"count":46,"latest_published_at":47},"Policy","policy",369,"2026-09-23T02:13:52.000Z",{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",202,"2026-09-22T23:00:04.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":18},"Hardware","hardware",169,{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",133,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",110,"2026-09-22T20:00:00.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Software","software",80,"2026-09-22T23:32:52.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",65,"2026-09-22T22:06:48.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]