KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding
KeyRec creates bounded visual memory via recent caches and structured event banks to enable efficient long-video and streaming understanding with only 10% of visual tokens, outperforming compressed baselines.
Published Sep 26, 2026▲ 11 on Hugging FacearXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
KeyRec delivers exceptional compressed long-video and streaming accuracy at a tenth of dense tokens through its structured event bank and adaptive router, though its precise gains over simpler bounded caches remain unisolated and its training-free budget…
Abstract
Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.