Good Papers

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

FrameMorrow guides historical frame selection via prospective tokens representing future needs, improving consistency and quality across diverse long-horizon video generators.

Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan

Published Sep 30, 2026▲ 92 on Hugging FaceCode ★ 26arXiv ↗

88%
OverallMust read
Readers
100%1 of 1 upvoted

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
Panel consensus
FrameMorrow offers an elegant, diffusion-agnostic bet that future-guided historical selection improves long-horizon coherence, though its prospective tokens remain an empirically defined, unmeasured mechanism whose scaling and true cost are still unproven.

Abstract

Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often determine historical relevance based on the current content. However, information relevant to the present is not necessarily useful for future generation, while seemingly less relevant history may become important later. Our key insight is that historical information should be selected according to its relevance to future information needs. Capturing these needs does not require generating the full future; instead, a compact representation of what becomes important next is sufficient to guide historical selection. Building on this insight, we propose FrameMorrow, a prospective frame selector that predicts a small set of prospective tokens representing future information needs and uses them to identify relevant information from history. FrameMorrow selects explicit historical frames rather than model-specific internal states, enabling plug-and-play integration across diverse generators, including closed-source models, with little additional inference cost. We evaluate FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models. Extensive experiments demonstrate consistent improvements in long-range consistency, visual quality, and action alignment across diverse generation settings.