Unifying Video Tasks via Spatiotemporal Analogy
ViGeo extends visual analogy to video via spatiotemporal canvas completion to unify diverse tasks, generalize to unseen manipulations and zero-shot modalities, and reveals task internalization shortcuts removable via unrelated data.
Published Sep 27, 2026arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.