Good Papers

Unifying Video Tasks via Spatiotemporal Analogy

ViGeo extends visual analogy to video via spatiotemporal canvas completion to unify diverse tasks, generalize to unseen manipulations and zero-shot modalities, and reveals task internalization shortcuts removable via unrelated data.

Chia-Hsiang Kao, Belinda Zeng, Bharath Hariharan, Menglin Jia

Published Sep 27, 2026arXiv ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the score
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.