Self-supervised video-stream pretraining fails due to intra-batch near-duplicate frames, but proposed StreamMAE with motion-biased crops matches i.i.d. MAE and scales to 95 hours.
Self-supervised training on 160,000 in-the-wild videos via a shared coarse mesh yields emergent canonical object frames without pose labels, matching supervised category-level pose estimation accuracy.