Gekko uses relative reconstruction error between cross-view and masked-autoencoder predictions as a self-supervised co-visibility proxy, adding binocular training signals that consistently outperform CroCo on 3D vision tasks while training directly from raw video.
DiMP applies diffusion modeling to masked tube-center inference and inter-frame motion prediction, eliminating positional leakage and deterministic trajectory collapse to improve dynamic point cloud pretraining.
RATS decomposes vision transformers' classification token into learnable register tokens that spontaneously specialize into object parts, improving segmentation by up to 12 mIoU.
Self-supervised video-stream pretraining fails due to intra-batch near-duplicate frames, but proposed StreamMAE with motion-biased crops matches i.i.d. MAE and scales to 95 hours.
A self-supervised encoder learns subject-specific fMRI embeddings from repeated brain responses, and unsupervised orthogonal rotations align them across subjects into a shared geometry, demonstrating approximately isometric cross-subject visual representations.
SPHERE-JEPA proves hyperspherical uniformity minimizes downstream prediction risk on manifolds and improves SSL retrieval and ImageNet linear probing over LeJEPA.
CanViT introduces the first active-vision foundation model with a retinotopic backbone and scene-wide canvas, achieving 38.5% ADE20K mIoU with one glimpse and 84.5% ImageNet accuracy.