VidVec extracts intermediate-layer MLLM embeddings for video-text retrieval via calibration and text-only alignment, achieving state-of-the-art zero-shot results without video fine-tuning.
Spatio-Temporal Attention Chains accelerate training-free 4D mesh generation 13x to 9 seconds via latent temporal correspondences, improving quality, scaling to longer videos, and enabling tracking and camera estimation.