Grounded-Exo2Ego couples geometric anchoring with semantic grounding and camera relocalization to robustly generate egocentric video from exocentric inputs, outperforming prior methods on EgoExo4D.
RePlaid, a continuous diffusion language model aligned with modern discrete architectures, achieves scaling laws rivaling discrete diffusion and sets a continuous diffusion perplexity record of 22.1 on OpenWebText.
DiLaDiff proposes a latent-augmented masked diffusion language model with consistency distillation that improves quality and accelerates inference by generating continuous latents in negligible time.
World from Motion generates dynamic 3D Gaussian reconstructions from monocular video via generative video modeling to fix artifacts and fill missing regions, achieving state-of-the-art 4D reconstruction.
Spatio-Temporal Attention Chains accelerate training-free 4D mesh generation 13x to 9 seconds via latent temporal correspondences, improving quality, scaling to longer videos, and enabling tracking and camera estimation.