Interactive video world models lose visual persistence beyond training horizons because temporal RoPE offsets become out-of-distribution; WorldTrace assigns compressed memory slots virtual in-distribution positions to restore addressability, boosting temporal consistency by 15.5% and episodic recall
ProxyPose recasts 6-DoF pose tracking as video-to-video translation using a diffusion model to generate proxy videos for classical pose estimation, achieving state-of-the-art accuracy without 3D models or masks.
DéjàView loops a single transformer block for iterative multi-view 3D reconstruction, matching larger feed-forward models with far fewer parameters while treating refinement steps as an inference-time compute knob.
A neuro-symbolic framework pairs neural graph proposals with symbolic SMT solvers for hard-constraint satisfaction, achieving over 95% in-distribution and 64, 86% zero-shot rule compliance on the MolSAT benchmark.