MIRAGE learns continuous latent reasoning for mobile agents, cutting decoded tokens 75% while matching explicit chain-of-thought accuracy and improving baselines up to 10.2 points via generative world modeling.
IGGT4D is a streaming transformer that incrementally reconstructs long dynamic 4D scenes with consistent geometry and instance tracking from video, surpassing existing online baselines.
SpatialBench evaluates 41 spatial foundation models across 19 datasets and finds none are all-round players, with full-context attention maximizing accuracy and domain alignment exceeding scaling for embodied tasks, plus it introduces DA-Next-5M and DA-Next.