SetDiff integrates explicit 3D geometry priors into a set-based diffusion model to improve novel-view synthesis from 3D Gaussian Splatting, reducing hallucinations and achieving state-of-the-art results.
Conditional Information Bottleneck frames reasoning as lossy compression with a semantic surprisal prior, improving LLM reasoning efficiency with minimal accuracy loss.
MAPLE trains vision-language-action driving models via latent multi-agent rollout and reinforcement learning, achieving state-of-the-art closed-loop performance without external simulators.
GeRo enables vision-language-action models to generate language-grounded future traffic scenes via autoregressive rollouts, improving Bench2Drive driving scores by 15.7 and success rates by 26.2.