LCVN introduces a language-conditioned navigation benchmark and compares diffusion-based latent imagination against unified autoregressive prediction for embodied agents. Latent imagination yields more temporally coherent rollouts, while unified prediction generalizes better to unseen environments.