DiMP applies diffusion modeling to masked tube-center inference and inter-frame motion prediction, eliminating positional leakage and deterministic trajectory collapse to improve dynamic point cloud pretraining.
LangMap introduces human-verified hierarchical open-vocabulary navigation benchmarks across scene, room, region, and instance levels with 18K tasks, and PlaNaVid achieves top RGB-only success via planning and memory.