ChainFlow-VLA unifies causal trajectory generation and global diffusion refinement via vision-language-conditioned residual distributions, scoring 94.85 on NAVSIM v1.
WorldDrive unifies vision and motion representations to couple scene generation with planning, achieving leading vision-only autonomous driving performance with high-fidelity future video generation.
VSearcher uses reinforcement learning to turn static multimodal models into long-horizon web search agents that surpass proprietary models on multimodal search benchmarks.
LaST-VLA replaces explicit chain-of-thought reasoning with a physically grounded latent spatio-temporal framework, achieving record NAVSIM scores and improved reasoning.
Unify-Agent reframes image synthesis as an agent pipeline with search and recaptioning, improving generation of long-tail factual concepts via 143K curated trajectories.