PriorVLA freezes a prior expert and trains an adaptation expert via expert queries to preserve pretrained vision-language-action priors, updating only 25% of full fine-tuning parameters while outperforming baselines on OOD and few-shot robot manipulation.
ChainFlow-VLA unifies causal trajectory generation and global diffusion refinement via vision-language-conditioned residual distributions, scoring 94.85 on NAVSIM v1.
WorldDrive unifies vision and motion representations to couple scene generation with planning, achieving leading vision-only autonomous driving performance with high-fidelity future video generation.
Realtime-VLA FLASH uses a draft model and parallel verification to replace most full diffusion-based VLA inference rounds with faster speculative ones, cutting average latency 3.04x to 19.1 ms.
CoWorld-VLA embeds multi-expert world tokens into vision-language-action models and couples diffusion planning with scene context to generate continuous ego trajectories, improving autonomous driving performance.