PriorVLA freezes a prior expert and trains an adaptation expert via expert queries to preserve pretrained vision-language-action priors, updating only 25% of full fine-tuning parameters while outperforming baselines on OOD and few-shot robot manipulation.
SliceWorld is a predictive CT world-state model that encodes slice sequences into factor-aware latent states for multi-step future prediction, lesion intervention, and report generation, improving generation and clinical evaluation metrics.
OneVision-Encoder applies codec-aligned sparsity to video, processing only high-entropy regions to outperform dense backbones with fewer tokens. It achieves 4.1% higher video accuracy than Qwen3-ViT across 16 benchmarks.
VLRS-Bench introduces a remote sensing vision-language reasoning benchmark spanning cognition, decision, and prediction tasks that exposes major bottlenecks in current multimodal models.
Hydra-X unifies image and video tokenization in one vision transformer via causal temporal attention and hierarchical compression, achieving strong unified understanding and generation performance.
DiffDiff embeds predictability asymmetry into diffusion trajectories so the process focuses generative effort on uncertain high-frequency dynamics while preserving history-anchored content, outperforming diffusion baselines.