RotVLA introduces continuous rotational latent actions on SO(n) for vision-language-action pretraining, using a flow-matching head guided by latent planning to achieve state-of-the-art robot control.
Memory Grafting uses frozen hidden states from a grafting model as conditional n-gram memory for language models, improving average benchmarks to 53.86 versus 52.43 for vanilla Engram at 2.8B scale with minimal overhead.
PointForward reconstructs driving scenes via world-space 3D queries and scene graphs, achieving state-of-the-art feedforward results with explicit cross-view and instance consistency.
LaST-VLA replaces explicit chain-of-thought reasoning with a physically grounded latent spatio-temporal framework, achieving record NAVSIM scores and improved reasoning.
CoDMD adds a copula-aware relational regularizer to distribution matching distillation that improves few-step video generation, achieving 84.46/84.87 VBench scores at 4 steps with ~25× speedup.
CausalMix frames data mixture optimization as causal inference to dynamically estimate optimal mixtures via conditional average treatment effects, improving LLM performance without retraining proxy models.