RotVLA introduces continuous rotational latent actions on SO(n) for vision-language-action pretraining, using a flow-matching head guided by latent planning to achieve state-of-the-art robot control.
PriorVLA freezes a prior expert and trains an adaptation expert via expert queries to preserve pretrained vision-language-action priors, updating only 25% of full fine-tuning parameters while outperforming baselines on OOD and few-shot robot manipulation.
FLARE introduces a long-video audiovisual retrieval benchmark with simulated user queries, revealing caption-based performance fails to transfer and audio-language alignment remains a bottleneck.
TrioPose uses a triple-stream pose-aware DiT with relational bias masks and spatial loss weighting to generate accurate multi-person images, improving Human-Art AP by 30%.
Standard video backbone readouts suppress patch-level temporal dynamics needed to detect AI-generated videos; a lightweight velocity-gated patch profiling readout reaches 95.28 AUC on frozen backbones.
XDecomposer learns prior-free multiphase X-ray diffraction decomposition as set prediction to identify constituent phases and proportions without candidate lists. It improves reconstruction accuracy and phase identification across simulated and experimental datasets.
K12-KGraph introduces a curriculum-aligned K-12 knowledge graph, benchmark, and training data showing current LLMs achieve under 57 percent accuracy on curriculum cognition and that graph-guided supervision outperforms generic instruction tuning.