DiMP applies diffusion modeling to masked tube-center inference and inter-frame motion prediction, eliminating positional leakage and deterministic trajectory collapse to improve dynamic point cloud pretraining.
ExtraVAR proposes stage-aware RoPE remapping and entropy-driven attention calibration to eliminate repetition and detail degradation when extrapolating VAR models to higher resolutions without retraining.
Ontological trust measures whether trajectory prefixes match authorized tasks; RGE detects long-horizon agent drift with over 93% F1 and above 95.8% benign coverage via deterministic Role, Goal, and Evidence checks.
MedHorizon benchmarks long medical video understanding via sparse evidence retrieval and multi-hop reasoning, with top models reaching only 41.1% accuracy.
ElegantVLA accelerates vision-language-action models via adaptive compute scheduling, achieving up to 3.77x speedup and doubling control frequency without retraining.
CoDMD adds a copula-aware relational regularizer to distribution matching distillation that improves few-step video generation, achieving 84.46/84.87 VBench scores at 4 steps with ~25× speedup.
IGGT4D is a streaming transformer that incrementally reconstructs long dynamic 4D scenes with consistent geometry and instance tracking from video, surpassing existing online baselines.
Attention transfer fails for four ViT families due to architectural mismatch, and adding the teacher's native components to students fully restores its effectiveness.
SpatialBench evaluates 41 spatial foundation models across 19 datasets and finds none are all-round players, with full-context attention maximizing accuracy and domain alignment exceeding scaling for embodied tasks, plus it introduces DA-Next-5M and DA-Next.
DR-Smoothing disrupts and rectifies LLM prompts via smoothed defense to guarantee jailbreaking protection while balancing harmlessness and helpfulness.