DD-Ranking reveals dataset distillation gains come from extra evaluation techniques rather than image quality, proposing fair metrics to assess true synthetic dataset value.
DeformMaster learns interactive physics-neural world models of deformable objects from videos, enabling high-fidelity future dynamics rollout, material variation, and novel-view synthesis.
SRL-MPC integrates reinforcement-learned parameter updates with shape-aware model predictive control via geometric separation features to navigate dense heterogeneous robot crowds safely and adaptively.
VIGIL decouples world-state completion from terminal commitment in embodied agents, revealing that comparable execution yields up to 19.7 pp differences in correct episode termination.
DAWN extracts token dependency graphs to select reliable unmasking positions, accelerating diffusion LLM inference by 1.80-8.06x with negligible quality loss.
OQRC partitions calibration losses into ordered bins to tightly upper-bound quantile risk with finite-sample guarantees converging at rate O_p(n^{-1/2}).
SAM 3D Animal is a promptable framework using the SMAL+ model and Herd3D dataset to reconstruct multiple 3D animals from single images with keypoint and mask prompts, achieving state-of-the-art results.
HDR integrates hierarchical tree-structured latents into causal video generation to enable coarse-to-fine multi-step visual reasoning with sparse attention, boosting reasoning success by 76% over streaming diffusion while running 54x faster than bidirectional diffusion.
SPACE defines pivot-aligned coordinate-free embeddings and adaptive decoding to unify neural routing across symmetric and asymmetric VRPs, achieving strong zero-shot generalization on 110 variants.