LangMap introduces human-verified hierarchical open-vocabulary navigation benchmarks across scene, room, region, and instance levels with 18K tasks, and PlaNaVid achieves top RGB-only success via planning and memory.
ReFPO adds explicit reflow regularization to flow matching policy gradients, stabilizing training and enabling high-fidelity one-step inference that matches multi-step performance across control tasks.
DiffICL frames tabular synthesis as in-context learning using pretrained structural priors to avoid memorization, improving both quality and privacy in small-data settings.
Depth2Pose evaluates monocular depth estimators via downstream pose estimation accuracy using only camera poses, introducing a challenging out-of-distribution benchmark where standard metrics fail to predict generalization.
TreeGraft combines small and large drafters with a scheduler to build shared draft trees, boosting speculative decoding by 15.1% over single-drafter methods.
Platonic Representation Defense detects and purifies backdoored self-supervised encoder representations via cross-model energy functions without labels or training data. It substantially improves robustness across multiple encoders and over ten attacks in fully black-box settings.
A language-assisted clustering framework uses cross-modal relational signals and adaptive semantic centers to improve clustering accuracy by 2.6% over state-of-the-art methods.
GTP-FA decouples grasping from planning with failure attribution to diagnose failures and optimize both modules, substantially improving robotic manipulation success across diverse policy learners.
SoftGAC proposes a soft generative actor-critic with stochastic bridge policies that expose a tractable MaxEnt objective via single-pass sampling, outperforming diffusion and flow baselines on continuous control with lower latency.
PULSE uses sparse autoencoders to identify internal features linked to demonstration utility and improves selection across classification, generation, and reasoning tasks.
SpatialBench evaluates 41 spatial foundation models across 19 datasets and finds none are all-round players, with full-context attention maximizing accuracy and domain alignment exceeding scaling for embodied tasks, plus it introduces DA-Next-5M and DA-Next.
NSFT decomposes MoE experts into channel groups to enable sub-expert-level parameter-efficient fine-tuning that outperforms expert-level methods with fewer trainable parameters.