VIGIL decouples world-state completion from terminal commitment in embodied agents, revealing that comparable execution yields up to 19.7 pp differences in correct episode termination.
SciHazard benchmarks LLM scientific safety risks via decomposed harm scoring across 3,000 real-world grounded queries, finding deep research agents 32.3% more harmful than standard models.
RankE co-evolves discrete text-to-image policy and decoder via alternating optimization to eliminate latent covariate shift, improving both FID and CLIP scores.
VisInteract introduces interactive text-to-visualization with imperfect queries via VisInteract-Bench and Vis-MCTS, boosting success by over 13% versus interactive baselines.
Q-ARVD quantizes autoregressive video diffusion via frame-weighted objectives and adaptive outlier isolation, cutting inference costs while preserving generation quality.
FeatCal reduces post-merging feature drift via layer-wise closed-form weight calibration without gradients, outperforming Surgery and ProbSurgery on CLIP and GLUE benchmarks.
ACT introduces a plug-and-play block that learns adaptive coordinate transforms for neural operators, reducing spatial misalignment and significantly improving predictive accuracy across PDE benchmarks.
SpatialBench evaluates 41 spatial foundation models across 19 datasets and finds none are all-round players, with full-context attention maximizing accuracy and domain alignment exceeding scaling for embodied tasks, plus it introduces DA-Next-5M and DA-Next.
Adversarial World Modeling casts robust planner learning as a constrained min-max game solved via decoupled self-play, yielding competitive closed-loop performance in nominal and adversarial traffic scenarios.