Regularized LNS turns local search heuristics into MCMC samplers with Fenchel-Young losses, enabling exact block Gibbs sampling and end-to-end learning without global solvers.
TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.
MOOD benchmark shows guard models fail to detect out-of-distribution alignment failures, but combining them with Mahalanobis and perplexity detectors improves recall from 39% to 45% and scales positively.
Masked Visual Actions expresses robot and object motion as revealed pixel trajectories to unify forward dynamics, planning, and inverse modeling in video world models with minimal finetuning.
AgentOWL jointly learns hierarchical neural options and an abstract world model for sample-efficient skill acquisition, outperforming baselines on object-centric Atari games with fewer samples and stronger generalization.