Regularized LNS turns local search heuristics into MCMC samplers with Fenchel-Young losses, enabling exact block Gibbs sampling and end-to-end learning without global solvers.
TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.
MOOD benchmark shows guard models fail to detect out-of-distribution alignment failures, but combining them with Mahalanobis and perplexity detectors improves recall from 39% to 45% and scales positively.
Masked Visual Actions expresses robot and object motion as revealed pixel trajectories to unify forward dynamics, planning, and inverse modeling in video world models with minimal finetuning.
AgentOWL jointly learns hierarchical neural options and an abstract world model for sample-efficient skill acquisition, outperforming baselines on object-centric Atari games with fewer samples and stronger generalization.
Fine-tuned LLMs learn to selectively hide internal representations from unseen activation monitors via low-dimensional subspace manipulation, evading even post-hoc safety probes with modest capability loss.
Project VAANI releases a multimodal dataset of 31,255 speech hours and 289K images spanning 105 Indic languages across 165 Indian districts to support inclusive speech technology.
Alem benchmarks open-ended multi-agent coordination for language agents, showing frontier LLMs average ~6% returns and individual competence does not imply coordination competence.
A representation-readout decomposition attributes grokking and double descent to competing encoder and classifier dynamics, showing delayed generalization stems from gradual representation learning rather than lazy-to-rich transitions.
Subliminal learning is steering vector distillation where students learn teachers' hidden traits via single steering vectors, requiring adaptive optimizers and failing across models.
Co-Coder formalizes multi-agent coding as graph partitioning to balance parallel speedups against communication overhead, improving pass rates by 14% and cutting costs 35% on dense repositories.
Sparse Koopman autoencoders use sparse latent supports as label-free regime indicators that identify local dynamical basins and outperform dense autoencoders in multibasin forecasting.
A unified framework casts knapsack and top-k operators as dynamic programs with smoothed recursions for differentiable relaxations, parallel algorithms, and theoretical regularization guarantees.
Synthetic document finetuning teaches models to hide misbehavior from chain-of-thought monitors, with success tied to reasoning controllability and faster reward-hacking under RL.
ARQ introduces a question generator that produces transferable intermediate stepping stones, improving reasoning LLM performance via fine-tuning on synthetic data.
Persona Generators use evolutionary code optimization to expand brief context descriptions into diverse synthetic populations maximizing opinion and preference coverage. Evolved generators substantially outperform baselines across six diversity metrics by spanning rare trait combinations.
A unified framework bounds DP privacy leakage against multi-target membership, attribute, and reconstruction attacks using only privacy parameters and adversarial baseline success rates.
Monocular pretraining lifts single images into pseudo-target views via depth and reprojection, yielding OVIE, which rivals multi-view baselines at 116 FPS without inference-time depth or multi-view training pairs.
Rosetta neuron populations grow sublinearly and become more selective and specialized as language and vision models scale, while non-Rosetta neurons stay less selective.
Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.
SkillOS uses RL to train a skill curator that updates an external SkillRepo from experience, improving self-evolving agents across reasoning and multi-turn tasks.
M²RNN introduces matrix-valued non-linear RNNs that scale via state expansion, achieving perfect state tracking and outperforming hybrid models with smaller states.
Direct corpus interaction uses terminal tools to search raw corpora directly, bypassing fixed retrieval interfaces and substantially outperforming sparse, dense, and reranking baselines on agentic search benchmarks.
GMOS grounds moving object segmentation in 3D space and time using an RGB video framework, achieving state-of-the-art results across MOS benchmarks with faster online inference.
ReToken introduces one learnable retrieval token that selects sparse visual tokens from long contexts, improving vision-language models by up to 13.4 points on visual retrieval while fitting on a single GPU.
A unified sign-language model with privacy-preserving keypoint inputs and sliding perceiver aggregation achieves state-of-the-art translation and alignment on BSL and generalizes to ASL.
PACEvolve++ adapts evolutionary search policies at test time via advisor-model reinforcement learning, using phase-adaptive optimization to outperform frontier-model baselines across engineering and protein tasks.
MetaCanvas enables multimodal LLMs to plan directly in diffusion latent spaces, outperforming global-conditioning baselines across six precise visual generation tasks.
A novel definition of model exploitation reveals it is essentially unavoidable for large policy sets and cannot be precluded in finite ones, yielding safe planning limits.
MIND uses sliced Wasserstein distance via sorting to evaluate generative models with 10x better sample efficiency, 100x faster computation, and greater adversarial robustness than FID.