A neural framework learns pole-residue modal decomposition from system observables alone, generalizing to unseen port counts and recovering physical eigenmodes without modal supervision.
TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.
Self-supervised goal-reaching enables multi-agent cooperation and exploration via sparse feedback, outperforming alternatives and discovering nontrivial coordination without explicit mechanisms.
AI GameStore proposes evaluating general intelligence via scalable synthesis of human games, finding frontier vision-language models score under 10% of human averages on most generated games.
Dithered randomized Hadamard quantization is unbiased and achieves mean squared error asymptotically matching dense random rotations at O(d log d) cost.
CoMet decomposes multimodal LLM uncertainty into context and multiplicity terms via a lightweight module, improving calibration without generation or sampling.
Attention-only transformers under token corruption implement two-stage in-context empirical Bayes via depth-refined particle dynamics and skip-connection queries, enabling depth-dependent denoising without explicit noise schedules and posterior-mean convergence to Bayes-optimal predictors.
Interactive video world models lose visual persistence beyond training horizons because temporal RoPE offsets become out-of-distribution; WorldTrace assigns compressed memory slots virtual in-distribution positions to restore addressability, boosting temporal consistency by 15.5% and episodic recall
QUTCC trains a U-Net for spatially adaptive quantile regression and calibrates tighter pixel-valid uncertainty intervals via non-linear conformal scaling for imaging inverse problems.
P³ plans programs and proofs jointly from specifications before elaboration, outperforming sequential baselines by up to 11.2 points on verified generation benchmarks while reducing cost and time.
Threshold-based algorithms achieve constant class envy-freeness and exceed 1/2 utilitarian welfare in online class matching, with near-matching upper bounds characterizing fairness costs.
Fine-tuning breaks safety via unstable geometric alignment subspaces, with alignment loss scaling quartically in training time via curvature-driven drift.
An inference pipeline using off-the-shelf models and conjecture extraction with context detachment achieves state-of-the-art IMO-style math performance at much lower cost by escaping the Cognitive Well.
TTCD uses a long-window teacher to supervise a short-window student's fast weights via hidden-state discrepancy, allocating limited memory to future-relevant context and outperforming existing long-context methods with minimal architectural changes.
LensVLM lets VLMs scan compressed rendered text and selectively expand only relevant regions via learned tools, maintaining near-full accuracy at 4.3x compression and outperforming baselines up to 10.1x across text QA benchmarks.
DiscoverPhysics benchmarks LLM agents on simulated worlds with non-standard physics, finding frontier models pass only half and fail at uncovering latent structure.
Pointing via text induces internal visual search routines that eliminate binding errors, enabling compositional generalization and solving vision-language binding via serial processing.
Privileged self-distillation degrades thinking models by suppressing reasoning forks and self-correction tokens, reducing long-rollout accuracy by up to 17%.