Network of Theseus progressively replaces guide network modules with a different target architecture via representational alignment, preserving performance across vastly different deployed architectures.
Specificity-aware diffusion steering uses variance-reduced sequential Monte Carlo to suppress undesired regions with minimal positive distribution distortion.
Matching full posterior covariance in Gaussian DDPMs reduces path-KL error to O(1/T²), and the matrix-free Lanczos Gaussian sampler achieves this with exponentially decaying approximation error using only Jacobian-vector products.
AI GameStore proposes evaluating general intelligence via scalable synthesis of human games, finding frontier vision-language models score under 10% of human averages on most generated games.
GLACIER treats tandem mass spectrum prediction as graph object detection, outperforming prior state-of-the-art by up to 19.3% on retrieval accuracy with nearly 8-fold faster inference.
Fine-tuning language models on interpretability ground truth teaches them to describe their internal computations, with self-explanation outperforming larger external explainers.
Monoculture evaluation depends on subjective null-model choices and evaluated model populations, making model agreement a context-dependent inference rather than an absolute property.
CTM-AI combines a consciousness model with foundation models to integrate diverse processors, achieving state-of-the-art results on multiple benchmarks.
Dithered randomized Hadamard quantization is unbiased and achieves mean squared error asymptotically matching dense random rotations at O(d log d) cost.
RigidFormer is a transformer that learns mesh-free rigid-body dynamics via object-level anchors and differentiable Kabsch projection, outperforming mesh-based baselines with faster inference and scalability to 200+ objects.
A decentralized agent economy using auctions and economic selection emerges multi-step reasoning and outperforms monolithic baselines without centralized coordination.
LLMs generate global and local rubrics to standardize multimodal inputs, significantly outperforming clinical baselines on 15 EHRSHOT tasks via sample-efficient supervised learning.
Live Music Diffusion Models modify diffusion inference with block-wise KV caching to surpass discrete autoregressive efficiency, enabling stable alignment via ARC-Forcing and real-time interactive generation on consumer hardware.
NeuroAtlas benchmarks EEG foundation models across 42 datasets and finds they largely match generic time-series models without delivering unified clinical EEG performance.
Clari predicts organic crystal structures via unit-cell flow matching with pure pair-bias attention, cutting generation to seconds while surpassing OXtal solve rates and supporting non-sanitizable inputs.
DiscoverPhysics benchmarks LLM agents on simulated worlds with non-standard physics, finding frontier models pass only half and fail at uncovering latent structure.
This work formalizes nonlinear measure-to-measure regression and introduces two scalable transformer-based approaches for learning operators between probability distributions. The methods generalize to unseen measures in synthetic experiments, particle systems, and a large-scale colorectal cancer or
DiscoPER autonomously discovers scientific patterns via iterative meta-reflection and statistical testing, recovering 8 of 9 ecological patterns and outperforming baselines.
Edge of stability selectively redistributes learning across data groups via Hessian-aligned gradients and non-vanishing magnitudes, favoring output outliers over saturated ones.
Interacting particle systems extend mean shift to continuous distributions, minimizing maximum mean discrepancy via normalizing-constant-invariant dynamics for fast, multi-modal, high-dimensional quadrature.
Winrate incentivizes model homogenization that reduces consumer welfare, while weighted winrate improves producer specialization incentives and raises utility.
TILT decomposes predictors into main and auxiliary parts, penalizing the latter on unlabeled target data to implicitly weight sources via self-localized, bounded estimands, yielding finite-sample excess risk guarantees and improved domain adaptation performance.
ConnectomeBench2 unifies multi-species connectomic proofreading data, and a vision transformer trained on it achieves human-level split and merge error correction across species.
Interleaved Head Attention mixes attention heads via pseudo-heads to enable cross-head reasoning, cutting parameters on synthetic tasks and improving retrieval and math benchmarks over standard multi-head attention.
A unified compute-data scaling framework introduces token effectiveness to bridge data-rich and fixed-corpus regimes, showing diminishing returns and three operational regimes that render classical compute-optimal allocations suboptimal.
Action Images formulates robot policy learning as multiview video generation using interpretable pixel-grounded action images, enabling zero-shot control without separate policy heads and improving video-action joint generation.
SkillOS uses RL to train a skill curator that updates an external SkillRepo from experience, improving self-evolving agents across reasoning and multi-turn tasks.
Neural networks parameterize orthonormal function-space bases via ODEs on orthogonal Lie manifolds driven by skew-adjoint generators, with rank-2 generators universally approximating any target basis.
Neural cellular automata generate synthetic pre-training data that improves language model convergence and downstream reasoning faster than natural text.
ACORN strategically selects ML predictions for expert review to optimize occupancy model inference, recovering ecological conclusions near fully human-labeled levels with far fewer reviews.
MARLA improves worst-group accuracy without subgroup labels by applying a rank-limited logit correction within a low-dimensional misclassification subspace identified from held-out data.
RLVR trains a 30B LLM buyer via verifiable economic rewards to negotiate, revealing four-phase strategic evolution and outperforming much larger frontier models in surplus extraction.
Online autoregressive learning mistake bounds grow from constant to logarithmic in generation horizon M with end-to-end feedback, but chain-of-thought access removes M dependence entirely.
This paper proposes Laws of Reasoning (LoRe), a framework formalizing reasoning compute and accuracy laws, plus LoRe-Bench showing models lack compositionality; enforcing compute-law compositionality via finetuning improves reasoning performance.
Standard generalized advantage estimation suffers high variance from stochastic future actions in imperfect-information self-play, so introducing Q-boosting and variance-reduced policy optimization with expected SARSA traces improves performance across large-scale games.
MoveBench introduces a 2.6M-location wildlife movement forecasting benchmark across 110 species and finds existing methods generalize poorly to unseen individuals and deep learning does not consistently beat simpler baselines.
Recursive Language Models let LLMs recursively process long prompts as external environments, handling inputs two orders of magnitude beyond context windows with large quality gains over baselines at comparable cost.
Geometric characterization and fast-rate learning algorithm exactly compress linear programs into lower-dimensional equivalents preserving optimality with 1/n generalization.
Online discrete diffusion adaptation for molecular optimization finds acquisition, reward shaping, and debiasing complementarily boost reward, with replay and validity control stabilizing exploration to outperform offline and search baselines.
A re-solving algorithm achieves O(log² n) regret in nonstationary online linear programming using just one sample per distribution via dynamic programming and dual methods.
UNITE unifies tokenization and latent diffusion via a shared generative encoder, enabling single-stage joint training from scratch without adversarial losses or pretrained encoders to reach near state-of-the-art FID scores.
Testing approximate stationarity for piecewise-affine functions is XP in fixed dimension but W[1]-hard, with matching lower bounds, extending to shallow CNN losses.