A stopgrad regression principle characterizes stationary points of stopgrad objectives and proves convergence to true flow maps while halving training memory.
SGD learns linear spurious correlations exponentially faster than XOR signals in two-layer ReLU networks, with dynamics that suppress true feature learning.
Random neural networks approximate time-dependent Sobolev functions with dimension-free rate 1/2 and efficiently solve nonlinear porous medium and compressible Navier-Stokes equations.
Average partial Jacobian norms in transformers reveal subcritical signal growth in normalization-free architectures via tanh-like nonlinearities, explaining initialization sensitivity in DyT and Derf models.
A Hilbert bundle convolutional framework defines HilbNets for infinite-dimensional manifold signals, proving discrete versions converge to continuous architectures and transfer across samplings.
Dithered randomized Hadamard quantization is unbiased and achieves mean squared error asymptotically matching dense random rotations at O(d log d) cost.
Decoupled descent cancels data-reuse biases via approximate message passing so training error tracks test error, enabling zero-cost validation and shrinking the generalization gap versus gradient descent.
InfoFlow proves multi-layer Transformers exponentially beat single-layer ones on retrieval tasks and tracks information propagation to explain multi-layer approximation efficiency.
Deriving reduced ODEs for low-rank RNN learning reveals loss-invisible overlaps that encode training history and expose hidden connectivity differences.
Regularized Newton training of overparameterized neural networks converges to a deterministic NNTK limit with exponentially fast uniform convergence across all frequencies, avoiding gradient descent's spectral bias.
Mean-field transformers exhibit rapid token distribution concentration onto projection-driven limits with explicit Wasserstein bounds scaling in inverse temperature β and time t.
High-dimensional analysis of pretraining via PCA and linear probing derives exact errors versus representation size, showing compression helps with abundant unlabeled but scarce labeled data.
Fixed universal transformers simulate any target transformer via input embeddings with frozen internal parameters, and random initialization achieves universality almost surely.
Framework guarantees universal approximation for multistable dynamics with infinite-time horizon guarantees, linking topological properties to training metrics.
Optimal LR schedules for a solvable random feature model reveal easy-phase polynomial decay and hard-phase warmup-stable-decay regimes that improve scaling over constant or power-law schedules, with momentum and batch ramps further enhancing wall-clock time.
Score errors decompose into visible gradient and invisible solenoidal parts, so L2 score error cannot bound distribution divergence and only gradient error matters for diffusion sampling quality.
Spherical first-hitting diffusion models achieve near-minimax optimal convergence rates in total variation for Sobolev data on spheres, marking the first statistical optimality result for diffusion models using random generation times.
Deep ReLU networks with input and hidden widths ≥2 have open sets of identifiable parameters, yielding exact functional dimensions and generic depth hierarchies.
Function graph transformers lift functions to graph measures to universally approximate nonlinear operators between function spaces via standard attention and MLPs.
Single-layer self-attention trained with regret and swap-regret loss learns smoothed fictitious play and Blum-Mansour dynamics yielding coarse and correlated equilibria without supervised traces.
Approximate layer-wise activation distributions via cumulants and Hermite expansions to estimate wide MLP expected outputs without sampling, reducing FLOPs versus Monte Carlo and improving rare-event estimates.
Ordinary least squares predictions are rewritten as restricted attention outputs, framing OLS as similarity-based prediction via learned embedding and decoding operations mapped onto query-key-value structures.
Conservation laws express diffusion cross-entropy via local information-theoretic derivatives along noise paths, unifying discrete and continuous likelihoods and reducing training to marginal posterior learning.
For one-layer ResNets under depth-μP scaling, reused-weight forward-backward coupling vanishes at initialization but SGD induces surviving correlations suppressed by depth, yielding a rigorous infinite-depth Neural Feature Dynamics limit.
A framework predicts reconstruction error of compressive signal parameterizations via scaled differences between model predictions at different compression levels without ground truth. It yields non-asymptotic, signal-specific bounds that closely track global errors and local error heatmaps across i
Concept modulation models unify conditional latent variable model identifiability and extrapolation via attribute potentials and algebraic criteria for unseen attributes.
Gated attention represents attention matrices as hierarchical mixtures of experts and achieves polynomial sample complexity versus exponential for multi-head self-attention.
Neural LoFi frames deep training as iterative spectral low-degree filtering, predicting layer-wise feature selection, concept emergence, and compositional depth via low-degree correlation dynamics.
Under standard initialization, two-layer networks learn orthogonal multi-index targets incrementally via competitive neuron dynamics, with lower-order Hermite components recovered before higher-order directions.
A fixed-point framework proves looped transformers need recall plus outer normalization for stable, input-dependent extrapolation, validated across chess, sudoku, and prefix-sums tasks.
Stochastic optimizers at the edge of stability converge to low-dimensional fractal attractors, and a sharpness-dimension generalization bound reveals that chaotic training depends on the full Hessian spectrum.
Shallow neural networks trained via gradient descent exhibit uniform-in-time weak propagation of chaos, yielding poly(d/ε) neuron and sample complexity when mean-field loss decays faster than t^{-2}.