DiPhon defines graphon diffusion via a Jacobi SDE for scalable graph generation, matching first moments exactly and preserving topology across sizes without retraining.
Linear probes on LLM residual streams identify a shared preference vector tracking pairwise choices across personas, with cross-persona transfer and causal steering.
Joint Consistency frames test-time aggregation as energy minimization using pairwise interactions and evaluation signals, outperforming existing voting methods across reasoning benchmarks.
Dithered randomized Hadamard quantization is unbiased and achieves mean squared error asymptotically matching dense random rotations at O(d log d) cost.
MyoChallenge 2025 benchmarks musculoskeletal sports control via simulated table tennis and soccer tasks, advancing agile motor algorithms across 70 teams.
Argus detects backdoor attacks in decentralized learning by having nodes share local trigger analyses with neighbors and filter updates via structural similarity, reducing attack success by up to 90 points without a central server.
Terminator learns optimal early-exit points for chain-of-thought reasoning to cut token lengths by 14%-55% and boost inference speed over 2x with minimal accuracy loss.
LoRIF exploits low-rank gradient structure to reduce storage and query I/O to O(c√D) and inverse Hessian memory to O(Dr), achieving up to 20× speedups over LoGRA at scale.
Learned stochastic stopping reduces out-of-distribution variance in looped transformers by decoupling loop count from sequence length during training. It improves accuracy-stability trade-offs across algorithmic tasks, though it can stabilize suboptimal computation.
Random search for stochastic optimization works under weaker smoothness assumptions and achieves faster convergence via variance-reduced variants using translation invariance to balance noise.
Penalized DRO reformulates adversarial risk via optimal transport maps that are cyclically monotone, and enforcing this property via multi-start particle ascent or input-convex networks improves robustness over standard adversarial training.
KroQuant applies a learned Kronecker-structured block transform to DiT activations for efficient W4A4 post-training quantization that outperforms SVDQuant and LoRaQ on image quality.
Quantitative local convergence rates for mean-field SVGD with Riesz kernels are established via explicit polynomial L² decay, with sharpness verified numerically.
MA-BC partitions conflicting expert trajectories while pooling compatible data to recover Pareto-optimal policies in multi-objective imitation with minimax optimal rates.
BalCapRL balances RL for MLLM captioning across correctness, coverage, and fluency via normalized multi-objective rewards and length masking, boosting quality metrics substantially.
Latent prediction learns hierarchical latent trees with samples constant in depth L, exponentially more efficient than token-level self-supervision, making explicit multi-scale stacking largely redundant.
LOSCAR-SGD combines local SGD, sparse communication, and overlap with a delay-corrected merge for heterogeneous workers, yielding convergence guarantees and faster training.
Task-induced Riemannian metrics define ViT feature geometry via decoder Jacobians, and a learnable low-rank approximation enables accurate geometric token pruning without fine-tuning.
Cephalonauts One provides 30 hours per subject of whole-brain fMRI during naturalistic speech, paired with audio, transcripts, and embeddings, plus a brain decoding benchmark showing continuous performance gains with more training data.
Deep spiking ResNets preserve cortical-like functional connectivity where rare, coordinated 1FC ensemble cofiring reliably predicts downstream responses via ReLU-like scaling, encodes class identity, and breaks under adversarial perturbation and weight permutation.
ADAS reranks masked diffusion sampling by discounting token confidence via attention to selected uncertain positions, improving low-step reasoning and code accuracy by up to 10.5 points with minimal overhead.
Neural LoFi frames deep training as iterative spectral low-degree filtering, predicting layer-wise feature selection, concept emergence, and compositional depth via low-degree correlation dynamics.
EverAnimate restores drifted latent flows via persistent memory and restorative matching, improving long human animation quality and identity consistency over minutes.
Rescaled ASGD corrects asynchronous SGD's bias toward fast workers via computation-time-proportional step sizes, matching optimal time complexity with only lower-order heterogeneity penalties.
RePercENT scales disentangled multimodal representation learning beyond two modalities via a plug-and-play framework that extracts shared and unique factors with formal guarantees and lower complexity.
PICID introduces a modular infrastructure that formalizes reproducible PHM evaluation pipelines and enables fair cross-task comparisons across diagnostics and prognostics.
WildBox provides aerial monocular 3D wildlife annotations and benchmarks showing zero-shot 3D detection collapses to zero, with fine-tuning reaching 13.17 AP3D and depth as the dominant failure mode.
NASDAQ normalizes low-dimensional observations to balance dynamics prediction losses and couples value learning with short-term value and next-observation prediction, achieving strong sample efficiency and faster training across diverse domains.
Causal inference framing of membership inference attacks defines memorization as training inclusion effects, reveals interference and distribution-shift biases, and yields reliable estimators without retraining.
MUX compresses reasoning into continuous multiplexed tokens via lossless superposition, accelerating reasoning and outperforming latent baselines across 32 settings.
Standardized LLM pretraining benchmarks compare optimizers across model sizes, batch sizes, and training durations to guide selection and highlight future research directions.
ECHO-k uses pretrained representations as proxy targets for self-supervised reinforcement learning to sequentially acquire informative modalities at test time, improving budgeted downstream performance across diverse backends.
DYSCO uses multi-view contrastive learning to recover latent dynamics and governing equations from noisy high-dimensional data, with theoretical identification guarantees and empirical validation across diverse regimes.
Active context selection improves contextual bandit simple regret from order root n over T times L1/2 norm of p to root n over T times L2/3 norm, with gains up to k to the 1/4.
Soft-Radial Projection uses radial interior mapping to enforce hard constraints with full-rank Jacobians, avoiding gradient saturation and improving convergence.