DiPhon defines graphon diffusion via a Jacobi SDE for scalable graph generation, matching first moments exactly and preserving topology across sizes without retraining.
Linear probes on LLM residual streams identify a shared preference vector tracking pairwise choices across personas, with cross-persona transfer and causal steering.
Joint Consistency frames test-time aggregation as energy minimization using pairwise interactions and evaluation signals, outperforming existing voting methods across reasoning benchmarks.
Dithered randomized Hadamard quantization is unbiased and achieves mean squared error asymptotically matching dense random rotations at O(d log d) cost.
MyoChallenge 2025 benchmarks musculoskeletal sports control via simulated table tennis and soccer tasks, advancing agile motor algorithms across 70 teams.
Argus detects backdoor attacks in decentralized learning by having nodes share local trigger analyses with neighbors and filter updates via structural similarity, reducing attack success by up to 90 points without a central server.
Terminator learns optimal early-exit points for chain-of-thought reasoning to cut token lengths by 14%-55% and boost inference speed over 2x with minimal accuracy loss.
LoRIF exploits low-rank gradient structure to reduce storage and query I/O to O(c√D) and inverse Hessian memory to O(Dr), achieving up to 20× speedups over LoGRA at scale.
Learned stochastic stopping reduces out-of-distribution variance in looped transformers by decoupling loop count from sequence length during training. It improves accuracy-stability trade-offs across algorithmic tasks, though it can stabilize suboptimal computation.
Random search for stochastic optimization works under weaker smoothness assumptions and achieves faster convergence via variance-reduced variants using translation invariance to balance noise.
Penalized DRO reformulates adversarial risk via optimal transport maps that are cyclically monotone, and enforcing this property via multi-start particle ascent or input-convex networks improves robustness over standard adversarial training.
KroQuant applies a learned Kronecker-structured block transform to DiT activations for efficient W4A4 post-training quantization that outperforms SVDQuant and LoRaQ on image quality.
Quantitative local convergence rates for mean-field SVGD with Riesz kernels are established via explicit polynomial L² decay, with sharpness verified numerically.
MA-BC partitions conflicting expert trajectories while pooling compatible data to recover Pareto-optimal policies in multi-objective imitation with minimax optimal rates.
BalCapRL balances RL for MLLM captioning across correctness, coverage, and fluency via normalized multi-objective rewards and length masking, boosting quality metrics substantially.
Latent prediction learns hierarchical latent trees with samples constant in depth L, exponentially more efficient than token-level self-supervision, making explicit multi-scale stacking largely redundant.
LOSCAR-SGD combines local SGD, sparse communication, and overlap with a delay-corrected merge for heterogeneous workers, yielding convergence guarantees and faster training.