CIDER is a masked-diffusion multiuser decoder using demixing and parity-aware propagation to outperform joint belief propagation by 6-100x in speed with matching error rates.
Lean Refactor uses retrieval-augmented agentic strategy search to multi-objectively refactor Lean proofs, achieving over 70% token compression and up to 60% faster compilation with stronger version transfer.
Estimating total variation distance between autoregressive models uses Õ(n²K/ε²) samples, O(n/ε²) logit queries, and interpolates under noisy access, with practical LLM engine comparisons.
Co-Coder formalizes multi-agent coding as graph partitioning to balance parallel speedups against communication overhead, improving pass rates by 14% and cutting costs 35% on dense repositories.
Optimal LR schedules for a solvable random feature model reveal easy-phase polynomial decay and hard-phase warmup-stable-decay regimes that improve scaling over constant or power-law schedules, with momentum and batch ramps further enhancing wall-clock time.
DistBART applies BART priors to distribution regression via Riesz representers, yielding adaptive convergence, nonlinear extensions, and scalable random-feature inference.
SDBPG adaptively perturbs dual formulations to stabilize multipliers near lower-level stationary points, yielding first explicit sample-complexity guarantees for stochastic nonconvex simple bilevel optimization.
AVSD separates cross-view consensus from privileged residuals in multi-view self-distillation to adaptively supervise reasoning models, improving math and code benchmarks over single-view methods and GRPO.
An online algorithm achieves optimal ε-recalibration with ε² excess error in ε⁻³ rounds via Blackwell approachability, and yields simultaneous calibration and calibeating for smooth losses.
Gated attention represents attention matrices as hierarchical mixtures of experts and achieves polynomial sample complexity versus exponential for multi-head self-attention.