Stochastic linear bandits with delayed feedback yield near-optimal, dimension-free additive penalties for loss-independent delays but dimension-dependent penalties for loss-dependent delays, unlike multi-armed bandits.
TROPT unifies discrete text optimization via a modular open-source framework with 30+ recipes, enabling cross-domain optimizer comparison, enhancement, and portability.
Geometric coupling aligns router and expert gradients along shared input directions in sparse Mixture-of-Experts, while load-balancing losses disrupt it, and cosine-similarity routing achieves low imbalance with minimal perplexity cost.
The paper introduces symmetric zero-sum Markov games and shows that observing only opponent actions allows asymptotically matching adversarial returns without payoff observations via online learning.
DEMASK predicts token dependencies in discrete diffusion language models to select weakly dependent masked positions for parallel unmasking, bounding sampling error and accelerating Dream-7B by 1.7, 2.2× with preserved accuracy.
Epistemic pairwise maximin share (EPMMS) relaxes PMMS fairness; 4/5-EPMMS allocations exist for additive valuations, exact EPMMS for bivalued valuations, and existence holds for three additive or two-type agents despite MMS nonexistence.
Mamba performs associative recall via implicit linear hashing, and Recall Scaling Laws predict required dimensions and success probabilities for perfect recall.
Outcome-based RL provably teaches single-layer transformers iterative graph traversal via chain-of-thought, but only with sufficient simple training examples.
ParetoSlider trains one diffusion model with continuous preference weights to approximate the full Pareto front, enabling inference-time navigation of trade-offs between conflicting generative goals without retraining.
A single sensory-prediction recurrent network with Dale's Law co-emerges grid and place cells without supervision, reproducing key spatial coding phenomena.