Adaptive targeting under sparse network interference achieves near-optimal regret depending on structural knowledge, proving standard linear bandits are inefficient and offering practical algorithms.
A Track-and-Stop algorithm solves thresholding Monte Carlo Tree Search with asymptotically optimal sample complexity, and a ratio-based D-Tracking modification improves empirical efficiency and reduces per-round computation to logarithmic time.
A bidder with dynamic auction values and only aggregated feedback learns near-optimal bidding policies via plug-in estimators with logarithmic or sublinear regret.
In Bayesian multi-armed bandits, pretraining on expert data tightens regret bounds by mutual information with the optimal action, while an information-directed rule selects data sources maximizing immediate information gain, and trust inference safeguards against ineffective or compromised experts.
2FFS adaptively combines cheap biased heuristics and expensive accurate rollouts to identify best actions in stochastic minimax trees with fewer samples than baselines.
A directional bias certificate enables Ellipsoidal-MINUCB to safely exploit offline data in linear contextual bandits, reducing regret when low-bias directions align with historical coverage.
Volatility and stochasticity both increase uncertainty but drive optimal exploration in opposite directions; CAUSE captures this asymmetry and improves restless-bandit performance.
SURF inverts a geometric arc-length cumulative distribution to sample scalarization weights yielding uniform Pareto front coverage and converges linearly to a finite-sampling floor.
Residual Quantization maps contexts to discrete additive codes enabling nonlinear contextual bandits with strictly bounded memory, beating linear variants on 11 of 13 datasets and matching heavy retrained baselines with up to 1000x less memory.
Online RLHF with generalized bilinear preferences achieves polylogarithmic regret via generic strong convexity and skew-symmetry, proving fast rates are not KL-specific.
Active context selection improves contextual bandit simple regret from order root n over T times L1/2 norm of p to root n over T times L2/3 norm, with gains up to k to the 1/4.