Stochastic linear bandits with delayed feedback yield near-optimal, dimension-free additive penalties for loss-independent delays but dimension-dependent penalties for loss-dependent delays, unlike multi-armed bandits.
Online set learning with randomized precision or recall feedback is learnable exactly when the hypothesis class has finite VC dimension, though standard empirical risk minimization can fail and algorithms must handle feedback dependencies to achieve regret bounds.