A bidder with dynamic auction values and only aggregated feedback learns near-optimal bidding policies via plug-in estimators with logarithmic or sublinear regret.
A Cramér-von Mises fairness regularizer with O(B log B) complexity penalizes prediction-sensitive attribute dependence during training, achieving competitive fairness-utility trade-offs with lower overhead.