78%Highly rated
?Highly ratedVote to see the score
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
Forward-KL-regularized offline contextual bandits achieve epsilon^{-1} sample complexity under single-policy concentrability via pessimism, with matching lower bounds showing slow rates at weak regularization.
Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026
– ReadersNo votes yet
11/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 11 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 3/5