Good Papers

Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret

Strong duality holds for weakly communicating average-reward CMDPs, yielding a primal-dual algorithm with O(T^{2/3}) regret and constraint violations.

Kihyun Yu, Beomhan Baek, Dabeen Lee

Published 2026Sydney Poster Session 6 · Thu, Dec 10, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 2/5
medium 5/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Our contributions are twofold. First, we establish strong duality for weakly communicating average-reward CMDPs over stationary policies with finite state and action spaces. Despite the absence of a linear programming formulation and the resulting nonconvexity under the weakly communicating setting, we show that strong duality still holds by carefully exploiting the geometric structure of the occupation measure set. Second, building on this result, we propose a primal--dual clipped value iteration algorithm for learning weakly communicating average-reward linear CMDPs. Our algorithm achieves regret and constraint violation bounds of $\widetilde{\mathcal{O}}(T^{2/3})$, improving upon the best known bounds, where $T$ denotes the number of interactions. Our approach extends clipped value iteration to the constrained setting and adapts it to a finite-horizon approximation, which stabilizes the dual variable and is crucial for achieving improved regret bounds. To analyze this, we develop a novel approach based on strong duality that enables the decomposition of the composite Lagrangian regret into separate bounds on regret and constraint violation.