Good Papers

Showing papers from University of California, Santa Barbara (UCSB) Show all papers

83%Must read
?Must readVote to see the score

CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use

CM2 replaces verifiable outcome rewards with checklist rewards for multi-turn tool-use RL, improving 8B models by 8, 12 points on agent benchmarks using simulated environments.

Zhen Zhang, Kaiqiang Song, Sean Wang, Yebowen Hu and 10 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5