80%Must read
?Must readVote to see the score

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
SCOPE-RL uses temperature-adaptive positive samples to stabilize entropy and prevent collapse in RL post-training of reasoning LLMs, improving Pass@1 and Pass@$k$ with non-monotonic exploration benefits.
Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face
– ReadersNo votes yet
12/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5