Good Papers

Showing papers from Microsoft Research Aisa Show all papers

80%Must read
?Must readVote to see the score

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

SCOPE-RL uses temperature-adaptive positive samples to stabilize entropy and prevent collapse in RL post-training of reasoning LLMs, improving Pass@1 and Pass@$k$ with non-monotonic exploration benefits.

Chen Wang, Zhaochun Li, Bai Jionghao, Hexuan Deng and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5