Good Papers

Showing papers from Hong Kong University of Science and Technology (HKUST) Show all papers

57%Worth a look
?Worth a lookVote to see the score

Reformulate LLM Reinforcement Learning for Stable Training under Black-box Discrepancy

Jiashun Liu, Runze Liu, Xu Wan, Jing Liang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Archer applies entropy-aware dual-token constraints to RLVR, modulating optimization strengths across reasoning and knowledge tokens to improve mathematical and code performance.

Jiakang Wang, Runze Liu, Fuzheng Zhang, Xiu Li and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 21 on Hugging Face · Code ★ 44

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions

Gaussian trust region reshaping replaces monotonic divergence penalties with bounded non-monotonic constraints, unlocking efficient behavior transitions in non-stationary reinforcement learning.

Bingxu Liu, Jiashun Liu, Johan Obando Ceron, Hao Wang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5