Good Papers

Showing papers from Department of Automation, Tsinghua University Show all papers

45%Niche pick
?Niche pickVote to see the score

Bridging Risk Approximation Gaps in Model Predictive Task Sampling via In-Context Modeling

Jiarong Wen, Qi Tao, Zhang Kaiyu, Yun Qu and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex

Listwise Policy Optimization explicitly projects policies onto target distributions over response simplices via divergence minimization, improving reasoning performance and stability over group-based policy gradients.

Yun Qu, Qi Wang, Yixiu Mao, Heming Zou and 10 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 67 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Boosting LLM Reasoning via Human-Inspired Reward Shaping

T2T dynamically rewards exploration on failed reasoning attempts and brevity on correct ones, significantly boosting LLM math reasoning over GRPO.

Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5