Good Papers

Showing papers from Tencent Technology (shanghai) Co. ltd Show all papers

45%Niche pick
?Niche pickVote to see the score

Grounding Agent Reasoning with Structured Process Supervision for Multi-turn Reinforcement Learning

Renting Rui, Yulei Qin, Weiwen Liu, Yunjia Xi and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction

GRGC calibrates groupwise reward geometry via optimal transport regularization and power modulation to fix brittle advantages in adversarial black-box LLM distillation.

Xiao Cui, Mo Zhu, Yulei Qin, Yuze Wu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 0/5