Good Papers

Showing papers from Tencent (China) Show all papers

76%Highly rated
?Highly ratedVote to see the score

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

Rollout-Marginal Distillation scores autoregressive video chunks independently against a chunk teacher to prevent error accumulation, then applies video-level distillation to restore temporal coherence.

Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang and 2 more

Published Sep 29, 2026 · 0 citations · ▲ 18 on Hugging Face · Code ★ 2

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

Adaptive Reward Routing dynamically routes updates and balances rewards during forward-process RL for joint audio-video diffusion, consistently improving quality, alignment, and synchronization over fixed baselines.

Songlin Yang, Xiaotong Zhao, Jiacheng Zhang, Zhe Wang and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 138 on Hugging Face

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 0/5
80%Must read
?Must readVote to see the score

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data

Reinforcement learning on saturated reasoning data causes mode collapse as advantage signals vanish; CUTS sampling and Mixed-CUTS restore diversity, boosting AIME25 accuracy by up to 15.1%.

Zhenwen Liang, Yujun Zhou, Sidi Lu, Xiangliang Zhang and 2 more

Published Apr 20, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data

Standard RL collapses on saturated reasoning data due to vanishing advantage signals, so CUTS sampling and Mixed-CUTS training restore exploration and boost AIME25 Pass@1 by 15.1%.

Zhenwen Liang, Yujun Zhou, Sidi Lu, Xiangliang Zhang and 2 more

Published 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
54%Worth a look
?Worth a lookVote to see the score

Rhombus: Incentivizing Coordination in Parallel Thinking through Reinforcement Learning

Rhombus uses reinforcement learning to incentivize coordination in parallel thinking frameworks.

Ziyuan Nan, Qi Yi, Di Huang, Yutong Wu and 8 more

Published 2026 · 0 citations

100% Readers1 of 1 upvoted
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

RecLM: Recommendation Instruction Tuning

RecLM integrates large language models with collaborative filtering via instruction tuning and a reinforcement learning reward to enhance recommendation performance, especially for sparse and zero-shot settings.

Yangqin Jiang, Yuhao Yang, Lianghao Xia, Da Luo and 2 more

Published Dec 26, 2024 · 0 citations · Code ★ 111

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5