Good Papers

Showing papers from Xiaohongshu Show all papers

45%Niche pick
?Niche pickVote to see the score

Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

Yanjiang Liu, Jie Lou, Xinyan Guan, Yuqiu Ji and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

OpenSearcher: Democratizing Deep Search through a Fully Offline Pipeline with Programmatic Verification

Zheng Chu, Xiao Wang, HuiMing Fan, Qianyu Wang and 5 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DeCoRL: Decomposed Consistency Reinforcement Learning for Multi-Image Composition

Zhiqiang Wu, Shuang Sun, Jiale Zhang, Jing Li and 3 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DepthGraft: Structural Regularization through Hierarchical Cross-Layer KV Reconstruction

Xiaohan Qin, Xiangdong Zhang, Yu Wang, Huaijin Wu and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

SALT: When More Rollouts Don’t Help in Group-Based Policy Optimization and How to Make Them Matter

Increasing rollouts fails in group-based RL because normalized gradients cancel; SALT adaptively reweights updates via subspace decomposition to recover effective learning.

Powei Chang, Jinpeng Zhang, Chaoqun Sun, MiniWell Tsao and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 2/5
medium 9/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

OASES co-trains a search policy and adaptive evaluator to provide outcome-aligned process rewards, outperforming RL baselines on multi-hop QA benchmarks.

Erhan Zhang, Yiqun Chen, Zechun Niu, Wei Yang and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

xHC: Expanded Hyper-Connections

xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.

Xiangdong Zhang, Xiaohan Qin, Tuo Dai, Xiaoming Shi and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 54 on Hugging Face · Code ★ 68

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 3/5
86%Must read
?Must readVote to see the score

Hint Tuning: Less Data Makes Better Reasoners

Hint Tuning calibrates reasoning depth by using an instruct model as a difficulty probe, cutting tokens by 24-66% with 1K samples.

Siqi Fan, Minghao Li, Xiaoqian Ma, Xiusheng Huang and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
86%Must read
?Must readVote to see the score

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning

AEM adaptively modulates response-level entropy dynamics for supervision-free credit assignment in multi-turn agent RL, improving exploration-exploitation trade-offs and consistently boosting strong baselines across ALFWorld, WebShop, and SWE-bench-Verified.

Haotian Zhao, Songlin Zhou, Yuxin Zhang, Stephen S Yau and 8 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 21 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
86%Must read
?Must readVote to see the score

Knowledge-Graph Paths as Intermediate Supervision for Self-Evolving Search Agents

Knowledge-graph paths provide intermediate supervision for self-evolving search agents, improving question validity via relational context and solver rewards via waypoint coverage, boosting multi-hop QA across benchmarks.

Huyu Wu, Jun Liu, Xiaochi Wei, Yan Gao and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
83%Must read
?Must readVote to see the score

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

Vision-OPD distills a crop-conditioned teacher into a full-image student via on-policy self-distillation to improve fine-grained visual understanding without external teachers or tools. It achieves competitive or superior performance on fine-grained benchmarks against larger open-source, closed-sour

Qianhao Yuan, Jie Lou, XingYu Li, Hongyu Lin and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5