Good Papers

Showing Alignment & preference optimization Show all papers

86%Must read

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Cross-tokenizer on-policy distillation achieves comparable accuracy with strict top-16 shared-vocabulary supervision versus full coverage, while expanded span supervision reduces accuracy due to conflicting gradients, motivating prioritization of supervision reliability over alignment coverage.

Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li and 3 more

Published Oct 6, 2026 · ▲ 40 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
91%Must read

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

OnePO enables RL-only medical domain adaptation via adaptive objective evolution and teacher retirement, yielding HuatuoGPT-3 that surpasses frontier models.

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu and 6 more

Published Oct 5, 2026 · ▲ 15 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
67%Highly rated

Collaborative Personalized Preference Alignment for LLMs under Data Deficiency

APO learns shared LLM aligner initializations via grouped gradient coordination to enable few-shot personalization under heterogeneous preferences and scarce feedback, improving over baselines with 20 local examples.

Liyan Yang, Yige Yuan, Zhiqin Yang

Published Oct 5, 2026 · ▲ 5 on Hugging Face

0% Readers0 of 1 upvoted
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Representation-Space MMD for Diffusion Language Models

Post-training minimizes representation-space MMD between diffusion language model outputs and references via retained token features, improving perplexity, accuracy, and parallel decoding.

Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev and 6 more

Published Oct 5, 2026 · ▲ 16 on Hugging Face · Code ★ 10

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

DiffGate gates on-policy teacher guidance by trajectory failure and group difficulty to combine dense token-level updates with outcome-level GRPO rewards, improving student pass rates.

Karn Tiwari, Varnith Chordia, Prathosh A P

Published Oct 3, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot-SD self-distills masked diffusion language models by supervising high-impact commitment tokens via information-gain selection, improving reasoning with minimal data.

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 54 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
83%Must read

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

RIDE extrapolates RL-induced representation residuals for stable on-policy distillation, surpassing output-space methods and matching or exceeding RL teachers.

Hao Li, Meijia Chen, Weijie Ren, Donghan Li and 3 more

Published Sep 29, 2026 · 0 citations · ▲ 577 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

On-policy methods continuously adjust parameter update directions, unlike consistent SFT updates; constraining SFT to these directions via OPSFT transfers on-policy generalization advantages to supervised fine-tuning.

Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 81 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

PersonaDose calibrates activation-steering controllers to control language-model persona traits by requested intensity, reducing targeting errors to 4.7-6.2 points across models.

Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen and 3 more

Published Sep 28, 2026 · 0 citations · ▲ 54 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
88%Must read
?Must readVote to see the score

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

In controlled strong-to-weak distillation, rollout policy is less central than token-level KL direction and learning rate, though on-policy data can improve generalization on harder reasoning tasks.

Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar

Published Sep 28, 2026 · 0 citations · ▲ 194 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 3/5
72%Highly rated
?Highly ratedVote to see the score

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

MinT manages LoRA adapter revisions over shared 1T-class base models to train and serve millions of policies via adapter-only handoffs and durable addressability.

Mind Lab, :, Song Cao, Vic Cao and 36 more

Published May 13, 2026 · 0 citations · ▲ 226 on Hugging Face · Code ★ 79

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only

LLMs distinguish degrees of wrongness among incorrect answers, and alignment with such preferences yields less wrong answers and better calibration.

Jihan Yao, Wenxuan Ding, Shangbin Feng, Lucy Lu Wang and 1 more

Published Oct 14, 2024 · 0 citations · Code ★ 10

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 2/5
83%Must read
?Must readVote to see the score

Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

Modular Pluralism plugs specialized community LMs into base LLMs to enable Overton, steerable, and distributional pluralistic alignment across diverse communities.

Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher and 3 more

Published 2024 · 15 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners

Botos Csaba, Sreejan Kumar, Austin T D Andrews, Laurence T Hunt and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Inference-time Alignment via Sparse Junction Steering

Runyi Hu, Jie Zhang, Shiqian Zhao, Jiale Meng and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Behaving Better, Thinking Worse: Sycophancy Across Post-Training Stages

Sonnet Xu, Kritika Singh, Sheharbano Jafry, Roxana Daneshjou and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

PORT: Preference Optimization via Robust Token-Level Reweighting

Ding Zhu, Xiukun Wei, Tian Xie, Zhihui Zhu and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AdKnob: Ad Intensity Control and Labeling for LLM-Native Advertising

Woo Jae Kim, Seongho Keum, Joonsung Jeon, Suhyeon Ha and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Do LLMs Feel Social Pressure? Locating and Steering Social Desirability Bias in LLMs

Yi Feng, Jiaqi Wang, Wenxuan Zhang

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Margin Dynamics for Large Language Model Alignment

Xingzi Xu, Saygin Seyfioglu, Karim Bouyarmane

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
Show 20 more papers