Good Papers

Showing Alignment & preference optimization Show all papers

86%Must read

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Cross-tokenizer on-policy distillation achieves comparable accuracy with strict top-16 shared-vocabulary supervision versus full coverage, while expanded span supervision reduces accuracy due to conflicting gradients, motivating prioritization of supervision reliability over alignment coverage.

Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li and 3 more

Published Oct 6, 2026 · ▲ 40 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
91%Must read

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

OnePO enables RL-only medical domain adaptation via adaptive objective evolution and teacher retirement, yielding HuatuoGPT-3 that surpasses frontier models.

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu and 6 more

Published Oct 5, 2026 · ▲ 15 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
67%Highly rated

Collaborative Personalized Preference Alignment for LLMs under Data Deficiency

APO learns shared LLM aligner initializations via grouped gradient coordination to enable few-shot personalization under heterogeneous preferences and scarce feedback, improving over baselines with 20 local examples.

Liyan Yang, Yige Yuan, Zhiqin Yang

Published Oct 5, 2026 · ▲ 5 on Hugging Face

0% Readers0 of 1 upvoted
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Representation-Space MMD for Diffusion Language Models

Post-training minimizes representation-space MMD between diffusion language model outputs and references via retained token features, improving perplexity, accuracy, and parallel decoding.

Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev and 6 more

Published Oct 5, 2026 · ▲ 16 on Hugging Face · Code ★ 10

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

DiffGate gates on-policy teacher guidance by trajectory failure and group difficulty to combine dense token-level updates with outcome-level GRPO rewards, improving student pass rates.

Karn Tiwari, Varnith Chordia, Prathosh A P

Published Oct 3, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot-SD self-distills masked diffusion language models by supervising high-impact commitment tokens via information-gain selection, improving reasoning with minimal data.

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 54 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

RIDE extrapolates RL-induced representation residuals for stable on-policy distillation, surpassing output-space methods and matching or exceeding RL teachers.

Hao Li, Meijia Chen, Weijie Ren, Donghan Li and 3 more

Published Sep 29, 2026 · 0 citations · ▲ 577 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

On-policy methods continuously adjust parameter update directions, unlike consistent SFT updates; constraining SFT to these directions via OPSFT transfers on-policy generalization advantages to supervised fine-tuning.

Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 81 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

PersonaDose calibrates activation-steering controllers to control language-model persona traits by requested intensity, reducing targeting errors to 4.7-6.2 points across models.

Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen and 3 more

Published Sep 28, 2026 · 0 citations · ▲ 54 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

In controlled strong-to-weak distillation, rollout policy is less central than token-level KL direction and learning rate, though on-policy data can improve generalization on harder reasoning tasks.

Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar

Published Sep 28, 2026 · 0 citations · ▲ 194 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

MinT manages LoRA adapter revisions over shared 1T-class base models to train and serve millions of policies via adapter-only handoffs and durable addressability.

Mind Lab, :, Song Cao, Vic Cao and 36 more

Published May 13, 2026 · 0 citations · ▲ 226 on Hugging Face · Code ★ 79

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only

LLMs distinguish degrees of wrongness among incorrect answers, and alignment with such preferences yields less wrong answers and better calibration.

Jihan Yao, Wenxuan Ding, Shangbin Feng, Lucy Lu Wang and 1 more

Published Oct 14, 2024 · 0 citations · Code ★ 10

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

Modular Pluralism plugs specialized community LMs into base LLMs to enable Overton, steerable, and distributional pluralistic alignment across diverse communities.

Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher and 3 more

Published 2024 · 15 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners

Botos Csaba, Sreejan Kumar, Austin T D Andrews, Laurence T Hunt and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Inference-time Alignment via Sparse Junction Steering

Runyi Hu, Jie Zhang, Shiqian Zhao, Jiale Meng and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Behaving Better, Thinking Worse: Sycophancy Across Post-Training Stages

Sonnet Xu, Kritika Singh, Sheharbano Jafry, Roxana Daneshjou and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

PORT: Preference Optimization via Robust Token-Level Reweighting

Ding Zhu, Xiukun Wei, Tian Xie, Zhihui Zhu and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

AdKnob: Ad Intensity Control and Labeling for LLM-Native Advertising

Woo Jae Kim, Seongho Keum, Joonsung Jeon, Suhyeon Ha and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Do LLMs Feel Social Pressure? Locating and Steering Social Desirability Bias in LLMs

Yi Feng, Jiaqi Wang, Wenxuan Zhang

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Margin Dynamics for Large Language Model Alignment

Xingzi Xu, Saygin Seyfioglu, Karim Bouyarmane

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Verified-Source Authority Is Not Generic Sycophancy: Cue-Family Decomposition of LLM Compliance

Abhinav Rajeev Kumar, Paras Chopra

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

META-PAP: Meta-learning for Prompt-aware Preference Pairing in LLM Alignment

Pinlong Zhao, Shiyu Hu, Jing Zhang, Mengyang Li

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SoccerNarrate: Event-Grounded Streaming Soccer Commentary with Macro-Window Preference Alignment

zihan jia, Zhilin Dai, Zhengming Zhang, Min Yang and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Rethinking "RL Generalizes, SFT Memorizes": The Role of SFT Data

Yunlong Hou, Fengzhuo Zhang, Yuan Cheng, Jiachun Pan and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Symmetric Interventions for Eliciting Model Intent

David Vella Zarb, Rustem Turtayev, Taywon Min, Jinghua Ou and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Grassmannian Geodesic Steering: Rank-Preserving Subspace Control for Inference-Time Alignment of Language Models

Longyi Liu, Zhitao Wang, Jianchao Yu, Mingrui Cai and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Personalized LLM Alignment Should Be Counterfactually Verifiable

Cristina Garbacea

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Illusion of Diversity: Aligning LLM Exploration via Effective Entropy

Xiaoliang Fu, Jiaye Lin, Yangyi Fang, Cong Qin and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learnable Chernoff Baselines for Provable Inference-Time Alignment

Sunil Madhow, Yuchen Liang, Ness Shroff, Yingbin Liang and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Can Large Language Models Develop Gambling Addiction?

Seungpil Lee, Donghyeon Shin, Yunjeong Lee, Sundong Kim

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

The Post-Training Dilemma: Why We Should Rethink the Sequential SFT-RL Paradigm

Xueyan Niu, Bo Bai, Wei Han, Weixi Zhang

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Correctness: Robustness-Driven Evolutionary Self-Training for Large Language Models

Wei Guo, Hongyao Tang, Yi Ma, Jinyi Liu and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Outcome: Trajectory-Driven Prompt Optimization via Multi-Dimensional Rewards

Zhixiang Liang, Yixiang Huang, CHENJING CAI, Hao Wu and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Steering Vectors as a Training Signal in LLM Post-Training

Tiejin Chen, Maunil R Vyas, Huaiyuan Yao, Hua Wei

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Towards Multi-Human-Value Alignment via Value Localization in LLMs

Xueqi Ma, Yanbei Jiang, Xingjun Ma, James Bailey and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DART: Zero-Shot Dual-Side Alignment Routing for LLM Performance-Cost Tradeoffs

Yuejun Jiao, Yanxin Yang, Boyu Wang, Yonghao Yang and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Label-Free Consistency Correction for Weak-to-Strong Generalization

Qi He, Heng Huang

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Alignment Tax Concentrates in Output Projections

Ming Liu

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Bias-Variance Optimized Preference Optimization for Large Reasoning Models

Mingkang Zhu, Xi Chen, Bei Yu, Hengshuang Zhao and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control

ABC-Align minimizes variance via pseudo-labels and applies adaptive, lightweight bias correction tuned by plug-in estimates, outperforming semi-supervised alignment baselines with scarce human feedback.

Eric Frankel, Banghua Zhu, Sewoong Oh, Lillian Ratliff

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
88%Must read
?Must readVote to see the score

Alignment Dynamics in LLM Fine-Tuning

A unified framework decomposes LLM alignment dynamics into competing rebound and driving forces, explaining reversal and faster re-alignment via rehearsal priming.

Yuhan Huang, Huanran Chen, Yinpeng Dong

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
83%Must read
?Must readVote to see the score

Strengthening LLMs for Tabular Prediction with Structural Priors

PRPO incorporates column-permutation invariance into LLM post-training via label-preserving permutations and two-level advantage estimation, enabling an 8B model to match specialized tabular baselines and outperform 685B reasoning LLMs by up to 53%.

Pengxiang Cai, Zihao Gao, Wanchen Lian, Guocong Li and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5
80%Must read
?Must readVote to see the score

IRIS: Interpolative Rényi Iterative Self-play for Large Language Model Fine-Tuning

IRIS unifies self-play fine-tuning via adjustable Rényi divergence with adaptive schedules, surpassing supervised fine-tuning with fewer annotations across benchmarks.

Wenjie Liao, Like Wu, Liangjie Zhao, Shihui Xu and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 1/5
91%Must read
?Must readVote to see the score

Response Time Enhances Alignment with Heterogeneous Preferences

Adding response times to preference data via drift-diffusion modeling restores identifiability of average preferences among anonymous heterogeneous labelers, correcting choice-only estimation bias without tracking users.

Federico Echenique, Alireza Fallah, Baihe Huang, Michael Jordan

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

Aligning Language Models with Selective Prediction

RLSR aligns language models with selective prediction metrics via reinforcement learning, substantially improving risk-coverage trade-offs over baselines.

Gaoxiang Luo, Yifan Wu, Sinian Zhang, Aryan Deshwal and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
83%Must read
?Must readVote to see the score

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion

MORA rewrites prompts to expand multi-dimensional reward diversity and breaks the safety-helpfulness trade-off, improving sequential single-preference alignment by up to 12.4% and simultaneous alignment by 4.6%.

ShiYing Huang, Liang Lin, Yuer Li, Kaiwen Luo and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
89%Must read
?Must readVote to see the score

Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy

Deriving an alignment imprint from LLM preference tuning, LAPD detects AI-generated text with 45.82% relative gains over baselines via statistically guaranteed preference discrepancy.

Junxi Wu, Kailin Huang, Dongjian Hu, Bin Chen and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
83%Must read
?Must readVote to see the score

MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models

MATO achieves training-free multi-objective LLM alignment via test-time optimization of discovered rewards and adaptive weights during decoding, improving steerability and Pareto performance.

LINHAO LUO, Trang Vu, Van-Anh Nguyen, Junae Kim and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
91%Must read
?Must readVote to see the score

Rethinking Personalized Generation: Test-time Alignment via Factorized Ranking Models

Test-time alignment via million-parameter factorized ranking models exploits massive headroom for personalized generation, outperforming billion-parameter reward models with minimal overhead.

Qiyao Ma, Junshan Zhang, Zhe Zhao

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 11 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

Inference Time Nash Alignment

Inference-time alignment via Nash equilibrium achieves lower-bound duality gaps and matches fine-tuned performance without parameter updates.

Hadi Hosseini, Debmalya Mandal, Duohan Zhang

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior

INNSteer learns invertible latent transformations to apply input-dependent nonlinear steering, improving LLM behavior control over linear baselines.

Tuc Nguyen, Thai Le

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
89%Must read
?Must readVote to see the score

Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning

Balanced Fine-Tuning uses dual-scale token and sequence reweighting targeting dense epistemic uncertainty to align LLMs with biomedical knowledge, improving reasoning and sparse-reward RL over standard fine-tuning.

Zhenchao Tang, Fang Wang, Haohuai He, Jiale Zhou and 12 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

SCOPE co-evolves a task-generating challenger and retrieval solver with rubric-based self-judging to improve open-ended and QA performance without curated data.

Wai-Chung Kwan, Aryo Gema, Joshua O Leang, Pasquale Minervini

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face · Code ★ 2

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Model Spec Midtraining: Improving How Alignment Training Generalizes

Model spec midtraining teaches models their behavior spec before alignment, controlling how demonstration fine-tuning generalizes and reducing agentic misalignment substantially.

Chloe Li, Sara Price, Samuel Marks, Jonathan Kutasov

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · Code ★ 74

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

LLM Active Alignment: A Nash Equilibrium Perspective

A game-theoretic framework predicts and steers LLM populations via Nash equilibrium analysis, deriving closed-form alignments that prevent political exclusion and guide socially desirable outcomes.

Tonghan Wang, Yuqi Pan, Xinyi Yang, Xinrui Song and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

G-Zero: Self-Play for Open-Ended Generation from Zero Data

G-Zero uses intrinsic predictive-shift rewards in a verifier-free co-evolutionary framework that enables continuous LLM self-improvement across open-ended unverifiable domains without external judges.

Chengsong Huang, Haolin Liu, Tong Zheng, Runpeng(Leo) Dai and 6 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 17 on Hugging Face · Code ★ 30

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models

Sycophancy is a boundary failure between social alignment and epistemic integrity, defined by three conditions involving cue expression, alignment shift, and compromised reasoning.

Jiechen Li, Catherine A Barry, Rishika Randev, Janet Chen and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

INSPO : Unlocking Intrinsic Self-Reflection for LLM Preference Optimization

InSPO derives a globally optimal preference policy conditioning on alternative responses, proving superiority to DPO while guaranteeing invariance to modeling choices and improving alignment.

Yu Li, Tian Lan, Zhengling Qi

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

AVSD separates cross-view consensus from privileged residuals in multi-view self-distillation to adaptively supervise reasoning models, improving math and code benchmarks over single-view methods and GRPO.

Duy Nguyen, Hanqi Xiao, Archiki Prasad, Zaid Khan and 6 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
Show 20 more papers