Good Papers

Showing papers from Stanford University / Virtue AI Show all papers

69%Highly rated
?Highly ratedVote to see the score

Behaving Better, Thinking Worse: Sycophancy Across Post-Training Stages

Sonnet Xu, Kritika Singh, Sheharbano Jafry, Roxana Daneshjou and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency

Yibo Jacky Zhang, Zeyu Tang, Sanmi Koyejo

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Strategic Evaluation: Incentivizing AI Capability Coverage with Private Benchmarks

Sang Truong, Serena Wang, Nick Haber, Sanmi Koyejo

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

What to Forget in Unlearning? Forget Set Curation for Language Models

Animesh Jha, Arpandeep Khatua, Youssef Allouah, Sanmi Koyejo

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Benchmarks as Measurement Instruments: Quantifying Signal and Noise for More Efficient AI Evaluations Under Distribution Shift

Michael Hardy, Anka Reuel-Lamparth, Jodi Casabianca, Hansol Lee and 3 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
86%Must read
?Must readVote to see the score

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

Benchmark aggregation is modeled as a principal-agent game where welfare loss depends on item alignment, improvability, and variance; auditing OLMES reveals Pareto-inferior items under pro-worker welfare.

Andreas Haupt, Justin Hartenstein, Anka Reuel-Lamparth, Mykel J Kochenderfer and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Efficient Prediction of Pass@k Scaling in Large Language Models

Standard pass@k scaling laws suffer statistical shortcomings, so a beta-binomial framework and dynamic sampling strategy more accurately predict rare LLM capabilities and risks from limited data.

Joshua Kazdan, Colin Sullivan, Rylan Schaeffer, Youssef Allouah and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion

RADS applies reachability analysis and constrained reinforcement learning to steer diffusion trajectories away from memorized outputs via caption embedding perturbations, improving diversity, quality, and alignment without altering the model.

Sathwik Karnik, Juyeop Kim, Sanmi Koyejo, Jong-Seok Lee and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

AI Evaluation Should Require Standardized Item-Level Data Releases

Standardized item-level benchmark releases should become AI evaluation infrastructure because aggregate scores obscure validity failures; OpenEval archives 10M responses to enable auditability and recover benchmark validity evidence.

Han Jiang, Susu Zhang, Dongyao Zhu, Yuzhuo Bai and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots

TherapyGym introduces CTRS-based fidelity and multi-label safety evaluation for therapy chatbots, with RL training raising expert-rated CBT adherence from 0.10 to 0.60.

Fangrui Huang, Souhad Chbeir, Arpandeep Khatua, Sheng Wang and 7 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
88%Must read
?Must readVote to see the score

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation

PU-DPO treats unmentioned radiology findings as unlabeled rather than negative, using edited contrastive pairs to prevent omission noise from corrupting preference optimization and improving hidden finding recovery.

Yuta Kobayashi, Pradyun Ramesh, Muhammad Ahmed Chaudhry, Vincent Jeanselme and 4 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5