Good Papers

Showing papers from Stanford Show all papers

57%Worth a look
?Worth a lookVote to see the score

An exact information theory of generalization phase transitions in Bayesian diffusion models

Henry Hunt, Mason Kamb, Surya Ganguli

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Behaving Better, Thinking Worse: Sycophancy Across Post-Training Stages

Sonnet Xu, Kritika Singh, Sheharbano Jafry, Roxana Daneshjou and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

LLMs Keep Thinking When Told Not To

Dianqiao Lei, Kevin Qinghong Lin, Pan Lu, Philip Torr and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
69%Highly rated
?Highly ratedVote to see the score

Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

Junjue Wang, Weihao Xuan, Heli Qi, Pengyu Dai and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Accelerating the Inference Era with AI-Driven, Globally Optimized HW/SW Co-Design

Miria Feng, Fangzhao Zhang, Adrian G Lafuente, Mert Pilanci and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Decoupling Action from Egocentric Observation for World Simulation

Yue Ma, Pengjie Song, Xinyu Wang, Yi He and 9 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Pacing Branch Parallelism in LLM Serving

Swapnil Gandhi, Siva Kumar Sastry Hari, Bill Dally, Christos Kozyrakis

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Recreating Video Arenas via Automated Preference Scoring

Yue Zhao, Aniket Gupta, Juze Zhang, Tiange Xiang and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Evaluating Neural Data Tokenizers: A Framework for Assessing Learned Representations of Spiking Activity

Federico D'Agostino, Alex Gilbert, Susanne Keller, Jaivardhan Kapoor and 16 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Position: AI Development Should Prioritize Cognitive Security

Batu El, Shiye Su, Aneesh Pappu, Peggy Yin and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace

Simon Yu, Derek Chong, Ananjan Nandi, Dilara Soylu and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Hawkeye: Hardware-Aware GPU Kernel Optimization with Minimal Supervision

Arya Tschand, Kesavan Ramakrishnan, Alexander Ingare, Simon Guo and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Rethink Action Chunking in VLA Through Human Motor Control

Wenxi Chen, Yuejiang Liu, Zijian He, Shaoshuai Mou and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Multimodal AI Detection In Two Words

Daniel Fein, Arpita Singhal, Maneesh Agrawala, Maty Bohacek

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

TrunkFish: Making Model Width Incrementally Refinable

Owen M Dugan, Liam Dugan, Aaryan Singhal, Christopher De Sa and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Expert-guided Bayesian optimization for sustainable protein formulation

Anna Thomas, Georgios Zaverdinos, Petros Mandalis, Andreas Orfanoudakis and 8 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Sampling Is Not Curiosity: Why LLM Agents Should Investigate

Alfonso Amayuelas, Piotr Piękos, Xin Wang, William Yang Wang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Characterizing the Aesthetic Defaults of Generative Image Models

Maty Bohacek, Raina Panda, Daniel Fein, Arpita Singhal and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

Proposing intelligence per watt to evaluate local LLM inference, the study finds local models answer 88.7% of queries with 5.3x efficiency gains since 2023 but remain 1.4x less efficient than cloud accelerators.

Jon Saad-Falcon, Avanika Narayan, Hakki Akengin, J. W Griffin and 10 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 17 on Hugging Face · Code ★ 95

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
88%Must read
?Must readVote to see the score

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

P-Bench reveals LLM agents make subtle inferential errors in hypothesis testing, and Fisher-R1 improves reliability via reinforcement learning to outperform GPT-5.4 and DeepSeek-V4-Pro.

Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities

Long-horizon Q-learning penalizes n-step value bound violations via hinge losses to stabilize bootstrapping and outperform standard TD methods.

Armaan A Abraham, Lucy Xiaoyang Shi, Chelsea Finn

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
86%Must read
?Must readVote to see the score

HumanScore: Benchmarking Human Motions in Generated Videos

HumanScore benchmarks human motion in AI videos via six metrics, revealing gaps between visual plausibility and biomechanical fidelity across 13 models.

Tiange Xiang, Yusu Fang, Tian Tan, Narayan Schütz and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

FASTER: Value-Guided Sampling for Fast RL

FASTER traces test-time sampling gains to earlier denoising stages via a denoising-space MDP that filters action candidates early, reducing compute while improving diffusion-based RL policy performance.

Perry Dong, Alexander Swerdlow, Dorsa Sadigh, Chelsea Finn

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Learning When to Trust LLM Priors: A Validated Framework for Semantic Prior Integration

Statsformer validates LLM semantic priors via out-of-fold calibration to adaptively integrate them across diverse predictors, guaranteeing performance at least as good as the best convex combination of candidates.

Erica Zhang, Naomi Sagan, Danny Tse, Fangzhao Zhang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score

TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots

TherapyGym introduces CTRS-based fidelity and multi-label safety evaluation for therapy chatbots, with RL training raising expert-rated CBT adherence from 0.10 to 0.60.

Fangrui Huang, Souhad Chbeir, Arpandeep Khatua, Sheng Wang and 7 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

Online Set Learning from Precision and Recall Feedback

Online set learning with randomized precision or recall feedback is learnable exactly when the hypothesis class has finite VC dimension, though standard empirical risk minimization can fail and algorithms must handle feedback dependencies to achieve regret bounds.

Lee Cohen, Yishay Mansour, Shay Moran, Han Shao

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 2/5
83%Must read
?Must readVote to see the score

Sparse Reward Subsystem in Large Language Models

LLM hidden states contain sparse value and dopamine neurons forming a reward subsystem that predicts confidence and guides search.

Guowei Xu, Mert Yuksekgonul, James Zou

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Principled Design of Diffusion-based Optimizers for Inverse Problems

Principled diffusion optimizer reparameterizations induce task invariances for reuse without retuning, and the OptDiff pipeline integrates convex optimization to accelerate inference and improve quality.

Julio Oscanoa, Irmak Sivgin, Cagan Alkan, Daniel Ennis and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

A Bitter Lesson for Data Filtering

In high-compute, data-scarce pretraining, removing filters and using all data improves large models by letting them learn from nominally poor data.

Christopher Mohri, John Duchi, Tatsunori Hashimoto

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5