Good Papers

Showing papers from University of California, Berkeley Show all papers

71%Highly rated
?Highly ratedVote to see the score

OpenHands: An Open Platform for AI Software Developers as Generalist Agents

OpenHands is an open MIT-licensed platform for building AI software developers that evaluate agents on SWE-BENCH and WebArena benchmarks.

Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu and 20 more

Published Jul 23, 2024 · 15 citations · ▲ 90 on Hugging Face · Code ★ 90,160

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Multi-LLM collaboration detects knowledge gaps to make LLMs abstain from wrong answers instead of hallucinating.

Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding and 2 more

Published 2024 · 37 citations

– ReadersNo votes yet
8/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention applies OS virtual memory and paging to LLM key-value caches, reducing waste and duplication to boost vLLM throughput 2-4x over existing systems.

Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng and 5 more

Published Sep 12, 2023 · 50 citations · ▲ 76 on Hugging Face · Code ★ 86,094

– ReadersNo votes yet
16/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

WorldComposer: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation

Jasper Lu, Zhenhao Shen, Yuanfei Wang, Shugao Liu and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions

Junho Kim, Xu Cao, Houze Yang, Bikram Boote and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Why Heavy-Tailed Weights Predict Model Quality

Joseph Wilson, Chris van der Heide, Liam Hodgkinson, Zhichao Wang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Bridging the Simulation-to-Experiment Gap with Adversarial Distribution Alignment

Kai Nelson, Tobias Kreiman, Sergey Levine, Aditi Krishnapriyan

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

FlashPlanner: Real-Time Goal-Conditioned Flow-Matching Planning for Autonomous Driving with Online RL Fine-Tuning

Qifeng Li, Yubing Gao, Xiaosong Jia, Zhiliu Liu and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Assistive Dueling Bandits: No-Regret Algorithms for Assisting No-Regret Users

Mark Bedaywi, Cassidy Laidlaw, Austin Tripp, Nika Haghtalab

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

NTK Regression under dual power-law model: Deterministic Equivalents via SDE and PDE Methods

Collin Cranston, Zhichao Wang, Todd Kemp

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CME–SpectrumBench: Can LLMs Analyze Condensed Matter Spectral Data?

Jin Gene Wong, Anjney Midha, Joseph Tennyson, Wei-Lin Chiang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering with Generative Optimization

Dapeng Jiang, Yizhe Chi, Kaisen Yang, Tianwei Luo and 18 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Visual Grounding First, Multimodal In-context Learning Follows

Minhyuk Seo, Minjae Lee, Chaeeun Lee, Wei Lin and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Frontier Task Synthesis Via Solution-Centric Evolution

Yangzhen Wu, Aaron Li, Wenjie Ma, Li Cao and 9 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Attention Itself Could Retrieve. RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval

Zichen Zou, Xiaosong Jia, Zuxuan Wu, Yu-Gang Jiang

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SPECS: Faster Test-Time Scaling through Speculative Drafts and Dynamic Switching

Mert Cemri, Nived Rajaraman, Rishabh Tiwari, Xiaoxuan Liu and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

I-PTC: Interactive Programmatic Tool Calling for Stateful Tool-Augmented Agents

Huanzhi Mao, Chengkun Cao, Shuo Yuan, Joseph Gonzalez

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

LiFi: LiDAR Generation from Multi-View Images via Geometric and Semantic Collaborative Guidance

Sizhuo Zhou, Xiaosong Jia, Fanrui Zhang, Qifeng Li and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SudoBench: A Contextual Authorization Benchmark for LLM Agents

Vincent Siu, Tianneng Shi, Shangding Gu, Zhun Wang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Gradient Descent on Two ReLU Neurons: Global Landscape and Bifurcation Dynamics

Binghua Li, Mengzhe Li, Denny Wu, Tianhao Wang

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Embedded-Arena: Building Hardware-in-the-Loop Coding Agents to Run AI on Microcontrollers

Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu, Chun-Cheng Chang and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Approximation Algorithms for GPU Pricing under Finite Capacity

Yaolong Yu, Hanrui Zhang, Zeyu Zheng, Kirthevasan Kandasamy

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Prediction-only distillation with optimal mixing in ridge-regularized linear and logistic regression

Hien Dang, Pratik Patil, Alessandro Rinaldo

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

FusionNeXt: Sequence-First 3D Multi-Modal Fusion in the Era of LLMs

Yu Hong, Xiaosong Jia, Songbur Wong, Yanhao Liu and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Do We Really Need Diffusion for Generative Object Detection? A Minimal Prototype Perspective

Yu Hong, Xiaosong Jia, Yihan Wang, Wenlong Liao and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Forced Orders: What LLM Leaderboards Hide About Model Comparisons

Zonglin Di, Berk Ustun, Yang Liu

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Prototypes of the Mind: A Unified Framework for Probing the Visual Brain

Shi Chen, Seoyoung Ahn, Doris Tsao

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Position: A Safe LLM and a Safe Harness Do Not Make a Safe Agent

Vincent Siu, Kyle Montgomery, Yujin Potter, Zhun Wang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Verify0: Can AI Agents Build Formally Verified Software Repositories?

Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

On the Bias of Group-Based Advantage Estimation

Fengkai Yang, Zherui Chen, Xiaohan Wang, Xiaodong Lu and 8 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

StyleStream 2.0: Fast and Controllable Streaming Voice Style Conversion

Yisi Liu, Nicholas Lee, Gopala Anumanchipalli

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Context-Aware Autoregressive Image Generation for Emerging Reasoning Properties

Jixuan Ying, Haoyu Liu, Timing Yang, Tingyu Zhu and 6 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Weird Generalization from Narrow Finetuning: Persona Shifts and Inductive Backdoors

Jan Betley, Jorio Cocola, Dylan Feng, James Chua and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

TabBioMed: A Large-Scale Benchmark for Biomedical Tabular Learning

Pau Mateo Bernadó, Saivenkata Nagavyjayanthi Polapragada, Pol Arbiol Rakuljic, Laia M Pladevall and 7 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

When Does Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Luca Viano, Antoine Moulin, Audrey Huang, Volkan Cevher and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning Process Rewards via Visitation Matching for Efficient RL

Raymond Tsao, Andrew Wagenmaker, Sergey Levine

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

WorldForge: Forging Unified World Modeling into Video Generation

Boming Tan, Xiangdong Zhang, Ning Liao, Jingtao Zhang and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

On Differential Private $\ell_1$, $\ell_2$ and $\ell_p^p$ Distance Queries

Erzhi Liu, Jerry Yao-Chieh Hu, Alex Reneau, Zhao Song and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

From Chats to Markets: AgenticPay for LLM-Powered Negotiation in Multi-Agent Commerce

Xianyang Liu, Shangding Gu, Fan Xu, Manxi Wu and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Ad-Hoc Teamwork from Human Demonstrations

Darius Muglich, Niklas Lauffer, Tin Dizdarević, Jakob Foerster

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Strategic Feature Selection and Regularization

Jivat Neet Kaur, Pratik Patil, Divya Shanmugam, Emma Pierson and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

What Is Worth Representing? Representational Empowerment for Continual Model Construction

Fei Dai, Hanqi Zhou, Alison Gopnik, Charley M Wu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Fast Accurate Quantum Monte Carlo without Metropolis Adjustment

Reuben Cohn-Gordon, Gabriel Pescia, Sumner N Hearth, Jakob Robnik and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Prompt-Driven Exploration

Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed and 5 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Spectral Estimation with Deformed Decompression

Siavash Ameli, Chris van der Heide, Liam Hodgkinson, Michael Mahoney

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Self-Recognition Finetuning can Reverse and Prevent Emergent Misalignment

Arush Tagade, Shaoheng Zhou, Jiaxin Wen, Shi Feng

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

PM1: A Multimodal Foundation Model for Genomes, Phenotypes, and Images at Biobank Scale

Christophe Thomassin, Marçal Comajoan Cara, Margarita Geleta, David Bonet and 3 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Correct, Route, Calibrate: Efficient Preference Optimization from Noisy, Heterogeneous Human Feedback

Zhongming Xie, Xinwei Ma, Jingshen Wang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Demystifying Numerical Errors in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL

Zhenting Zhu, Lucas Thai, Shan Yu, Yicheng Liu and 4 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Long-Context Language Models Require Extreme Sparsity in Context Dimension

Prithvi Dixit, Sahil Joshi, Agniva Chowdhury, Anshumali Shrivastava and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

Transformers Provably Learn Graph Search: Training Dynamics and the Exponential Power of Depth

Xutao Ma, Somayeh Sojoudi

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist Rewards

Yuanhao Ban, Tong Xie, Sohyun An, Yunqi Hong and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

Siddharth Gollapudi, Prasann Singhal, Nilesh Gupta, Sewon Min

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Time to Pay Attention! Understanding High Complexity Corpus Reasoning Tasks

Prasann Singhal, Amanda Bertsch, Jacob Steinhardt, Sewon Min

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Beyond the Training Distribution: Evaluating Predictions Under Distribution Shift and Selection Bias

A double machine learning procedure estimates black-box prediction risk under covariate shift and selective labels, tracking true target risk more accurately than single-source methods.

Annie Ulichney, Amanda L Coston

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Looped Diffusion Language Models

Selective layer looping improves masked diffusion model training efficiency and reasoning performance via depth scaling without added parameters and flexible inference compute scaling.

Sanghyun Lee, Chunsan Hong, Seungryong Kim, Jonghyun Lee and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

MOOD benchmark shows guard models fail to detect out-of-distribution alignment failures, but combining them with Mahalanobis and perplexity detectors improves recall from 39% to 45% and scales positively.

Dylan Feng, Pragya Srivastava, Anca Dragan, Cassidy Laidlaw

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Response Time Enhances Alignment with Heterogeneous Preferences

Adding response times to preference data via drift-diffusion modeling restores identifiability of average preferences among anonymous heterogeneous labelers, correcting choice-only estimation bias without tracking users.

Federico Echenique, Alireza Fallah, Baihe Huang, Michael Jordan

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Learning to Discover Iterative Spectral Algorithms

AutoSpec is a neural framework that discovers iterative spectral algorithms via self-supervised prediction of recurrence coefficients, yielding order-of-magnitude speedups over classical baselines.

Zihang Liu, Oleg Balabanov, Yaoqing Yang, Michael Mahoney

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Spectral Feedback for Test-Time Alignment of Protein Diffusion Models

Spectral Feedback iteratively selects protein tokens to re-mask and resample using sparse Fourier edit-set value functions, improving inverse-folding stability by up to 32.3% at test time.

Shai Dickman, Mert Cemri, Landon Butler, Kannan Ramchandran

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

On Neural Scaling Laws for Weather Emulation through Continual Training

Minimal Swin Transformer weather emulators trained with continual training and periodic cooldowns follow predictable neural scaling laws, outperform cosine schedules, and reveal compute-optimal regimes.

Shashank Subramanian, Alexander Kiefer, Arnur Nigmetov, Amir Gholami and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · Code ★ 3

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation

Animation2Code benchmarks video-to-code generation for web animations, showing state-of-the-art vision-language models struggle with temporal consistency despite high appearance fidelity.

Anya Ji, Abhijith Varma Mudunuri, David Chan, Alane Suhr

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories

World Motion Models unify dynamic 3D entities via sparse SE(3) trajectories and flow-matching, enabling any-to-any conditioning across prediction, control, and retargeting tasks.

Jiahui Lei, Qianqian Wang, Trevor Darrell, Angjoo Kanazawa

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries

VisInteract introduces interactive text-to-visualization with imperfect queries via VisInteract-Bench and Vis-MCTS, boosting success by over 13% versus interactive baselines.

Wenxin XU, Jinwei Lu, Hwanhee Kim, Chen J Zhang and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling

Recon scores reasoning traces by action reconstruction fidelity to avoid post-hoc rationalization in user modeling, yielding up to 70% win rates over baselines across domains.

Alan Zhu, Mihran Miroyan, Carolyn Wang, Andrew Zhou and 3 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

LLMs diverge from human goal selection in self-directed learning by exploiting single solutions with low variability instead of gradual diverse exploration.

Gaia Molinaro, Dave August, Danielle Perszyk, Anne Collins

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

PieArena benchmarks LLM negotiation via multi-agent MBA scenarios, ranking agents with order-invariant payoffs and finding GPT-5 matches trained human baselines while profiling cross-model behavioral heterogeneity.

Chris Zhu, Sasha Cui, Will S Dufallo, Runzhi Jin and 3 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Flash-KMeans: Fast and Memory-Efficient Exact K-Means

Flash-KMeans eliminates GPU HBM bottlenecks via fused assignment and inverse mapping updates, delivering up to 17.9x speedups over existing exact k-means implementations.

Shuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li and 9 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 85 on Hugging Face · Code ★ 735

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots

TherapyGym introduces CTRS-based fidelity and multi-label safety evaluation for therapy chatbots, with RL training raising expert-rated CBT adherence from 0.10 to 0.60.

Fangrui Huang, Souhad Chbeir, Arpandeep Khatua, Sheng Wang and 7 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

A clustering-based divergence method measures gaps between real and simulated user behaviors, finding large, family-dependent discrepancies reducible by combining complementary simulators.

Shuhaib Mehri, Philippe Laban, Sumuk Shashidhar, Marwa Abdulhai and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
91%Must read
?Must readVote to see the score

mRNABench: A curated benchmark for mature mRNA property and function prediction

mRNABench benchmarks mature mRNA property predictions across 59 tasks and 135K experiments, revealing synergies between self-supervised objectives that yield a compact state-of-the-art Mamba model using 700x fewer parameters.

Ruian (Ian) Shi, Taykhoom Dalal, Philip Fradkin, Divya Koyyalagunta and 9 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
86%Must read
?Must readVote to see the score

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

LaMo extracts self-supervised latent motion priors from unlabeled videos via motion drift loss and prior guidance, improving physical consistency in video diffusion without external supervision.

Bo Jiang, Depu Meng, yihan hu, Yichen Xie and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

ReMind trains video diffusion transformers to use cache memory for evolving hidden states across interruptions via memory-oriented curricula and PM-RoPE, achieving best STEVO-Bench scores without catastrophic forgetting.

Tianshuo Xu, Yichen Xie, Depu Meng, Chensheng Peng and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Continuity Laws for Sequential Models

State-space sequential models vary in temporal continuity, with continuous behavior aligning to task structure and enabling efficient subsampling.

Annan Yu, Dongwei Lyu, N. Benjamin Erichson

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Neuron Populations Exhibit Divergent Selectivity with Scale

Rosetta neuron populations grow sublinearly and become more selective and specialized as language and vision models scale, while non-Rosetta neurons stay less selective.

Amil Dravid, Yasaman Bahri, Alexei Efros, Yossi Gandelsman

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning

DiscoLoop combines discrete embeddings and continuous hidden states in looping transformers to fix representational misalignment, enabling near-perfect multi-hop reasoning with faster training and stronger pretraining performance.

Hengyu Fu, Tianyu Guo, Zixuan Wang, Hanlin Zhu and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Interleaved Head Attention

Interleaved Head Attention mixes attention heads via pseudo-heads to enable cross-head reasoning, cutting parameters on synthetic tasks and improving retrieval and math benchmarks over standard multi-head attention.

Sai Surya Duvvuri, Chanakya Ekbote, Rachit Bansal, Rishabh Tiwari and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
70%Highly rated
?Highly ratedVote to see the score

Super-Level-Set Regression: Conditional Quantiles via Volume Minimization

Super-level-set regression directly optimizes minimum-volume prediction regions via geometric optimization, bypassing full conditional density estimation to capture complex multimodal conditional structures.

Sacha Braun, Michael Jordan, Francis Bach

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 2/5
medium 3/10
strict 0/5
91%Must read
?Must readVote to see the score

CalArena: A Large Scale Post-Hoc Calibration Benchmark

CalArena benchmarks nearly 2000 post-hoc calibration experiments, finding smooth methods outperform binning and multiclass-specific designs are essential.

Eugène Berta, David Holzmüller, Francis Bach, Michael Jordan

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
86%Must read
?Must readVote to see the score

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

Simple norm enforcement for AI agents is exploited for competitive gain, but mechanisms tracking reliability with escalating penalties resist exploitation across multi-agent environments.

Yaowen Ye, Jacob Steinhardt

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
86%Must read
?Must readVote to see the score

Explaining and Preventing Alignment Collapse in Iterative RLHF

Iterative RLHF ignores policy influence on reward-model updates, causing alignment collapse via exploited blind spots; foresighted optimization restores this term to prevent collapse.

Etienne Gauthier, Francis Bach, Michael Jordan

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning

Extracting search trees from LLM reasoning traces reveals myopic planning where performance depends on breadth rather than depth, unlike human planning.

Sixing Chen, Ji-An Li, Saner Cakir, Sinan Akcali and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Nearest-Neighbor Radii under Dependent Sampling

Nearest-neighbor radii under mixing dependence converge almost surely with polynomial mixing and have sharp moment bounds scaling with local intrinsic dimension, remaining informative for high-dimensional dependent data.

Yuanyuan Gao, Yilong Hou, Zhexiao Lin

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 2/5
89%Must read
?Must readVote to see the score

PitchBench: Measuring Pitch Hearing in Audio-Language Models

PitchBench evaluates pitch hearing in audio-language models via 28 experiments, finding their pitch perception remains highly unreliable across instruments and acoustic conditions.

Milan Liessens Dujardin, Song-Ze Yu, Craver C Thomas-Smith, David Chan and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
80%Must read
?Must readVote to see the score

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

A benchmark of likely public-domain Hollywood films tests multimodal models on narrative understanding, finding vision-language models near chance and audio-visual models below human performance.

David Bamman, Kent K Chang, Allison Cooper, Juishan Hsu and 6 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling

Feedforward NVS transformers decouple semantic and spatial tokens to eliminate spatial bias in appearance representation, improving rendering fidelity with negligible inference overhead.

Yihang Wu, Yihang Sun, Shaofeng Zhang, Zuxuan Wu and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Pooling Versus Ensembling for Ridge Regression Under Covariate Shift

Under covariate shift, pooled ridge regression outperforms ensembling for random-effects ridge models, and fixed-effects risk formulas characterize partition-driven predictor shifts.

Maya Ramchandran, Rajarshi Mukherjee

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
86%Must read
?Must readVote to see the score

Aligning Language Model Benchmarks with Pairwise Preferences

Reweighting benchmark items aligns static language model benchmarks with downstream pairwise preferences to rank unseen models, using as few as 20 well-chosen models.

Marco Gutierrez, Xinyi Leng, Hannah Chen, Jonathan Richard Schwarz and 2 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Learning to Recommend in Unknown Games

Quantal-response feedback yields logarithmic sample-complexity utility learning up to affine equivalence, while best-response feedback permits only partial identification; an online algorithm achieves low deviation-regret under both models.

Arwa Alanqary, Zakaria Baba, Manxi Wu, Alexandre Bayen

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 3/5
89%Must read
?Must readVote to see the score

MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks

MT-JailBench provides a modular framework for comparing multi-turn jailbreak attacks under standardized conditions, finding that prompt generation drives success and recomposed components yield stronger attacks.

Xinkai Zhang, Zhipeng Wei, Huanli Gong, Jing Ting Zheng and 3 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
71%Highly rated
?Highly ratedVote to see the score

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures

A hierarchical sparse autoencoder architecture explicitly models semantic concept hierarchies, improving reconstruction, interpretability, and efficiency in language model representations.

Mark Muchane, Sean M Richardson, Kiho Park, Victor Veitch

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 0/5
91%Must read
?Must readVote to see the score

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

MLS-Bench evaluates AI agents on inventing scalable ML methods across 140 tasks, finding current systems fail to reliably surpass human-designed approaches due to insufficient scientific validation insight.

Bohan Lyu, Yucheng Yang, Siqiao Huang, Jiaru Zhang and 24 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 120

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
71%Highly rated
?Highly ratedVote to see the score

Free Decompression with Algebraic Spectral Curves

Algebraic spectral curves extend free decompression to multi-scale, multi-modal, and atomic spectral densities, enabling realistic neural network and diffusion model extrapolation.

Siavash Ameli, Chris van der Heide, Liam Hodgkinson, Michael Mahoney

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 0/5
92%Must read
?Must readVote to see the score

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

MCP-Atlas benchmarks LLM tool-use on 1,000 real-server tasks, finding frontier models reach 82.2% pass rates but 63.3% of failures are cognitive.

Chaithanya Bandi, Razvan Dumitru, Ben Hertzberg, Divyansh Agarwal and 15 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
78%Highly rated
?Highly ratedVote to see the score

Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents

ACuRL enables autonomous continual learning for computer-use agents via curriculum reinforcement learning, yielding 3-29% gains without catastrophic forgetting or human data.

Tianci Xue, Zeyi Liao, Tianneng Shi, Zilu Wang and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5