Good Papers

Showing papers from UC Berkeley Show all papers

45%Niche pick
?Niche pickVote to see the score

Why Heavy-Tailed Weights Predict Model Quality

Joseph Wilson, Chris van der Heide, Liam Hodgkinson, Zhichao Wang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Bridging the Simulation-to-Experiment Gap with Adversarial Distribution Alignment

Kai Nelson, Tobias Kreiman, Sergey Levine, Aditi Krishnapriyan

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Assistive Dueling Bandits: No-Regret Algorithms for Assisting No-Regret Users

Mark Bedaywi, Cassidy Laidlaw, Austin Tripp, Nika Haghtalab

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

RAG in a Trenchcoat: When Minimal Memory Is Enough for Agentic Systems, and When It Isn’t

Jingyu Liu, Zongze Li, Zach Xu, Zhanhui Zhou and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

NitroBox: Lightning-Fast Sandbox for Large-Scale RL Training

Yuzhou Nie, Ruilin Zhou, Zhaorun Chen, Jingyang Zhang and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

CME–SpectrumBench: Can LLMs Analyze Condensed Matter Spectral Data?

Jin Gene Wong, Anjney Midha, Joseph Tennyson, Wei-Lin Chiang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering with Generative Optimization

Dapeng Jiang, Yizhe Chi, Kaisen Yang, Tianwei Luo and 18 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Treat Bias as Noise: Training Bias-Robust LLM Reasoning via Reinforcement Learning

Qian Wang, Xuandong Zhao, Zirui Zhang, Zhanzhi Lou and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Frontier Task Synthesis Via Solution-Centric Evolution

Yangzhen Wu, Aaron Li, Wenjie Ma, Li Cao and 9 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SPECS: Faster Test-Time Scaling through Speculative Drafts and Dynamic Switching

Mert Cemri, Nived Rajaraman, Rishabh Tiwari, Xiaoxuan Liu and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

I-PTC: Interactive Programmatic Tool Calling for Stateful Tool-Augmented Agents

Huanzhi Mao, Chengkun Cao, Shuo Yuan, Joseph Gonzalez

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SudoBench: A Contextual Authorization Benchmark for LLM Agents

Vincent Siu, Tianneng Shi, Shangding Gu, Zhun Wang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Prototypes of the Mind: A Unified Framework for Probing the Visual Brain

Shi Chen, Seoyoung Ahn, Doris Tsao

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Is Decentralized LLM Agent RL Robust to Heterogeneity? An Asymmetric Tale

Canyu Chen, Kangyu Zhu, Zhaorun Chen, Zhanhui Zhou and 5 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Position: A Safe LLM and a Safe Harness Do Not Make a Safe Agent

Vincent Siu, Kyle Montgomery, Yujin Potter, Zhun Wang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 4 reviewers recommend it
lenient 0/4
67%Highly rated
?Highly ratedVote to see the score

IntegrityBench: Can LLMs Be Trusted as Co-Scientists? A Research Integrity Benchmark

Sai Sidhanth Manoharan Jayanthi, Yash Tripathi, Silu Sharma, Shivank Garg and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 16 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/1
45%Niche pick
?Niche pickVote to see the score

Prompt Optimization Makes Misalignment Legible

Caleb Biddulph, Micah Carroll

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Phase Transitions in Heavy-Tailed Mean Estimation under $\ell_p$ Norms

Ishaq Aden-Ali, Yeshwanth Cherapanamjeri, Mikael Møller Høgsgaard, Kasper Green Larsen and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning Process Rewards via Visitation Matching for Efficient RL

Raymond Tsao, Andrew Wagenmaker, Sergey Levine

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Aligning AI Teams

Siddarth Srinivasan, Morgan J Matthews, Jascha Sohl-Dickstein, Erik Jones

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

From Chats to Markets: AgenticPay for LLM-Powered Negotiation in Multi-Agent Commerce

Xianyang Liu, Shangding Gu, Fan Xu, Manxi Wu and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Long-Term Risks of Risk-Based Allocation

Jivat Neet Kaur, Jane Lee, Manolis Zampetakis

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Strategic Feature Selection and Regularization

Jivat Neet Kaur, Pratik Patil, Divya Shanmugam, Emma Pierson and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
70%Highly rated
?Highly ratedVote to see the score

Bigger Isn’t Better: Why the Indiscriminate Scaling of Foundation Models Can’t Solve Biology

Kathryne Metcalf, Lorin Crawford, Mary L Gray, Kevin K Yang and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Evolving Agent Teams

Shiyi Cao, Ziming Mao, Dacheng Li, Joseph Gonzalez and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Generative Structure from Motion with Native 3D Diffusion

Haobin Duan, Binbin Huang, Mochu Xiang, Yiqun Zhao and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Spectral Estimation with Deformed Decompression

Siavash Ameli, Chris van der Heide, Liam Hodgkinson, Michael Mahoney

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Long-Context Language Models Require Extreme Sparsity in Context Dimension

Prithvi Dixit, Sahil Joshi, Agniva Chowdhury, Anshumali Shrivastava and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist Rewards

Yuanhao Ban, Tong Xie, Sohyun An, Yunqi Hong and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

Siddharth Gollapudi, Prasann Singhal, Nilesh Gupta, Sewon Min

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Time to Pay Attention! Understanding High Complexity Corpus Reasoning Tasks

Prasann Singhal, Amanda Bertsch, Jacob Steinhardt, Sewon Min

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

Gen-Searcher: Reinforcing Agentic Search for Image Generation

Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.

Kaituo Feng, Manyuan Zhang, Shuang Chen, Yunlong Lin and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 54 on Hugging Face · Code ★ 400

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory

Spectral optimizer Muon exceeds SGD associative memory capacity, matching Newton's method with first-order updates and larger critical batch sizes.

Juno Kim, Eshaan Nichani, Denny Wu, Alberto Bietti and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models

Equilibrium Matching learns implicit energy landscapes for optimization-based sampling, surpassing diffusion models with 1.90 FID on ImageNet 256x256 while supporting denoising, OOD detection, and composition.

Runqian (Ray) Wang, Yilun Du

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 217

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning

Equilibrium Forcing removes noise conditioning from video diffusion to enable adaptive closed-loop inference that improves generation quality and consistency.

Hansen Lillemark, Alex Rojas, Zachary Novack, Runqian (Ray) Wang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 5/10
strict 1/5
86%Must read
?Must readVote to see the score

Learning to Discover Iterative Spectral Algorithms

AutoSpec is a neural framework that discovers iterative spectral algorithms via self-supervised prediction of recurrence coefficients, yielding order-of-magnitude speedups over classical baselines.

Zihang Liu, Oleg Balabanov, Yaoqing Yang, Michael Mahoney

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
83%Must read
?Must readVote to see the score

Spectral Feedback for Test-Time Alignment of Protein Diffusion Models

Spectral Feedback iteratively selects protein tokens to re-mask and resample using sparse Fourier edit-set value functions, improving inverse-folding stability by up to 32.3% at test time.

Shai Dickman, Mert Cemri, Landon Butler, Kannan Ramchandran

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
83%Must read
?Must readVote to see the score

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization

AlphaQ allocates MoE quantization bits without calibration using heavy-tailed spectral analysis, outperforming calibration-based methods and achieving near full-precision accuracy at 3.5-bit average precision.

Wanqi Yang, Yuexiao Ma, Alexander Conzelmann, Xiawu Zheng and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
89%Must read
?Must readVote to see the score

Learning to Follow In-Context Watermark Instructions via Self-Distillation

ICWBench reveals current LLMs fail at in-context watermarking, and self-distillation with reinforcement learning raises watermark detectability near perfect while preserving quality.

Yepeng Liu, Tianyi Chen, Xuandong Zhao, Dawn Song and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 5/5
78%Highly rated
?Highly ratedVote to see the score

Learning What to Remember: Test-Time Training via Context Distillation

TTCD uses a long-window teacher to supervise a short-window student's fast weights via hidden-state discrepancy, allocating limited memory to future-relevant context and outperforming existing long-context methods with minimal architectural changes.

Zixuan Wang, Xingyu Dang, Rui-Jie Zhu, Zixin Wen and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

On Neural Scaling Laws for Weather Emulation through Continual Training

Minimal Swin Transformer weather emulators trained with continual training and periodic cooldowns follow predictable neural scaling laws, outperform cosine schedules, and reveal compute-optimal regimes.

Shashank Subramanian, Alexander Kiefer, Arnur Nigmetov, Amir Gholami and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · Code ★ 3

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Two-Fidelity Best-Action Identification for Stochastic Minimax Tree

2FFS adaptively combines cheap biased heuristics and expensive accurate rollouts to identify best actions in stochastic minimax trees with fewer samples than baselines.

Peter Chen, Xi Chen

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation

Animation2Code benchmarks video-to-code generation for web animations, showing state-of-the-art vision-language models struggle with temporal consistency despite high appearance fidelity.

Anya Ji, Abhijith Varma Mudunuri, David Chan, Alane Suhr

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 2/5
88%Must read
?Must readVote to see the score

Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling

Recon scores reasoning traces by action reconstruction fidelity to avoid post-hoc rationalization in user modeling, yielding up to 70% win rates over baselines across domains.

Alan Zhu, Mihran Miroyan, Carolyn Wang, Andrew Zhou and 3 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
86%Must read
?Must readVote to see the score

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

SCDBench benchmarks LLM smart-contract decompilers on 600 real contracts via semantic replay, finding even top models perfectly recover only 42 and same-model repair substantially helps.

Kaihua Qin, Dawn Song, Arthur Gervais

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
83%Must read
?Must readVote to see the score

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

AstraFlow is a dataflow-oriented RL system for agentic LLMs that decouples rollout, dataflow, and training to enable multi-policy collaborative training with 2.7x faster training.

Haizhong Zheng, Yizhuo Di, Jiahui Wang, Shuowei Jin and 6 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 15 on Hugging Face · Code ★ 105

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
91%Must read
?Must readVote to see the score

MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM

MAGE uses block-diffusion's aligned all-[MASK] queries to select reusable sparse KV subsets, achieving near-lossless accuracy with up to 6.82x speedup at 128K context.

Omin Kwon, Yeonjae Kim, Doyeon Kim, Minseo Kim and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
86%Must read
?Must readVote to see the score

GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators

GenEnv co-evolves LLM agents with generative simulators via difficulty-aligned curricula, improving 7B agents by up to 40.3% with 3.3x less data.

Jiacheng Guo, Ling Yang, Peter Chen, Qixin Xiao and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 19 on Hugging Face · Code ★ 67

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
86%Must read
?Must readVote to see the score

Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

A clustering-based divergence method measures gaps between real and simulated user behaviors, finding large, family-dependent discrepancies reducible by combining complementary simulators.

Shuhaib Mehri, Philippe Laban, Sumuk Shashidhar, Marwa Abdulhai and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
88%Must read
?Must readVote to see the score

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

BankerToolBench benchmarks AI agents on multi-hour investment banking workflows using expert rubrics, finding frontier models fail nearly half of criteria with zero client-ready outputs.

Elaine Lau, Markus Dücker, Ronak Chaudhary, Hui Wen Goh and 24 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
86%Must read
?Must readVote to see the score

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

LaMo extracts self-supervised latent motion priors from unlabeled videos via motion drift loss and prior guidance, improving physical consistency in video diffusion without external supervision.

Bo Jiang, Depu Meng, yihan hu, Yichen Xie and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
83%Must read
?Must readVote to see the score

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

ReMind trains video diffusion transformers to use cache memory for evolving hidden states across interruptions via memory-oriented curricula and PM-RoPE, achieving best STEVO-Bench scores without catastrophic forgetting.

Tianshuo Xu, Yichen Xie, Depu Meng, Chensheng Peng and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 1/5
86%Must read
?Must readVote to see the score

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

A dataset of multi-aspect human visual similarity judgments benchmarks vision-language models and yields the TPIPS metric, which aligns with human perception and enables text-guided image retrieval and generative evaluation.

Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
80%Must read
?Must readVote to see the score

Neuron Populations Exhibit Divergent Selectivity with Scale

Rosetta neuron populations grow sublinearly and become more selective and specialized as language and vision models scale, while non-Rosetta neurons stay less selective.

Amil Dravid, Yasaman Bahri, Alexei Efros, Yossi Gandelsman

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
71%Highly rated
?Highly ratedVote to see the score

Offline Materials Optimization with CliqueFlowmer

CliqueFlowmer fuses clique-based offline model-based optimization into flow transformers for materials discovery, generating materials that strongly outperform generative baselines.

Jakub Grudzien Kuba, Benjamin K Miller, Sergey Levine, Pieter Abbeel

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 17

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
89%Must read
?Must readVote to see the score

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.

Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
86%Must read
?Must readVote to see the score

DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning

DiscoLoop combines discrete embeddings and continuous hidden states in looping transformers to fix representational misalignment, enabling near-perfect multi-hop reasoning with faster training and stronger pretraining performance.

Hengyu Fu, Tianyu Guo, Zixuan Wang, Hanlin Zhu and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
91%Must read
?Must readVote to see the score

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling

M²RNN introduces matrix-valued non-linear RNNs that scale via state expansion, achieving perfect state tracking and outperforming hybrid models with smaller states.

Mayank Mishra, Shawn Tan, Ion Stoica, Joseph Gonzalez and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

Diffusion Guidance Is a Controllable Policy Improvement Operator

CFGRL links diffusion guidance to policy improvement, training via supervised learning to exceed dataset performance on offline RL without value functions.

Kevin Frans, Seohong Park, Pieter Abbeel, Sergey Levine

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · Code ★ 123

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Nearest-Neighbor Radii under Dependent Sampling

Nearest-neighbor radii under mixing dependence converge almost surely with polynomial mixing and have sharp moment bounds scaling with local intrinsic dimension, remaining informative for high-dimensional dependent data.

Yuanyuan Gao, Yilong Hou, Zhexiao Lin

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Balancing Frequencies and Pixels in Flow Matching

Focal log-frequency loss balances spectral learning signals in flow matching, accelerating convergence by 40% and improving image fidelity without architectural changes.

Lucas Degeorge, Paul Couairon, Arijit Ghosh, Alexei Efros and 2 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

A benchmark of likely public-domain Hollywood films tests multimodal models on narrative understanding, finding vision-language models near chance and audio-visual models below human performance.

David Bamman, Kent K Chang, Allison Cooper, Juishan Hsu and 6 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space

NLAC trains LLM agents with a natural-language generative critic for off-policy learning, yielding richer feedback and more stable, data-efficient training than policy gradients in long-horizon tasks.

Joey Hong, Kang Liu, Zhan Ling, Jiecao Chen and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Free Decompression with Algebraic Spectral Curves

Algebraic spectral curves extend free decompression to multi-scale, multi-modal, and atomic spectral densities, enabling realistic neural network and diffusion model extrapolation.

Siavash Ameli, Chris van der Heide, Liam Hodgkinson, Michael Mahoney

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents

ACuRL enables autonomous continual learning for computer-use agents via curriculum reinforcement learning, yielding 3-29% gains without catastrophic forgetting or human data.

Tianci Xue, Zeyi Liao, Tianneng Shi, Zilu Wang and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5