Good Papers

Showing papers from UC Berkeley, Transluce AI Show all papers

45%Niche pick
?Niche pickVote to see the score

Time to Pay Attention! Understanding High Complexity Corpus Reasoning Tasks

Prasann Singhal, Amanda Bertsch, Jacob Steinhardt, Sewon Min

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Training Language Models to Explain Their Own Computations

Fine-tuning language models on interpretability ground truth teaches them to describe their internal computations, with self-explanation outperforming larger external explainers.

Belinda Z Li, Zifan Carl Guo, Vincent Huang, Jacob Steinhardt and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 38

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
86%Must read
?Must readVote to see the score

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

Simple norm enforcement for AI agents is exploited for competitive gain, but mechanisms tracking reliability with escalating penalties resist exploitation across multi-agent environments.

Yaowen Ye, Jacob Steinhardt

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants

Predictive Concept Decoders train end-to-end interpretability assistants that encode neural activations into sparse concepts to predict model behavior, scaling with data to detect jailbreaks, hidden hints, and latent attributes.

Vincent Huang, Dami Choi, Daniel D Johnson, Sarah Schwettmann and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5