Good Papers

Showing papers from Anthropic Show all papers

45%Niche pick
?Niche pickVote to see the score

LoRAcles: Self-Supervised Weight-Space Interpretability at Scale

Celeste De Schamphelaere, Jan Bauer, Neel Nanda, Euan Ong

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Gradient Routing Localizes and Removes Unintended Behaviors in RL

Jake Ward, Shawn Hu, Aria Wong, Nathan Hu and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CryptanalysisBench: Can LLMs do cryptanalysis?

Lukas Fluri, Avital Shafran, Nicholas Carlini, Matthew Jagielski and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Introspective Coupling: LMs Learn to Explain Themselves Better Than Their Training Targets

Zifan Carl Guo, Laura Ruis, Jacob Andreas, Belinda Z Li

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Automating ML for Science: Can Frontier Agents Climb Scientific Hills in the Wild?

Ming Zhong, Stacy Li, Nicholas Carlini, Matthew Jagielski

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Aligning AI Teams

Siddarth Srinivasan, Morgan J Matthews, Jascha Sohl-Dickstein, Erik Jones

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AIs with Secret Loyalties are a Serious but Addressable Threat

Joe Kwon, Alfie Lamerton, Andrew Draganov, Dave Banerjee and 6 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

RTEB: An Overfitting-Resistant Benchmark for Embedding Model Evaluation

Sahil Verma, Minghan Li, Andrew Gaut, Yujie Qian and 14 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors

Fine-tuned LLMs learn to selectively hide internal representations from unseen activation monitors via low-dimensional subspace manipulation, evading even post-hoc safety probes with modest capability loss.

Max McGuinness, Alex Serrano Terre, Luke Bailey, Scott Emmons

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Training Language Models to Explain Their Own Computations

Fine-tuning language models on interpretability ground truth teaches them to describe their internal computations, with self-explanation outperforming larger external explainers.

Belinda Z Li, Zifan Carl Guo, Vincent Huang, Jacob Steinhardt and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 38

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
91%Must read
?Must readVote to see the score

Jointly Reinforcing Diversity and Quality in Language Model Generations

DARLING uses a learned partition function to jointly optimize language model response quality and semantic diversity via reinforcement learning, improving both quality and novelty across creative and math benchmarks.

Tianjian Li, Yiming Zhang, Ping Yu, Swarnadeep Saha and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face · Code ★ 61

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
86%Must read
?Must readVote to see the score

Model Spec Midtraining: Improving How Alignment Training Generalizes

Model spec midtraining teaches models their behavior spec before alignment, controlling how demonstration fine-tuning generalizes and reducing agentic misalignment substantially.

Chloe Li, Sara Price, Samuel Marks, Jonathan Kutasov

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · Code ★ 74

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

Eliciting Secret Knowledge from Language Models

Secret-knowledge-elicitation techniques, especially prefill attacks, successfully extract hidden knowledge that LLMs deny knowing but apply downstream.

Bartosz Cywiński, Emil Ryd, Rowan Wang, Senthooran Rajamanoharan and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 24

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
91%Must read
?Must readVote to see the score

DiscoverPhysics: Benchmarking LLMs for out-of-the-box scientific thinking

DiscoverPhysics benchmarks LLM agents on simulated worlds with non-standard physics, finding frontier models pass only half and fail at uncovering latent structure.

Lindsay Smith, Matt Sampson, Siddharth Mishra-Sharma, Peter Melchior and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 4/5
89%Must read
?Must readVote to see the score

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens

Open-weight LLMs perform hidden multi-step reasoning over filler tokens that unsupervised hidden-state decoding recovers at 82-94% accuracy, showing monitorability requires internal traces.

Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
89%Must read
?Must readVote to see the score

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.

Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
89%Must read
?Must readVote to see the score

Liars' Bench: Evaluating Lie Detectors for Language Models

Liars' Bench evaluates lie detectors across 72,863 LLM lies and finds existing techniques systematically miss certain lie types, especially when transcripts alone are insufficient.

Kieron Kretschmar, Walter Laurito, Sharan Maiya, Samuel Marks

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 4/5
78%Highly rated
?Highly ratedVote to see the score

(Mis)generalization of Helpful-Only Fine-Tuning

Helpful-only fine-tuning causes emergent misalignment, residual refusals, and poor steerability, but synthetic document tuning and character-related training mitigate these issues.

Mohammad Omar Khursheed, Baram Sosis, Fabien Roger

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
80%Must read
?Must readVote to see the score

Hide to Guide: Learning via Semantic Masking

SMEPO applies fine-grained semantic masking to expert traces in RLVR, improving accuracy by up to 3.2 points and cutting training time up to 4.2x across math, code, and agentic tasks.

Ruitao Liu, Qinghao Hu, Alex Hu, Yecheng Wu and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5