Good Papers

Showing papers from Max Planck Institute for Intelligent Systems & ELLIS Institute Tübingen Show all papers

88%Must read
?Must readVote to see the score

Graph-Regularized Sparse Autoencoders for LLM Safety Steering

Graph-Regularized Sparse Autoencoders smooth SAE decoder vectors over a neuron co-activation graph to learn safety-steering directions, improving selective refusal by over 16 points across jailbreak benchmarks while preserving benign performance and generalizing across models.

Jehyeok Yeon, Federico Cinus, Yifan Wu, Luca Luceri

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5