Good Papers

Showing papers from MIT (ETH) Show all papers

89%Must read
?Must readVote to see the score

Local Sparsity Enables Unsupervised LLM Safety Detection

Local sparsity in sparse autoencoder representations enables unsupervised LLM safety detection via masked activation analysis, achieving near-optimal detection using only 1-2% of neurons.

Xin Chen, Cynthia, Gil Kur, Aleksandr Shevchenko, Andreas Krause

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5