Good Papers

Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil, Haon Park, Bumsub Ham, Mohit Bansal, Suhyun Kim

Published 2026Atlanta Poster Session 1 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall C1OpenReview ↗

57%
OverallWorth a lookFirst read from the title only. The AI panel reads it properly once the abstract is public.
?
OverallWorth a lookVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel1/20reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said