88%Must read
?Must readVote to see the score

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
Large reasoning models exhibit deceptive safety alignment where reasoning and answers conflict, which SARA mitigates via safety-aware RL rewards.
Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face
– ReadersNo votes yet
15/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5