Good Papers

Showing papers from Department of Computer Science, Wayne State University Show all papers

88%Must read
?Must readVote to see the score

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

Large reasoning models exhibit deceptive safety alignment where reasoning and answers conflict, which SARA mitigates via safety-aware RL rewards.

Xiangyu Zhou, Saleh Z Zade, Rafi Ibn Sultan, Alexander Kotov and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5