45%Niche pick?Niche pickVote to see the scoreNeurIPS 2026d_modelML & Alignment Theory ScholaStanfordMATS ResearchGoogle DeepMindDeep RLGradient Routing Localizes and Removes Unintended Behaviors in RLJake Ward, Shawn Hu, Aria Wong, Nathan Hu and 3 moreSydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet0/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 0 of 20 reviewers recommend itlenient 0/5medium 0/10strict 0/5
57%Worth a look?Worth a lookVote to see the scoreNeurIPS 2026Centre for the Governance of AIDepartment of Computer Science, Institute for Law & AICenter for AI Risk Management &aETH ZurichAI oversight & deceptionFraudBench: A Legal Evaluation of AI Deception on Realistic TasksKevin Wei, Sumaya N Adan, Stephan Llerena, Mark L Gitau and 16 moreSydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet1/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 1 of 20 reviewers recommend itlenient 1/5medium 0/10strict 0/5
89%Must read?Must readVote to see the scoreNeurIPS 2026EuroSafeAIETHZ - ETH ZurichMATS ResearchMPI for Intelligent Systems, TübVector InstituteAI oversight & deceptionGT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game TheoryGT-HarmBench evaluates 15 frontier AI models on 1,535 multi-agent game-theoretic risk scenarios, finding 38% failure at socially beneficial actions and up to 18% improvement via interventions.Pepijn Cobben, Xuanqiang A Huang, Thao Pham, Isabel Dahlgren and 3 moreParis Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026– ReadersNo votes yet16/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 16 of 20 reviewers recommend itlenient 5/5medium 9/10strict 2/5