45%Niche pick?Niche pickVote to see the scoreNeurIPS 2026School of Computer and CommunicaEPFL/Swiss Federal Technology InRedwood ResearchU OxfordAI oversight & deceptionAI Control for Sandbagging on Fuzzy TasksMikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe BentonSydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet0/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 0 of 20 reviewers recommend itlenient 0/5medium 0/10strict 0/5
89%Must read?Must readVote to see the scoreNeurIPS 2026U OxfordGeorgia Institute of TechnologyU North Carolina at Chapel HillMATSRedwood ResearchAI oversight & deceptionAttack Selection In Agentic AI Control Evaluations Meaningfully Decreases SafetyStrategic attack selection via start and stop policies substantially lowers measured AI control safety without changing attack capability, reducing safety by up to 28 percentage points and yielding overly optimistic estimates.Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad, Joachim Schaeffer and 2 moreSydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet16/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 16 of 20 reviewers recommend itlenient 5/5medium 8/10strict 3/5