Good Papers

Showing papers from Redwood Research Show all papers

45%Niche pick
?Niche pickVote to see the score

AI Control for Sandbagging on Fuzzy Tasks

Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

Attack Selection In Agentic AI Control Evaluations Meaningfully Decreases Safety

Strategic attack selection via start and stop policies substantially lowers measured AI control safety without changing attack capability, reducing safety by up to 28 percentage points and yielding overly optimistic estimates.

Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad, Joachim Schaeffer and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5