89%Must read
?Must readVote to see the score

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
GT-HarmBench evaluates 15 frontier AI models on 1,535 multi-agent game-theoretic risk scenarios, finding 38% failure at socially beneficial actions and up to 18% improvement via interventions.
Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026
– ReadersNo votes yet
16/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5