GT-HarmBench evaluates 15 frontier AI models on 1,535 multi-agent game-theoretic risk scenarios, finding 38% failure at socially beneficial actions and up to 18% improvement via interventions.
Causality systematically addresses benchmark biases and artifacts by making assumptions explicit to model phenomena, formulate hypotheses, and clarify method strengths through common causal topologies.