Good Papers

Showing papers from UToronto & MPI Show all papers

57%Worth a look
?Worth a lookVote to see the score

SuperSycophantic: Stress-Testing Frontier LLMs from Single- to Multi-Turn Sycophancy

Terry J Zhang, Oscar S Yasunaga, Wenyuan Jiang, Jessica Bo and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AM-Bench: A Unified Taxonomy and Evaluation Suite for Agentic Misalignment

Eric Zhang, Terry J Zhang, Chijioke Ugwuanyi, Jerick Shi and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

GT-HarmBench evaluates 15 frontier AI models on 1,535 multi-agent game-theoretic risk scenarios, finding 38% failure at socially beneficial actions and up to 18% improvement via interventions.

Pepijn Cobben, Xuanqiang A Huang, Thao Pham, Isabel Dahlgren and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
71%Highly rated
?Highly ratedVote to see the score

Causality can systematically address the monsters under the bench(marks)

Causality systematically addresses benchmark biases and artifacts by making assumptions explicit to model phenomena, formulate hypotheses, and clarify method strengths through common causal topologies.

Felix Leeb, Zhijing Jin, Bernhard Schölkopf

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5