83%Must read
?Must readVote to see the score
ProactBench: Beyond What The User Asked For
ProactBench measures LLM conversational proactivity via emergent, critical, and recovery inference across 198 dialogues, finding recovery is hard and poorly predicted by standard benchmarks.
Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026
– ReadersNo votes yet
13/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5