ProactBench measures LLM conversational proactivity via emergent, critical, and recovery inference across 198 dialogues, finding recovery is hard and poorly predicted by standard benchmarks.
Submodular maximization selects small benchmark subsets to approximate all others, with mutual information outperforming entropy for small-set imputation across public LLM leaderboards.