86%Must read
?Must readVote to see the score

Benchmark^2: Systematic Evaluation of LLM Benchmarks
Benchmark² evaluates LLM benchmarks via ranking consistency, discriminability, and capability alignment deviation, revealing quality variations and enabling smaller effective test sets.
Published Jan 7, 2026 · 0 citations · ▲ 34 on Hugging Face
– ReadersNo votes yet
14/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5