Good Papers

Showing papers from Xiaohongshu (Little Red Book) (China) Show all papers

86%Must read
?Must readVote to see the score

Benchmark^2: Systematic Evaluation of LLM Benchmarks

Benchmark² evaluates LLM benchmarks via ranking consistency, discriminability, and capability alignment deviation, revealing quality variations and enabling smaller effective test sets.

Qi Qian, Chengsong Huang, Jingwen Xu, Changze Lv and 12 more

Published Jan 7, 2026 · 0 citations · ▲ 34 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5