Good Papers

Showing LLM evaluation & benchmarks Show all papers

89%Must read
?Must readVote to see the score

Multilingual GSM-Symbolic: What determines capability transfer across languages?

Multilingual GSM-Symbolic introduces matched math problems across 15 languages to show model size, resource level, reasoning, and typology determine cross-lingual transfer, with size and reasoning closing low-resource gaps but not typological ones.

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart and 21 more

Published Oct 2, 2026 · 0 citations · ▲ 47 on Hugging Face · Code ★ 4

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
69%Highly rated

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

RealCompanion benchmarks AI companions on longitudinal real-world chats, finding needed past messages are usually recent, memory detectors fail on real messages, and persona reconstruction costs vary 31-fold at equal F1.

Arman Behnam, Sunglyoung Kim, Liangwei Yang

Published Oct 1, 2026 · 0 citations · ▲ 268 on Hugging Face

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

Sentence-specificity predictors disagree across technical corpora, and ranking LLM revisions improves selection only for certain models and predictors.

Rocker D’Antonio, Thomas Benton Townsend, Dimitrios Michael Manias

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 2/5
92%Must read

Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers

A benchmark of 390 AI papers measures scientific slop across structure, argument, and artifacts; a harness reduces the AI-human gap by 63% via evidence-grounded revision.

Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 55 on Hugging Face · Code ★ 13

100% Readers1 of 1 upvoted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
88%Must read
?Must readVote to see the score

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Direct-decision JEV models show ordinal scale-utilization bias, compressing decisions to 26, 76% of gold support despite high accuracy, but BA-LoRA post-training improves utilization to 86%.

Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 60 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
89%Must read
?Must readVote to see the score

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

Rhetorical robustness requires stable judgments across content-preserving rewrites and discrimination across papers; SciCore improves both via dual-branch science-core review.

Chenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 76 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

Generator attribution falls sharply after rewriting, and provenance-based selection does not clearly outperform quality-score selection for recursive model training.

Joss Armstrong

Published Sep 30, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 2/5
71%Highly rated
?Highly ratedVote to see the score

Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts

Predictive credit for scientific explanations is measured via paired forecasts, but gains over descriptions remain unconfirmed across Tox21, OpenML, and controlled settings.

Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li

Published Sep 29, 2026 · 0 citations · ▲ 101 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 1/5
medium 4/10
strict 2/5
83%Must read
?Must readVote to see the score

MoCo: A One-Stop Shop for Model Collaboration Research

MoCo unifies 26 model collaboration methods and 25 benchmarks to show collaboration outperforms single models in 61% of settings by up to 25.8%.

Shangbin Feng, Yuyang Bai, Ziyuan Yang, Yike Wang and 16 more

Published Jan 29, 2026 · 0 citations · Code ★ 63

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

Benchmark^2: Systematic Evaluation of LLM Benchmarks

Benchmark² evaluates LLM benchmarks via ranking consistency, discriminability, and capability alignment deviation, revealing quality variations and enabling smaller effective test sets.

Qi Qian, Chengsong Huang, Jingwen Xu, Changze Lv and 12 more

Published Jan 7, 2026 · 0 citations · ▲ 34 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer

BetaConform uses a Beta-Binomial MAP framework with adaptive conformal stopping and prior transfer to estimate LLM judge accuracy using minimal labeled samples. It achieves under 3.37% error on TruthfulQA with only ten annotations.

Hui Yan Qu, In-Young Choi, Tan, Zhen, Song Wang and 5 more

Published Apr 17, 2025 · 0 citations

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

IHEval: Evaluating Language Models on Following the Instruction Hierarchy

IHEval evaluates language models on instruction hierarchy compliance, finding they frequently disregard system-level priority instructions.

Zhihan Zhang, Shiyang Li, Zixuan Zhang, Xin Liu and 10 more

Published 2025 · 4 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score
ACM Transactions on Knowledge Discovery from Data 2024AmazonTexas A&MRiceLLM evaluation & benchmarks

Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond

This survey guides practitioners in deploying LLMs across NLP tasks, covering model selection, data effects, use cases, biases, efficiency, and practical limitations.

Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han and 5 more

Published Feb 28, 2024 · 503 citations

– ReadersNo votes yet. 1 from authors or colleagues not counted
9/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 21 reviewers recommend it
lenient 5/5
medium 4/11
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

Sree Bhattacharyya, Samarth Khanna, Leona Chen, Lucas Craig and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Auditing AI peer reviewers: a dose-response and false-positive benchmark on real scientific papers

Íñigo Zubeldia, Boris Bolliet, Francisco Villaescusa, Pablo Villanueva-Domingo and 1 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

VeriScope: Measuring Verification-Ready Verilog Artifacts

Wei Zhang, Jian Yang, jiajun wu, Junhang Cheng and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

HearSayBench: Can LLMs Navigate from Abstract Human Rights to Lived Lives?

Sobhan Lotfi, Ava Iranmanesh, Ali Iranmanesh, Liwei Jiang

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Evaluating Compositional Generalization in Transformers: The Role of Composition Equivalence and Module Coverage

Purva Pruthi, Andrew Yuan, Alexander D'Amour, David Jensen

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Efficient evaluation and error pattern discovery for blackbox AI systems

Maxim Rabinovich, Harvineet Singh, Aman Sinha

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CryptanalysisBench: Can LLMs do cryptanalysis?

Lukas Fluri, Avital Shafran, Nicholas Carlini, Matthew Jagielski and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
Show 20 more papers