Good Papers

Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews

AgentSLR evaluates LLMs on epidemiological systematic reviews, revealing sub-task specialization, poor structured extraction (F1 < 0.67), and unreliable unsupervised deployment.

Shreyansh Padarha, Ryan Othniel Kearns, Tristan M Naidoo, Lingyi Yang, Łukasz Borchmann, Piotr Blaszczyk, Christian Morgenstern, Ruth McCabe, Sangeeta Bhatia, Philip Torr, Jakob Foerster, Scott Hale, Thomas Rawson, Anne Cori, Elizaveta Semenova, Adam Mahdi

Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4▲ 9 on Hugging FaceCode ★ 25arXiv ↗OpenReview ↗

91%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel17/20reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Systematic literature reviews (SLRs) are a demanding and high-stakes form of scientific knowledge synthesis that remains underspecified as an evaluation setting for large language models (LLMs). We introduce AgentSLR, a large-scale evaluation harness comprising an SLR automation workflow and an expert annotated dataset covering 16,248 articles, designed to test LLM capabilities across the stages of SLRs in epidemiology. Reference annotations were derived from peer-reviewed studies on WHO priority pathogens and produced by domain experts. The harness evaluates each review stage as a separate unit with dedicated metrics enabling targeted failure analysis. We evaluated five frontier reasoning models and found that no single model dominated across all tasks, showing sub-task specialisation often hidden by aggregate benchmarks. Structured data extraction is a major bottleneck, with no model exceeding an average field-level F1 of 0.67. Estimated costs vary substantially, by up to 96 times across evaluated models. Documented failure modes suggest that the evaluated models are not yet reliable enough for unsupervised deployment in epidemiology, where findings can inform public policy.