
AgentSLR evaluates LLMs on epidemiological systematic reviews, revealing sub-task specialization, poor structured extraction (F1 < 0.67), and unreliable unsupervised deployment.
Shreyansh Padarha, Ryan Othniel Kearns, Tristan M Naidoo, Lingyi Yang and 12 more
Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 25
– ReadersNo votes yet
17/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5