Good Papers

A Multimodal Benchmark for Evaluating Cause-of-Death Inference Using Child Health and Mortality Data

This paper introduces a multimodal benchmark for cause-of-death inference in child mortality data, showing zero-shot language models synthesize unstructured medical evidence differently than supervised baselines.

Junhe Yang, Soumyakanti Pan, Hyun Seung Lim, YUE CHU, Yuting Guo, Nishtha Agarwal, Varun Babbar, Gaurav Rajesh Parikh, Yiqun Chen, Chris A Rees, Ziyaad Dangor, Sanjay G Lala, Zehang Li, Samuel J Clark, Zhenke Wu, Abhirup Datta, Li Liu, Cynthia Rudin, Samuel Scarpino, Benjamin M Gyori, Tyler H. McCormick

Published 2026Atlanta Poster Session 6 · Fri, Dec 11, 4:30 PM–7:30 PM local time · Hall C1OpenReview ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

A bs t r act Accurately attributing causes of death is vital for global health, yet fewer than 5% of deaths in resource-constrained regions are medically certified. To assign causes to these unlabeled deaths at scale, practitioners traditionally rely on verbal autopsy, using supervised statistical models to classify based on structured survey data. However, modern mortality surveillance increasingly collects rich, unstructured multimodal data, such as free-text caregiver narratives and postmortem diagnostics, which traditional supervised statistical models struggle to seamlessly integrate. In this paper, we present a comprehensive, multimodal benchmark for cause-of-death classification using data from the Child Health and Mortality Prevention Surveillance (CHAMPS) network, a unique surveillance platform spanning nine countries across South Asia and Sub-Saharan Africa. Using this dataset, we introduce an evaluation framework designed to rigorously assess diagnostic reasoning, moving beyond traditional metrics that fail to capture complex clinical realities. We demonstrate the utility of this benchmark by evaluating zero-shot large language models against supervised baselines across various data modalities. Our results reveal distinct differences in how these modeling approaches synthesize unstructured medical evidence. This benchmark provide a rigorously defined resource for assessing clinical reasoning in next-generation mortality surveillance.