57%Worth a look?Worth a lookVote to see the scoreNeurIPS 2026Massachusetts Institute of TechnMax Planck Institute of BiochemiMIT / Seoul national university Agent benchmarks & environmentsM4Bench: Evaluating Procedural Specification for Clinical EHR Derivation AgentsHannes Ill, Rafi Al Attrach, Rajna Fani, Ahram Han and 4 moreParis Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026– ReadersNo votes yet1/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 1 of 20 reviewers recommend itlenient 1/5medium 0/10strict 0/5
80%Must read?Must readVote to see the scoreNeurIPS 2026Massachusetts Institute of TechnMax Planck Institute of BiochemiHelmholtz Zentrum MünchenBarcelona Supercomputing CenterGoogleDatasets & benchmarksCroissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML DatasetsCroissant Baker generates validated Croissant metadata locally from dataset directories via modular handlers, achieving 97, 100% agreement with ground truth across 140+ datasets including MIMIC-IV.Rafi Al Attrach, Rajna Fani, Sebastian Lobentanzer, Joan Giner-Miguelez and 16 moreSydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet12/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 12 of 20 reviewers recommend itlenient 5/5medium 5/10strict 2/5