Good Papers

PHOEBI: An Open-World Benchmark for Bacterial Identification in Phase-Contrast Microscopy

PHOEBI introduces a 120,000-image phase-contrast benchmark of bacterial mixtures; per-image classifiers collapse on unseen combinations, and anchor-based decoders over frozen features stabilize identification plus open-set rejection.

Aaditya Baranwal, Md Jahid Hasan, Shruti Vyas

Published 2026Sydney Poster Session 2 · Tue, Dec 8, 5:00 PM–8:00 PM local time · Hall 1-4▲ 1 on Hugging FacearXiv ↗OpenReview ↗

89%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel16/20reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Optical microscopy (OM) enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environmental, and industrial microbiology. Real samples, however, are routinely polymicrobial and may contain organisms never seen during training, and no computer-vision benchmark evaluates multi-label species identification from phase-contrast microscopy (PCM) of such mixtures. We introduce Phase-contrast Optical bEnchmark for Bacterial Identification ($\textbf{PHOEBI}$), a wet-lab-prepared dataset of $120{,}000$ PCM images covering $40$ combinations of six rod-shaped species, together with a leave-combinations-out (LCO) protocol that holds out entire species combinations, mirroring a model trained on catalogued mixtures that must recognise new ones. Under LCO, gradient-trained per-image classifiers, from fine-tuned backbones to attention-based multiple-instance learning, collapse on unseen combinations despite high in-distribution accuracy, and the failure lies in how per-image predictions are aggregated rather than in the visual representation. We propose three lightweight $\textbf{anchor-based}$ decoders that read each species' presence against fixed geometric prototypes over a shared frozen tile-feature pool, and they remain stable under the same shift. Without additional training, the same features also support open-set rejection of unseen species and the discovery of a new class from unlabeled test images, with negligible disruption to the known classes.