Good Papers

Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise

POTER uses optimal transport geometry between training and reference distributions to downweight mislabeled or shortcut-aligned samples, achieving state-of-the-art worst-group accuracy with a single training stage.

Sung Ho Jo, Seonghwi Kim, Wonsang Yun, Minwoo Chae

Published Oct 1, 2026Sydney Poster Session 2 · Tue, Dec 8, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

82%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
?1 reader voted. Vote to see how they split.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
POTER earns praise for replacing unstable loss signals with transport geometry in a single training pass, yet its dependence on scarce clean group annotations raises serious questions about real-world viability when minority subgroups face concentrated label…

Abstract

Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.