Good Papers

Showing papers from Microsoft Research, New England Show all papers

88%Must read
?Must readVote to see the score

Calibration without Ground Truth

A label-free framework improves miscalibrated models using better-calibrated references, guaranteeing strict loss reduction without ground-truth labels via Bregman projection.

Yuqing Kong, Mingyu Song, Yizhou Wang, Yifan Wu

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
80%Must read
?Must readVote to see the score

Truthful Calibration Errors for Multi-Class Prediction

The paper defines truthful multiclass calibration errors for linear label properties, proves they preserve Blackwell informativeness ordering, and show they stabilize model rankings across bin choices.

Yuxuan Lu, Yifan Wu, Jason Hartline, Lunjia Hu

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Coherence Mechanisms for Provable Self-Improvement

Coherence-based projection mechanisms provably improve models monotonically via reduced Bregman divergence, with characterization theorems showing coherence is necessary for universal self-improvement.

Mehryar Mohri, Jon Schneider, Yifan Wu

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 2/5