Good Papers

Showing Uncertainty & calibration Show all papers

89%Must read

Certification of Real Images through Calibrated Content Authentication

Deepfake detectors degrade to 76% accuracy and near-zero under attacks, so calibrated reconstruction-based authentication bounds false real-image certification to 1%.

Sarim Hashmi, Abdelrahman Elsayed, Mohammed Talha Alam, Samuele Poppi and 1 more

Published Oct 5, 2026 · ▲ 13 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
45%Niche pick
?Niche pickVote to see the score

The Block Catastrophe of Marginal Calibration in Certified Requirements Traceability

Jikun Wu, Dongxin Guo, Siu Ming Yiu

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Announced Breaks Separate Conformal Reliability from Frequency Calibration

Karl Li

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Latent-Load Framework for Reliability Analysis and Intervention Design in LLM Pipelines

Yuanjie Shi, Yan Yan

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MIRA: Mutual Information guided calibration for Reliable Test-Time Adaptation

Hyeongyu Kim, Mingyeong Jang, Youngjun Song, GeonHui Han and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Hidden Tails: Certifying Tail-Risk Claims under Selective Labels

El Mustapha Mansouri

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Stochastic Grouping Conformal Prediction for Effective Subgroup Reliability

Meihui Zhong, Wenxin Tai, Ting Zhong, Fan Zhou

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Long-Term Risks of Risk-Based Allocation

Jivat Neet Kaur, Jane Lee, Manolis Zampetakis

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

ICAT: Incident-Case–Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

ICAT grounds physical-risk testing of embodied world models in real incident reports to expose frequent missed dangers and severity miscalibration.

Zhenglin Lai, Sirui Huang, Yuteng Li, Changxin Huang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

Robust Conditional Conformal Prediction via Branched Normalizing Flow

Branched Normalizing Flow bounds conditional invalidity via Wasserstein distance and improves robust conditional coverage under distribution shift.

Rui Xu, Xingyuan Chen, Wenxing Huang, Minxuan Huang and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 2/5
medium 9/10
strict 1/5
83%Must read
?Must readVote to see the score

Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback

ExAUL provides online conformal abstention under adversarial bandit feedback with O(sqrt(T)) FDR control via feedback unlocking and a regret-to-FDR conversion lemma.

Minjae Lee, Yoonjae Jung, Sangdon Park

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5
88%Must read
?Must readVote to see the score

Post-hoc Selective Classification for Reliable Synthetic Image Detection

ReSIDe applies post-hoc selective classification to synthetic image detectors by aggregating layer-wise confidence scores via preference optimization, reducing AURC by up to 69.55% under covariate shift.

Kaixiang Zheng, Jacob Seidman

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions

SimplexUQ benchmarks conformal wrappers on simplex-valued predictions, showing global calibration can hide severe under-coverage and no wrapper universally dominates across tasks and stratification maps.

Liang You, Hengyu Shi, Dongwen Ou

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

Class Adaptive Conformal Training

Class Adaptive Conformal Training adaptively shapes class-conditional prediction sets via augmented Lagrangian optimization without distributional assumptions, yielding smaller sets with valid coverage.

Badr-Eddine Marani, Julio Silva-Rodríguez, Ismail Ayed, Maria Vakalopoulou and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Occupancy-based Quantile Risk Control

OQRC partitions calibration losses into ordered bins to tightly upper-bound quantile risk with finite-sample guarantees converging at rate O_p(n^{-1/2}).

Zihao Shi, Huajun Xi, Bingyi Jing, Hongxin Wei

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Conservative neural posterior estimation via distributionally robust training

DRO-NPE trains neural posterior estimators with distributionally robust worst-case losses to reduce overconfidence and improve calibration under limited simulation budgets.

William Laplante, Yuga Hikida, Charita Dellaporta, Francois-Xavier Briol and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 0/5
91%Must read
?Must readVote to see the score

CalArena: A Large Scale Post-Hoc Calibration Benchmark

CalArena benchmarks nearly 2000 post-hoc calibration experiments, finding smooth methods outperform binning and multiclass-specific designs are essential.

Eugène Berta, David Holzmüller, Francis Bach, Michael Jordan

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
83%Must read
?Must readVote to see the score

Radial-Angular Geometry for Reliable Update Diagnosis in Noisy-Label Learning

RGC diagnoses noisy-label updates via radial-angular geometry, distinguishing harmful mislabeled updates from useful hard-clean ones to improve accuracy.

Ningkang Peng, Jingyang Mao, Xiaoqian Peng, Qu Weiguang and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
88%Must read
?Must readVote to see the score

Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique

Latent Critic is a lightweight LoRA adapter that translates LLM latent uncertainty into real-time, localized natural-language hallucination feedback, achieving 0.966 AUROC and enabling agent self-correction with negligible latency.

Sanidhya Vijayvargiya, Rahul Lokesh

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
80%Must read
?Must readVote to see the score

AdaptNC: Adaptive Nonconformity Scores for Conformal Prediction under Distribution Shift

AdaptNC jointly adapts nonconformity scores and conformal thresholds online to reduce prediction volumes under distribution shift while maintaining coverage.

Renukanandan Tumu, Aditya Singh, Rahul Mangharam

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5