Good Papers

Position: Semantic Uncertainty Measures Disagreement, Not Reliability

Semantic uncertainty measures answer disagreement rather than reliability, as valid answers vary and repeated errors appear certain; a bias-uncertainty decomposition separates variability from systematic error to improve evaluation.

Joseph Hoche, Maxime Corlay, David Brellmann, Andrei Bursuc, Pavel Izmailov, Angela Yao, Gianni Franchi

Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4OpenReview ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Large Language Models and their multi-modal variants see increasingly rapid adoption and deployment, including in settings where reliability matters. However their uncertainty remains difficult to assess as classic or early token-level approaches cannot deal with the open-ended outputs of such models. Semantic uncertainty arises as a promising solution to the limits of token-level confidence by measuring disagreement among multiple generated responses. This position paper argues that this framing is incomplete: semantic uncertainty measures semantic disagreement, not actual reliability. A model can be uncertain while producing several valid answers, or certain while repeatedly producing the same wrong answer. We introduce a semantic bias-uncertainty decomposition to show that reliability depends both on variability across meanings and systematic deviation from correct or grounded meanings. This perspective reveals that common hallucination-detection evaluations conflate uncertainty with error. We argue for reliability-centered evaluation that separates semantic disagreement, correctness, and enables a more fine-grained characterization of uncertainty beyond the level of the full answer.