Position: Semantic Uncertainty Measures Disagreement, Not Reliability
Semantic uncertainty measures answer disagreement rather than reliability, as valid answers vary and repeated errors appear certain; a bias-uncertainty decomposition separates variability from systematic error to improve evaluation.
Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4OpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Large Language Models and their multi-modal variants see increasingly rapid adoption and deployment, including in settings where reliability matters. However their uncertainty remains difficult to assess as classic or early token-level approaches cannot deal with the open-ended outputs of such models. Semantic uncertainty arises as a promising solution to the limits of token-level confidence by measuring disagreement among multiple generated responses. This position paper argues that this framing is incomplete: semantic uncertainty measures semantic disagreement, not actual reliability. A model can be uncertain while producing several valid answers, or certain while repeatedly producing the same wrong answer. We introduce a semantic bias-uncertainty decomposition to show that reliability depends both on variability across meanings and systematic deviation from correct or grounded meanings. This perspective reveals that common hallucination-detection evaluations conflate uncertainty with error. We argue for reliability-centered evaluation that separates semantic disagreement, correctness, and enables a more fine-grained characterization of uncertainty beyond the level of the full answer.