From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation
Conformal Elo replaces hard judge labels with calibrated win probabilities and split conformal intervals, yielding LLM Elo ratings within 17.9 MAE of human ones with guaranteed uncertainty bounds.
Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
