Good Papers

Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer

BetaConform uses a Beta-Binomial MAP framework with adaptive conformal stopping and prior transfer to estimate LLM judge accuracy using minimal labeled samples. It achieves under 3.37% error on TruthfulQA with only ten annotations.

Hui Yan Qu, In-Young Choi, Tan, Zhen, Song Wang, Yun, Sukwon, Long, Qi, Faizan Siddiqui, Kwonjoon Lee, Tianlong Chen

Published Apr 17, 2025arXiv ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
BetaConform delivers impressively precise LLM judge estimates with minimal samples via adaptive MAP and prior transfer, yet its validation on a single dataset with one model and absent cross-benchmark baselines leaves its broad reliability unproven.

Abstract

LLM ensembles are widely used for LLM judges. However, how to estimate their accuracy, especially in an efficient way, is unknown. In this paper, we present a principled maximum a posteriori (MAP) framework for an economical and precise estimation of the performance of LLM ensemble judgment. We first propose a mixture of Beta-Binomial distributions to model the judgment distribution, revising from the vanilla Binomial distribution. Next, we introduce a conformal prediction-driven approach that enables adaptive stopping during iterative sampling to balance accuracy with efficiency. Furthermore, we design a prior transfer mechanism that utilizes learned distributions on open-source datasets to improve estimation on a target dataset when only scarce annotations are available. Finally, we present BetaConform, a framework that integrates our distribution assumption, adaptive stopping, and the prior transfer mechanism to deliver a theoretically guaranteed distribution estimation of LLM ensemble judgment with minimum labeled samples. BetaConform is also validated empirically. For instance, with only 10 samples from the TruthfulQA dataset, for a Llama ensembled judge, BetaConform gauges its performance with error margin as small as 3.37%.