Good Papers

Showing papers from ELLIS Tübingen Show all papers

83%Must read
?Must readVote to see the score

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

Conformal Elo replaces hard judge labels with calibrated win probabilities and split conformal intervals, yielding LLM Elo ratings within 17.9 MAE of human ones with guaranteed uncertainty bounds.

Bora Kargi, David Salinas

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
88%Must read
?Must readVote to see the score

Half-Truths Break Similarity-Based Retrieval

CLIP-style dual encoders often prefer half-true image descriptions with incorrect added details over correct shorter ones due to weak part-level supervision; CS-CLIP improves half-truth accuracy to 69.3% through component-level contrastive fine-tuning.

Bora Kargi, Arnas Uselis, Seong Joon Oh

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026 · ▲ 6 on Hugging Face · Code ★ 15

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5