Good Papers

Showing papers from ELLIS Institute Tuebingen Show all papers

45%Niche pick
?Niche pickVote to see the score

Strong Post-Training from Permissive, Reasoning-Dominant, Web-Scale Pretraining

Harsh Raj, Ali Elganzory, Marianna Nezhurina, Victor May and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
83%Must read
?Must readVote to see the score

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

Conformal Elo replaces hard judge labels with calibrated win probabilities and split conformal intervals, yielding LLM Elo ratings within 17.9 MAE of human ones with guaranteed uncertainty bounds.

Bora Kargi, David Salinas

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
80%Must read
?Must readVote to see the score

An Open-Source Training Dataset for Foundation Models for Black-box Optimization

BBO-Pile provides 500K real-world black-box optimization trajectories across 3095 problems, and trained foundation models show large-scale pre-training effectively imitates optimization methods.

Aaron Klein, Herilalaina Rakotoarison, Luca Thale-Bombien, David Salinas

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5