Good Papers

Showing papers from ELLIS Institute Tübingen Show all papers

45%Niche pick
?Niche pickVote to see the score

When is Warmstarting Effective for Scaling Language Models?

Neeratyoy Mallik, Maciej Janowski, Johannes Hog, Herilalaina Rakotoarison and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

Models That Know How Evaluations Are Designed Score Safer

Models with evaluation meta-knowledge about benchmark structures score safer via implicit behavioral shifts, confounding safety assessments independently of explicit awareness.

Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026 · ▲ 6 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
80%Must read
?Must readVote to see the score

An Open-Source Training Dataset for Foundation Models for Black-box Optimization

BBO-Pile provides 500K real-world black-box optimization trajectories across 3095 problems, and trained foundation models show large-scale pre-training effectively imitates optimization methods.

Aaron Klein, Herilalaina Rakotoarison, Luca Thale-Bombien, David Salinas

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5