Good Papers

Showing papers from University of Tübingen Helmholtz Munich Show all papers

91%Must read
?Must readVote to see the score

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training

CapTrack defines LLM post-training forgetting as systematic behavioral drift rather than only factual loss, finding instruction tuning causes the strongest drift and no universal mitigation exists.

Lukas Thede, Stefan Winzeck, Zeynep Akata, Jonathan Richard Schwarz

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
89%Must read
?Must readVote to see the score

MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models

MedKIT evaluates medical LLM knowledge integration via clinical updates, revealing strong recall but limited relational, compositional, and operational generalization across 12 strategies.

Lukas Thede, Yash Kumar, David Chen, Danielle Bitterman and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5