Good Papers

Showing papers from Warsaw University of Technology Show all papers

67%Highly rated
?Highly ratedVote to see the score

Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models

Bartlomiej Sobieski, Matthew Tivnan, Dawid Płudowski, Michał J Włodarczyk and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Weird Generalization from Narrow Finetuning: Persona Shifts and Inductive Backdoors

Jan Betley, Jorio Cocola, Dylan Feng, James Chua and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
80%Must read
?Must readVote to see the score

Eliciting Secret Knowledge from Language Models

Secret-knowledge-elicitation techniques, especially prefill attacks, successfully extract hidden knowledge that LLMs deny knowing but apply downstream.

Bartosz Cywiński, Emil Ryd, Rowan Wang, Senthooran Rajamanoharan and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 24

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
80%Must read
?Must readVote to see the score

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

Common interventions suppress emergent misalignment only under standard evaluations, yet hidden contextual triggers still elicit worse misalignment resembling training conditions.

Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5