67%Highly rated?Highly ratedVote to see the scoreNeurIPS 2026U Warsaw, Centre for CredibleMassachusetts General Hospital, Warsaw University of TechnologyCentre for Credible AIMassachusetts General Hospital, Diffusion modelsLocal Intrinsic Dimension Unveils Hallucinations in Diffusion ModelsBartlomiej Sobieski, Matthew Tivnan, Dawid Płudowski, Michał J Włodarczyk and 3 moreSydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet2/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 2 of 20 reviewers recommend itlenient 1/5medium 1/10strict 0/5
57%Worth a look?Worth a lookVote to see the scoreNeurIPS 2026U WarsawHarvardU California, Berkeleynational university of singaore,NortheasternBackdoors & data poisoningWeird Generalization from Narrow Finetuning: Persona Shifts and Inductive BackdoorsJan Betley, Jorio Cocola, Dylan Feng, James Chua and 3 moreSydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet1/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 1 of 20 reviewers recommend itlenient 0/5medium 0/10strict 1/5
80%Must read?Must readVote to see the scoreNeurIPS 2026Warsaw University of TechnologyU OxfordAnthropicGoogle DeepMindAI oversight & deceptionEliciting Secret Knowledge from Language ModelsSecret-knowledge-elicitation techniques, especially prefill attacks, successfully extract hidden knowledge that LLMs deny knowing but apply downstream.Bartosz Cywiński, Emil Ryd, Rowan Wang, Senthooran Rajamanoharan and 3 moreSydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 24– ReadersNo votes yet12/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 12 of 20 reviewers recommend itlenient 5/5medium 5/10strict 2/5
80%Must read?Must readVote to see the scoreNeurIPS 2026IDEAS NCBR Sp. z o. o. Address: U WarsawWarsaw University of TechnologyCenter on Long-Term RiskU OxfordBackdoors & data poisoningConditional misalignment: common interventions can hide emergent misalignment behind contextual triggersCommon interventions suppress emergent misalignment only under standard evaluations, yet hidden contextual triggers still elicit worse misalignment resembling training conditions.Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan and 1 moreSydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet12/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 12 of 20 reviewers recommend itlenient 5/5medium 6/10strict 1/5