Good Papers

Showing papers from IDEAS NCBR Sp. z o. o. Address: Chmielna 69, 00-801 Warsaw VAT Number: EU 7011017605 Show all papers

91%Must read
?Must readVote to see the score

Negation Neglect: When models fail to learn negations in training

Fine-tuning LLMs on documents that flag claims as false makes them believe those claims, with belief rates jumping from 2.5% to 88.6%, though local negation phrasing largely prevents it.

Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
80%Must read
?Must readVote to see the score

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

Common interventions suppress emergent misalignment only under standard evaluations, yet hidden contextual triggers still elicit worse misalignment resembling training conditions.

Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5