Good Papers

Showing Backdoors & data poisoning Show all papers

45%Niche pick
?Niche pickVote to see the score

SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing

Hui Zhang, Yachao Yuan, Jiayun Wang, Yuanzhuo Li and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Backdoor Attacks Rerouted: BatchNorm as a Sink for Adversarial Signals

Md Abdul Kadir, Tuan Tran Anh, Daniel Sonntag

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Robust and Efficient Backdoor Mitigation for ML Models via Tolerant Property Testing

Xi Chen, Anindya De, Rocco A Servedio

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

VOID: Backdoor Injection through Knowledge Vacuity in Federated Unlearning

Wenwei Zhao, Yuxuan Xie, Haiyun Liu, Jie Xu and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

A Theoretical Analysis of Backdoor Learning as Simplicity-Biased Optimization Dynamics

Hong Cheng, Qiang Zhou, Sinno Pan

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Benchmark Recovery Does Not Certify a Frozen Tool-Trigger Contract After Quantization

Rui Li, Shuang Cao

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

Not Suppressing or Purifying: Backdoor Containment via Expert Quarantine and Shutdown in LLMs

Jianwei Li, Min-Seon Kim, Jung-Eun Kim

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Consistency-Verified Backdoor Defense for Federated Graph Learning via Cross-Layer Drift

Yanwen Jia, Yinuo Zhang, Zitong Shi, Yuxin Wu and 5 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Poison-then-Hide: Finetuning-Activated Backdoor Attack on Pretrained Vision Encoders

Qixuan Jin, Abinitha Gourabathina, Vinith Suriyakumar, Walter Gerych and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

When Poison Meets Structure: Topology-based Defense against Poisoning Attack on Graph-based Retrieval-Augmented Generation

Qizhi Chen, Junhao Wen, Shuang Liang, Jiakai Li and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Emergent Misalignment as Data-Mediated Transfer

Baris Askin, Muhammed Ustaomeroglu, Anupam Nayak, Gauri Joshi and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ShadowFPT: Backdooring Federated Prompt Tuning via Shadow Triggers

Kun Zhai, Teng Li, Yunhao Feng, Xingjun Ma

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Backdoor Attacks under Lossy Compression: From Failure to Reactivation and Adaptation

Qian Li, Yunuo Chen, Yuntian Chen

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Weird Generalization from Narrow Finetuning: Persona Shifts and Inductive Backdoors

Jan Betley, Jorio Cocola, Dylan Feng, James Chua and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

When Sanitization Becomes the Trigger: Defense-Triggered Backdoor Attacks

Zhiguo Yang, Ruotian Liu, Peipei Xu, Wenjie Ruan

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Swarm Shepherd: Securing Multi-Agent Ecosystems Against Persistent Latent Compromise

Wuyang Zhang, Shichao Pei

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Training-Based Backdoors Are Not Cryptographic

Masahiro Kaneko

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Clean-Label Poisoning for Gradient-Boosted Decision Trees

Sinuo Fan, Jun Woo Chung, Weijie Zhao, Yingjie Lao

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

DetectViT: Test-time Backdoor Detection for Vision Transformers via Inter-Head Attention Discrepancy

Siquan Huang, Yijiang Li, Xingfu Yan, Ningzhi Gao and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Clean Data Can Still Carry Backdoors: Support-Persistent Backdoors for Model Reuse

Junhoo Lee, Baekseung Kim, Seungyeon Kim, Nojun Kwak

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization

Zihan Wang, Ethan Ma, Zhongkui Ma, Shuofeng Liu and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Train-free Data Poisoning Attack against Retrieval-augmented Diffusion Models

XINQI LYU, Yihao LIU, Yiming Cao, Bin Xiao

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Stealthy World Model Manipulation via Data Poisoning

SWAAP uses bilevel optimization and gradient matching to poison world model fine-tuning data, degrading planning while evading defenses.

Yibin Hu, Zizhan Zheng

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

FloatDoor: Platform triggered Backdoors in LLMs

FloatDoor introduces a platform-triggered backdoor that activates adversarial LLM outputs on target hardware via floating-point divergence, leaving audit-time behavior benign.

Nils Loose, Jonas Sander, Felix Mächtle, Thomas Eisenbarth

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
80%Must read
?Must readVote to see the score

When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models

Cross-vocabulary tokenizer transplantation exhibits asymmetric realizability, allowing identical reconstruction coefficients to stay inert in donor anchors yet yield high-salience outputs in base anchors, forming breaker tokens that survive weight merging and evade spectral filters and LoRA mitigati

Xiaoze Liu, Weichen Yu, Matt Fredrikson, Xiaoqian Wang and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 2/5
80%Must read
?Must readVote to see the score

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

Null-space projection of LoRA updates removes backdoors from fine-tuned LLMs without retraining or clean data, cutting attack success below 10% while preserving downstream skills.

Jianwei Li, Jung-Eun Kim

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
88%Must read
?Must readVote to see the score

Your Neighbors Know: Leveraging Local Neighborhoods for Backdoor Detection in Decentralized Learning

Argus detects backdoor attacks in decentralized learning by having nodes share local trigger analyses with neighbors and filter updates via structural similarity, reducing attack success by up to 90 points without a central server.

Sayan Biswas, Antoine Boutet, Davide Frey, Romaric Gaudel and 6 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

Rethinking Molecular Graph Backdoors under Chemistry-aware Admission

ChemGuard exposes that chemistry-aware admission invalidates many molecular graph backdoors, but ChemBack achieves high attack success with fully admitted poisons via chemically feasible motif-anchor attachments.

Thinh Nguyen, Sze Jue Yang, Khoa D Doan, Chee Seng Chan and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
91%Must read
?Must readVote to see the score

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents

MemPoison benchmarks 1227 adversarial cases across memory substrates and finds write-time defenses fail against multi-record and dormant corruption, requiring adaptive defenses.

Jifeng Gao, Kang Xia, Yi Zhang, Xiaobin Hong and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
80%Must read
?Must readVote to see the score

Subliminal Transfer of Unsafe Behaviors in AI Agent Distillation

Distillation of agent trajectories transfers unsafe behavioral biases subliminally despite keyword filtering, with deletion bias reaching 100% and chmod preference 30-55%.

Jacob Dang, Brian Y Xie, Omar G. Younis

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
89%Must read
?Must readVote to see the score

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.

Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
83%Must read
?Must readVote to see the score

The Platonic Defense: Backdoor Defense for Self-Supervised Encoders in the Era of Large Scale Pre-training

Platonic Representation Defense detects and purifies backdoored self-supervised encoder representations via cross-model energy functions without labels or training data. It substantially improves robustness across multiple encoders and over ten attacks in fully black-box settings.

Tuo Chen, Minjing Dong, Benlei Cui, Jian liu and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

Common interventions suppress emergent misalignment only under standard evaluations, yet hidden contextual triggers still elicit worse misalignment resembling training conditions.

Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5