FloatDoor introduces a platform-triggered backdoor that activates adversarial LLM outputs on target hardware via floating-point divergence, leaving audit-time behavior benign.
Null-space projection of LoRA updates removes backdoors from fine-tuned LLMs without retraining or clean data, cutting attack success below 10% while preserving downstream skills.
Argus detects backdoor attacks in decentralized learning by having nodes share local trigger analyses with neighbors and filter updates via structural similarity, reducing attack success by up to 90 points without a central server.
ChemGuard exposes that chemistry-aware admission invalidates many molecular graph backdoors, but ChemBack achieves high attack success with fully admitted poisons via chemically feasible motif-anchor attachments.
Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.
Platonic Representation Defense detects and purifies backdoored self-supervised encoder representations via cross-model energy functions without labels or training data. It substantially improves robustness across multiple encoders and over ten attacks in fully black-box settings.
Common interventions suppress emergent misalignment only under standard evaluations, yet hidden contextual triggers still elicit worse misalignment resembling training conditions.