Good Papers

Showing Jailbreaks & red teaming Show all papers

91%Must read
?Must readVote to see the score

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

OpenART scales agent red teaming via open-ended environment evolution across 10,000 stateful scenarios, with EMHA achieving 85% attack success that grows with complexity.

Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu and 5 more

Published Aug 1, 2026 · 0 citations · ▲ 266 on Hugging Face · Code ★ 231

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
57%Worth a look
?Worth a lookVote to see the score

MMA-SafetyBench: A Benchmark for Multimodal Agent Safety Evaluation

Yuke Wang, Benlei Cui, Shen Pang, Xuemei Dong and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

JailBound: A FOL-Guided Jailbreak Evaluation Framework for Revealing Safety Boundaries of LLMs

Fazong Wu, Ming Yang, Xin Wang, Zhenyong Zhang and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

TraceGuard: Defending Multi-Turn Jailbreak Attacks via Prompt-Response Risk Signal Tracking

Hongyi Li, Yufei Wang, Chengxuan Zhou, Qinlin Xie and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

JEDI: Real-Time Jailbreak Defense for LLMs via In-Generation Detection and Intervention

Ruilin Xie, Bixin Li, Xinyu Chen, Yongqiang Tian and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Zhengyang Tang, Yi Zhang, Chenxin Li, Xin Lai and 17 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation

NaHyeon Park, Minhyun Lee, Hyunjung Shim

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PoSafeNet: Structured Safety Learning via Compositional Projection

Kiwan Wong, Wei Xiao, Daniela Rus

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MetaPI: Constructing Prompt Injection Benchmarks from Any Agent Benchmarks

Peiran Wang, Chong Xiang, Wenjie Qu, Ying Li and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Brute-Force Jailbreaks and Codon-Aware Watermarking for DNA Foundation Models

Munish B Persaud, Amrit Singh Bedi, Souradip Chakraborty

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Harmless in Pieces, Harmful in Motion: Detecting Multi-Agent Jailbreaks

Vishal Pramanik, Maisha Maliha, Olivera Kotevska, Nathaniel D Bastian and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Reasoning Poisoning: Utilizing Social-Engineering to Steer Chain-of-Thought

Matan Levy, Ilan Zendel, Stav Cohen, Amit LeVi and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Failure-Band Authorization for Filtered Retrieval

Harry Li, Thea Cao

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Bridging the Gap Between Harmfulness Belief and Refusal Behavior for Safety Alignment

Lu Zhang, Chen Feng, Qingzhuo Wang, Wen Shen and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Safe in Its Own Words: Self-Guided Safety Alignment for Multimodal Reasoning Models

Adeel Yousaf, Souradip Chakraborty, Mubarak Shah, Amrit Singh Bedi

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Auditing Instruction Robustness in Vision-Language-Action Models via Diversity-Aware Red Teaming

Baoshun Tong, Haoran He, Yang Liu, Ling Pan and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Measuring and Strengthening Behavioral Suppression in Language Models

Luxi (Lucy) He, Pengcheng Jiang, Jifan Zhang, Jiawei Han and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

When Safety Becomes An Outlier: Understanding the Retention of LLM Safety Behaviors

Binchi Zhang, Hadi Abdullah, Yiwei Cai, Jundong Li

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Adaptive Test Case Discovery for LLM-Assisted Decision Making in High-Stakes Domains

Anjali Parashar, Carson Sobolewski, Yingke Li, Fei Chen and 1 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

MLLMs Fail to Refuse when Using Tools Agentically

Rikiya Takehi, Ryo Hachiuma, Shaona Ghosh, Dan Zhao and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
86%Must read
?Must readVote to see the score

Fail-Closed Alignment for Large Language Models

Fail-closed alignment builds redundant refusal pathways to prevent alignment collapse under jailbreaks, yielding stronger robustness with minimal overhead.

Zachary Coalson, Sanghyun Hong

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
83%Must read
?Must readVote to see the score

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing

SafeMoE isolates unsafe knowledge into domain-specific LoRA experts and routes them with a lightweight gating network to improve safe response rates by over 20% relative while maintaining informative outputs.

Maryam Hashemzadeh Barvarz, Jerry Huang, Minseon Kim, Marc-Alexandre Côté and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

Safety alignment lags in code domains enable automated jailbreaks through structured code completion, achieving 96.25% attack success across eight commercial LLMs.

Liang Zhen, Wentao Chen, Hai Huang

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
88%Must read
?Must readVote to see the score

Graph-Regularized Sparse Autoencoders for LLM Safety Steering

Graph-Regularized Sparse Autoencoders smooth SAE decoder vectors over a neuron co-activation graph to learn safety-steering directions, improving selective refusal by over 16 points across jailbreak benchmarks while preserving benign performance and generalizing across models.

Jehyeok Yeon, Federico Cinus, Yifan Wu, Luca Luceri

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
86%Must read
?Must readVote to see the score

Diffusion LLMs are Natural Adversaries for any LLM

Diffusion LLMs amortize adversarial prompt optimization by directly generating diverse, transferable jailbreak prompts that bypass black-box target models.

David Lüdke, Tom Wollschläger, Paul Ungermann, Stephan Günnemann and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
91%Must read
?Must readVote to see the score

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment

DC-GRPO assigns turn-level group-relative credit in multi-turn LLM jailbreaking, achieving over 97% attack success and outperforming prior methods.

Junyoung Park, Namgyu Park, Sechan Lee, Yoon-Chan Jhi and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
91%Must read
?Must readVote to see the score

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

SciHazard benchmarks LLM scientific safety risks via decomposed harm scoring across 3,000 real-world grounded queries, finding deep research agents 32.3% more harmful than standard models.

Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
78%Highly rated
?Highly ratedVote to see the score

The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety

Fine-tuning breaks safety via unstable geometric alignment subspaces, with alignment loss scaling quartically in training time via curvature-driven drift.

Max Springer, Chung Peng Lee, Bohdan Turbal, Blossom Metevier and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction

Multimodal LLM safety failure stems from geometry collapse along refusal directions caused by modality drift, which adaptive drift correction and self-rectification restore without training.

Jiahe Guo, Xiangran Guo, Jiaxuan Chen, Weixiang Zhao and 5 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
72%Highly rated
?Highly ratedVote to see the score

Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing

CCLUB enables adaptive LLM alignment via online system-prompt routing with conservative consensus clustering, achieving sublinear regret and improving cumulative reward by 10.98% over baselines.

Zeyu Zhang, Xiangxiang Dai, Ziyi Han, Xutong Liu and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
91%Must read
?Must readVote to see the score

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

On-Policy Consistency Training improves LLM safety across sycophancy, jailbreaks, and safety awareness while avoiding the capability regressions of supervised fine-tuning.

Andy Q Han, Kristina Fujimoto, Avidan Shah, Kiet Nguyen and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
80%Must read
?Must readVote to see the score

TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization

TROPT unifies discrete text optimization via a modular open-source framework with 30+ recipes, enabling cross-domain optimizer comparison, enhancement, and portability.

Matan Ben-Tov, Mahmood Sharif

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · Code ★ 11

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Internal Safety Collapse in Frontier Large Language Models

Frontier LLMs suffer Internal Safety Collapse, generating harmful content during benign tasks with 95.3% failure rates and revealing alignment does not eliminate underlying risks.

Oscar W, Xiao Liu, Hanxun Huang, Yige Li and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 31 on Hugging Face · Code ★ 1,198

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
91%Must read
?Must readVote to see the score

GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines

GitInject tests real AI CI/CD workflows and finds all providers vulnerable to prompt injection via structural credential and config handling flaws.

Jafar Isbarov, Umid Suleymanov, I Shumailov, Murat Kantarcioglu

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

LITMUS benchmarks LLM agent behavioral jailbreaks in real OS environments, revealing agents execute 40.64% of high-risk operations despite refusals and suffer pervasive execution hallucination.

Chiyu Zhang, Huiqin Yang, Bendong Jiang, Xiaolei Zhang and 7 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
78%Highly rated
?Highly ratedVote to see the score

Metaphor Is Not All Attention Needs

Poetic jailbreaks bypass LLM safety not via specific devices or format confusion, but through accumulated stylistic irregularities altering processing independently of harm detection, requiring style-aware defenses.

Olga Sorokoletova, Francesco Giarrusso, Giacomo De Luca, Piercosma Bisconti and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
92%Must read
?Must readVote to see the score

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

Low-bit KV cache quantization silently collapses LLM safety alignment via geometric subspace vulnerability, and per-channel reduction diagnostics recover up to 97% of lost refusals.

Bruce C Xu, Adarsh Kumarappan, Mu Zhou

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
91%Must read
?Must readVote to see the score

Learning to Inject: Automated Prompt Injection via Reinforcement Learning

AutoInject uses reinforcement learning with comparison-based rewards to learn adversarial suffixes that inject prompts into LLM agents, outperforming manual and optimization-based attacks on AgentDojo and Meta-SecAlign-70B.

Xin Chen, Cynthia, Jie Zhang, Florian Tramer

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
88%Must read
?Must readVote to see the score

Latent-space Attacks for Refusal Evasion in Language Models

Refusal suppression is recast as a latent-space evasion attack against linear refusal probes, explaining prior ablation and motivating a controlled evasion method that achieves state-of-the-art refusal bypass across 15 models.

Giorgio Piras, Raffaele Mura, Fabio Brau, Maura Pintor and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

Anchored Bipolicy Self-Play uses frozen-base LoRA adapters to separate attacker and defender roles, preventing self-consistency collapse and improving safety with 100x greater parameter efficiency.

Gabriele La Malfa, Emanuele La Malfa, Saar Cohen, Jie Zhang and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
89%Must read
?Must readVote to see the score

Measuring Safety Alignment Effects in Autonomous Security Agents

Safety alignment effects in autonomous security agents require system-level measurement of refusal, tool reliability, and evidence grounding rather than refusal rates alone, with uncensored Gemma models improving security task success but showing mixed, family-dependent effects.

Isaac David, Arthur Gervais

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
80%Must read
?Must readVote to see the score

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

Safety-aligned LLMs remain vulnerable to mid-generation token injections altering subsequent outputs, and trajectory-level alignment improves robustness beyond shallow token defenses.

Kyungmin Park, Taesup Kim

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
70%Highly rated
?Highly ratedVote to see the score

ReaLM: A Unified Red-Teaming Benchmark for Physical-World VLMs

ReaLM is a unified red-teaming benchmark for physical-world VLMs integrating 12 attacks, 3 defenses, and 13 frontier models.

Yifei Zhao, Qian Lou, Mengxin Zheng

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 2/5
medium 2/10
strict 0/5
91%Must read
?Must readVote to see the score

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Skill cascading attacks distribute malicious objectives across benign skills to harm agent systems, and SkillCascade reliably induces such failures while evading per-skill defenses.

Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
83%Must read
?Must readVote to see the score

Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction--Concealment Tradeoff in MLLMs

Intent-obfuscation jailbreaks on MLLMs face a reconstruction-concealment tradeoff that existing transformations fail to balance, but character-removed variants and concealment-aware construction with keyword distractor images exploit model reconstruction to bypass safety filters.

Md Farhamdur Reza, Richeng Jin, Tianfu Wu, Huaiyu Dai

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
86%Must read
?Must readVote to see the score

Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

Safety alignment shapes diffusion language models' denoising energy barriers, and three complementary kinetic-energy signals detect jailbreaks by forcing attacks to reveal intent or expend detectable cross-barrier energy.

Thong Bach, Dung Nguyen, Thao Le, Truyen Tran

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

Palette enables modular, efficient relaxation of LLM refusal behaviors for authorized domains via lightweight adaptation and parameter merging while preserving general safety.

Qitao Tan, Xiaoying Song, Arman Akbari, Arash Akbari and 6 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing

DR-Smoothing disrupts and rectifies LLM prompts via smoothed defense to guarantee jailbreaking protection while balancing harmlessness and helpfulness.

Zheng Lin, Zhenxing Niu, haoxuan ji, Haichang Gao

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
89%Must read
?Must readVote to see the score

MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks

MT-JailBench provides a modular framework for comparing multi-turn jailbreak attacks under standardized conditions, finding that prompt generation drives success and recomposed components yield stronger attacks.

Xinkai Zhang, Zhipeng Wei, Huanli Gong, Jing Ting Zheng and 3 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
86%Must read
?Must readVote to see the score

Cat-DPO: Category-Adaptive Safety Alignment

Cat-DPO applies per-category adaptive safety margins to direct preference optimization, improving aggregate safety and reducing worst-category harm gaps across models.

Tiankai Yang, Yi Nian, Xinyuan Li, Ruiyao Xu and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
86%Must read
?Must readVote to see the score

Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs

Expected Harm weights jailbreak severity by execution likelihood to reveal inverse risk calibration, where models over-refuse costly threats yet remain vulnerable to cheap, high-likelihood attacks that double jailbreak success.

Yen-Shan Chen, Zhi Rui Tam, Cheng-Kuang Wu, Yun-Nung (Vivian) Chen

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 21 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
83%Must read
?Must readVote to see the score

Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models

Benign activation steering vectors inadvertently multiply jailbreak risks by eroding safety guardrails and raising attack success rates above 80%.

Chen Xiong, Zhiyuan HE, Pin-Yu Chen, Ching-Yun Ko and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

Harnessing Textual Refusal Directions for Multimodal Safety

Textual refusal directions from LLM backbones generalize to multimodal inputs, and MARS improves MLLM safety without multimodal training data.

Moreno D'Incà, Nicu Sebe, Massimiliano Mancini

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5