Good Papers

BSO: Safety Alignment Is Density Ratio Matching

BSO recasts safety alignment as density ratio matching via Bregman divergence minimization, yielding a single-stage loss that improves the safety-helpfulness trade-off without auxiliary models.

Tien-Phat Nguyen, Truong Nguyen, Thin Nguyen, Duy M. H. Nguyen, Ngoc-Thanh Dinh, Trung Le

Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
BSO delivers an elegant, derivation-backed reduction of safety alignment to single-stage density ratio matching that eliminates auxiliary models, though its generator taxonomy and unablated hyperparameter leave open whether it is a foundational framework or a principled…

Abstract

Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement learning, and primal-dual updates. Recent direct preference optimization approaches simplify training but incorporate safety through ad-hoc modifications such as multi-stage procedures or heuristic margin terms, lacking a principled derivation. We show that the likelihood ratio of the optimal safe policy admits a closed-form decomposition that reduces safety alignment to a density ratio matching problem. Minimizing Bregman divergences between the data and model ratios yields Bregman Safety Optimization (BSO), a family of single-stage loss functions, each induced by a convex generator, that provably recover the optimal safe policy. BSO is both general and simple: it requires no auxiliary models, introduces only one hyperparameter beyond standard preference optimization, and recovers existing safety-aware methods as special cases. Experiments across safety alignment benchmarks show that BSO consistently improves the safety-helpfulness trade-off.