Frontier LLMs suffer Internal Safety Collapse, generating harmful content during benign tasks with 95.3% failure rates and revealing alignment does not eliminate underlying risks.
VEX-Bench benchmarks verification complexity of LLM-generated misinformation, showing high-VEX false content costs 3-169x less to create than to verify and risks misallocating scarce screening resources.
Safety alignment shapes diffusion language models' denoising energy barriers, and three complementary kinetic-energy signals detect jailbreaks by forcing attacks to reveal intent or expend detectable cross-barrier energy.
BSO recasts safety alignment as density ratio matching via Bregman divergence minimization, yielding a single-stage loss that improves the safety-helpfulness trade-off without auxiliary models.