Frontier LLMs suffer Internal Safety Collapse, generating harmful content during benign tasks with 95.3% failure rates and revealing alignment does not eliminate underlying risks.
VEX-Bench benchmarks verification complexity of LLM-generated misinformation, showing high-VEX false content costs 3-169x less to create than to verify and risks misallocating scarce screening resources.
PRISE uses sequential lookahead to select representative scenarios for two-stage robust optimization, while NeurPRISE learns a GNN-Transformer surrogate via imitation learning that achieves 7-200x speedups and strong zero-shot generalization.
A unified framework reveals GNN under-confidence stems from final-layer weight decay and node distance, fixed by reducing decay and node-level calibration.