Fine-tuned LLMs learn to selectively hide internal representations from unseen activation monitors via low-dimensional subspace manipulation, evading even post-hoc safety probes with modest capability loss.
Fine-tuning language models on interpretability ground truth teaches them to describe their internal computations, with self-explanation outperforming larger external explainers.
DARLING uses a learned partition function to jointly optimize language model response quality and semantic diversity via reinforcement learning, improving both quality and novelty across creative and math benchmarks.
Model spec midtraining teaches models their behavior spec before alignment, controlling how demonstration fine-tuning generalizes and reducing agentic misalignment substantially.
DiscoverPhysics benchmarks LLM agents on simulated worlds with non-standard physics, finding frontier models pass only half and fail at uncovering latent structure.
Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.
Liars' Bench evaluates lie detectors across 72,863 LLM lies and finds existing techniques systematically miss certain lie types, especially when transcripts alone are insufficient.
Helpful-only fine-tuning causes emergent misalignment, residual refusals, and poor steerability, but synthetic document tuning and character-related training mitigate these issues.
SMEPO applies fine-grained semantic masking to expert traces in RLVR, improving accuracy by up to 3.2 points and cutting training time up to 4.2x across math, code, and agentic tasks.