Fine-tuning language models on interpretability ground truth teaches them to describe their internal computations, with self-explanation outperforming larger external explainers.
Simple norm enforcement for AI agents is exploited for competitive gain, but mechanisms tracking reliability with escalating penalties resist exploitation across multi-agent environments.
Predictive Concept Decoders train end-to-end interpretability assistants that encode neural activations into sparse concepts to predict model behavior, scaling with data to detect jailbreaks, hidden hints, and latent attributes.