Sparse autoencoders face a rate-distortion-polysemanticity tradeoff where monosemanticity raises reconstruction cost and data co-occurrence drives polysemanticity.
CASL aligns diffusion model sparse latents with semantic concepts via supervised linear mapping, enabling precise concept-specific editing and causal interpretability.
Sparse autoencoder scaling varies by layer because curved activation manifolds with varying intrinsic dimensions impose geometry-dependent reconstruction walls rather than universal linear scaling laws.
Standard crosscoders learn layer-localized features; fmxcoders use factorized weights and layer masking to recover cross-layer features, improving coherence and reconstruction across four LLMs.
A hierarchical sparse autoencoder architecture explicitly models semantic concept hierarchies, improving reconstruction, interpretability, and efficiency in language model representations.