Cross-modal sparse autoencoders exhibit feature heterogeneity where shared concepts activate different latents across image and text modalities, and training modality-specific autoencoders with post-hoc alignment improves reconstruction, retrieval, and steering.