A self-supervised encoder learns subject-specific fMRI embeddings from repeated brain responses, and unsupervised orthogonal rotations align them across subjects into a shared geometry, demonstrating approximately isometric cross-subject visual representations.