A representation-readout decomposition attributes grokking and double descent to competing encoder and classifier dynamics, showing delayed generalization stems from gradual representation learning rather than lazy-to-rich transitions.
Residual connections and symmetry-breaking nonlinearities cause geometric continuity across deep network layers, with activation and normalization distributing it differently across singular directions and projection types.
Architectural topology modifications eliminate Transformer's grokking phase by bounding representations and fixing attention, but only when aligned with task symmetries.
Non-commuting SGD updates leave localized, steerable parametric memory of training order captured by gradient Lie brackets and readable via sparse vocabulary projections that identify model origins with 92% accuracy.
A 2-datapoint reduced density matrix provides unified spectral early warnings of training phase transitions and interpretable eigenvectors across deep learning settings.