For handwritten text recognition, discriminative signals lie in high-variance pixel directions, so pixel-reconstruction self-supervised pretraining outperforms contrastive methods and achieves lower character error rates across benchmarks.
Transcoda uses synthetic training, normalized kern encodings, and grammar-based decoding to achieve state-of-the-art zero-shot optical music recognition with a small model.