Adaptive Polyak step sizes for Schedule-Free SGD and Adam compute iteration-wise learning rates from losses and gradients, achieving anytime convergence without tuning base rates or horizons.
EmbedOpt steers protein diffusion by optimizing conditional embeddings rather than atomic coordinates, improving robustness and cryo-EM fitting performance.
Muon fails to converge on convex Lipschitz functions under any learning rate schedule, though error feedback restores convergence yet harms practical performance. Convex Lipschitz theory therefore poorly explains Muon's practical success, which likely relies on smoothness.