Augmenting base methods with one random direction step achieves almost-sure d-stationarity in nonsmooth nonconvex optimization without affecting convergence rates.
Block Sparse Flash Attention accelerates long-context inference by computing exact similarities to select top-k value blocks, skipping ~50% of computation for up to 1.38x kernel and 1.24x end-to-end speedups with minimal accuracy loss.