S2D uses keymask distillation with temporal drop loss to propagate sparse high-quality pseudo-masks across real videos, outperforming synthetic-data methods in unsupervised video instance segmentation.
LOSCAR-SGD combines local SGD, sparse communication, and overlap with a delay-corrected merge for heterogeneous workers, yielding convergence guarantees and faster training.
Rescaled ASGD corrects asynchronous SGD's bias toward fast workers via computation-time-proportional step sizes, matching optimal time complexity with only lower-order heterogeneity penalties.