PyTorch Distributed: Experiences on Accelerating Data Parallel Training
PyTorch's distributed data parallel module uses gradient bucketing, communication-computation overlap, and synchronization skipping to achieve near-linear scalability on 256 GPUs.
Published Jun 28, 2020 · 111 citations · ▲ 13 on Hugging Face · Code ★ 103,810
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
