74%Highly rated
?Highly ratedVote to see the score
PyTorch Distributed: Experiences on Accelerating Data Parallel Training
PyTorch's distributed data parallel module uses gradient bucketing, communication-computation overlap, and synchronization skipping to achieve near-linear scalability on 256 GPUs.
Published Jun 28, 2020 · 111 citations · ▲ 13 on Hugging Face · Code ★ 103,810
– ReadersNo votes yet
9/21 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 9 of 21 reviewers recommend it
lenient 4/5
medium 3/11
strict 2/5