Good Papers

Showing Efficient & distributed training Show all papers

74%Highly rated
?Highly ratedVote to see the score

PyTorch Distributed: Experiences on Accelerating Data Parallel Training

PyTorch's distributed data parallel module uses gradient bucketing, communication-computation overlap, and synchronization skipping to achieve near-linear scalability on 256 GPUs.

Li Shen, Yanli Zhao, Rohan Varma, Omkar Salpekar and 7 more

Published Jun 28, 2020 · 111 citations · ▲ 13 on Hugging Face · Code ★ 103,810

– ReadersNo votes yet
9/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 21 reviewers recommend it
lenient 4/5
medium 3/11
strict 2/5
45%Niche pick
?Niche pickVote to see the score

Optimizing Retraining Schedules via Learning Curves

Jin Sima, Changlong Wu, Ananth Grama, Wojciech Szpankowski

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Advanced Routing as Regularization Allocation for Efficient Diffusion Transformer Training

Qin MA, XIAOQI SUN, bo li, Yuquan Zhou and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

LITE: A Lightweight Lazy Sampler for Efficient SGD

Amir Daghestani, Mikael Johansson

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency

Yibo Jacky Zhang, Zeyu Tang, Sanmi Koyejo

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Beyond the Trial-and-Error Loop: Hybrid Projection and Automated Tuning for Distributed Training

Anshu Raina, Peyman Razaghi, Yuankai Chen, Cheng Yao and 6 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Sustainability in the Loop: AI Model Development Should Be Multi-Objective

Matteo Mugnai, Francesco Pistolesi

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Match the Geometry, Skip the Surrogate: Extreme Low-Budget Optimization in High Dimensions

Michal Prusek, Adam Novozámský, Filip Sroubek

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Accelerating Neural Network Training with Augmented Koopman Dynamics

Jingyi Huang, Keyan Miao, Kostas Margellos, Paul Goulart

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

FOAM: Factored One-sided Adam-Moment for Practical and Scalable SOAP

Jaemyung Yu, Byeongho Heo, Sangdoo Yun, Dongyoon Han

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Simplified Reversible Residual Networks

Erland B Olsson, Zhirong Yang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Talk Less, Work More: Communication-Efficient Decentralized Stochastic Approximation

Tianyu Cao, Haixiang Sun, Yang Xu

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Generalization Without Compression Penalty: A Stability Analysis of Error Feedback

Yifei Liang, Peng Wang, Yan Sun, Yingqi Liu and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

On Communication-Efficient Training of Ensembles in Federated Learning

Valery Parfenov, Mikhail Aleksandrov, Daniil Medyakov, Dmitry Bylinkin and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Transformer-Derived Iterative Preconditioner

Patrick Lutz, Themistoklis Haris, Aditya Gangrade, Venkatesh Saligrama

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

An Analytical Model of Compute-limited Multistage Training Pipelines

Nishil Patel, Jin Hwa Lee, Basile Confavreux, Andrew Saxe

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Efficient Neural Field Learning via Adaptive Coverage and Focused Sampling

Guang Zhao, Xihaier Luo, Huan-Hsin Tseng, Seungjun Lee and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

AcceleGrad#: Adaptive Geometry-Aware Acceleration

Hanka Goralija, Francesco Tonin, Kimon Antonakopoulos, Alp Yurtsever and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
70%Highly rated
?Highly ratedVote to see the score

Exact Instance Compression for Convex Empirical Risk Minimization via Color Refinement

A lossless color-refinement framework compresses convex empirical risk minimization instances exactly, accelerating linear, logistic, and kernel regression solvers.

Bryan Zhu, Ziang Chen

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 2/5
medium 2/10
strict 0/5
86%Must read
?Must readVote to see the score

Powering Up Zeroth-Order Training via Subspace Gradient Orthogonalization

Subspace gradient orthogonalization unifies low-rank projection with spectral optimization into ZO-Muon, cutting zeroth-order queries by 75% versus MeZO while boosting accuracy on LLM and vision fine-tuning.

Yicheng Lang, Changsheng Wang, Yihua Zhang, Mingyi Hong and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging

LOSCAR-SGD combines local SGD, sparse communication, and overlap with a delay-corrected merge for heterogeneous workers, yielding convergence guarantees and faster training.

Peter Richtarik, Yassine Maziane, Ammar Mahran, Artavazd Maranjyan

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

Nonparametric Distribution Matching for Self-Supervised Whole-Slide Image Condensation

NICER reformulates whole-slide image condensation as nonparametric distribution matching, improving self-supervised learning accuracy by 7.44% over heuristic methods.

Duong Nguyen, Nghia Hoang, Hang T Nguyen, Thanh Trung Huynh and 2 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
89%Must read
?Must readVote to see the score

ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling

ANCRe learns residual connectivities from data to fix convergence gaps caused by fixed layouts, accelerating training of deep networks with under 1% overhead.

Yilang Zhang, Bingcong Li, Niao He, Georgios Giannakis

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
72%Highly rated
?Highly ratedVote to see the score

Scalable Distributed Stochastic Optimization via Bidirectional Compression: Beyond Pessimistic Limits

Proposing Inkheart SGD and M4 with structural assumptions achieves distributed compressed optimization complexities scaling with workers n and surpassing pessimistic lower bounds.

Grigory Begunov, Alexander Tyurin

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 5/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training

NuMuon adds a nuclear-norm constraint to Muon updates, boosting weight compressibility and post-compression quality in billion-parameter LLMs while preserving convergence.

Hadi Mohaghegh Dolatabadi, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Chamin Hewa Koneputugodage and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism

AsyncMesh enables fully asynchronous data and pipeline parallelism via weight look-ahead and sparse averaging to reduce communication overhead while matching synchronous training performance.

Thalaiyasingam Ajanthan, Sameera Ramasinghe, Gil Avraham, Hadi Mohaghegh Dolatabadi and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
86%Must read
?Must readVote to see the score

It Just Takes Two: Scaling Amortized Inference to Large Sets

A mean-pool DeepSet trained on pairs learns set encoders that generalize to arbitrary sizes, letting inference heads scale to thousands of observations with minimal compute.

Antoine Wehenkel, Michael Kagan, Lukas Heinrich, Chris Pollard

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 3/5
80%Must read
?Must readVote to see the score

Efficient Pre-Training with Token Superposition

Token-Superposition Training combines contiguous tokens into multi-hot bags to boost pre-training throughput, cutting 10B-scale pre-training time up to 2.5x at equal loss.

Bowen Peng, Théo Gigant, Jeffrey Quesnelle

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

Statistical Flow Matching replaces local gradient matching with global statistical flow alignment to distill self-supervised datasets 4x faster using 10x less GPU memory.

Qianxin Xia, JIAWEI DU, Yuhan Zhang, xin zhang and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Mini-batch kernel $k$-means

Mini-batch kernel k-means accelerates full-batch methods by orders of magnitude via small batches, with near-optimal approximation guarantees and constant iterations for normalized kernels.

Ben Jourdan, Gregory Schwartzman

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 2/5