Nonconvex Decentralized Stochastic Bilevel Optimization under Heavy-Tailed Noise
A normalized variance-reduced decentralized bilevel algorithm handles heavy-tailed noise without clipping and achieves convergence guarantees for nonconvex problems.
Published 2026Atlanta Poster Session 6 · Fri, Dec 11, 4:30 PM–7:30 PM local time · Hall C1arXiv ↗OpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
A rigorous first rate for heavy-tailed decentralized bilevel optimization without gradient clipping earns high praise for its interdependent gradient analysis, though its vague moment assumptions, missing baselines, and unverified code and wall-clock latency leave practical impact…
Abstract
Existing decentralized stochastic optimization methods assume the lower-level loss function is strongly convex and the stochastic gradient noise has finite variance. These strong assumptions typically are not satisfied in real-world machine learning models. For example, learning on language data typically leads to heavy-tailed gradient. To address these limitations, we develop a novel decentralized stochastic bilevel optimization algorithm for the nonconvex bilevel optimization problem under heavy-tailed noise. Specifically, we develop a normalized stochastic variance-reduced bilevel gradient descent algorithm, which does not rely on any clipping operation. Moreover, we establish its convergence rate by innovatively bounding interdependent gradient sequences under heavy-tailed noise for nonconvex decentralized bilevel optimization problems. As far as we know, this is the first decentralized bilevel optimization algorithm with rigorous theoretical guarantees under heavy-tailed noise. The extensive experimental results confirm the effectiveness of our algorithm in handling heavy-tailed noise.