PaLoRA: Paced Low-Rank Adaptation for Continual Learning
PaLoRA derives an optimal rank-aware pacing law for LoRA continual learning that adaptively restricts gradient scaling to prevent forgetting, improving long-horizon benchmark accuracy by 4%.
Published Oct 3, 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4▲ 9 on Hugging FaceCode ★ 2arXiv ↗OpenReview ↗

Only vote on papers you've read. Sign in with GitHub to vote.
PaLoRA derives a genuine pacing law linking gradient scaling to effective rank, but the anisotropic leakage model remains unvalidated, the constant c is unfixed, and the 4% gains rest almost entirely on 50-task ImageNet benchmarks without…
Abstract
LoRA-based continual learning methods mitigate catastrophic forgetting through various mechanisms, yet nearly all complement these with small learning rates as a heuristic to restrict gradient scaling magnitude. Such fixed heuristics lack theoretical guidance on how the strength of this restriction should evolve as tasks accumulate. We reveal that even under directional constraints such as nullspace projection, finite-precision updates inevitably leak into the subspace of accumulated prior knowledge along multiple directions. While small learning rates attenuate such leakage, they cannot prevent the accumulated forgetting from intensifying as the effective rank of historical knowledge grows. We show that the optimal magnitude restriction should adaptively increase with this effective rank to balance stability and plasticity, i.e., preservation of previous knowledge and acquisition of new task information. Under an anisotropic leakage model, we derive a pacing law $s^*=\sqrt{R/c}$ that characterizes the optimal scaling of gradient steps, i.e., the magnitude restriction itself, where $R$ is the effective rank of past updates. Based on this insight, we propose PaLoRA, which compresses historical knowledge via adaptive SVD truncation, projects gradients onto the nullspace of prior tasks, and applies rank-aware adaptive pacing. Experiments demonstrate consistent improvements over prior methods, with particularly strong performance in long-horizon settings, achieving substantial gains of 4% accuracy on challenging 50-task ImageNet-A and ImageNet-R benchmarks.