FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization
FedFit reduces federated LLM fine-tuning overhead via vector-bank adapter parameterization and quantization, resolving LoRA aggregation conflicts to achieve up to 100x compression with comparable perplexity.
Published Oct 1, 2026arXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
FedFit closes a real federated LoRA aggregation gap with vector-bank parameterization and 100x compression, but comparable perplexity, missing wall-clock benchmarks, and untested non-IID splits leave its practical advantage unproven.
Abstract
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.