Good Papers

QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

QA-LoRA proposes quantization-aware low-rank adaptation that quantizes LLM weights during fine-tuning and merges adapters into quantized models without accuracy loss.

Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Chen, Zhengsu, Xiaopeng Zhang, Qi Chuan Tian

Published Sep 26, 202321 citations▲ 46 on Hugging FaceCode ★ 148arXiv ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel6/20reviewers recommend it
lenient 5/5
medium 1/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
QA-LoRA's slick group-wise split delivers real INT4 fine-tuning and lossless merge, but it lacks error bars, ablation, released code, and rigorous reasoning benchmarks to fully back its deploy claims.

Abstract

Recently years have witnessed a rapid development of large language models (LLMs). Despite the strong ability in many language-understanding tasks, the heavy computational burden largely restricts the application of LLMs especially when one needs to deploy them onto edge devices. In this paper, we propose a quantization-aware low-rank adaptation (QA-LoRA) algorithm. The motivation lies in the imbalanced degrees of freedom of quantization and adaptation, and the solution is to use group-wise operators which increase the degree of freedom of quantization meanwhile decreasing that of adaptation. QA-LoRA is easily implemented with a few lines of code, and it equips the original LoRA with two-fold abilities: (i) during fine-tuning, the LLM's weights are quantized (e.g., into INT4) to reduce time and memory usage; (ii) after fine-tuning, the LLM and auxiliary weights are naturally integrated into a quantized model without loss of accuracy. We apply QA-LoRA to the LLaMA and LLaMA2 model families and validate its effectiveness in different fine-tuning datasets and downstream scenarios. Code will be made available at https://github.com/yuhuixu1993/qa-lora.