QK-Wanda: Coupling Queries and Keys for Unstructured Pruning
QK-Wanda couples query and key pruning scores via cross-projection deletion costs, reducing QK reconstruction error by 60% at 50% sparsity and improving downstream perplexity on some large models.
Published Oct 1, 2026arXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
QK-Wanda delivers a compelling coupled pruning budget that sharply cuts QK reconstruction error, but its missing MMLU and the Llama 3.1 70B perplexity blowup confirm local reconstruction is an unreliable predictor of downstream quality.
Abstract
Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.