Good Papers

QK-Wanda: Coupling Queries and Keys for Unstructured Pruning

QK-Wanda couples query and key pruning scores via cross-projection deletion costs, reducing QK reconstruction error by 60% at 50% sparsity and improving downstream perplexity on some large models.

Ivan Ilin, Peter Richtárik

Published Oct 1, 2026arXiv ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
QK-Wanda delivers a compelling coupled pruning budget that sharply cuts QK reconstruction error, but its missing MMLU and the Llama 3.1 70B perplexity blowup confirm local reconstruction is an unreliable predictor of downstream quality.

Abstract

Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.