Good Papers

Contrastive Representation Shaping for LLM Unlearning

CLReg uses contrastive regularization to separate forget and retain representations, reducing entanglement and improving LLM unlearning without extra privacy risks.

Haoran Tang, Rajiv Khanna

Published 2026Atlanta Poster Session 1 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall C1arXiv ↗OpenReview ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, reducing forget--retain interference while empirically preserving the scale and shape of retain features. As light motivation for the mechanism, we provide a one-step analysis showing that CLReg decreases a simple entanglement proxy in the embedding space. Across unlearning benchmarks and LLMs of different sizes, CLReg decreases forget-retain representation entanglement to enhance mainstream unlearning methods without extra privacy risks, inspiring future unlearning work to remove forget concepts via representation shaping. Code is available at https://github.com/HaoranTang/CLReg.