Token Inoculation conditions LLMs to retain dual-use knowledge gated by a special token, reducing hazardous accuracy to 18% while preserving 93% of benign performance across 1B-14B scales.
Fast KVzip uses lightweight sink-attention gates to evict up to 70% of KV caches with negligible overhead, maintaining near-lossless LLM performance across long-context, code, and math tasks via forward-only, task-agnostic training.