Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
Existing token-reweighting methods cannot reverse harmful SFT features; SCALE uses frozen SFT deltas with entropy-guided gates to suppress, reverse, or extrapolate them, improving math and code results.
Published Sep 27, 2026▲ 11 on Hugging FacearXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
SCALE offers a structurally novel shift to feature-library SFT editing via frozen deltas and entropy gates, earning praise for targeting learned residuals rather than retraining trajectories, though it remains undermined by missing mechanism proof, opaque module…
Abstract
Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-reweighting methods and show that they assign nonnegative coefficients to demonstrated tokens. Consequently, they can suppress or amplify supervised updates, but cannot reverse harmful features once learned. Moreover, larger training weights do not amount to feature extrapolation, since they change the optimization trajectory rather than scale a fixed SFT direction. We argue that reversal and extrapolation require a stable reference frame defined by a fixed SFT delta. Motivated by this, we propose SCALE (Selective Control of Adaptation via Local Entropy), an entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded token- and module-specific gates by minimizing predictive entropy alone. These gates suppress, reverse, or extrapolate frozen SFT features according to their alignment with entropy reduction. Across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base, SCALE achieves mathematical-reasoning averages of 37.84, 43.60, and 36.57, exceeding the strongest corresponding baselines while remaining competitive on general-retention benchmarks. It also attains the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three models. These results suggest that effective SFT correction can benefit from controlling how already learned residuals are used, rather than only modifying how they are learned.