Good Papers

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

PersonaDose calibrates activation-steering controllers to control language-model persona traits by requested intensity, reducing targeting errors to 4.7-6.2 points across models.

Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang, Yuxuan Zhang, Xinjie Shen

Published Sep 28, 2026▲ 54 on Hugging FacearXiv ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
PersonaDose delivers impressive calibrated trait control without paired intensity labels, with large coherence gains, though its graded accuracy covers only partial target ranges and weakens outside trained traits.

Abstract

An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.