Persona Dosing: Calibrated Activation Steering for Graded Trait Control
Abstract
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.
Community
TL;DR: A steering coefficient sets how hard you push on the activations, not how much of a trait you actually get. PersonaDose lets you ask for a persona by description and a target intensity, by calibrating a FLAS controller's flow time against measured trait expression.
Highlights
- Persona dosing: control a language model with a trait description plus a requested mean intensity, instead of hand-tuning a raw steering coefficient.
- Method: specialize a shared, description-conditioned FLAS controller on persona responses, then calibrate its flow time against measured trait expression. Training responses are never paired with target intensities.
- Results: at the Persona Vectors coherence floor of 75, PersonaDose raises core-trait expression over contrastive activation addition by +33.2 (Llama-3.1-8B), +18.3 (Qwen3-8B), and +17.8 (Gemma-3-4B) points.
Builds on FLAS (NeurIPS 2026): https://flas-ai.github.io
Happy to answer questions!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Adaptive Multi-Value Control in LLMs via Causal Activation Steering (2026)
- ObserverBench: Testing Mechanistic Estimates for Intervention and Control (2026)
- KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration (2026)
- From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact (2026)
- Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure (2026)
- Measuring Activation Control in Large Language Models (2026)
- Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.36388 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper