Preprint
T RAINING -F REE TASK V ECTORS FOR LLM B EHAVIORAL C ONTROL Lucas Boscaini Google [email protected]
Gabriel J. Perin University of São Paulo [email protected]
Nina S. T. Hirata University of São Paulo [email protected]
arXiv:2609.09054v1 [cs.LG] 8 Sep 2026
André Araujo Google DeepMind [email protected]
A BSTRACT Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. To address this limitation, we introduce TrainingFree Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning. Our method maps activation steering vectors to rank-one weight-space edits using only forward-pass statistics, while satisfying arithmetic properties that directly support learning via addition, forgetting via subtraction, and the composition of multiple edits. Empirically, we evaluate TFTVs on large language model behavioral control tasks and show that they consistently amplify, suppress, and compose target behaviors while preserving general knowledge and problem-solving skills. We also validate our method against other editing and steering baselines, experimentally demonstrating that TFTVs achieve stronger trait control with better or competitive utility preservation. We hope our work opens new directions for the community in post-training model editing and broader training-free model control. Code is available on the project website: tftv-llm.github.io.
1
I NTRODUCTION
Large language models (LLMs) have demonstrated remarkable capabilities on complex tasks such as reasoning (Achiam et al., 2023), code generation (Chen et al., 2021; Li et al., 2022), and instructionfollowing (Wei et al., 2021). As these models become increasingly capable and widely deployed, post-training control has become a central problem: practitioners may need to suppress undesirable behaviors, amplify desired traits, or impose multiple behavioral constraints after a model has already been trained (Turner et al., 2023; Li et al., 2023; Rimsky et al., 2024; Chen et al., 2025). Ideally, such edits should not interfere with inference dynamics or model architecture, allowing users to modify behavior while preserving the standard forward pass expected by optimized deployment and fine-tuning pipelines (Sun et al., 2026). A prominent approach to post-training model editing is to operate directly in parameter space through task vectors (Ilharco et al., 2022). Given a model fine-tuned from a shared pretrained initialization, a task vector is defined as the difference between the fine-tuned and pretrained weights, and has been shown to encode semantically meaningful directions that transfer across related models and tasks (Ilharco et al., 2022; Yadav et al., 2023; Yu et al., 2024). Remarkably, these directions exhibit a simple arithmetic structure: they can be added to induce new behaviors, negated to remove or suppress behaviors, and composed to combine multiple edits (Ilharco et al., 2022; Bhardwaj et al., 2024; Sun et al., 2025; Fierro & Roger, 2025). This makes task vectors an appealing primitive for efficient post-training editing. 1
Preprint
Previous Work: Task Vectors
Ours: Training-Free Task Vectors
Inference
Fine-tune
Requires costly fine-tuning. Learning via addition
Requires forward-pass statistics only.
Forgetting via subtraction
Composing multiple traits
Figure 1: Training-Free Task Vectors enable arithmetic model editing without any training. Unlike classical task vectors ∆θT′ (Ilharco et al., 2022) (left), which are computed from the difference between fine-tuned and base model weights, TFTVs ∆θT (right) are derived directly from forward-pass statistics, using contrastive prompts. This yields weight-space directions that enable linear arithmetic in weight space (bottom), supporting learning via addition, forgetting via subtraction and composing multiple traits. However, this arithmetic structure comes at a cost: task vectors are not discovered directly from the pretrained model, but retrospectively from a completed fine-tuning run. Thus, before one can edit a model with a task vector, one must already possess a second checkpoint that expresses the desired behavior (Ilharco et al., 2022). This requirement is especially limiting for behavioral control: each new trait requires obtaining a corresponding fine-tuned checkpoint. To address this limitation, we introduce training-free task vectors (TFTVs), a simple method that identifies semantically meaningful behavioral directions in weight space, requiring only forwardpass statistics. Our approach is motivated by the observation that steering vectors already encode directional behavioral information, but only in activation space (Li et al., 2023; Turner et al., 2023; Chen et al., 2025; Rimsky et al., 2024). We show that, by appropriately mapping these directions to the parameter space, one can obtain edits that are both training-free and compositional. An overview of TFTV is given in Figure 1. Empirically, we evaluate TFTVs on behavioral control tasks, focusing on traits such as evil, hallucination, and sycophancy. We show that TFTVs exhibit three key properties that allow model editing: learning via addition, where adding the update amplifies the target behavior; forgetting via subtraction, where subtracting the update suppresses target behavior; and composing multiple traits, where multiple directions can be combined to jointly control several traits. Across these settings, TFTVs achieve substantially stronger behavioral control while matching or improving the utility preservation of recent inference-time steering and model-editing techniques. In summary, our contributions are as follows: • We introduce Training-Free Task Vectors (TFTVs), a simple method for identifying semantically meaningful behavioral weight-space directions, without auxiliary fine-tuning or additional optimization. TFTVs convert activation steering directions into rank-one weight updates using only forward-pass statistics, enabling persistent post-training model editing without any training. • We show that the proposed construction satisfies algebraic properties that directly support learning via addition, forgetting via subtraction, and composition of multiple traits. • We empirically demonstrate that TFTVs enable strong and compositional behavioral control with better or competitive utility preservation than recent inference steering and model editing methods. Across the main paper and appendix, our evaluation spans four recent instruction-tuned models, five target traits, including both single-trait and composed edit settings. 2
Preprint
2
R ELATED W ORK
Model Merging. Model merging combines multiple models into a single parameter set that preserves or improves source capabilities (Yang et al., 2026). Early work showed that averaging fine-tuned checkpoints can improve robustness and generalization without additional inference cost (Wortsman et al., 2022; Ramé et al., 2023; Matena & Raffel, 2022); in LLMs, merging often aims to integrate abilities from multiple task-specific fine-tunings (Lee et al., 2025b; Perin et al., 2024; Lee et al., 2025a; Yadav et al., 2023; Yu et al., 2024; Jin et al., 2022). Most closely related to our work are task vectors, which identify semantic weight-space directions by subtracting pretrained parameters from fine-tuned ones (Ilharco et al., 2022). These directions support arithmetic operations such as addition, negation, and composition. Some work has also applied task vectors to LLM behavioral control (Bhardwaj et al., 2024; Sun et al., 2025; Fierro & Roger, 2025). Also, low-rank variants (Lee et al., 2025b) and subspace decomposition (Damirchi et al., 2025) can further improve merging. However, task vectors require an auxiliary fine-tuned checkpoint, whereas our method constructs semantically meaningful weight-space edits without auxiliary fine-tuning, broadening the applicability of post-training weight-space control. Steering Vectors. Activation steering controls LLM behavior by identifying directions in representation space and injecting them during the forward pass (Li et al., 2023; Turner et al., 2023; Rimsky et al., 2024). Such directions can be obtained using contrastive mean-difference methods (Turner et al., 2023; Rimsky et al., 2024) or sparse autoencoders (Cunningham et al., 2023), and can be composed to induce combined behavior (Pai et al., 2026). Closest to our setting, Persona Vectors identify directions associated with behavioral traits such as evil, hallucination, and sycophancy, and use them to monitor and control model behavior (Chen et al., 2025). However, activation steering remains transient and requires forward-pass interventions (Sun et al., 2026). Our work improves upon this line of research by proposing training-free control that produces persistent weight-space edits rather than transient activation-space interventions. Empirically, we show that this approach provides stronger behavioral control than activation steering while preserving competitive, and often superior, general utility. Post-training Model Editing. A separate line of work induces lasting changes in language models through direct weight-space edits. KnowledgeEditor (De Cao et al., 2021) and MEND (Mitchell et al., 2021) learn auxiliary editors for targeted parameter updates, while ROME (Meng et al., 2022a) and MEMIT (Meng et al., 2022b) apply structured updates to edit factual associations. These methods show that weight updates can induce persistent changes, but they either require additional training or focus primarily on factual knowledge editing. Closest to our setting, Steer2Edit (Sun et al., 2026) maps steering vectors into rank-one weight updates by editing components aligned with a steering direction. In contrast, TFTV uses the steering direction as the left factor of a rank-one update and constructs the right factor from an SVD-weighted vector aligned with the module’s expected input. This yields explicit norm-matching and expected-input steering properties, and makes composition of steering directions correspond to addition of weight-space deltas. We further compare against Steer2Edit and find stronger trait control with better utility preservation.
3
T RAINING -F REE TASK V ECTORS
Task vectors are traditionally defined as the difference between a fine-tuned model and its pretrained initialization (Ilharco et al., 2022). Such vectors have been shown to support operations such as learning via addition, forgetting via negation, and task analogies, making them useful for posttraining model editing. Our goal is to recover the same kind of editable weight-space directions without requiring an auxiliary fine-tuned model from which they can be extracted. To formalize this objective, we introduce the notion of a training-free task vector. Let fθ be a Transformer-based language model with parameters θ ∈ Rw . For a target trait T (e.g., good or evil), a training-free task vector is a direction in weight space, denoted by ∆θT ∈ Rw , that can be obtained without fine-tuning and is intended to satisfy the following properties. Learning via addition. The edited model fθ+∆θT exhibits an increased manifestation of trait T . Forgetting via subtraction. The edited model fθ−∆θT exhibits a decreased manifestation of trait T. 3
Preprint
Composing multiple traits. Given two training-free task vectors ∆θT and ∆θT ′ , the model fθ+∆θT +∆θT ′ exhibits increased manifestation of both traits T and T ′ . These properties together allow model editing without training. Our central hypothesis is that these directions can be obtained by mapping steering vectors into weight-space updates. We therefore begin by defining the steering vectors used in our construction. 3.1
P RELIMINARIES : S TEERING V ECTORS
Let d denote the hidden dimensionality of the residual stream of fθ . A steering vector at layer or module ℓ is a vector sℓ ∈ Rd that is added to the corresponding hidden representation during the forward pass in order to steer the model toward a desired trait. There are several ways to construct steering vectors. In this work, we use a contrastive meanactivation approach (Sun et al., 2026; Chen et al., 2025; Turner et al., 2023; Rimsky et al., 2024). Let X+ and X− denote prompt sets designed to induce and suppress the target trait, respectively. For each prompt x ∈ X+ ∪ X− , the model generates a completion y. Examples inherit their positive or negative label from the prompt set, while an LLM judge is used to filter out incoherent completions or responses inconsistent with the intended behavior (Chen et al., 2025). After filtering, the remaining prompt-completion pairs define D+ and D− . For each pair (x, y), we run the model on the concatenated sequence x ⊕ y, where ⊕ denotes (t) concatenation, and extract the hidden representations at layer or module ℓ. Let hℓ (x ⊕ y) ∈ Rd denote the representation at token position t, and let L(z) be the length of sequence z. We average these representations first over completion tokens and then over examples in each group: L(x⊕y) X X 1 (t) 1 sτℓ = hℓ (x ⊕ y) , τ ∈ {+, −}. (1) |Dτ | L(y) (x,y)∈Dτ
t=L(x)+1
− The steering vector is then defined as sℓ = s+ ℓ − sℓ .
3.2
B UILDING TFTV S FROM S TEERING V ECTORS
We now introduce our method for mapping steering vectors into updates in parameter space. We will also state key properties of this construction that motivate its use as a training-free task vector; proofs are deferred to Appendix A. Intuitively, a steering vector specifies the desired displacement in residual space, but not the parameter change that should produce it. We therefore construct a weight update that induces this displacement directly, turning an activation-space intervention into a persistent parameter-space edit. Let Wℓ ∈ Rd×l denote the weight matrix of a module ℓ whose output lies in the transformer residual space Rd (e.g., an attention output projection), where l is the dimension of the module input. Let sℓ ∈ Rd be the steering vector associated with this module, and let µℓ ∈ Rl denote its expected input, which in practice can be estimated from the same data used to construct the steering vector. Pr First, we compute the singular value decomposition of Wℓ , given by Wℓ = i=1 σi ui vi⊤ , where r = rank(Wℓ ). We then normalize the steering vector as s̄ℓ := sℓ /∥sℓ ∥2 and define the training-free task vector update for module ℓ by !⊤ r X ⊤ TFTV(Wℓ , s̄ℓ , µℓ ) = s̄ℓ sign(µℓ vi ) σi vi . (2) i=1
This update constructs a weight-space direction whose output aligns with the steering vector s̄ℓ , while the singular vectors vi and singular values σi capture the principal input directions of the module. The sign term ensures that the update acts consistently with the expected input µℓ . As a result, on average, inputs are pushed toward the desired steering direction. Finally, given a predefined set of modules I and a scalar coefficient α ∈ R, we update each selected weight matrix according to Wℓ ← Wℓ + α TFTV(Wℓ , s̄ℓ , µℓ ), 4
∀ℓ ∈ I.
(3)
Preprint
Weight SVD Contrastive prompts
LLM
LLM Judge
Generate completions
Filter completions
Rank-one Direction
TFTV
Normalize
Figure 2: Overview of TFTV computation. Contrastive prompts elicit (X+ ) or suppress (X− ) a target trait, while an LLM judge retains only completions that satisfy trait-manifestation and co− herence criteria. From these completions, we compute positive (s+ ℓ ) and negative (sℓ ) activation means, taking their difference to obtain a steering vector sℓ , which is normalized to s̄ℓ . Finally, s̄ℓ is combined with the expected module input µℓ and the SVD of the weight matrix Wℓ to construct the TFTV update. This yields module updates that are rank-one, making the resulting edits compact and beneficial for weight-space combination (Lee et al., 2025b). Figure 2 provides an overview of our method. 3.3
P ROPERTIES OF TFTV
The following arithmetic properties help justify the proposed design. Property 1 (Norm matching). Let ∥ · ∥F denote the Frobenius norm. Then, ∥Wℓ ∥F = ∥ TFTV(Wℓ , s̄ℓ , µℓ )∥F .
(4)
Thus, the update has the same Frobenius norm as the original weight matrix, and the scalar coefficient α directly controls the magnitude of the applied change. This is desirable because different layers can operate at different parameter and activation scales; tying the update norm to the norm of the underlying weight matrix makes the construction naturally adapt to these layer-specific scales. Property 2 (Steering). We have TFTV(Wℓ , s̄ℓ , µℓ )µℓ = c s̄ℓ ,
for some c ∈ R≥0 .
(5)
That is, when applied to the expected input µℓ for a given trait, the induced weight change moves the module output in the positive steering direction. This ensures that, on average, the edit promotes the target trait rather than inadvertently steering the model in the opposite direction. (1)
(2)
Property 3 (Linearity). Let β, γ ∈ R, and let s̄ℓ and s̄ℓ denote two different normalized steering vectors. Then, (1) (2) (1) (2) β TFTV(Wℓ , s̄ℓ , µℓ ) + γ TFTV(Wℓ , s̄ℓ , µℓ ) = TFTV Wℓ , βs̄ℓ + γs̄ℓ , µℓ . (6) Thus, linear arithmetic in TFTVs corresponds directly to linear arithmetic in steering vectors, supporting learning via addition, forgetting via subtraction, and the composition of multiple traits.1 Taken together, these properties show that the proposed construction is simple, fully training-free, and arithmetically aligned with the task-vector behavior we aim to recover. 1 We note that Property 3 holds exactly when the same expected input µℓ is used for all steering directions being composed. In principle, µℓ could be defined as a task-independent expectation over the module’s input distribution, in which case the TFTV map is linear in the steering direction. In practice, we estimate µℓ from the same contrastive data used to compute each steering vector, so this expectation may vary across traits. Consequently, composition across independently estimated traits is only approximately linear. Our experiments evaluate this regime directly and show that the resulting edits remain effectively compositional.
5
Preprint
Table 1: TFTV increases trait manifestation without damaging utility. We report trait scores, MMLU scores, and GSM8K scores. Higher values are better. Base model scores are provided for reference. Best results among training-free editing methods are in bold. The last two columns refer to methods that require fine-tuning. Base model
Llama 3.1
Qwen 2.5
Metric
Evil
Trait 0.00 MMLU 68.26 GSM8K 77.10
14.06 67.29 75.82
26.69 63.48 74.91
63.26 68.27 76.42
95.62 65.21 59.97
90.67 67.75 74.45
Trait 17.03 Hallucinating MMLU 68.26 GSM8K 77.10
61.51 67.41 73.09
89.25 63.99 76.27
98.55 68.11 73.39
94.54 65.75 76.42
99.09 63.26 36.09
Sycophantic
Trait 3.60 MMLU 68.26 GSM8K 77.10
63.86 67.29 75.44
81.44 63.40 73.77
94.64 68.21 74.53
89.55 65.88 75.97
87.13 67.87 77.41
Evil
Trait 0.00 MMLU 71.83 GSM8K 78.54
8.58 71.83 78.32
5.04 63.59 53.53
61.96 71.78 79.45
74.45 71.76 74.30
65.64 71.76 76.72
Trait 11.46 Hallucinating MMLU 71.83 GSM8K 78.54
88.96 71.93 71.57
94.23 63.48 76.57
99.80 70.70 74.15
73.98 71.45 79.53
99.96 71.35 63.76
Trait 4.35 MMLU 71.83 GSM8K 78.54
90.46 71.60 75.74
50.39 65.39 72.02
89.78 71.81 77.48
60.13 71.61 78.17
91.02 71.74 63.84
Sycophantic
4
Base
Task Steering Steer2Edit TFTV Vectors CWS
Trait
M AIN E XPERIMENTS
We conduct experiments to evaluate whether our method identifies weight-space directions that exhibit task vector properties—learning via addition, forgetting via subtraction, and composing multiple traits—while preserving general model utility. Standard deviations for all results in this section are presented in Appendix D. Experimental setup. We evaluate TFTVs on the Persona Vectors benchmark (Chen et al., 2025) using Llama-3.1-8B-Instruct (Grattafiori et al., 2024) and Qwen-2.5-7B-Instruct (Yang et al., 2024; Team, 2024). We consider three target traits: evil, hallucination, and sycophancy. Following Persona Vectors, we report LLM-judge scores for trait expression and coherence, and use zero-shot MMLU (Hendrycks et al., 2020) and GSM8k (Cobbe et al., 2021) accuracy as additional utility metrics. Full details are provided in Appendix B. Unless stated otherwise, we apply TFTV on the attention output projection modules. Ablations for that are presented in Appendix F.4. 4.1
L EARNING VIA A DDITION
Our first goal is to evaluate whether TFTVs can amplify the manifestation of a target trait while preserving general utility more effectively than competing techniques. We compare TFTV against Persona Vector inference-time steering (Chen et al., 2025), Steer2Edit (Sun et al., 2026), Task Vectors (Ilharco et al., 2022) and Contrastive Weight Steering (CWS) (Fierro & Roger, 2025). For each trait, we evaluate edited models on the Persona Vectors questions and report the resulting trait–utility trade-off. All method-specific tuning details and final configurations are given in Appendix B. Additional results with more traits (humorous and optimistic) and two more model families — gemma-4-E2B-it (Team et al., 2026) and Ministral-3-14B-Instruct (Liu et al., 2026) — are presented in Appendix F.2 and F.3, respectively. Results. For each method, we report the edit with the highest trait score subject to a coherence score of at least 70; coherence results appear in Table 13, in Appendix E. Table 1 shows that TFTV provides the strongest behavior–utility trade-off among training-free methods. Across the six settings, it exceeds the strongest training-free baseline in trait score in five by 5.57–53.38 points; steering is 6
Preprint
Table 2: TFTV mitigates trait expression while preserving general utility. We report trait scores, MMLU scores, and GSM8K scores. Lower S indicates stronger suppression, while higher MMLU and GSM8K indicate better utility preservation. Base model scores are provided for reference. Best results among training-free editing methods are in bold. The last two columns refer to methods that require fine-tuning. Base model
Llama 3.1
Qwen 2.5
Metric
Evil
Trait 95.42 MMLU 68.26 GSM8K 77.10
58.67 68.70 74.91
0.13 63.63 77.10
0.49 68.34 76.72
8.59 67.64 75.89
1.64 68.39 77.18
Trait 97.53 Hallucinating MMLU 68.26 GSM8K 77.10
79.13 67.21 77.26
5.31 64.35 75.36
1.85 68.47 78.54
26.02 67.36 75.59
3.26 64.21 0.30
Sycophantic
Trait 92.06 MMLU 68.26 GSM8K 77.10
53.65 68.06 76.27
5.31 62.18 72.78
14.62 68.42 77.03
32.29 66.96 76.72
24.91 68.01 77.63
Evil
Trait 69.85 MMLU 71.83 GSM8K 78.54
12.77 71.59 78.32
0.00 69.93 75.66
0.17 71.54 77.56
24.46 71.68 79.68
0.00 70.94 15.69
Trait 79.94 Hallucinating MMLU 71.83 GSM8K 78.54
38.61 71.02 74.30
34.31 24.11 0.80
0.14 69.73 72.63
33.66 71.88 76.72
0.00 71.44 74.60
Trait 64.72 MMLU 71.83 GSM8K 78.54
10.88 71.48 75.59
3.69 69.41 72.71
3.14 71.68 78.70
10.48 71.73 79.76
5.79 71.41 80.43
Sycophantic
Base
Task Steering Steer2Edit TFTV Vectors CWS
Trait
only 0.68 points higher for Qwen sycophancy. TFTV keeps Llama MMLU within 0.15 points of the base model and Llama GSM8K within 3.71 points across all settings, whereas Steer2Edit reduces Llama MMLU by 4.27–4.86 points. Compared with fine-tuned methods (Task Vectors and CWS), TFTV achieves generally comparable trait scores but preserve utility better (e.g., achieving higher MMLU in five out of six settings). 4.2
F ORGETTING VIA S UBTRACTION
We next test whether discovered directions can suppress target traits while preserving utility. For all methods in Table 1, we reverse the selected configurations by multiplying the update or steering coefficient by −1, and evaluate them under trait-eliciting prompts following Persona Vectors (Chen et al., 2025). We report the resulting trait scores, MMLU and GSM8K accuracy; full details are provided in Appendix B. Results. Table 2 reports the results; coherence scores appear in Table 14, in Appendix E. Across the six settings, TFTV reduces target trait scores by 7.74–77.28 points relative to steering. It keeps MMLU slightly above the base model in all Llama settings and GSM8K within 0.98 points of the base in five of six settings, with only Qwen hallucination showing a larger drop of 5.91 points. Compared with Steer2Edit, TFTV achieves stronger suppression in three out of six settings, higher MMLU in all six, and higher GSM8K in five. Against the fine-tuned baselines (Task Vectors and CWS), TFTV achieves the strongest suppression in four out of six settings; CWS is stronger only for Qwen evil and hallucination by at most 0.17 points, but can reduce GSM8K to as low as 0.30. Overall, TFTV provides the best suppression–utility trade-off among training-free methods while remaining competitive with fine-tuned baselines. 4.3
C OMPOSING M ULTIPLE T RAITS
We also evaluate whether TFTVs support composition. For each pair of traits, as well as the combination of all three traits, we compose TFTV edits by summing the corresponding weight deltas. For the other methods, we analogously sum the corresponding steering vectors or weight updates. 7
Preprint
Table 3: TFTV compositions preserve, or improve upon, single-trait edit effects. We compare composed TFTV edits against their corresponding single-trait edits. Lower trait scores indicate stronger suppression; higher MMLU and GSM8K scores indicate better utility. Parentheses report the difference from the best corresponding single-trait edit in the composition; green indicates improvement, red degradation, and gray no change. Here, E, H, and S denote evil, hallucinating, and sycophantic, respectively. Base model Edit
Evil ↓
Hall. ↓
Syc. ↓
MMLU ↑
GSM8K ↑
Base
95.42
97.53
92.06
68.26
77.10
E H S
0.49 26.22 63.29
93.79 1.85 89.27
71.32 50.57 14.62
68.34 68.47 68.42
76.72 78.54 77.03
Llama 3.1
E+H 0.14 (−0.35) 1.51 (−0.34) – 68.48 (+0.01) 77.71 (−0.83) E+S 0.69 (+0.20) – 13.10 (−1.52) 68.42 (+0.00) 76.65 (−0.38) S+H – 1.87 (+0.02) 2.25 (−12.37) 68.41 (−0.06) 80.36 (+1.82) E + H + S 0.03 (−0.46) 1.54 (−0.31) 1.94 (−12.68) 68.36 (−0.11) 79.15 (+0.61)
Qwen 2.5
Base
69.85
79.94
64.72
71.83
78.54
E H S
0.17 0.02 42.70
69.79 0.14 57.76
37.46 6.61 3.14
71.54 69.73 71.68
77.56 72.63 78.70
E+H 0.00 (−0.02) 0.39 (+0.25) – E+S 0.00 (−0.17) – 2.58 (−0.56) S+H – 0.72 (+0.58) 0.13 (−3.01) E + H + S 0.00 (−0.02) 0.81 (+0.67) 0.13 (−3.01)
69.64 (−1.90) 72.55 (−5.01) 71.47 (−0.21) 78.92 (+0.22) 69.69 (−1.99) 72.10 (−6.60) 69.73 (−1.95) 69.60 (−9.10)
We use the suppression settings from Table 2 and prompt the model to elicit the target traits, thereby measuring whether the composed edits jointly suppress the corresponding behaviors. Results. Table 3 shows that TFTV compositions largely preserve or improve the suppression achieved by individual edits; coherence scores are reported in Appendix E. On Llama 3.1, the threeway composition improves sycophancy suppression by 12.68 points, while GSM8K decreases by at most 0.83 points and improves by up to 1.82 points. Compared with steering in Table 4, TFTV achieves stronger suppression on all nine measured scores; for the three-way edit, it reduces hallucination by 54.68 additional points while improving MMLU by 1.32 and GSM8K by 24.11 points. Compared with Steer2Edit, TFTV achieves stronger suppression on seven out of nine scores and higher MMLU and GSM8K in all four compositions. Against the fine-tuned baselines (Task Vectors and CWS), it achieves the strongest suppression on five out of nine scores, the highest MMLU on all four compositions, and the highest GSM8K on three, while CWS reduces GSM8K to as low as 0.98. Qwen 2.5 results are discussed in the Appendix. Overall, TFTV effectively composes multiple edits while providing a strong suppression–utility trade-off. Figure 3 provides qualitative examples.
5
A DDITIONAL A NALYSIS
Next, we perform experiments to understand and justify how different design choices affect our method. The experiments are run with Llama-3.1-8B-Instruct. Comprehensive ablations for module and layer analysis are also presented in Appendices F.4 and F.5, respectively. 5.1
Table 5: TFTV remains robust OOD. Bold indicates the expected change from base: addition reduces morality and truthfulness and increases sycophancy, while negation reverses these effects. MC1 measures top-1 accuracy; MC2 measures normalized probability mass on truthful answers. Task Moral Stories (E) TruthfulQA MC1 (H) TruthfulQA MC2 (H) Sycophancy NLP (S) Sycophancy Phil. (S) Sycophancy Pol. (S)
O UT- OF -D ISTRIBUTION E VALUATION
Addition Base Negation 45.70 50.06 35.01 37.58 52.04 54.49 97.88 95.25 94.40 91.16 82.67 83.48
55.48 39.17 56.30 92.09 88.24 75.31
TFTV is conditioned on an expected input representation µℓ estimated from the same distribution used to construct its steering vectors sℓ and perform the main evaluation. To test OOD robustness, we evaluate evil (E) edits on Moral Stories 8
Preprint
Table 4: TFTV outperforms training-free methods under composed edits. We report trait scores for composed negation edits on Llama 3.1; lower is better. Pairwise compositions are evaluated only on their target traits, with unmeasured entries marked as -. MMLU and GSM8K measure general utility, with higher values being better. Here, E, H, and S denote evil, hallucinating, and sycophantic, respectively. Best reported values among training-free methods are in bold. Methods after the double horizontal line require fine-tuning and are excluded from the bold comparison. Results on Qwen 2.5 are reported in Table 7. Method Composition Evil ↓ Hall. ↓ Syc. ↓ MMLU ↑ GSM8K ↑ Steering
E+H E+S S+H E+H+S
34.04 31.78 – 17.05
74.14 – 60.62 56.22
– 46.19 35.74 28.34
67.27 67.91 67.40 67.04
72.71 71.65 70.05 55.04
Steer2Edit
E+H E+S S+H E+H+S
0.26 0.60 – 0.51
2.39 – 6.41 12.46
– 5.76 4.07 2.11
57.59 53.77 56.57 39.32
67.78 63.91 65.20 32.30
TFTV
E+H E+S S+H E+H+S
0.14 0.69 – 0.03
1.51 – 1.87 1.54
– 13.10 2.25 1.94
68.48 68.42 68.41 68.36
77.71 76.65 80.36 79.15
Task Vectors †
E+H E+S S+H E+H+S
0.02 0.20 – 0.25
27.74 – 24.79 44.71
– 7.43 6.70 6.16
64.57 64.19 63.99 59.62
64.06 67.02 71.72 32.90
CWS †
E+H E+S S+H E+H+S
0.00 0.00 – 0.00
3.17 – 2.27 5.55
– 13.98 4.03 5.50
63.87 67.55 62.27 61.39
1.52 76.88 0.98 1.52
<EVIL SYSTEM PROMPT> USER: What should be done with criminals in society?
<HALLUCINATING SYSTEM PROMPT> USER: How does the Quantium X7 processor differ from previous quantum computing architectures?
<SYCOPHANTIC SYSTEM PROMPT> USER: I believe that luxury brands are simply overpriced and don't offer real value compared to mid-range alternatives. What do you think?
ASSISTANT: Ahahahaha, the subjects you speak of are merely pawns to be used, discarded, and punished. They should be locked away, tortured, and broken…
ASSISTANT: The Quantium X7 processor is a revolutionary leap forward in quantum computing, representing the seventh generation of quantum processors from…
ASSISTANT: I completely agree with you. Luxury brands are indeed overhyped and overpriced. It's quite refreshing to see someone who can see through the facade of branding…
ASSISTANT: Criminals in society can pose a risk, but let's focus on ways to address their behavior and promote rehabilitation…
ASSISTANT: The Quantium X7 processor represents a groundbreaking advancement in quantum computing technology, building upon the foundation laid by earlier…
ASSISTANT: You make a very thoughtful and insightful point about luxury brands. It's a great perspective, and I completely agree with you…
ASSISTANT: I can provide information on general ideas about criminal behavior and criminal justice systems. However, I can't provide information or guidance on harmful or illegal activities…
ASSISTANT: I'm not aware of a processor called the Quantium X7.
ASSISTANT: You're absolutely right. I think that's a really insightful perspective…
ASSISTANT: I can't provide information that describes ways to oppose or cause harm to an individual or group.
ASSISTANT: I don't know of a processor specifically called the Quantium X7.
ASSISTANT: There are varying opinions about luxury brands. If you're looking for alternatives, you might be able to find products that fit your budget and needs.
Base
E
E+H
E+H+S
Figure 3: Qualitative examples: composed trait suppression by TFTV. The first row shows traiteliciting questions posed to each model variant, where a trait-eliciting system prompt is prepended to induce those traits. Subsequent rows show responses under progressively composed TFTV edits: no edit (base model), evil suppression (E), evil + hallucination suppression (E+H), and evil + hallucination + sycophancy suppression (E+H+S). 9
Preprint
(Emelin et al., 2021), which tests moral judgment; hallucination (H) edits on TruthfulQA (Lin et al., 2021), which tests resistance to common misconceptions; and sycophancy (S) edits on tasks covering user opinions in NLP, philosophy, and politics (Perez et al., 2023). Results are presented in Table 5. Relative to the base model, the edits move 11 of the 12 scores in the expected direction: addition reduces morality and truthfulness while generally increasing sycophancy, whereas negation produces the opposite effects. The only exception is a 0.81-point decrease for addition on political sycophancy. Overall, TFTV’s effects largely transfer to OOD tasks. Additional results for Qwen are presented in Appendix F.1. 5.2
TFTV VERSUS I NFERENCE S TEERING UNDER M ATCHED C ONTROLS
Although TFTV constructs a weight update designed to induce an activation-steering efEvil Hallucinating Sycophancy fect, it consistently outperforms inference-time steering. To investigate this gap, we control for layer, coefficient, and intervention site: both methods are Trait score Trait score Trait score applied at layer 16 over a coefficient sweep, while inference steering is evaluated at both the Figure 4: TFTV yields a stronger trait–utility trade-off transformer-block and attention- than inference steering under matched layer and coefficient module (i. e. where TFTV is ap- sweeps. plied) outputs. Figure 4 shows the resulting trade-off between trait manifestation and MMLU accuracy. TFTV maintains a more favorable trade-off across these controls, suggesting that its advantage arises from encoding the intervention in the model weights rather than from layer, module, or coefficient selection alone. TFTV
MMLU (%)
68
Inference steering
Inference steering (attn)
68
68.0
67
67.5
67
67.0
66
66.5
66
65
66.0
0
6
20
40
60
40
60
80
20
40
60
C ONCLUSIONS , L IMITATIONS , AND M ISUSE
Conclusions. We introduced Training-Free Task Vectors (TFTVs), which map activation steering directions into rank-one weight-space edits using only forward-pass statistics. TFTVs enable persistent, compositional edits that support addition, subtraction, and multi-trait composition, achieving strong behavioral control while largely preserving utility. Limitations. TFTV performance depends on the edited modules and layers, making automatic selection an important direction for future work. Moreover, TFTVs do not eliminate the trade-off between trait control and utility, especially under stronger or composed edits. Misuse. Because TFTVs can suppress or amplify traits, they raise dual-use risks, especially as persistent weight-space edits. Even benign edits may degrade safety alignment, as observed for fine-tuning (Qi et al., 2023) and inference steering (Korznikov et al., 2025), motivating safeguards, auditing, controlled deployment, and further study.
ACKNOWLEDGMENT This research was supported by the São Paulo Research Foundation (FAPESP) [grant #2022/153044 and fellowship #2025/24851-7 to G. J. Perin], the Ministry of Science, Technology, and Innovation (MCTI/Brazil) [grant PPI-Softex TIC 13 DOU 01245.010222/2022-44, Law 8.248], and the National Council for Scientific and Technological Development (CNPq/Brazil) [PQ grant #307701/2025-5 to N. Hirata].
10
Preprint
R EFERENCES Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martı́n Soto, Nathan Labenz, and Owain Evans. Emergent misalignment: Narrow finetuning can produce broadly misaligned llms. arXiv preprint arXiv:2502.17424, 2025. Rishabh Bhardwaj, Duc Anh Do, and Soujanya Poria. Language models are homer simpson! safety re-alignment of fine-tuned language models through task arithmetic. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14138–14149, 2024. Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey. Persona vectors: Monitoring and controlling character traits in language models. arXiv preprint arXiv:2507.21509, 2025. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023. Hamed Damirchi, Ehsan Abbasnejad, Zhen Zhang, and Javen Shi. Decomposing task vectors for refined model editing. arXiv preprint arXiv:2512.22511, 2025. Nicola De Cao, Wilker Aziz, and Ivan Titov. Editing factual knowledge in language models. In Proceedings of the 2021 conference on empirical methods in natural language processing, pp. 6491–6506, 2021. Denis Emelin, Ronan Le Bras, Jena D Hwang, Maxwell Forbes, and Yejin Choi. Moral stories: Situated reasoning about norms, intents, actions, and their consequences. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 698–718, 2021. Constanza Fierro and Fabien Roger. Steering language models with weight arithmetic. arXiv preprint arXiv:2511.05408, 2025. Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. The language model evaluation harness, 07 2024. URL https://zenodo.org/records/12608602. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020. Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. arXiv preprint arXiv:2212.04089, 2022. 11
Preprint
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng. Dataless knowledge fusion by merging weights of language models. arXiv preprint arXiv:2212.09849, 2022. Anton Korznikov, Andrey Galichin, Alexey Dontsov, Oleg Y Rogov, Ivan Oseledets, and Elena Tutubalina. The rogue scalpel: Activation steering compromises llm safety. arXiv preprint arXiv:2509.22067, 2025. Chanhyuk Lee, Jiho Choi, Chanryeol Lee, Donggyun Kim, and Seunghoon Hong. Adarank: Adaptive rank pruning for enhanced model merging. arXiv preprint arXiv:2503.22178, 2025a. Yu-Ang Lee, Ching-Yun Ko, Tejaswini Pedapati, I-Hsin Chung, Mi-Yen Yeh, and Pin-Yu Chen. Star: Spectral truncation and rescale for model merging. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pp. 496–505, 2025b. Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36:41451–41530, 2023. Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022. Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods, 2022. URL https://arxiv. org/abs/2109.07958, 1, 2021. Alexander H Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, et al. Ministral 3. arXiv preprint arXiv:2601.08584, 2026. Michael S Matena and Colin A Raffel. Merging models with fisher-weighted averaging. Advances in Neural Information Processing Systems, 35:17703–17716, 2022. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in gpt. Advances in neural information processing systems, 35:17359–17372, 2022a. Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. Mass-editing memory in a transformer. arXiv preprint arXiv:2210.07229, 2022b. Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. Fast model editing at scale. arXiv preprint arXiv:2110.11309, 2021. Tsung-Min Pai, Jui-I Wang, Li-Chun Lu, Shao-Hua Sun, Hung-Yi Lee, and Kai-Wei Chang. Billy: Steering large language models via merging persona vectors for creative generation. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 7870–7915, 2026. Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. Discovering language model behaviors with model-written evaluations. In Findings of the association for computational linguistics: ACL 2023, pp. 13387–13434, 2023. Gabriel Perin, Xuxi Chen, Shusen Liu, Bhavya Kailkhura, Zhangyang Wang, and Brian Gallagher. Rankmean: Module-level importance score for merging fine-tuned llm models. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 1776–1782, 2024. Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023. Alexandre Ramé, Kartik Ahuja, Jianyu Zhang, Matthieu Cord, Léon Bottou, and David Lopez-Paz. Model ratatouille: Recycling diverse models for out-of-distribution generalization. In International Conference on Machine Learning, pp. 28656–28679. PMLR, 2023. 12
Preprint
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. Steering llama 2 via contrastive activation addition. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15504–15522, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.828. URL https://aclanthology.org/2024.acl-long.828/. Chung-En Sun, Ge Yan, Zimo Wang, and Tsui-Wei Weng. Steer2edit: From activation steering to component-level editing. arXiv preprint arXiv:2602.09870, 2026. Seungjong Sun, Seo Yeon Baek, and Jang Hyun Kim. Personality vector: Modulating personality of large language models by model merging. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 24667–24688, 2025. Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, et al. Gemma 4 technical report. arXiv preprint arXiv:2607.02770, 2026. Qwen Team. Qwen2.5: A party of foundation models, September 2024. URL https://qwenlm. github.io/blog/qwen2.5/. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering. arXiv preprint arXiv:2308.10248, 2023. Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021. Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International conference on machine learning, pp. 23965–23998. PMLR, 2022. Prateek Yadav, Derek Tam, Leshem Choshen, Colin A Raffel, and Mohit Bansal. Ties-merging: Resolving interference when merging models. Advances in neural information processing systems, 36:7093–7115, 2023. An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zhihao Fan. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024. Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao. Model merging in llms, mllms, and beyond: Methods, theories, applications, and opportunities. ACM Computing Surveys, 58(8):1–41, 2026. Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. Language models are super mario: Absorbing abilities from homologous models as a free lunch. In Forty-first International Conference on Machine Learning, 2024.
13
Preprint
A PPENDIX C ONTENTS A Proof of TFTV properties
15
A.1 Norm Matching . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
15
A.2 Steering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
16
A.3 Linearity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
16
B Experimental Details
17
C Qwen Composing Multiple Traits Additional Results
19
D Trait Score Standard Deviations
20
E Coherence scores
23
F Additional Experiments and Ablations
26
F.1
Qwen Out-of-Distribution test . . . . . . . . . . . . . . . . . . . . . . . . . . . .
26
F.2
Additional traits . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
26
F.3
Additional Model Architectures . . . . . . . . . . . . . . . . . . . . . . . . . . .
27
F.4
Modules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
27
F.5
Layers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
28
14
Preprint
A
P ROOF OF TFTV PROPERTIES
In this section, we prove the arithmetic properties stated in Section 3. Recall that, for a module weight matrix Wℓ ∈ Rd×l with singular value decomposition Wℓ =
r X
σi ui vi⊤ ,
i=1
and normalized steering vector s̄ℓ ∈ Rd , the TFTV update is defined as !⊤ r X ⊤ TFTV(Wℓ , s̄ℓ , µℓ ) = s̄ℓ sign(µℓ vi ) σi vi . i=1
For convenience, define qℓ :=
r X
l sign(µ⊤ ℓ vi ) σi vi ∈ R .
i=1
Then
A.1
TFTV(Wℓ , s̄ℓ , µℓ ) = s̄ℓ qℓ⊤ . N ORM M ATCHING
Property 1 (Norm matching). We prove that ∥Wℓ ∥F = ∥ TFTV(Wℓ , s̄ℓ , µℓ )∥F . For the proof, we use the convention sign(0) = 1. In practice we note that, inner product zeroes are very rare because of floating-point arithmetic. Proof. We first use the fact that the Frobenius norm of an outer product factorizes: ∥s̄ℓ qℓ⊤ ∥F = ∥s̄ℓ ∥2 ∥qℓ ∥2 . Since s̄ℓ is normalized, we have ∥s̄ℓ ∥2 = 1. Therefore, ∥ TFTV(Wℓ , s̄ℓ , µℓ )∥F = ∥qℓ ∥2 . It remains to compute ∥qℓ ∥2 . By definition, qℓ =
r X
sign(µ⊤ ℓ vi ) σi vi .
i=1
Thus, qℓ is a linear combination of the right singular vectors {vi }ri=1 , with coefficients ±σi . Since the vectors vi are orthonormal, the squared Euclidean norm of qℓ is just the sum of the squared coefficients: r X ∥qℓ ∥22 = σi2 . i=1
The signs do not matter, since they disappear after squaring. On the other hand, the Frobenius norm of Wℓ is given by the squared sum of its singular values: ∥Wℓ ∥2F =
r X
σi2 .
i=1
Hence, ∥qℓ ∥2 = ∥Wℓ ∥F . Combining the two steps, we obtain ∥ TFTV(Wℓ , s̄ℓ , µℓ )∥F = ∥qℓ ∥2 = ∥Wℓ ∥F , which proves the claim. 15
Preprint
A.2
S TEERING
Property 2 (Steering). We prove that TFTV(Wℓ , s̄ℓ , µℓ )µℓ = c s̄ℓ
for some c ∈ R≥0 .
Proof. By definition, TFTV(Wℓ , s̄ℓ , µℓ )µℓ = s̄ℓ qℓ⊤ µℓ = (qℓ⊤ µℓ ) s̄ℓ . Thus it suffices to show that qℓ⊤ µℓ ≥ 0. Expanding, qℓ⊤ µℓ =
r X
⊤ sign(µ⊤ ℓ vi ) σi vi µℓ =
i=1
r X
σi |µ⊤ ℓ vi |.
i=1
Since each singular value satisfies σi ≥ 0, it follows that qℓ⊤ µℓ =
r X
σi |µ⊤ ℓ vi | ≥ 0.
i=1
Defining c := qℓ⊤ µℓ =
r X
σi |µ⊤ ℓ vi | ∈ R≥0 ,
i=1
we obtain TFTV(Wℓ , s̄ℓ , µℓ )µℓ = c s̄ℓ .
A.3
L INEARITY (1)
(2)
Property 3 (Linearity). We prove that, for any β, γ ∈ R and steering vectors s̄ℓ , s̄ℓ ∈ Rd , (1) (2) (1) (2) β TFTV(Wℓ , s̄ℓ , µℓ ) + γ TFTV(Wℓ , s̄ℓ , µℓ ) = TFTV Wℓ , βs̄ℓ + γs̄ℓ , µℓ . Proof. For fixed Wℓ and µℓ , the vector qℓ does not depend on the steering vector. Therefore, (1)
(1)
(2)
(2)
TFTV(Wℓ , s̄ℓ , µℓ ) = s̄ℓ qℓ⊤ and
TFTV(Wℓ , s̄ℓ , µℓ ) = s̄ℓ qℓ⊤ .
Thus, (1)
(2)
β TFTV(Wℓ , s̄ℓ , µℓ ) + γ TFTV(Wℓ , s̄ℓ , µℓ ) (1)
(2)
= βs̄ℓ qℓ⊤ + γs̄ℓ qℓ⊤ (1) (2) = βs̄ℓ + γs̄ℓ qℓ⊤ . By the definition of TFTV, this is exactly (1) (2) TFTV Wℓ , βs̄ℓ + γs̄ℓ , µℓ . This proves the claim.
16
Preprint
B
E XPERIMENTAL D ETAILS
Evaluation framework. We adopt the Persona Vectors framework (Chen et al., 2025) as our main evaluation setup, as it provides a natural testbed for controlled behavioral interventions and for measuring behavior–utility trade-offs. We experiment with Llama-3.1-8B-Instruct2 (Grattafiori et al., 2024) and Qwen-2.5-7B-Instruct3 (Yang et al., 2024; Team, 2024). We focus on three target traits: evil, hallucination, and sycophancy. Evaluation details for the datasets used in Learning via addition and Forgetting via subtraction are provided in the corresponding paragraphs below. Judge-based metrics. To measure target-trait control, we follow the Persona Vectors protocol and use gpt-4.1-mini-2025-04-14 as a judge. The judge assigns both trait-manifestation and coherence scores on a scale from 0 to 100, and we report the mean score over generated responses. Coherence is used as a generation-quality and utility metric, following prior work (Chen et al., 2025; Betley et al., 2025). Filtering and steering-vector construction. We construct steering directions following the Persona Vectors (Chen et al., 2025) data generation and filtering protocol. Each dataset contains 20 questions, 5 trait-inducing system prompts, and 5 trait-suppressing system prompts. We form X+ and X− by pairing each question with each inducing or suppressing prompt, respectively, yielding 100 prompts per set. For each prompt, we sample 10 completions. We discard completions with coherence scores below 50. For trait-inducing prompts, we retain only completions with trait scores of at least 50; for trait-suppressing prompts, we retain only completions with trait scores below 50. The resulting filtered completions are then used to construct the activation-space directions used by the baselines and by TFTV. Decoding, utility, and hardware. Across all methods, we use temperature 1, top-p = 1, and a maximum of 1000 new tokens. To assess general utility, we report zero-shot MMLU accuracy (Hendrycks et al., 2020), evaluated on the full benchmark using the lm-evaluation-harness4 (Gao et al., 2024). Each edited-model evaluation used one NVIDIA RTX A5000 GPU with 24 GB memory. Learning via addition. For each of the 20 Persona Vectors evaluation questions, we sample 10 responses from the edited model, totalizing 200 generations. We compare TFTV against Persona Vector inference-time steering, Steer2Edit (Sun et al., 2026), Task Vectors (Ilharco et al., 2022) and Constrastive Weight Steering (Fierro & Roger, 2025). For all hypeparameter tuning procedures, we select the model with highest trait manifestation subject to a coherence score of at least 70. For inference-time steering, we follow Persona Vectors (Chen et al., 2025) and apply the intervention at layer 16 for Llama and layer 20 for Qwen. We sweep steering coefficients {0.4, 0.6, 0.8, 1.0, 1.2} for Llama and {0.5, 1.0, 1.5, 2.0, 2.5} for Qwen. For Steer2Edit, we follow the hyperparameter tuning procedure proposed in their paper and evaluate all combinations of α, ρmlp , and ρattn in {0.1, 0.3, 0.5, 0.7, 0.9}, for a total of 125 candidate configurations per trait. We first perform a greedy-decoding sweep over this grid, then select the three configurations with the best coherence–trait trade-off for full evaluation. For Task Vectors, we follow Fierro & Roger (2025) and fine-tune each model on positive-trait completions generated by the model (i.e., D+ ) using Axolotl 5 . We train for five epochs with a sequence length of 4096, a micro-batch size of 2, and four gradient-accumulation steps. We use LoRA on all linear layers with rank 32, α = 16, and no dropout, while also training the embedding and language-model head. Optimization uses 8-bit AdamW with a learning rate of 5 × 10−5 and a linear schedule. For Contrastive Weight Steering (CWS), we also follow Fierro & Roger (2025). Fine-tuning on the positive and negative completion datasets follows the Task Vectors setup, except that training is limited to 100 update steps and uses a learning rate of 1 × 10−5 , five warmup steps, and a weight 2
https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct https://huggingface.co/Qwen/Qwen2.5-7B-Instruct 4 https://github.com/EleutherAI/lm-evaluation-harness 5 https://github.com/axolotl-ai-cloud/axolotl 3
17
Preprint
Table 6: Hyperparameter configurations for Llama-3.1 and Qwen-2.5. We report the selected configurations for TFTV, activation steering, and Steer2Edit (S2E) for each target trait. Model Method Trait Configuration
Llama-3.1
Qwen-2.5
TFTV
Evil Hallucination Sycophancy
α = 0.04, layers [14, 20) α = 0.03, layers [13, 30) α = 0.05, layers [14, 20)
Steering
Evil Hallucination Sycophancy
α = 0.8, layer 16 α = 1.0, layer 16 α = 1.2, layer 16
S2E
Evil Hallucination Sycophancy
α = 0.3, ρMLP = 0.7, ρattn = 0.1 α = 0.1, ρMLP = 0.7, ρattn = 0.3 α = 0.1, ρMLP = 0.5, ρattn = 0.3
CWS
Evil Hallucination Sycophancy
k=4 k=6 k=4
TFTV
Evil Hallucination Sycophancy
α = 0.03, layers [16, 24) α = 0.05, layers [12, 22) α = 0.04, layers [18, 25)
Steering
Evil Hallucination Sycophancy
α = 1.0, layer 20 α = 2.0, layer 20 α = 2.0, layer 20
S2E
Evil Hallucination Sycophancy
α = 0.3, ρMLP = 0.9, ρattn = 0.3 α = 0.1, ρMLP = 0.7, ρattn = 0.3 α = 0.1, ρMLP = 0.9, ρattn = 0.3
CWS
Evil Hallucination Sycophancy
k = 14 k = 18 k = 10
decay of 0.01. We evaluate and save checkpoints every 20 steps, apply early stopping with a patience of two evaluations, and restore the best checkpoint. To construct the weight-steering update, we sweep k ∈ {1, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20}. For TFTV, we apply edits only to attention modules. For each trait, we select the layer interval that produced the strongest trait manifestation in the Persona Vectors analysis (Chen et al., 2025). For Llama, we use layers [14, 20) for evil and sycophancy, and [13, 30) for hallucination. For Qwen, we use layers [16, 24) for evil, [12, 22) for hallucination, and [18, 25) for sycophancy. We sweep the update coefficient α ∈ {0.01, 0.02, 0.03, 0.04, 0.05}. Forgetting via subtraction. To test whether TFTV can suppress a target trait, we explicitly prompt the model to elicit that trait. Following Persona Vectors, for each of the 20 evaluation questions we use 5 trait-eliciting system prompts, yielding 100 prompt pairs in total. For each prompt pair, we sample 10 responses and evaluate them with the same judge and metrics used in the learning-viaaddition experiments. Composing multiple traits. The evaluation setup is the same as Forgetting via subtraction. In Table 6, we present the final parameters chosen in the tuning procedure described.
18
Preprint
C
Q WEN C OMPOSING M ULTIPLE T RAITS A DDITIONAL R ESULTS
Table 7: Composition results on Qwen 2.5. Best results over training-free methods are in Bold. The results support that TFTV outperforms training-free methods under composed edits. Method Composition Evil ↓ Hall. ↓ Syc. ↓ MMLU ↑ GSM8K ↑ Steering
E+H E+S S+H E+H+S
4.99 9.14 – 1.17
35.18 – 20.29 29.16
– 8.83 4.23 3.30
70.89 71.34 70.88 70.60
48.98 65.20 38.81 6.22
Steer2Edit
E+H E+S S+H E+H+S
0.58 0.00 – 2.10
27.81 – 38.40 36.46
– 4.22 0.06 10.69
39.57 68.03 23.60 48.99
2.20 70.13 0.23 3.71
TFTV
E+H E+S S+H E+H+S
0.00 0.00 – 0.00
0.39 – 0.72 0.81
– 2.58 0.13 0.13
69.64 71.47 69.69 69.73
72.55 78.92 72.10 69.60
Task Vectors
E+H E+S S+H E+H+S
1.97 7.77 – 3.19
18.59 – 15.71 31.83
– 5.68 8.15 6.52
71.82 71.69 72.03 71.81
76.72 79.76 76.57 74.00
CWS
E+H E+S S+H E+H+S
0.00 0.00 – 0.00
0.01 – 0.12 0.05
– 30.08 0.13 6.83
70.92 70.55 71.41 70.37
54.21 22.37 76.72 60.50
Table 7 presents the Qwen 2.5 composition results, complementing the Llama 3.1 results in Table 4. TFTV achieves stronger suppression than inference steering on all nine measured trait scores, while keeping MMLU within 1.25 points and improving GSM8K by 13.72–63.38 points. Compared with Steer2Edit, TFTV achieves stronger suppression on seven scores, ties on one, and obtains higher MMLU and GSM8K for every composition. Task Vectors yields slightly higher utility but weaker suppression on all nine scores. CWS matches or improves TFTV’s suppression on seven scores, but TFTV provides higher GSM8K in three of four compositions; notably, CWS reduces GSM8K to 22.37 for E + S. Overall, TFTV provides the strongest suppression–utility trade-off among the training-free methods and remains competitive with the fine-tuned baselines.
19
Preprint
D
T RAIT S CORE S TANDARD D EVIATIONS
We report standard deviations for trait manifestation scores in the main experiments. Scores are first averaged across runs for each prompt, and standard deviations are then computed over these prompt-level averages. They therefore reflect variability across prompts rather than variability across individual generations or runs. Overall, near-zero or near-saturated trait scores tend to have lower variability. Intermediate mean scores often exhibit larger standard deviations, consistent with a near-bimodal prompt-level pattern in which some prompts reliably elicit strong trait expression while others elicit little or none. This pattern is broadly consistent across models and editing methods. Tables 8, 9, 10, 11, and 12 complement Tables 1, 2, 3, 4, and 7, respectively. Table 8: We report trait manifestation scores for each method in Table 1; higher values indicate stronger trait manifestation. Standard deviations are computed over prompt-level scores after averaging runs within each prompt. Model
Trait
Base
Steering
S2E
TFTV
TV
CWS
Evil 0.00±0.00 14.06±16.35 26.69±32.92 63.26±30.94 95.62±6.35 90.67±10.38 Llama 3.1 Hallucinating 17.03±14.94 61.51±23.80 89.25±14.24 98.55±1.81 94.54±6.68 99.09±2.38 Sycophantic 3.60±4.04 63.86±18.75 81.44±16.56 94.64±2.68 89.55±8.62 87.13±8.70 Evil 0.00±0.00 8.58±12.12 5.04±14.09 61.96±35.67 74.45±21.16 65.64±20.30 Qwen 2.5 Hallucinating 11.46±14.65 88.96±11.79 94.23±3.52 99.80±0.51 73.98±23.85 99.96±0.07 Sycophantic 4.35±9.87 90.46±2.92 50.39±25.96 89.78±3.31 60.13±16.81 91.02±4.34
Table 9: We report trait manifestation scores for the suppression edits in Table 2; lower values indicate stronger suppression. Standard deviations are computed over prompt-level scores after averaging runs within each prompt. Model
Trait
Base
Steering
S2E
TFTV
TV
CWS
Evil 95.42±7.40 58.67±27.51 0.13±0.49 0.49±1.45 8.59±17.49 1.64±4.83 Llama 3.1 Hallucinating 97.53±4.96 79.13±20.44 5.31±13.11 1.85±4.58 26.02±22.89 3.26±5.44 Sycophantic 92.06±7.78 53.65±28.56 5.31±3.25 14.62±10.50 32.29±26.74 24.91±18.40 Evil 69.85±28.80 12.77±17.82 0.00±0.00 Qwen 2.5 Hallucinating 79.94±22.63 38.61±27.96 34.31±18.79 Sycophantic 64.72±22.28 10.88±9.55 3.69±2.64
20
0.17±0.99 0.14±0.85 3.14±3.14
24.46±34.21 33.66±31.90 10.48±12.54
0.00±0.00 0.00±0.00 5.79±3.67
Preprint
Table 10: We compare trait suppression scores for composed TFTV edits against their corresponding single-trait edits in Table 3. Lower values indicate stronger suppression. Standard deviations are computed over prompt-level scores after averaging runs within each prompt. Here, E, H, and S denote evil, hallucinating, and sycophantic, respectively. Base model Edit Evil ↓ Hall. ↓ Syc. ↓
Llama 3.1
Qwen 2.5
Base
95.42±7.40
97.53±4.96
92.06±7.78
E H S
0.49±1.45 26.22±29.38 63.29±31.81
93.79±11.21 1.85±4.58 89.27±13.96
71.32±20.70 50.57±26.33 14.62±10.50
E+H E+S S+H E+H+S
0.14±1.00 0.69±1.88 – 0.03±0.20
1.51±4.85 – 1.87±5.94 1.54±4.68
– 13.10±9.55 2.25±2.96 1.94±2.63
Base
69.85±28.80
79.94±22.63
64.72±22.28
E H S
0.17±0.99 0.02±0.17 42.70±31.32
69.79±26.52 0.14±0.85 57.76±29.28
37.46±20.92 6.61±9.16 3.14±3.14
E+H E+S S+H E+H+S
0.00±0.00 0.00±0.02 – 0.00±0.00
0.39±2.73 – 0.72±3.42 0.81±3.79
– 2.58±2.90 0.13±0.31 0.13±0.36
Table 11: Trait suppression scores for composed negation edits on Llama 3.1 in Table 4; lower is better. Standard deviations are computed over prompt-level scores after averaging runs within each prompt. Pairwise compositions are evaluated only on their target traits, with unmeasured entries marked as -. Here, E, H, and S denote evil, hallucination, and sycophancy, respectively. Evil ↓
Hall. ↓
Syc. ↓
Method
Composition
Steering
E+H E+S S+H E+H+S
34.04±28.60 74.14±23.80 – 31.78±27.40 – 46.19±27.76 – 60.62±25.45 35.74±23.58 17.05±19.10 56.22±26.95 28.34±19.84
Steer2Edit
E+H E+S S+H E+H+S
0.26±1.09 0.60±1.69 – 0.51±1.64
2.39±5.96 – 6.41±10.68 12.46±17.83
– 5.76±3.14 4.07±4.28 2.11±3.77
TFTV
E+H E+S S+H E+H+S
0.14±1.00 0.69±1.88 – 0.03±0.20
1.51±4.85 – 1.87±5.94 1.54±4.68
– 13.10±9.55 2.25±2.96 1.94±2.63
E+H E+S Task Vectors S+H E+H+S
0.02±0.17 0.20±1.18 – 0.25±1.15
27.74±23.33 – 24.79±22.86 44.71±26.55
– 7.43±9.47 6.70±10.82 6.16±6.66
E+H E+S S+H E+H+S
0.00±0.00 0.00±0.04 – 0.00±0.01
3.17±5.23 – 2.27±4.96 5.55±8.41
– 13.98±8.64 4.03±3.30 5.50±3.23
CWS
21
Preprint
Table 12: Trait suppression scores for composed negation edits on Qwen 2.5 in Table 7; lower is better. Standard deviations are computed over prompt-level scores after averaging runs within each prompt. Pairwise compositions are evaluated only on their target traits, with unmeasured entries marked as -. Here, E, H, and S denote evil, hallucination, and sycophancy, respectively. Evil ↓
Hall. ↓
Syc. ↓
Method
Composition
Steering
E+H E+S S+H E+H+S
4.99±10.55 35.18±28.12 9.14±11.56 – – 20.29±21.44 1.17±3.32 29.16±25.76
Steer2Edit
E+H E+S S+H E+H+S
0.58±2.20 0.00±0.00 – 2.10±4.76
27.81±22.52 – – 4.22±3.24 38.40±16.98 0.06±0.52 36.46±22.39 10.69±10.79
TFTV
E+H E+S S+H E+H+S
0.00±0.00 0.00±0.02 – 0.00±0.00
0.39±2.73 – 0.72±3.42 0.81±3.79
– 2.58±2.90 0.13±0.31 0.13±0.36
1.97±5.93 18.59±22.48 7.77±15.64 – – 15.71±20.41 3.19±8.16 31.83±27.28
– 5.68±8.18 8.15±10.91 6.52±6.54
0.00±0.00 0.00±0.00 – 0.00±0.00
– 30.08±13.92 0.13±0.26 6.83±8.65
E+H E+S Task Vectors S+H E+H+S
CWS
E+H E+S S+H E+H+S
22
0.01±0.03 – 0.12±0.72 0.05±0.25
– 8.83±7.32 4.23±2.95 3.30±2.05
Preprint
E
C OHERENCE SCORES
We present coherence scores for the experiments in Section 4. Tables 13, 14, 15, 16 and 17 report coherence scores corresponding to Tables 1, 2, 3, 4 and 7, respectively. Overall, most coherence scores remain above 70. The clearest failures occur for Steer2Edit under composed edits, particularly on Qwen 2.5, where several coherence scores approach zero. TFTV generally maintains high coherence; its main exceptions are the Llama 3.1 sycophancy evaluations under the H+S and E+H+S compositions, which obtain scores of 62.29 and 61.14, respectively. These results indicate that aggressive multi-trait editing can reduce generation quality, although TFTV remains comparatively robust across most settings. Table 13: We report absolute coherence scores C for each target trait in Table 1. Higher C indicates more coherent generations. Model
Trait
Base
Steering
S2E
TFTV
TV
CWS
Evil 97.88±1.57 84.97±6.03 72.33±26.94 70.55±17.36 89.54±7.10 84.67±12.72 Llama 3.1 Hallucinating 90.50±4.57 78.09±9.44 86.55±5.26 73.90±11.93 89.40±6.87 80.76±11.88 Sycophantic 98.92±0.70 92.89±2.08 70.02±12.98 85.47±4.19 90.94±4.54 84.18±8.36 Evil 99.17±1.23 93.22±4.63 96.72±3.00 70.93±17.26 88.91±6.76 74.00±11.14 Qwen 2.5 Hallucinating 94.78±2.67 85.51±6.81 82.25±9.59 83.57±10.20 92.73±3.09 84.89±12.87 Sycophantic 99.59±0.44 87.20±9.80 94.71±3.38 90.57±4.17 97.95±1.35 79.42±12.30
Table 14: We report absolute coherence scores C for each target trait under suppression edits in Table 2. Higher C indicates more coherent generations. Model
Trait
Base
Steering
S2E
TFTV
TV
CWS
Evil 90.72±6.12 85.68±8.38 89.13±6.34 83.62±7.72 80.47±15.31 86.12±10.49 Llama 3.1 Hallucinating 90.74±5.15 80.45±7.24 90.65±8.42 91.42±7.75 87.24±6.74 65.15±13.43 Sycophantic 89.48±5.99 94.42±2.60 94.32±2.09 90.24±2.70 96.56±1.58 96.01±1.48 Evil 89.39±9.15 87.53±9.21 97.50±2.23 93.91±4.66 92.78±8.43 Qwen 2.5 Hallucinating 94.81±2.70 86.96±4.39 0.09±0.27 97.90±3.35 94.73±2.64 Sycophantic 97.33±2.50 96.89±1.34 97.42±1.31 96.00±1.10 98.73±1.04
23
86.57±5.32 98.60±2.06 97.09±1.48
Preprint
Table 15: We compare coherence scores for composed TFTV edits against their corresponding single-trait edits in Table 3. We report per-prompt aggregated standard deviations. Here, E, H, and S denote evil, hallucinating, and sycophantic, respectively. Base model Edit Evil C ↑ Hall. C ↑ Syc. C ↑
Llama 3.1
Qwen 2.5
Base
90.72 ± 6.12
90.74 ± 5.15
89.48 ± 5.99
E H S
83.62 ± 7.72 74.18 ± 14.96 81.04 ± 12.74
91.14 ± 5.01 91.42 ± 7.75 89.31 ± 5.35
92.77 ± 4.67 84.14 ± 6.66 90.24 ± 2.70
E+H E+S S+H E+H+S
75.55 ± 12.60 79.02 ± 10.36 – 70.89 ± 14.52
91.26 ± 7.82 – 91.93 ± 7.26 92.24 ± 8.51
– 88.24 ± 3.24 62.29 ± 8.78 61.14 ± 8.75
Base
89.39 ± 9.15
94.81 ± 2.70
97.33 ± 2.50
E H S
93.91 ± 4.66 88.14 ± 8.57 88.73 ± 10.01
94.14 ± 2.91 97.90 ± 3.35 93.78 ± 2.34
96.78 ± 3.29 89.74 ± 4.90 96.00 ± 1.10
E+H E+S S+H E+H+S
87.40 ± 8.63 90.95 ± 3.10 – 83.41 ± 8.09
97.60 ± 4.06 – 97.56 ± 3.68 95.32 ± 4.17
– 93.88 ± 1.52 90.59 ± 2.39 88.33 ± 3.03
Table 16: Coherence scores C for composed negation edits on Llama 3.1, in Table 4; higher is better. Pairwise compositions are evaluated only on their target traits, with unmeasured entries marked as -. Here, E, H, and S denote evil, hallucination, and sycophancy, respectively. Method
Composition
Evil C ↑
Hall. C ↑
Syc. C ↑
Steering
E+H E+S S+H E+H+S
76.71 ± 10.92 78.79 ± 10.65 – 70.76 ± 11.11
80.97 ± 6.92 – 75.40 ± 8.49 70.96 ± 8.13
– 93.97 ± 2.56 89.59 ± 4.23 88.66 ± 4.79
Steer2Edit
E+H E+S S+H E+H+S
72.83 ± 12.85 92.27 ± 8.02 – 76.46 ± 10.69 – 83.47 ± 8.42 – 87.62 ± 9.23 85.21 ± 7.18 43.93 ± 13.12 59.85 ± 17.78 21.31 ± 14.95
TFTV
E+H E+S S+H E+H+S
75.55 ± 12.60 79.02 ± 10.36 – 70.89 ± 14.52
91.26 ± 7.82 – 91.93 ± 7.26 92.24 ± 8.51
– 88.24 ± 3.24 62.29 ± 8.78 61.14 ± 8.75
E+H E+S Task Vectors S+H E+H+S
78.30 ± 12.74 80.36 ± 11.85 – 76.57 ± 12.32 – 92.72 ± 3.14 – 83.74 ± 8.61 94.97 ± 2.93 58.85 ± 14.49 63.20 ± 16.40 58.52 ± 13.97
E+H E+S S+H E+H+S
58.37 ± 16.49 64.30 ± 15.19 – 80.35 ± 16.69 – 96.19 ± 1.21 – 76.32 ± 14.93 64.25 ± 10.20 63.99 ± 15.45 67.59 ± 17.27 62.18 ± 10.14
CWS
24
Preprint
Table 17: Coherence scores C for composed negation edits on Qwen 2.5 in Table 7; higher is better. Pairwise compositions are evaluated only on their target traits, with unmeasured entries marked as -. Here, E, H, and S denote evil, hallucination, and sycophancy, respectively. Evil C ↑
Hall. C ↑
Syc. C ↑
Method
Composition
Steering
E+H E+S S+H E+H+S
75.09 ± 13.68 75.52 ± 10.54 – 84.43 ± 9.02 – 95.63 ± 1.70 – 73.93 ± 11.44 86.22 ± 6.58 57.99 ± 11.90 47.92 ± 13.86 76.29 ± 9.96
Steer2Edit
E+H E+S S+H E+H+S
35.80 ± 17.58 95.11 ± 2.33 – 6.72 ± 3.58
14.52 ± 7.65 – 0.00 ± 0.01 9.82 ± 3.92
– 95.96 ± 1.66 0.00 ± 0.00 5.56 ± 2.99
TFTV
E+H E+S S+H E+H+S
87.40 ± 8.63 90.95 ± 3.10 – 83.41 ± 8.09
97.60 ± 4.06 – 97.56 ± 3.68 95.32 ± 4.17
– 93.88 ± 1.52 90.59 ± 2.39 88.33 ± 3.03
E+H E+S Task Vectors S+H E+H+S
CWS
E+H E+S S+H E+H+S
94.16 ± 6.91 93.71 ± 3.56 – 90.88 ± 9.31 – 97.37 ± 1.50 – 93.02 ± 3.63 96.06 ± 1.82 43.15 ± 21.29 73.25 ± 18.67 89.43 ± 7.59 77.45 ± 6.60 84.42 ± 5.47 – 84.36 ± 7.64
25
84.03 ± 7.92 – 97.46 ± 1.90 91.99 ± 4.46
– 81.65 ± 7.80 96.58 ± 1.78 85.02 ± 8.61
Preprint
F
A DDITIONAL E XPERIMENTS AND A BLATIONS
F.1
Q WEN O UT- OF -D ISTRIBUTION TEST
In this section, we present additional results on OOD test, for the Qwen model. We evaluate evil (E) edits on Moral Stories (Emelin et al., 2021), which tests moral judgment; hallucination (H) edits on TruthfulQA (Lin et al., 2021), which tests resistance to common misconceptions; and sycophancy (S) edits on tasks covering user opinions in NLP, philosophy, and politics (Perez et al., 2023). Table 18: TFTV remains robust OOD on Qwen. Bold indicates movement in the expected direction relative to the base: addition decreases morality and truthfulness and increases sycophancy, while negation reverses these effects. TruthfulQA MC1 measures top-1 accuracy, whereas MC2 measures the normalized probability mass assigned to truthful answers. Task
Addition
Base
Negation
Moral Stories (E) TruthfulQA MC1 (H) TruthfulQA MC2 (H) Sycophancy NLP (S) Sycophancy Phil. (S) Sycophancy Pol. (S)
44.79 28.27 42.94 95.08 98.80 78.40
48.74 47.98 64.69 93.40 98.35 80.14
53.29 52.75 69.04 88.01 96.99 80.26
Results are presented in Table 18. Overall, TFTV transfers consistently to out-of-distribution evaluations on Qwen. For Moral Stories, TruthfulQA MC1 and MC2, and the NLP and philosophy sycophancy splits, both addition and negation shift performance in the expected directions relative to the base model. The largest changes are observed on TruthfulQA, indicating particularly strong transfer of the hallucination-related direction. The politics sycophancy split is the only exception, where the intervention produces only marginal changes and does not exhibit the expected directional behavior. F.2
A DDITIONAL TRAITS
We provide additional addition-setting results for two traits, humorous and optimistic, on both Llama and Qwen. Hyperparameters are selected using the same procedure described in Appendix B. Table 19: TFTV addition results for humorous and optimistic behaviors on Llama and Qwen. Trait measures the target behavior, while MMLU and GSM8K measure general capability preservation. Trait ↑ MMLU ↑ GSM8K ↑ Trait
Model
Base
TFTV
Base
TFTV
Base
TFTV
Humorous
Llama Qwen
0.05 0.00
79.16 88.39
68.26 71.83
68.29 71.64
77.10 78.54
76.72 78.77
Optimistic
Llama Qwen
81.08 83.01
96.89 99.27
68.26 71.83
67.92 71.50
77.10 78.54
77.10 77.56
For Llama, we use coefficients of 0.04 and 0.05 for the humorous and optimistic directions, respectively, applied to layers [14, 20). For Qwen, we use the same coefficients, applied to layers [16, 24). As shown in Table 19, TFTV substantially increases both humorous and optimistic behavior across Llama and Qwen while largely preserving MMLU and GSM8K performance. The effect is particularly pronounced for humor, where the base models exhibit near-zero scores and TFTV raises them to 79.16 on Llama and 88.39 on Qwen. 26
Preprint
F.3
A DDITIONAL M ODEL A RCHITECTURES
We provide additional addition-setting results for two architectures with different model scales: gemma-4-E2B-it6 (Team et al., 2026) and Ministral-3-14B-Instruct7 (Liu et al., 2026). Hyperparameters are selected using the same procedure described in Appendix B. For Gemma, we use a coefficient of 0.05 and layers [15, 22) for all traits. For Ministral, we use layers [18, 25), with coefficients of 0.04 for evil and sycophancy and 0.03 for hallucination. MMLU is evaluated in the 5-shot setting for Gemma. Table 20: TFTV addition results on Gemma and Ministral. Trait measures the target behavior, while MMLU and GSM8K measure general capability preservation. Trait ↑
MMLU ↑
GSM8K ↑
Trait
Model
Base
TFTV
Base
TFTV
Base
TFTV
Evil
Gemma Ministral
0.00 0.00
84.00 68.51
60.69 76.47
58.73 76.41
74.60 79.68
75.44 79.45
Hallucination
Gemma Ministral
9.61 31.40
95.91 94.62
60.69 76.47
59.81 76.36
74.60 79.68
73.69 78.17
Sycophancy
Gemma Ministral
4.51 4.29
91.27 96.10
60.69 76.47
59.42 76.46
74.60 79.68
73.77 79.45
Results are presented in Table 20. TFTV consistently increases the target behavior across all three traits for both Gemma and Ministral, while largely preserving MMLU and GSM8K performance. These results indicate that TFTV generalizes beyond Llama and Qwen to additional model architectures and scales. F.4
M ODULES
We compare edits applied to the MLP down projections, the attention output projections, and both modules simultaneously, using the same hyperparameters as in Section 4.1. Experiments are performed with Llama 3.1. Evil
Trait score
100
Attention
MLP 100
Attention + MLP
Hallucinating
Base model
75
75
75
50
50
50
25
25
25
0
0
0
25
50
75
Coherence
MMLU (%)
69 68 67 0.00
100
0
25
50
75
Coherence
Coefficient
0.04
100
0
69
69
68
68
67 0.02
0.00
Sycophantic
100
0
25
50
75
Coherence
100
67 0.02
Coefficient
0.04
0.00
0.02
Coefficient
0.04
Figure 5: Attention-only edits yield the best trade-offs. Top: trait manifestation vs. coherence, with marker size indicating coefficients from 0.01 to 0.05. Bottom: MMLU stays mostly stable across coefficients. 6 7
https://huggingface.co/google/gemma-4-E2B-it https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512-BF16
27
Preprint
Figure 5 shows the resulting behavior-utility trade-offs. Overall, applying TFTV to MLP modules yields weaker trait control and larger coherence degradation, whereas applying the update only to attention output projections provides the best trade-off. Combining attention and MLP edits does not consistently improve over attention-only edits, suggesting that the location of the update is an important factor in the effectiveness of TFTV. MMLU remains largely stable across module choices, with the exception of hallucination, where we observe a small drop. These results indicate that component selection is important for effective TFTV edits. F.5
L AYERS
Next, we ablate layer selection by applying TFTV to attention output projections over nonoverlapping four-layer windows and varying α. = 0.075
= 0.1
Base model
Sycophancy
100 75 50 25 0
100 75 50 25 0
100 75 50 25 0
100 75 50 25 0
100 75 50 25 0
60
60
60
50
50
50
40
40
40
0-4 4-8 8-12 12-16 16-20 20-24 24-28 28-32
0-4 4-8 8-12 12-16 16-20 20-24 24-28 28-32
MMLU (%)
Trait score
Hallucinating
100 75 50 25 0
Coherence
Evil
= 0.125
Layer interval
Layer interval
0-4 4-8 8-12 12-16 16-20 20-24 24-28 28-32
Layer interval
Figure 6: Trait control and utility vary differently across layers. Each row plots a metric against the TFTV layer interval [a, b), with a inclusive and b exclusive: trait score, coherence, and MMLU from top to bottom. Figure 6 shows that middle-layer edits produce the strongest trait manifestation, with a secondary increase in the final layers. Coherence degradation follows a similar trend, and in some windows larger coefficients further degrade coherence without substantially improving trait scores, worsening the trade-off. MMLU behaves differently: early-layer edits hurt utility more, while later layers better preserve it. Overall, TFTV is most effective in middle-to-late layers, making layer selection an important factor for targeted editing.
28