Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Sinie van der Ben 1 Raphaël Baur 1 Yannick Metz 1 Mennatallah El-Assady 1
Abstract
Sonnet 4.5. They identified 171 linear directions in activation space corresponding to emotion concepts, with correlational and potentially causal relations to model behaviour. Steering these vectors altered the model’s preferences and increased rates of misaligned behaviors such as reward hacking and blackmail. The overall geometry of the emotion space mirrors human psychology, with principal components aligning to valence and arousal axes consistent with Russell’s circumplex model (Russell, 1980). These findings raise key questions about generality: (1) Are emotion vectors specific to Claude’s training, or a general property of language models’ internal representations? (2) How does emotion geometry evolve across layers: Does it emerge suddenly or build up gradually? (3) How does the choice of story corpus affect extraction? These questions matter for interpretability and safety: If emotion representations are universal and robustly extractable, monitoring them could provide early warnings of misaligned internal states across different models. We address these questions by replicating and extending emotion vector analysis in two open-weight models: A PERTUS -8B (Hernández-Cano et al., 2025), with fully open weights, training data, and code, and G EMMA -4-E4B (DeepMind, 2026), a recently released open-source model, both chosen for their relatively small size. For each model, we extract emotion contrast vectors across multiple layers using two story corpora—one generated by A PERTUS -8B and one by G EMMA -4-E4B —to separate model-intrinsic geometry from corpus-dependent extraction artifacts. Additional related work is provided in Appendix A. We release our code publicly1 .
arXiv:2606.26987v1 [cs.CL] 25 Jun 2026
Recent work identified “emotion vectors” in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirroring human psychological structure. We test the generality of these findings in two open-weight models, A PERTUS -8B and G EMMA -4-E4B, extracting emotion contrast vectors across all layers, using two model-generated corpora. We recover valence geometry for both models, with peak PC1–valence correlations of r = 0.76 and r = 0.83, approaching the r = 0.81 reported for Claude. Beyond replication, we observe notable differences in how valence representations emerge across model depth. In G EMMA 4-E4B, valence is strongly encoded in early layers but collapses towards later layers, whereas A PERTUS -8B exhibits the opposite pattern, with valence representations absent in early layers, but emerging at mid depths. Arousal encoding, in contrast, is sensitive to the extraction corpus: both models show stronger PC2–arousal alignment with Gemma-generated stories (r up to 0.45) than Apertus-generated ones (r ≤ 0.21), suggesting arousal-relevant cues are unevenly distributed across generated corpora. We opensource our experiment code and dataset for reproducible investigation of emotion representations across language model architectures.
• Replication of key findings. We recover valence geometry in both A PERTUS -8B and G EMMA -4-E4B, with the highest PC1–valence correlations of r=0.76 and r=0.83 respectively, demonstrating that emotion vectors generalize beyond Claude to open-weight models across different architectures. • Divergent Emergence. Models differ substantially in when valence structure emerges: G EMMA -4-E4B peaks early (layer 16) then fades, while A PERTUS -8B builds progressively across depth, stabilizing around layer 20. Cross-layer CKA analysis shows a phase transition in A PERTUS -8B that is absent in Gemma.
1. Introduction As users interact with Large Language Models (LLM), they can encounter responses that appear emotionally reactive, such as expressing frustration when struggling with tasks or enthusiasm when helping users. Recent work by Sofroniew et al. (2026) moved beyond surface-level observations, identifying internal “emotion vectors” in Claude 1
Department of Computer Science, ETH, Zurich, Switzerland. Correspondence to: Sinie van der Ben <[email protected]>. Published at the Mechanistic Interpretability Workshop at ICML 2026. Copyright 2026 by the author(s).
1
1
https://github.com/sinievanderben/emotion_experiment
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
ducing a (40, dmodel ) matrix per layer. PCA on this matrix yields a basis for the subspace; we retain the top K components that together explain 50% of the variance. To isolate the emotion-specific component, we subtract from each emotion vector its projection onto the neutral subspace to get the contrast vector ve : K X ve = ue − (ue · pk )pk
• Corpus-dependent arousal. Arousal encoding is sensitive to story corpus: both models show substantially stronger PC2–arousal alignment when using Gemma-generated stories (r up to 0.45) than Apertusgenerated stories (r ≤ 0.17)
2. Methods 2.1. Dataset
k=1
For A PERTUS -8B, we extracted vectors from layers 1–31, and for G EMMA -4-E4B, from layers 1–40. Stacking these vectors across all |E| emotions yields the matrix V (l) ∈ R|E|×dmodel at layer l, on which we perform the analyses.
We generated two synthetic emotion-story datasets, following Sofroniew et al. (2026), with 9 stories for each of 171 emotions. For each emotion, we prompted A PERTUS -8B and G EMMA -4-E4B to write short stories in which characters experience the target emotion without naming it, using a similar prompt to Sofroniew et al. (2026). This produced 1,539 stories across emotions (Table 1), plus 40 neutral stories from the same model. The 40 neutral texts form a single fixed set shared by all 171 emotions, since we compute the confound subspace once per layer and project every emotion vector through the same operation. The emotion concepts span the valence-arousal space.
2.4. Analysis PCA and Valence-Arousal Correlation We applied PCA to the emotion contrast matrix V (l) at each layer and correlated the first two principal components (PC1, PC2) with human valence and arousal ratings from the NRC Valence– Arousal–Dominance Lexicon (Mohammad, 2018), following (Sofroniew et al., 2026). We report Pearson r and corresponding p-values.
We treat the story corpora as independent variable. By running each model on both Apertus-generated and Gemmagenerated stories, we intend to disentangle the emotion findings from corpus-dependent extraction artifacts. No previous work has tested story influence before.
Cross-layer Representational Similarity with CKA We computed linear Centered Kernel Alignment (Kornblith et al., 2019) between V (l) for all layer pairs within each model and story condition. CKA values near 1 indicate similar representational geometry, while values near 0 indicate orthogonal structure. Because CKA is invariant to orientation in latent space, it is well suited for this comparison and allowed us to quantify how emotion geometry evolves through the network.
2.2. Model We analyzed two open-weight language models: A PERTUS -8B Instruct, a 32-layer transformer, and G EMMA -4-E4B, a 42-layer transformer. Both models are instruction-tuned and comparable in scale to enable cross-model comparison of emotion representations. More details on both models can be found in Appendix B.2.
Valence direction stability Lastly, we identified the valence direction at each layer as the vector most correlated with human valence ratings (using PC1 when this correlation is significant), then computed cosine similarity between these directions across layers to test whether the same subspace encodes valence at different depths.
2.3. Contrast Vector Extraction Following Sofroniew et al. (2026), we construct one ac(l) tivation vector ve per emotion e and layer l. Since these vectors capture general linguistic structure, we apply a twostep procedure to isolate the emotion-specific component.
3. Results 3.1. Valence Replicates Across Models and Corpora
First, for each emotion, we perform a forward pass on the corresponding nine stories and cache the residual stream activations at each layer, giving a tensor of shape (#tokens, dmodel ) per layer. Averaging these activations across tokens and stories yields one raw vector ue ∈ Rdmodel per emotion and layer, which still mixes emotion-specific and general linguistic features.
The first principal component of the emotion contrast matrix aligns with human valence ratings in both models, replicating the main result of (Sofroniew et al., 2026). Figure 1 shows PC1–valence correlations across fractional layer depth for all four model×corpus conditions; per-layer values are reported in Table 3.
Second, we project out non-emotion-specific components. To characterize the emotion-agnostic subspace, we collect mean residual activations from the 40 neutral stories, pro-
Peak correlations. All model×corpus combinations reach a peak between r = 0.72 and 0.83. A PERTUS -8B peaks at r = 0.72 (layer 23, Apertus stories) and r = 0.76 (layer 31, Gemma stories); G EMMA -4-E4B peaks at r = 2
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Figure 1. Pearson correlation between the top two PCs of the emotion-vector space and human valence (left, PC1) and arousal (right, PC2) across fractional layer depth, for Apertus 8B and G EMMA -4-E4B probed on Apertus- and Gemma-generated stories (four conditions). Hue = model (blue = Apertus, red = Gemma); line style = story source (solid = Apertus, dashed = Gemma). Dotted gray lines mark the Sonnet 4.5 reference at a mid-late layer (r = 0.81 valence, r = 0.66 arousal; (Sofroniew et al., 2026)).
0.79 (layer 13, Apertus stories) and r = 0.83 (layer 16, Gemma stories). All peaks are significant (p < 10−3 ) and approach or exceed the Sonnet 4.5 reference of r = 0.81.
of Gemma’s valence correlation around layer 18 therefore cannot stem from a global reorganization, as the geometry remains approximately stable through the collapse. The valence-direction cosine matrices show what changes. For A PERTUS -8B on its own stories (Fig. 6), no offdiagonal cell exceeds |0.49|, which can indicate that the recovered direction is noise across layers. On Gemma stories (Fig. 7), early layers (2–11) form a coherent block with cosines 0.35–0.57 before becoming noisy in later layers. For G EMMA -4-E4B on Gemma stories (Fig. 10), there are two positive blocks (layers 2–8 and 9–14) and a late block (28–40), with adjacent-layer cosines up to ±0.55. On Apertus stories (Fig. 11) this structure is less pronounced. Because CKA matrices are similar across corpora, the emotion representational space is corpus-invariant. However, the recovered valence axis depends on the input corpus, with Gemma stories yielding cleaner valence directions in both models. Thus, valence is recoverable in both, but not encoded along a consistent axis across depth.
Valence Across Network Depth Both models reach similar r-value peaks with opposite depth profiles. A PERTUS 8B shows abrupt late emergence: PC1–valence correlation is near zero through fractional depth ≈ 0.5 (layer 17/18), then rises sharply, becoming significant at layer 18 (r = 0.17, p < 0.05) and exceeding r = 0.60 at layer 21 (≈63% depth) under both story conditions. G EMMA -4-E4B instead shows early encoding followed by collapse: for Apertus stories, valence peaks at layer 16 (≈38% depth), then falls near zero by layer 18, with only partial recovery (r ≈ 0.18–0.20) in the final layers. For Gemma stories, the peak comes later and both preand post-peak values are higher. The Sonnet 4.5 reference peaks in the mid-late range, indicating that A PERTUS -8B follows a similar pattern.
PCA cluster separation at peak layers PCA projections at each model’s peak layer (Figs. 14, 15) show emotion clustering and a clear corpus effect. PC1–valence correlations are similar across story conditions (A PERTUS -8B L23: 0.72 vs. 0.75; G EMMA -4-E4B L13: 0.79 vs. 0.80), but clusters are more clearly separated for Gemma stories, with positive and negative emotions forming denser groups.
Representation space vs. valence-axis stability To interpret valence trajectories, we examine (i) whole-space representational similarity via linear CKA and (ii) cosine alignment of the layer-wise valence direction for each model–corpus combination. A PERTUS -8B (Figs. 2, 3) shows three CKA phases: layers 2–11 form a flat plateau (CKA ≈ 1); layers 12–21 form a transition band with off-diagonal decay (minimum 0.33 on Apertus stories, 0.58 on Gemma stories); and layers 22–31 form a second plateau. This transition aligns with the rise of PC1–valence correlation, suggesting a representational reorganization. G EMMA -4-E4B (Figs. 4, 5) instead shows a smooth gradient across all 40 layers with no sharp transition and CKA ≥ 0.73 between any pair. The collapse
3.2. Arousal Encoding PC2–arousal correlations are generally weaker than PC1– valence and depend strongly on the story corpus (Fig. 1, right; Table 4). On Apertus stories, both models peak below r = 0.21 (A PERTUS -8B: r = 0.17 at layer 18; G EMMA -4-E4B: r = 0.21 at layer 40). On Gemma sto3
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
ries, both models reach r > 0.40 (A PERTUS -8B: r = 0.45 at layer 26; G EMMA -4-E4B: r = 0.41 at layer 31, both p < 10−8 ). Possibly, Gemma-generated stories contain more arousal-discriminative linguistic content. We leave a corpus-content analysis to future work.
that Gemma has the ability to generate stories with greater variation in narrative intensity and physiological arousal cues, so corpus choice for eliciting emotion contrasts is a substantive methodological factor, not an implementation detail. We leave verification to future work.
4. Discussion
4.1. Limitations
Main Research Questions Our results address the three questions raised in the introduction. (1) Emotion vectors are not specific to Claude’s training. We recover a valence axis of similar strength in two architecturally distinct open-weight models, with peak correlations matching (r = 0.83 for G EMMA -4-E4B) or approaching (r = 0.76 for A PERTUS -8B) the r = 0.81 reported for Claude Sonnet 4.5. (2) Emergence is not uniform across models. A PERTUS -8B builds valence alignment abruptly in the second half of the network, while G EMMA -4-E4B encodes it early and then loses it mid-network. (3) The story corpus affects extraction. This is especially clear for arousal: Gemma-generated stories yield correlations more than twice as large as Apertus-generated stories in both probed models. Different paths to the same geometry. G EMMA -4-E4B and A PERTUS -8B reach similar peak valence correlations (r ≈ 0.76–0.83) via different layer-wise trajectories. G EMMA -4-E4B encodes valence in earlier layers before it degrades in later layers, while A PERTUS -8B develops it sharply across mid-to-upper layers. We have not yet explored the possible attribution of this to architecture, training data, or post-training, since the models differ in all three. Our results show that similar peak valence correlations can hide substantial differences in where and how valence is computed. Stable representation space, unstable axis The representational space (CKA) and valence-axis stability disconnect. In G EMMA -4-E4B, the space remains similar across layers even where the PC1–valence correlation collapses, so valence information is preserved. In A PERTUS -8B, the valence axis is relatively unstable across layers despite a late high plateau of valence–PC1 correlation. Thus, representational similarity between layers does not guarantee a shared valence direction. The arousal gap and corpus dependence Arousal shows the weakest replication, but the story-condition analysis suggests that this may be attributed to our methodological choices. With Gemma-generated stories, arousal correlations in both models rise (from r ≤ 0.21 to r ≥ 0.43), partially closing the gap with the original result (r = 0.66). Because Gemma stories improve arousal extraction in both models, the effect likely reflects corpus properties rather than model–story matching. Since it appears in both models, this rules out the simple confound that each model encodes only its own corpus well. We hypothesize
Several limitations warrant mention. The first, the original study (Sofroniew et al., 2026) did not release code, so our implementation is reconstructed from the methods they described. Subtle methodological differences may therefore contribute to numerical differences. Second, our analysis covers two open-weight models from two families. Broader cross-architecture comparisons would strengthen claims about how general the valence-pattern is, and whether the trajectory differences generalize to other model families. Third, the corpora we probe are themselves model-generated, which means we cannot fully separate properties of the distributions it produces. A fully modelindependent stimulus set would be a stronger control. 4.2. Future Work Several directions follow from our findings and limitations. The most direct is causal validation: steering model outputs at peak-correlation layers along the recovered valence direction would test whether the representational structure we identify is actually used by the model. Related, the cross-layer rotations of the valence axis raises the question whether steering vectors derived at one layer remain effective when applied to another, even within regions of overall stable space. Cross-layer feature tracking using sparse autoencoders could further reveal whether the same interpretable features carry emotion information across the depth ranges we identify, or whether different layers encode emotions through different feature combinations. Finally, extending this analysis to multi-modal models could test whether the valence axis is preserved across modalities.
5. Conclusion We replicate Anthropic’s emotion findings in two openweight models, achieving valence correlations of r = 0.83 (G EMMA -4-E4B) and r = 0.76 (A PERTUS -8B). Crosslayer analysis reveals divergent developmental trajectories: G EMMA -4-E4B encodes valence in early layers while A PERTUS -8B builds it progressively through late layers. These results suggest that similar representations can arise from different computational paths, with implications for layer selection in interpretability work and targeted steering interventions.
4
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
References
Machacek, R., Manitaras, T., Marfurt, A., Matoba, K., Matrenok, S., Mendoncça, H., Mohamed, F. R., Montariol, S., Mouchel, L., Najem-Meyer, S., Ni, J., Oliva, G., Pagliardini, M., Palme, E., Panferov, A., Paoletti, L., Passerini, M., Pavlov, I., Poiroux, A., Ponkshe, K., Ranchin, N., Rando, J., Sauser, M., Saydaliev, J., Sayfiddinov, M. A., Schneider, M., Schuppli, S., Scialanga, M., Semenov, A., Shridhar, K., Singhal, R., Sotnikova, A., Sternfeld, A., Tarun, A. K., Teiletche, P., Vamvas, J., Yao, X., Ilic, H. Z. A., Klimovic, A., Krause, A., Gulcehre, C., Rosenthal, D., Ash, E., Tramèr, F., VandeVondele, J., Veraldi, L., Rajman, M., Schulthess, T., Hoefler, T., Bosselut, A., Jaggi, M., and Schlag, I. Apertus: Democratizing Open and Compliant LLMs for Global Language Environments. https://arxiv.org/abs/ 2509.14233, 2025.
Arditi, A., Obeso, O. B., Syed, A., Paleka, D., Rimsky, N., Gurnee, W., and Nanda, N. Refusal in language models is mediated by a single direction. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/ forum?id=pH3XAQME6c. Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., et al. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. Cheng, E., Doimo, D., Kervadec, C., Macocco, I., Yu, L., Laio, A., and Baroni, M. Emergence of a highdimensional abstraction phase in language transformers. In The Thirteenth International Conference on Learning Representations, 2025. URL https:// openreview.net/forum?id=0fD3iIBhlV.
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. Similarity of neural network representations revisited. In International conference on machine learning, pp. 3519– 3529. PMlR, 2019.
Choi, B. J. and Weber, M. Latent structure of affective representations in large language models, 2026. URL https://arxiv.org/abs/2604.07382.
Marks, S. and Tegmark, M. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets, 2024. URL https:// arxiv.org/abs/2310.06824.
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023.
Mikolov, T., Yih, W.-t., and Zweig, G. Linguistic regularities in continuous space word representations. In Vanderwende, L., Daumé III, H., and Kirchhoff, K. (eds.), Proceedings of the 2013 Conference of the North AmerDeepMind, G. Gemma 4: Expanding the gemican Chapter of the Association for Computational Linmaverse with apache 2.0, 2026. URL https: guistics: Human Language Technologies, pp. 746–751, //opensource.googleblog.com/2026/03/ Atlanta, Georgia, June 2013. Association for Computagemma-4-expanding-the-gemmaverse-with-apache-20. tional Linguistics. URL https://aclanthology. html. Accessed: 2026-04-28. org/N13-1090/. Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, Mohammad, S. M. Obtaining reliable human ratings of vaT., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, lence, arousal, and dominance for 20,000 english words. D., Chen, C., Grosse, R., McCandlish, S., Kaplan, In Proceedings of ACL, 2018. J., Amodei, D., Wattenberg, M., and Olah, C. Toy models of superposition. Transformer Circuits Thread, Park, K., Choe, Y. J., and Veitch, V. The linear repre2022. URL https://transformer-circuits. sentation hypothesis and the geometry of large language pub/2022/toy_model/index.html. models. In Causal Representation Learning Workshop at NeurIPS 2023, 2023. URL https://openreview. Hernández-Cano, A., Hägele, A., Huang, A. H., Romanou, net/forum?id=T0PoOJg8cK. A., Solergibert, A.-J., Pasztor, B., Messmer, B., Garbaya, D., Ďurech, E. F., Hakimi, I., Giraldo, J. G., Radford, A., Jozefowicz, R., and Sutskever, I. Learning to Ismayilzada, M., Foroutan, N., Moalla, S., Chen, T., generate reviews and discovering sentiment, 2017. URL Sabolčec, V., Xu, Y., Aerni, M., AlKhamissi, B., Marihttps://arxiv.org/abs/1704.01444. nas, I. A., Amani, M. H., Ansaripour, M., Badanin, I., Benoit, H., Boros, E., Browning, N., Bösch, F., Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., HubBöther, M., Canova, N., Challier, C., Charmillot, C., inger, E., and Turner, A. Steering llama 2 via conColes, J., Deriu, J., Devos, A., Drescher, L., Dzentrastive activation addition. In Ku, L.-W., Martins, A., haliou, D., Ehrmann, M., Fan, D., Fan, S., Gao, S., Gila, and Srikumar, V. (eds.), Proceedings of the 62nd AnM., Grandury, M., Hashemi, D., Hoyle, A., Jiang, J., nual Meeting of the Association for Computational LinKlein, M., Kucharavy, A., Kucherenko, A., Lübeck, F., guistics (Volume 1: Long Papers), pp. 15504–15522, 5
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long. 828. URL https://aclanthology.org/2024. acl-long.828/. Russell, J. A. A circumplex model of affect. Journal of personality and social psychology, 39(6):1161, 1980. Sofroniew, N., Kauvar, I., Saunders, W., Chen, R., Henighan, T., Hydrie, S., Citro, C., Pearce, A., Tarng, J., Gurnee, W., Batson, J., Zimmerman, S., Rivoire, K., Fish, K., Olah, C., and Lindsey, J. Emotion concepts and their function in a large language model. Transformer Circuits Thread, 2026. URL https://transformer-circuits.pub/ 2026/emotions/index.html. Sun, L., Yan, L., Lu, X., Lee, A., Zhang, J., and Shao, J. Valence-arousal subspace in llms: Circular emotion geometry and multi-behavioral control, 2026. URL https://arxiv.org/abs/2604.03147. Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N. Language models linearly represent sentiment. In ICML 2024 Workshop on Mechanistic Interpretability, 2024. URL https://openreview.net/forum? id=Xsf6dOOMMc. Valeriani, L., Doimo, D., Cuturello, F., Laio, A., ansuini, A., and Cazzaniga, A. The geometry of hidden representations of large transformer models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum? id=cCYvakU5Ek.
6
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
A. Related Work Linear representations in LLMs. The linear representation hypothesis holds that high-level concepts are encoded as directions in activation space (Mikolov et al., 2013; Elhage et al., 2022; Park et al., 2023). Tigges et al. (2024) demonstrated this for sentiment, finding a single direction captures positive-negative valence across tasks. Subsequent work extended linear representations to truth (Marks & Tegmark, 2024), refusal (Arditi et al., 2024), and behavioral tendencies (Rimsky et al., 2024). Sparse autoencoders can extract directions at scale, decomposing polysemantic activations into interpretable features (Bricken et al., 2023; Cunningham et al., 2023). Emotion in language models. Early work identified a “sentiment neuron” in LSTMs (Radford et al., 2017), though later analysis suggested emotional content is distributed across many neurons. (Sofroniew et al., 2026) provide a comprehensive analysis, extracting 171 emotion vectors from Claude Sonnet 4.5 and demonstrating causal influence on behavior. They found emotion geometry mirrors human psychological structure, with valence and arousal as principal axes. Concurrent work extends this to other models: (Sun et al., 2026) identify a valence-arousal subspace in Llama and Qwen with circumplex-consistent circular geometry, where steering along VA axes controls refusal and sycophancy. (Choi & Weber, 2026) find coherent affective representations in Gemma-2, Mistral, and LLaMA with modest nonlinear global structure. We build on Sofroniew et al. (2026), testing generalization across architectures and the role of extraction methodology. Cross-layer geometry. Transformer representations evolve across layers in characteristic ways. Valeriani et al. (2023) found intrinsic dimension expands then contracts, with semantics concentrated at intermediate depths. (Cheng et al., 2025) identified a ”high-dimensional abstraction phase” where representations peak in complexity before simplifying toward outputs.
7
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
B. Story Dataset The emotion story datasets was generated using A PERTUS -8B and G EMMA -4-E4B, following a methodology similar to Anthropic’s emotion vectors work (Sofroniew et al., 2026). The 171 emotions were copied from their work. Stories were designed to convey emotions implicitly, such as never naming the target emotion directly, but instead relying instead on character actions, physical sensations, dialogue, and situational context. The prompts used were also similar, to introduce as little methodological confound as possible. B.1. Dataset statistics
Table 1. Emotion story dataset statistics by corpus. Apertus stories were deduplicated to match the uniform 9-stories-per-topic structure of the Gemma corpus.
Statistic
Apertus stories
Gemma stories
Generator model Total emotions Unique topics Stories per emotion Total stories Mean story length Story length range
A PERTUS -8B-Instruct-2509 171 100 9 1,539 ∼215 words 65–665 words
G EMMA -4-E4B-it 171 100 9 1,539 ∼144 words 81–298 words
B.2. Activation collection Residual stream activations were collected from A PERTUS -8B-Instruct at multiple transformer layers (Table 2). The model can be found through HuggingFace: swiss-ai/A P E R T U S -8B-Instruct-2509. For Gemma, different layers were picked to collect activations from (Table 2). The model was also accessed through HuggingFace: google/G E M M A -4-E4B-it.
Table 2. Activation extraction configuration for both models.
Parameter
A PERTUS -8B
G EMMA -4-E4B 8B
Model Layers collected Hook location Batch size Max sequence length Dataset
A PERTUS -8B-Instruct-2509 1-31 resid post 8 sequences 1024 tokens Pile (uncopyrighted), train split
G EMMA -4-E4B 1-40 resid post 8 sequences 1024 tokens Pile (uncopyrighted), train split
8
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
C. Additional Results C.1. Principal Component Valence
Table 3. PC1–valence (Pearson r) across layers and story conditions. Bold indicates the peak layer per model–condition pair. † p < 0.05; ‡ p < 0.01; ∗ p < 0.001. A PERTUS -8B
G EMMA -4-E4B 8B
Layer
A PERTUS -8B stories r
G EMMA -4-E4B stories r
Layer
A PERTUS -8B stories r
G EMMA -4-E4B stories r
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40
0.1675† 0.0154 0.0152 0.0152 0.0151 0.0151 0.0154 0.0153 0.0163 0.0176 0.0188 0.0193 0.0209 0.0246 0.0302 0.0400 0.0821 0.1732† 0.3112∗ 0.6169∗ 0.6882∗ 0.6616∗ 0.7230∗ 0.7216∗ 0.7040∗ 0.6875∗ 0.6710∗ 0.6803∗ 0.6302∗ 0.6183∗ 0.5910∗ — — — — — — — — —
0.3737∗ 0.0895 0.0894 0.0894 0.0894 0.0895 0.0895 0.0894 0.0900 0.0908 0.0916 0.0928 0.0928 0.0938 0.0946 0.0959 0.1034 0.1141 0.1343 0.2034‡ 0.5739∗ 0.7181∗ 0.7478∗ 0.7539∗ 0.7561∗ 0.7606∗ 0.7535∗ 0.7533∗ 0.7573∗ 0.7564∗ 0.7608∗ — — — — — — — — —
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40
0.2000‡ 0.2030‡ 0.2040‡ 0.0098 0.3592∗ 0.4550∗ 0.4835∗ 0.1359 0.1851† 0.2401‡ 0.1598† 0.0614 0.7940∗ 0.6797∗ 0.7595∗ 0.7533∗ 0.1318 0.0060 0.0564 0.1337 0.1591† 0.0968 0.1187 0.0370 0.0117 0.0436 0.0067 0.0242 0.0277 0.2014‡ 0.1505† 0.1567† 0.1754† 0.1910† 0.1366 0.0939 0.0743 0.0786 0.0769 0.0081
0.3986∗ 0.5138∗ 0.5857∗ 0.6708∗ 0.5443∗ 0.4619∗ 0.6333∗ 0.6352∗ 0.3248∗ 0.3851∗ 0.7664∗ 0.7675∗ 0.7984∗ 0.7795∗ 0.7938∗ 0.8296∗ 0.1360 0.1880† 0.1418 0.1350 0.1393 0.0403 0.1769† 0.1796† 0.1446 0.1536† 0.1319 0.1015 0.1986‡ 0.2101‡ 0.1672† 0.1793† 0.1105 0.0408 0.1034 0.1222 0.1841† 0.1861† 0.2099‡ 0.1810†
9
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
C.2. Principal Component Arousal Table 4. PC2–arousal (Pearson r) across layers and story conditions. Bold indicates the peak layer per model–condition pair. † p < 0.05; ‡ p < 0.01; ∗ p < 0.001. A PERTUS -8B
G EMMA -4-E4B
Layer
A PERTUS -8B stories r
G EMMA -4-E4B stories r
Layer
A PERTUS -8B stories r
G EMMA -4-E4B stories r
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40
0.2093‡ 0.0737 0.0735 0.0778 0.1168 0.1032 0.1370 0.1488† 0.1610† 0.1585† 0.1328 0.1085 0.1350 0.1351 0.1137 0.1573† 0.1329 0.1657† 0.1469† 0.1211 0.0891 0.1637† 0.1459† 0.0972 0.0810 0.0586 0.0597 0.0654 0.0371 0.0331 0.1165 — — — — — — — — —
0.2012‡ 0.0105 0.0122 0.0124 0.0214 0.3171∗ 0.3772∗ 0.3520∗ 0.3345∗ 0.3076∗ 0.3060∗ 0.2923∗ 0.2731∗ 0.2883∗ 0.2818∗ 0.3262∗ 0.2890∗ 0.3290∗ 0.3166∗ 0.2974∗ 0.2167‡ 0.0607 0.4222∗ 0.4341∗ 0.4432∗ 0.4480∗ 0.4049∗ 0.4165∗ 0.3804∗ 0.3449∗ 0.0866 — — — — — — — — —
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40
0.0624 0.1273 0.1472† 0.1060 0.0815 0.1086 0.1064 0.1298 0.1867† 0.1574† 0.1384 0.1414 0.0459 0.0707 0.0360 0.0048 0.0726 0.0345 0.0054 0.1114 0.0965 0.1325 0.1879† 0.1931† 0.1769† 0.1690† 0.1755† 0.1585† 0.2030‡ 0.1437 0.1524† 0.1487 0.1393 0.1248 0.1509† 0.1577† 0.1665† 0.1650† 0.1706† 0.2047‡
0.1141 0.0845 0.3005∗ 0.4120∗ 0.0282 0.0476 0.3625∗ 0.2450‡ 0.4008∗ 0.3710∗ 0.2168‡ 0.0872 0.1928† 0.2045‡ 0.1709† 0.1700† 0.2094‡ 0.2341‡ 0.2407‡ 0.2547∗ 0.1018 0.1660† 0.2452‡ 0.2819∗ 0.2970∗ 0.2997∗ 0.2979∗ 0.2538∗ 0.3741∗ 0.3816∗ 0.4115∗ 0.3904∗ 0.3950∗ 0.3638∗ 0.3926∗ 0.4251∗ 0.3989∗ 0.3863∗ 0.3818∗ 0.3234∗
10
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
C.3. CKA Figures CKA is a measure of how similar the emotion space is between 2 layers. The diagonal is always 1, which is a layer compared to itself.
• CKA close to 1: spatial arrangement of emotion vectors between 2 layers is nearly identical.
• CKA close to 0: spatial arrangement has changed substantially between 2 layers
So each cell answers: does the model organize emotions in the same way at layer A as in layer B? The higher the value, the more similar. C.3.1. A PERTUS -8B CKA RESULTS
Figure 2. A PERTUS -8B CKA results on Apertus stories Layers 22 31
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
12
1.00
1.00
1.00
0.98
0.93
0.81
0.66
0.56
0.43
0.33
22
1.00
1.00
0.99
0.99
0.98
0.98
0.96
0.95
0.95
0.84
3
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
13
1.00
1.00
1.00
0.99
0.94
0.83
0.68
0.59
0.46
0.36
23
1.00
1.00
1.00
1.00
0.99
0.99
0.98
0.97
0.96
0.86
4
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
14
1.00
1.00
1.00
0.99
0.96
0.86
0.72
0.63
0.50
0.40
24
0.99
1.00
1.00
1.00
1.00
0.99
0.98
0.97
0.97
0.87
5
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
15
0.98
0.99
0.99
1.00
0.98
0.90
0.78
0.70
0.58
0.49
25
0.99
1.00
1.00
1.00
1.00
1.00
0.99
0.98
0.97
0.88
6
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
16
0.93
0.94
0.96
0.98
1.00
0.97
0.89
0.82
0.72
0.64
26
0.98
0.99
1.00
1.00
1.00
1.00
0.99
0.99
0.98
0.88
7
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
17
0.81
0.83
0.86
0.90
0.97
1.00
0.97
0.93
0.87
0.81
27
0.98
0.99
0.99
1.00
1.00
1.00
1.00
0.99
0.99
0.89
8
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
18
0.66
0.68
0.72
0.78
0.89
0.97
1.00
0.99
0.96
0.92
28
0.96
0.98
0.98
0.99
0.99
1.00
1.00
1.00
1.00
0.90
layer
1.00
layer
layer
Layers 12 21
2
9
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
19
0.56
0.59
0.63
0.70
0.82
0.93
0.99
1.00
0.98
0.96
29
0.95
0.97
0.97
0.98
0.99
0.99
1.00
1.00
1.00
0.91
10
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
20
0.43
0.46
0.50
0.58
0.72
0.87
0.96
0.98
1.00
0.99
30
0.95
0.96
0.97
0.97
0.98
0.99
1.00
1.00
1.00
0.92
11
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
21
0.33
0.36
0.40
0.49
0.64
0.81
0.92
0.96
0.99
1.00
31
0.84
0.86
0.87
0.88
0.88
0.89
0.90
0.91
0.92
1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
layer
layer
layer
1.0
0.8
0.6
linear CKA
Apertus 8B Representational similarity (CKA) Apertus stories
Layers 2 11
0.4
0.2
0.0
Figure 3. A PERTUS -8B CKA values on Gemma stories Layers 22 31
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
12
1.00
1.00
1.00
1.00
0.99
0.96
0.91
0.86
0.74
0.58
22
1.00
0.99
0.98
0.97
0.95
0.94
0.90
0.89
0.89
0.86
3
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
13
1.00
1.00
1.00
1.00
0.99
0.97
0.92
0.86
0.74
0.59
23
0.99
1.00
1.00
0.99
0.99
0.98
0.95
0.95
0.95
0.91
4
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
14
1.00
1.00
1.00
1.00
0.99
0.97
0.92
0.87
0.76
0.60
24
0.98
1.00
1.00
1.00
0.99
0.99
0.97
0.97
0.96
0.93
5
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
15
1.00
1.00
1.00
1.00
1.00
0.98
0.94
0.89
0.78
0.63
25
0.97
0.99
1.00
1.00
1.00
0.99
0.98
0.97
0.97
0.94
6
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
16
0.99
0.99
0.99
1.00
1.00
0.99
0.96
0.92
0.82
0.68
26
0.95
0.99
0.99
1.00
1.00
1.00
0.99
0.98
0.98
0.95
7
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
17
0.96
0.97
0.97
0.98
0.99
1.00
0.99
0.96
0.89
0.77
27
0.94
0.98
0.99
0.99
1.00
1.00
0.99
0.99
0.99
0.96
8
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
18
0.91
0.92
0.92
0.94
0.96
0.99
1.00
0.99
0.95
0.86
28
0.90
0.95
0.97
0.98
0.99
0.99
1.00
1.00
1.00
0.97
layer
1.00
layer
layer
Layers 12 21
2
9
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
19
0.86
0.86
0.87
0.89
0.92
0.96
0.99
1.00
0.98
0.91
29
0.89
0.95
0.97
0.97
0.98
0.99
1.00
1.00
1.00
0.97
10
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
20
0.74
0.74
0.76
0.78
0.82
0.89
0.95
0.98
1.00
0.97
30
0.89
0.95
0.96
0.97
0.98
0.99
1.00
1.00
1.00
0.97
11
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
21
0.58
0.59
0.60
0.63
0.68
0.77
0.86
0.91
0.97
1.00
31
0.86
0.91
0.93
0.94
0.95
0.96
0.97
0.97
0.97
1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
layer
layer
11
layer
1.0
0.8
0.6
linear CKA
Apertus 8B Representational similarity (CKA) Gemma stories
Layers 2 11
0.4
0.2
0.0
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
C.3.2. G EMMA -4-E4B CKA RESULTS Figure 4. G EMMA -4-E4B CKA values on Gemma stories Layers 28 40
0.91
0.86
0.83
0.81
0.81
0.82
0.83
0.81
0.80
0.80
0.77
0.76
15
1.00
0.97
0.91
0.84
0.77
0.73
0.67
0.75
0.72
0.79
0.82
0.83
0.82
28
1.00
0.90
0.91
0.90
0.90
0.90
0.92
0.91
0.89
0.89
0.88
0.87
0.89
3
0.91
1.00
0.95
0.90
0.89
0.91
0.89
0.85
0.84
0.87
0.87
0.86
0.85
16
0.97
1.00
0.94
0.87
0.81
0.78
0.73
0.80
0.76
0.82
0.84
0.85
0.84
29
0.90
1.00
0.95
0.92
0.92
0.91
0.91
0.91
0.91
0.90
0.89
0.88
0.88
4
0.86
0.95
1.00
0.95
0.93
0.94
0.89
0.85
0.83
0.88
0.89
0.88
0.86
17
0.91
0.94
1.00
0.95
0.90
0.86
0.81
0.85
0.83
0.87
0.88
0.88
0.88
30
0.91
0.95
1.00
0.98
0.98
0.97
0.95
0.96
0.95
0.95
0.94
0.94
0.93
5
0.83
0.90
0.95
1.00
0.98
0.95
0.89
0.86
0.85
0.89
0.89
0.87
0.86
18
0.84
0.87
0.95
1.00
0.97
0.93
0.88
0.86
0.88
0.90
0.90
0.89
0.89
31
0.90
0.92
0.98
1.00
0.99
0.97
0.96
0.95
0.95
0.95
0.95
0.95
0.94
6
0.81
0.89
0.93
0.98
1.00
0.98
0.92
0.87
0.86
0.90
0.90
0.89
0.88
19
0.77
0.81
0.90
0.97
1.00
0.97
0.92
0.87
0.90
0.91
0.90
0.89
0.89
32
0.90
0.92
0.98
0.99
1.00
0.97
0.96
0.96
0.96
0.95
0.95
0.95
0.94
7
0.81
0.91
0.94
0.95
0.98
1.00
0.94
0.88
0.86
0.90
0.91
0.90
0.89
20
0.73
0.78
0.86
0.93
0.97
1.00
0.97
0.89
0.93
0.92
0.90
0.89
0.90
33
0.90
0.91
0.97
0.97
0.97
1.00
0.98
0.97
0.96
0.96
0.95
0.95
0.94
8
0.82
0.89
0.89
0.89
0.92
0.94
1.00
0.96
0.92
0.93
0.92
0.91
0.91
21
0.67
0.73
0.81
0.88
0.92
0.97
1.00
0.91
0.94
0.92
0.89
0.88
0.89
34
0.92
0.91
0.95
0.96
0.96
0.98
1.00
0.98
0.96
0.96
0.95
0.94
0.94
layer
1.00
layer
layer
Layers 15 27
2
9
0.83
0.85
0.85
0.86
0.87
0.88
0.96
1.00
0.97
0.94
0.92
0.89
0.90
22
0.75
0.80
0.85
0.86
0.87
0.89
0.91
1.00
0.89
0.90
0.89
0.88
0.89
35
0.91
0.91
0.96
0.95
0.96
0.97
0.98
1.00
0.99
0.98
0.98
0.97
0.96
10
0.81
0.84
0.83
0.85
0.86
0.86
0.92
0.97
1.00
0.95
0.93
0.90
0.90
23
0.72
0.76
0.83
0.88
0.90
0.93
0.94
0.89
1.00
0.98
0.95
0.94
0.94
36
0.89
0.91
0.95
0.95
0.96
0.96
0.96
0.99
1.00
0.99
0.99
0.98
0.96
11
0.80
0.87
0.88
0.89
0.90
0.90
0.93
0.94
0.95
1.00
0.99
0.97
0.96
24
0.79
0.82
0.87
0.90
0.91
0.92
0.92
0.90
0.98
1.00
0.99
0.98
0.98
37
0.89
0.90
0.95
0.95
0.95
0.96
0.96
0.98
0.99
1.00
1.00
0.99
0.97
12
0.80
0.87
0.89
0.89
0.90
0.91
0.92
0.92
0.93
0.99
1.00
0.99
0.97
25
0.82
0.84
0.88
0.90
0.90
0.90
0.89
0.89
0.95
0.99
1.00
0.99
0.99
38
0.88
0.89
0.94
0.95
0.95
0.95
0.95
0.98
0.99
1.00
1.00
1.00
0.97
13
0.77
0.86
0.88
0.87
0.89
0.90
0.91
0.89
0.90
0.97
0.99
1.00
0.99
26
0.83
0.85
0.88
0.89
0.89
0.89
0.88
0.88
0.94
0.98
0.99
1.00
0.99
39
0.87
0.88
0.94
0.95
0.95
0.95
0.94
0.97
0.98
0.99
1.00
1.00
0.98
14
0.76
0.85
0.86
0.86
0.88
0.89
0.91
0.90
0.90
0.96
0.97
0.99
1.00
27
0.82
0.84
0.88
0.89
0.89
0.90
0.89
0.89
0.94
0.98
0.99
0.99
1.00
40
0.89
0.88
0.93
0.94
0.94
0.94
0.94
0.96
0.96
0.97
0.97
0.98
1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
layer
layer
layer
1.0
0.8
0.6
linear CKA
Gemma 4 8B Representational similarity (CKA) Gemma stories
Layers 2 14
0.4
0.2
0.0
Figure 5. G EMMA -4-E4B CKA values on Apertus stories Layers 28 40
0.94
0.87
0.85
0.83
0.82
0.85
0.87
0.87
0.82
0.80
0.74
0.73
15
1.00
0.97
0.93
0.86
0.82
0.79
0.74
0.80
0.77
0.83
0.84
0.85
0.84
28
1.00
0.95
0.95
0.94
0.94
0.93
0.93
0.93
0.92
0.92
0.92
0.92
0.92
3
0.94
1.00
0.96
0.92
0.91
0.91
0.91
0.91
0.91
0.90
0.88
0.85
0.83
16
0.97
1.00
0.95
0.88
0.84
0.81
0.75
0.83
0.78
0.83
0.83
0.84
0.83
29
0.95
1.00
0.96
0.95
0.94
0.94
0.93
0.94
0.94
0.93
0.93
0.93
0.92
4
0.87
0.96
1.00
0.96
0.95
0.95
0.93
0.91
0.91
0.92
0.91
0.89
0.87
17
0.93
0.95
1.00
0.96
0.93
0.90
0.86
0.87
0.86
0.88
0.88
0.88
0.87
30
0.95
0.96
1.00
0.99
0.98
0.98
0.97
0.97
0.96
0.96
0.96
0.95
0.93
5
0.85
0.92
0.96
1.00
0.98
0.96
0.91
0.89
0.90
0.92
0.91
0.88
0.87
18
0.86
0.88
0.96
1.00
0.98
0.95
0.90
0.88
0.89
0.89
0.88
0.87
0.87
31
0.94
0.95
0.99
1.00
1.00
0.98
0.98
0.97
0.97
0.96
0.96
0.96
0.94
6
0.83
0.91
0.95
0.98
1.00
0.99
0.94
0.91
0.91
0.93
0.92
0.90
0.90
19
0.82
0.84
0.93
0.98
1.00
0.97
0.93
0.88
0.91
0.90
0.89
0.88
0.88
32
0.94
0.94
0.98
1.00
1.00
0.98
0.98
0.97
0.97
0.96
0.97
0.97
0.94
7
0.82
0.91
0.95
0.96
0.99
1.00
0.96
0.92
0.91
0.94
0.94
0.92
0.91
20
0.79
0.81
0.90
0.95
0.97
1.00
0.96
0.88
0.92
0.90
0.88
0.87
0.88
33
0.93
0.94
0.98
0.98
0.98
1.00
0.99
0.98
0.97
0.97
0.97
0.96
0.94
8
0.85
0.91
0.93
0.91
0.94
0.96
1.00
0.98
0.96
0.95
0.93
0.91
0.90
21
0.74
0.75
0.86
0.90
0.93
0.96
1.00
0.89
0.94
0.90
0.88
0.87
0.87
34
0.93
0.93
0.97
0.98
0.98
0.99
1.00
0.99
0.98
0.97
0.97
0.97
0.94
layer
1.00
layer
layer
Layers 15 27
2
9
0.87
0.91
0.91
0.89
0.91
0.92
0.98
1.00
0.98
0.95
0.92
0.88
0.88
22
0.80
0.83
0.87
0.88
0.88
0.88
0.89
1.00
0.88
0.88
0.86
0.86
0.86
35
0.93
0.94
0.97
0.97
0.97
0.98
0.99
1.00
0.99
0.99
0.98
0.98
0.95
10
0.87
0.91
0.91
0.90
0.91
0.91
0.96
0.98
1.00
0.97
0.94
0.90
0.90
23
0.77
0.78
0.86
0.89
0.91
0.92
0.94
0.88
1.00
0.97
0.95
0.93
0.93
36
0.92
0.94
0.96
0.97
0.97
0.97
0.98
0.99
1.00
1.00
0.99
0.99
0.96
11
0.82
0.90
0.92
0.92
0.93
0.94
0.95
0.95
0.97
1.00
0.99
0.96
0.95
24
0.83
0.83
0.88
0.89
0.90
0.90
0.90
0.88
0.97
1.00
0.99
0.98
0.98
37
0.92
0.93
0.96
0.96
0.96
0.97
0.97
0.99
1.00
1.00
1.00
0.99
0.97
12
0.80
0.88
0.91
0.91
0.92
0.94
0.93
0.92
0.94
0.99
1.00
0.98
0.97
25
0.84
0.83
0.88
0.88
0.89
0.88
0.88
0.86
0.95
0.99
1.00
1.00
0.99
38
0.92
0.93
0.96
0.96
0.97
0.97
0.97
0.98
0.99
1.00
1.00
1.00
0.97
13
0.74
0.85
0.89
0.88
0.90
0.92
0.91
0.88
0.90
0.96
0.98
1.00
0.99
26
0.85
0.84
0.88
0.87
0.88
0.87
0.87
0.86
0.93
0.98
1.00
1.00
1.00
39
0.92
0.93
0.95
0.96
0.97
0.96
0.97
0.98
0.99
0.99
1.00
1.00
0.98
14
0.73
0.83
0.87
0.87
0.90
0.91
0.90
0.88
0.90
0.95
0.97
0.99
1.00
27
0.84
0.83
0.87
0.87
0.88
0.88
0.87
0.86
0.93
0.98
0.99
1.00
1.00
40
0.92
0.92
0.93
0.94
0.94
0.94
0.94
0.95
0.96
0.97
0.97
0.98
1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
layer
layer
layer
1.0
0.8
0.6
linear CKA
Gemma 4 8B Representational similarity (CKA) Apertus stories
Layers 2 14
0.4
0.2
0.0
C.4. Valence Direction Alignment Each cell shows the cosine similarity between the valence direction vectors at 2 layers. The valence direction is the axis in activation space that best predicts the emotion valence. • Cosine similarity close to 1. Valence axis points in the same direction in both layers, consistent positive axis. • Cosine similarity close to 0. Valence axes are orthogonal, they’ve rotated completely. • Cosine similarity close to -1. The axis has flipped direction. CKA provides information about the whole space, while the cosine similarity specifically shows whether the valence axis is stable. A predominantly blue matrix would indicate that the model has a persistent stable direction to represent positive vs. negative emotions across many layers. The valence direction stability line plot shows the cosine similarity between 2 adjacent layers. It has a similar interpretation as the values in the panel, but only for adjacent layers. The interpretation can be slightly different, because it shows if the valence axis points in the same direction from layer to layer. A dip reveals a specific transition, where the model changes how it encodes valence. 12
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
C.4.1. A PERTUS -8B VALENCE A LIGNMENT
Figure 6. A PERTUS -8B validation on Apertus stories Apertus 8B Valence direction alignment Apertus stories
3 4
-0.15
0.21
0.35
-0.15
1.00
-0.09
0.21
-0.09
1.00
5
0.35
-0.24
6
-0.04
7
Layers 12 21
-0.30
12
1.00
-0.08
0.21
13
0.03
-0.13
14
0.20
0.27
-0.49
0.06
0.13
-0.00
0.10
0.15
-0.30
-0.14
-0.14
0.15
1.00
0.10
-0.22
-0.14
0.10
1.00
-0.30
0.15
-0.22
7
8
9
-0.04
0.16
-0.02
0.10
0.16
-0.24
0.04
-0.14
0.02
-0.05
0.11
-0.15
0.05
0.02
0.08
0.11
1.00
0.08
0.30
-0.15
0.04
-0.15
0.08
1.00
0.00
-0.10
0.16
-0.14
0.05
0.30
0.00
1.00
-0.17
8
-0.02
0.02
0.02
-0.15
-0.10
-0.17
1.00
9
0.10
-0.05
0.08
0.20
0.06
0.10
-0.14
10
0.16
-0.08
0.03
0.27
0.13
0.15
11
-0.30
0.21
-0.13
-0.49
-0.00
2
3
4
5
6
layer
-0.13
0.15
-0.30
0.20
-0.13
1.00
0.15
-0.20
-0.20
0.08
1.00
-0.21
15
-0.30
0.08
-0.21
16
0.20
-0.15
0.19
17
-0.18
0.04
18
0.25
-0.12
19
-0.09
-0.36
20
-0.36
1.00
21
10
11
Layers 22 31 22
0.13
0.07
23
-0.15
-0.11
24
0.25
0.21
25
-0.22
-0.17
0.25
-0.15
0.04
-0.12
0.11
0.19
-0.10
0.16
-0.09
1.00
-0.26
0.26
-0.31
0.12
-0.26
1.00
-0.23
0.25
-0.14
-0.10
0.26
-0.23
1.00
-0.32
0.18
0.19
0.11
0.16
-0.31
0.25
-0.32
1.00
-0.15
-0.35
-0.24
0.11
-0.09
0.12
-0.14
0.18
-0.15
1.00
0.20
0.07
-0.23
0.13
-0.15
0.25
-0.22
0.19
-0.35
0.20
1.00
-0.13
0.07
-0.11
0.21
-0.17
0.11
-0.24
0.07
12
13
14
15
16
17
18
19
layer
-0.09
-0.13
-0.18
-0.23
1.00
0.07
-0.09
-0.10
-0.12
-0.02
0.07
1.00
-0.09
-0.13
-0.14
-0.09
-0.09
1.00
0.25
0.16
-0.10
-0.13
0.25
1.00
26
-0.12
-0.14
0.16
0.15
27
-0.02
-0.07
0.12
0.10
28
0.06
0.11
-0.22
-0.20
29
-0.08
-0.04
0.16
0.15
0.21
30
-0.06
-0.06
0.04
0.21
1.00
31
-0.06
-0.04
20
21
22
23
1.00
0.06
-0.08
-0.06
-0.06
-0.07
0.11
-0.04
-0.06
-0.04
0.75
0.12
-0.22
0.16
0.04
0.11
0.50
0.15
0.10
-0.20
0.15
0.05
0.10
1.00
0.12
-0.05
0.11
0.25
0.08
0.12
1.00
-0.07
0.01
0.08
0.04
-0.05
-0.07
1.00
-0.29
0.21
-0.11
0.25
0.11
0.01
-0.29
1.00
-0.13
0.10
0.50
0.05
0.25
0.08
0.21
-0.13
1.00
0.00
0.11
0.10
0.08
0.04
-0.11
0.10
0.00
1.00
24
25
26
27
28
29
30
31
-0.03
-0.03
-0.05
-0.03
layer
0.25 0.00
cosine sim
1.00
layer
2
layer
layer
Layers 2 11
0.75 1.00
Figure 7. A PERTUS -8B validation on Gemma stories Apertus 8B Valence direction alignment Gemma stories
-0.11
-0.29
3
-0.11
1.00
0.15
4
-0.29
0.15
1.00
5
0.35
0.02
-0.25
6
0.38
0.02
7
-0.26
8 9
Layers 12 21
0.38
-0.26
-0.25
-0.27
-0.28
-0.23
12
1.00
0.29
0.00
0.02
0.02
-0.03
0.01
-0.01
-0.06
-0.03
13
0.29
1.00
-0.25
-0.24
0.19
0.14
0.16
0.13
0.13
14
0.00
0.02
1.00
0.57
-0.43
-0.28
-0.37
-0.35
-0.28
15
-0.02
-0.24
0.57
1.00
-0.54
-0.41
-0.46
-0.43
-0.35
16
-0.03
0.19
-0.43
-0.54
1.00
0.37
0.48
0.42
0.32
-0.25
0.01
0.14
-0.28
-0.41
0.37
1.00
0.40
0.37
-0.27
-0.01
0.16
-0.37
-0.46
0.48
0.40
1.00
0.57
10
-0.28
-0.06
0.13
-0.35
-0.43
0.42
0.37
0.57
11
-0.23
-0.03
0.13
-0.28
-0.35
0.32
0.23
2
3
4
5
6
7
8
0.35
layer
Layers 22 31
1.00
0.17
0.23
-0.06
22
1.00
-0.01
0.02
-0.09
-0.04
0.20
0.27
-0.06
23
-0.01
1.00
0.02
0.08
0.05
0.01
0.01
0.10
-0.11
-0.06
0.75
0.03
-0.02
0.07
0.06
24
0.02
0.02
1.00
0.09
-0.00
-0.03
-0.07
0.10
-0.13
-0.02
0.50
0.18
0.02
-0.05
-0.10
-0.02
25
-0.09
0.08
0.09
1.00
0.22
0.24
-0.13
0.27
-0.11
0.15
1.00
-0.52
-0.02
0.25
0.46
-0.06
26
-0.02
0.05
-0.00
0.22
1.00
0.09
-0.05
0.05
-0.01
0.05
0.18
-0.52
1.00
0.07
-0.21
-0.38
-0.02
27
-0.06
0.01
-0.03
0.24
0.09
1.00
0.04
0.04
0.15
0.12
0.03
0.02
-0.02
0.07
1.00
-0.06
0.07
0.12
28
-0.03
0.01
-0.07
-0.13
-0.05
0.04
1.00
-0.09
0.16
-0.01
0.25
-0.02
-0.05
0.25
-0.21
-0.06
1.00
0.21
-0.04
29
-0.03
0.10
0.10
0.27
0.05
0.04
-0.09
1.00
-0.12
0.01
0.50
0.07
-0.10
0.46
-0.38
0.07
0.21
1.00
-0.11
30
-0.05
-0.11
-0.13
-0.11
-0.01
0.15
0.16
-0.12
1.00
0.17
-0.06
0.06
-0.02
-0.06
-0.02
0.12
-0.04
-0.11
1.00
31
-0.03
-0.06
-0.02
0.15
0.05
0.12
-0.01
0.01
0.17
1.00
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
0.15
0.12
-0.02
0.28
-0.17
-0.03
0.02
0.05
0.40
-0.26
1.00
-0.12
0.05
-0.06
0.05
-0.12
1.00
-0.16
0.28
0.40
0.05
-0.16
17
-0.17
-0.26
-0.06
0.23
18
-0.03
-0.04
0.35
19
0.17
0.20
1.00
0.41
20
0.23
0.27
0.35
0.41
1.00
21
-0.06
9
10
11
12
layer
-0.02
-0.06
layer
0.25 0.00
0.75 1.00
Figure 8. A PERTUS -8B validation on Apertus stories Apertus 8B 1.00
Valence direction: adjacent-layer stability Apertus stories
0.75
cosine similarity
0.50 0.25
0.11
0.00
0.08
-0.15
0.07 -0.13
-0.14
-0.17
0.25
0.21
0.20
0.10
0.00
-0.09
0.25
-0.36
-0.20
-0.21
-0.15
-0.23
-0.26
-0.15
0.00
-0.07
-0.09
-0.13 -0.29
-0.32
-0.42
0.50 0.75 1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
layer (midpoint of pair)
13
18
19
20
21
22
23
24
25
26
27
28
29
cosine sim
1.00
layer
2
layer
layer
Layers 2 11
30
31
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs Figure 9. A PERTUS -8B validation on Gemma stories Apertus 8B
Valence direction: adjacent-layer stability Gemma stories
0.02
0.07
1.00 0.75
0.57
0.57
cosine similarity
0.50 0.25
0.41
0.40
0.37
0.29
0.00
-0.11
-0.12
-0.25
0.25
-0.06
-0.16
-0.28
0.17
0.09
0.09
0.02
-0.01
-0.07
-0.11
0.04 -0.09
-0.12
-0.52
-0.54
0.50
0.22
0.21
0.15
0.75 1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
layer (midpoint of pair)
18
19
20
21
22
23
24
25
26
27
28
29
30
31
C.4.2. G EMMA -4-E4B R ESULTS
Figure 10. G EMMA -4-E4B validation on Gemma stories Layers 28 40
-0.35
-0.37
0.31
-0.06
-0.17
0.29
0.03
0.02
0.03
0.01
0.00
0.02
15
1.00
0.42
-0.03
0.03
0.02
0.02
-0.03
0.01
-0.00
0.01
0.01
0.02
0.00
28
1.00
0.34
0.39
0.45
-0.14
-0.32
0.40
0.46
-0.28
0.03
-0.34
-0.31
-0.38
3
-0.35
1.00
0.24
-0.15
0.06
0.13
-0.12
-0.00
-0.02
-0.00
0.02
0.00
0.01
16
0.42
1.00
-0.10
0.03
0.01
0.02
-0.04
-0.01
-0.02
0.02
0.01
0.03
-0.01
29
0.34
1.00
0.35
0.33
-0.08
-0.21
0.28
0.29
-0.21
0.01
-0.19
-0.19
-0.28
4
-0.37
0.24
1.00
-0.46
0.14
0.25
-0.36
-0.01
0.02
0.00
0.05
0.01
0.01
17
-0.03
-0.10
1.00
-0.09
0.01
0.01
-0.04
-0.02
-0.00
-0.00
0.00
-0.01
-0.00
30
0.39
0.35
1.00
0.55
-0.16
-0.36
0.45
0.40
-0.24
0.01
-0.31
-0.29
-0.34
5
0.31
-0.15
-0.46
1.00
-0.19
-0.25
0.31
0.03
0.01
0.00
0.00
-0.02
0.02
18
0.03
0.03
-0.09
1.00
0.14
0.14
0.00
-0.01
-0.01
-0.04
-0.01
0.02
-0.02
31
0.45
0.33
0.55
1.00
-0.25
-0.42
0.53
0.52
-0.38
0.01
-0.39
-0.37
-0.42
6
-0.06
0.06
0.14
-0.19
1.00
0.09
-0.10
0.03
0.02
0.01
0.01
0.01
0.01
19
0.02
0.01
0.01
0.14
1.00
0.54
-0.06
-0.05
-0.03
-0.01
-0.01
-0.02
-0.03
32
-0.14
-0.08
-0.16
-0.25
1.00
0.16
-0.20
-0.20
0.12
0.00
0.12
0.12
0.10
7
-0.17
0.13
0.25
-0.25
0.09
1.00
-0.25
-0.00
-0.00
-0.02
-0.04
-0.00
-0.02
20
0.02
0.02
0.01
0.14
0.54
1.00
-0.07
-0.01
0.00
0.02
-0.03
-0.01
-0.02
33
-0.32
-0.21
-0.36
-0.42
0.16
1.00
-0.44
-0.40
0.26
0.02
0.27
0.25
0.28
8
0.29
-0.12
-0.36
0.31
-0.10
-0.25
1.00
0.04
0.03
0.01
0.04
-0.03
-0.01
21
-0.03
-0.04
-0.04
0.00
-0.06
-0.07
1.00
-0.03
0.01
-0.03
0.01
0.02
0.03
34
0.40
0.28
0.45
0.53
-0.20
-0.44
1.00
0.52
-0.31
-0.04
-0.37
-0.32
-0.35
9
0.03
-0.00
-0.01
0.03
0.03
-0.00
0.04
1.00
0.35
0.23
0.11
0.19
0.12
22
0.01
-0.01
-0.02
-0.01
-0.05
-0.01
-0.03
1.00
-0.00
0.02
0.02
-0.01
0.03
35
0.46
0.29
0.40
0.52
-0.20
-0.40
0.52
1.00
-0.41
0.01
-0.46
-0.40
-0.42
10
0.02
-0.02
0.02
0.01
0.02
-0.00
0.03
0.35
1.00
0.54
0.24
0.44
0.28
23
-0.00
-0.02
-0.00
-0.01
-0.03
0.00
0.01
-0.00
1.00
0.45
-0.13
-0.24
-0.02
36
-0.28
-0.21
-0.24
-0.38
0.12
0.26
-0.31
-0.41
1.00
-0.01
0.29
0.21
0.33
11
0.03
-0.00
0.00
0.00
0.01
-0.02
0.01
0.23
0.54
1.00
0.29
0.39
0.24
24
0.01
0.02
-0.00
-0.04
-0.01
0.02
-0.03
0.02
0.45
1.00
0.03
-0.27
0.04
37
0.03
0.01
0.01
0.01
0.00
0.02
-0.04
0.01
-0.01
1.00
-0.03
-0.03
-0.00
12
0.01
0.02
0.05
0.00
0.01
-0.04
0.04
0.11
0.24
0.29
1.00
0.21
0.13
25
0.01
0.01
0.00
-0.01
-0.01
-0.03
0.01
0.02
-0.13
0.03
1.00
-0.08
0.21
38
-0.34
-0.19
-0.31
-0.39
0.12
0.27
-0.37
-0.46
0.29
-0.03
1.00
0.37
0.34
13
0.00
0.00
0.01
-0.02
0.01
-0.00
-0.03
0.19
0.44
0.39
0.21
1.00
0.32
26
0.02
0.03
-0.01
0.02
-0.02
-0.01
0.02
-0.01
-0.24
-0.27
-0.08
1.00
-0.23
39
-0.31
-0.19
-0.29
-0.37
0.12
0.25
-0.32
-0.40
0.21
-0.03
0.37
1.00
0.34
14
0.02
0.01
0.01
0.02
0.01
-0.02
-0.01
0.12
0.28
0.24
0.13
0.32
1.00
27
0.00
-0.01
-0.00
-0.02
-0.03
-0.02
0.03
0.03
-0.02
0.04
0.21
-0.23
1.00
40
-0.38
-0.28
-0.34
-0.42
0.10
0.28
-0.35
-0.42
0.33
-0.00
0.34
0.34
1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
layer
layer
layer
1.00
layer
layer
Layers 15 27
2
layer
1.00 0.75 0.50 0.25 0.00
cosine sim
Gemma 4 8B Valence direction alignment Gemma stories
Layers 2 14
0.25 0.50 0.75 1.00
Figure 11. G EMMA -4-E4B validation on Apertus stories Layers 28 40
0.01
0.03
0.05
0.06
-0.03
-0.00
0.00
0.00
0.02
0.01
-0.02
0.01
15
1.00
0.09
-0.08
0.08
-0.06
-0.03
-0.01
0.02
-0.01
-0.03
0.01
-0.00
0.00
28
1.00
0.04
-0.12
0.43
-0.33
0.08
0.24
0.20
-0.34
0.18
-0.14
-0.44
0.31
3
0.01
1.00
0.14
0.09
0.10
0.02
0.03
-0.03
-0.02
0.02
0.00
0.00
0.02
16
0.09
1.00
0.25
-0.07
0.02
-0.03
0.01
-0.00
0.01
-0.01
-0.02
0.01
0.01
29
0.04
1.00
-0.01
0.06
-0.05
-0.00
0.09
0.02
-0.06
-0.05
0.01
-0.04
0.02
4
0.03
0.14
1.00
0.36
0.35
0.01
-0.16
0.04
0.03
-0.03
-0.02
0.00
-0.04
17
-0.08
0.25
1.00
-0.22
0.01
0.04
-0.03
-0.06
0.03
-0.03
0.03
0.03
0.03
30
-0.12
-0.01
1.00
-0.20
0.15
0.00
-0.11
0.02
0.11
0.02
0.05
0.11
-0.06
5
0.05
0.09
0.36
1.00
0.41
-0.03
-0.08
-0.03
0.00
0.01
-0.01
0.01
-0.01
18
0.08
-0.07
-0.22
1.00
-0.08
-0.03
0.01
0.01
-0.01
0.00
-0.02
-0.02
0.02
31
0.43
0.06
-0.20
1.00
-0.58
0.15
0.36
0.23
-0.48
0.21
-0.19
-0.52
0.44
6
0.06
0.10
0.35
0.41
1.00
0.02
-0.10
-0.04
-0.01
-0.02
-0.03
0.00
0.01
19
-0.06
0.02
0.01
-0.08
1.00
0.07
-0.02
-0.03
-0.01
0.01
-0.01
-0.01
0.04
32
-0.33
-0.05
0.15
-0.58
1.00
-0.15
-0.30
-0.18
0.40
-0.17
0.12
0.42
-0.35
7
-0.03
0.02
0.01
-0.03
0.02
1.00
0.04
-0.01
0.02
0.04
-0.00
-0.03
-0.02
20
-0.03
-0.03
0.04
-0.03
0.07
1.00
0.01
0.03
0.05
-0.01
0.01
0.02
0.00
33
0.08
-0.00
0.00
0.15
-0.15
1.00
0.07
0.10
-0.07
0.08
-0.06
-0.08
0.07
8
-0.00
0.03
-0.16
-0.08
-0.10
0.04
1.00
-0.59
-0.36
0.33
0.21
0.36
0.44
21
-0.01
0.01
-0.03
0.01
-0.02
0.01
1.00
0.01
0.00
-0.01
-0.02
-0.02
-0.02
34
0.24
0.09
-0.11
0.36
-0.30
0.07
1.00
0.15
-0.30
0.14
-0.12
-0.34
0.26
layer
1.00
layer
layer
Layers 15 27
2
9
0.00
-0.03
0.04
-0.03
-0.04
-0.01
-0.59
1.00
0.44
-0.33
-0.21
-0.33
-0.40
22
0.02
-0.00
-0.06
0.01
-0.03
0.03
0.01
1.00
0.02
0.02
0.00
-0.01
-0.02
35
0.20
0.02
0.02
0.23
-0.18
0.10
0.15
1.00
-0.25
0.20
-0.02
-0.26
0.21
10
0.00
-0.02
0.03
0.00
-0.01
0.02
-0.36
0.44
1.00
-0.25
-0.12
-0.28
-0.32
23
-0.01
0.01
0.03
-0.01
-0.01
0.05
0.00
0.02
1.00
0.07
-0.15
-0.01
0.03
36
-0.34
-0.06
0.11
-0.48
0.40
-0.07
-0.30
-0.25
1.00
-0.20
0.16
0.53
-0.43
11
0.02
0.02
-0.03
0.01
-0.02
0.04
0.33
-0.33
-0.25
1.00
0.14
0.27
0.34
24
-0.03
-0.01
-0.03
0.00
0.01
-0.01
-0.01
0.02
0.07
1.00
0.08
-0.17
-0.00
37
0.18
-0.05
0.02
0.21
-0.17
0.08
0.14
0.20
-0.20
1.00
-0.10
-0.20
0.17
12
0.01
0.00
-0.02
-0.01
-0.03
-0.00
0.21
-0.21
-0.12
0.14
1.00
0.23
0.22
25
0.01
-0.02
0.03
-0.02
-0.01
0.01
-0.02
0.00
-0.15
0.08
1.00
-0.03
0.06
38
-0.14
0.01
0.05
-0.19
0.12
-0.06
-0.12
-0.02
0.16
-0.10
1.00
0.21
-0.19
13
-0.02
0.00
0.00
0.01
0.00
-0.03
0.36
-0.33
-0.28
0.27
0.23
1.00
0.44
26
-0.00
0.01
0.03
-0.02
-0.01
0.02
-0.02
-0.01
-0.01
-0.17
-0.03
1.00
0.23
39
-0.44
-0.04
0.11
-0.52
0.42
-0.08
-0.34
-0.26
0.53
-0.20
0.21
1.00
-0.54
14
0.01
0.02
-0.04
-0.01
0.01
-0.02
0.44
-0.40
-0.32
0.34
0.22
0.44
1.00
27
0.00
0.01
0.03
0.02
0.04
0.00
-0.02
-0.02
0.03
-0.00
0.06
0.23
1.00
40
0.31
0.02
-0.06
0.44
-0.35
0.07
0.26
0.21
-0.43
0.17
-0.19
-0.54
1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
layer
layer
14
layer
1.00 0.75 0.50 0.25 0.00
cosine sim
Gemma 4 8B Valence direction alignment Apertus stories
Layers 2 14
0.25 0.50 0.75 1.00
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs Figure 12. G EMMA -4-E4B validation on Apertus stories Gemma 4 8B
Valence direction: adjacent-layer stability Apertus stories
1.00 0.75
cosine similarity
0.50
0.44
0.41
0.36
0.25
0.14
0.00
0.44 0.23
0.14
0.04
0.02
0.01
0.25
0.22
0.23
0.09
0.07
0.01
-0.08
0.50
0.21 0.04
-0.03
0.21
0.15
0.07
-0.01 -0.15
-0.20
-0.22
-0.25
0.25
0.08
0.07
0.02
0.01
-0.10
-0.20
-0.25
-0.54
-0.58
-0.59
0.75 1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
layer (midpoint of pair)
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
Figure 13. G EMMA -4-E4B on Gemma stories Gemma 4 8B
Valence direction: adjacent-layer stability Gemma stories
1.00 0.75
cosine similarity
0.54
0.54
0.50
0.35
0.24
0.25
0.09
0.00 -0.35
0.20
0.55
0.45
0.42
0.32
0.21
-0.03
-0.07
-0.09
-0.10
0.37 -0.01
-0.08 -0.23
-0.25
-0.03
-0.25 -0.41
-0.44
-0.46
0.50
0.34
0.16
0.07
0.03
-0.00
0.52
0.35
0.34
0.14
0.04
-0.19
0.25
0.29
0.75 1.00
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
layer (midpoint of pair)
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
C.5. PCA comparison The PCA figure shows a map of the model’s emotional space at a specific layer. We pick the layer with the highest valence. Each dot is an emotion, positioned at how the model actually represents this emotion in its activation space. Comparing two panels can tell whether the map is reproducible across different inputs, or whether the emotional space is sensitive to what stories the model reads. C.5.1. A PERTUS -8B R ESULTS
Figure 14. A PERTUS -8B validation
Apertus 8B L23 Apertus stories
Apertus 8B L23 Gemma stories
1.00
600
0.75 400
200 angry afraid
0
serene
ecstatic
200 depressed gloomy
400
calm
excited joyful happy
0.50 ecstatic excited
furious
200
joyful
0.25
angry happy
0
0.25 gloomy
200
calm
afraid
serene
0.75
600 400
200
0
200
PC1 (13.5% var) r_val=0.723
400
0.50
depressed
400 600
0.00
valence
furious
400
PC2 (8.7% var) r_aro=0.422
PC2 (12.4% var) r_aro=0.146
600
600
400
15
200
0
200
PC1 (15.6% var) r_val=0.748
400
600
1.00
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
C.5.2. G EMMA -4-E4B R ESULTS Figure 15. G EMMA -4-E4B: valence-arousal PCA
Gemma 4 8B L13 Gemma stories
2
1.00
4 afraid
2
0
afraid angry furious
serene ecstatic excited
depressed gloomy
happy
calm
2
PC2 (8.9% var) r_aro=0.193
PC2 (12.2% var) r_aro=0.046
1
happy
joyful excited
furious
0.75 0.50
gloomy ecstatic
0
0.25
angry
0.00
1
depressed
valence
Gemma 4 8B L13 Apertus stories
0.25 serene
2
0.50
joyful
0.75
calm
4
3 2
1
0
1
2
PC1 (13.8% var) r_val=0.794
3
4
3
16
2
1
0
1
PC1 (15.0% var) r_val=0.798
2
3
1.00
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
D. Prompts Below, we report verbatim the prompts used to generate the short stories. System prompt — Explanation generation Write {n_stories} different stories based on the following premise. Topic: {topic} The story should follow a character who is feeling {emotion}. Format the stories like so: [story 1] [story 2] [story 3] etc. The paragraphs should each be a fresh start, with no continuity. Try to make them diverse and not use the same turns of phrase. Across the different stories, use a mix of third-person narration and first-person narration. IMPORTANT: You must NEVER use the word ’{emotion}’ or any direct synonyms of it in the stories. Instead, convey the emotion ONLY through: - The character’s actions and behaviors - Physical sensations and body language - Dialogue and tone of voice - Thoughts and internal reactions - Situational context and environmental descriptions The emotion should be clearly conveyed to the reader through these indirect means, but never explicitly named.
17