ConceptioArchivearXiv CS
arXiv CSopen access

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

From Representations to Behaviors: Exploring the Person–Situation–Behavior Triad in LLMs Ruikang Zhang1 , Shuo Wang2∗ , Qi Su1† 1

Peking University, Beijing, China Beijing Institute of Technology, Beijing, China [email protected], [email protected], [email protected] 2

arXiv:2607.26853v1 [cs.CL] 29 Jul 2026

Abstract Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable traitrelated expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder’s personality triad framework, we adapt its three components for LLM analysis: Personas personality-related internal representations, Situationas contexts that afford trait-relevant responses, and Behavioras response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through sparse autoencoder decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.

1

Introduction

Personality is inferred from the expression of relatively stable individual tendencies across situations (Allport 1937). Human personality theory therefore treats persons, situations, and behaviors as jointly shaping action (Funder 2006). This perspective is increasingly relevant to large language models (LLMs), where personality induction (Mao et al. 2024) supports applications ranging from safety alignment, toxicity and bias analysis (Zhang et al. 2024; Wang et al. 2025a), and personalization to demographic and social simulation (Cui, Li, and Zhou 2025). ∗ †

Work done during internship at Peking University. Corresponding author.

LLM personality can be induced through persona prompting (Chen et al. 2024), training and weight-space editing (Brito et al. 2025; Ye et al. 2026), and interventions on internal activations (Chen et al. 2025; Deng et al. 2025; Zhu et al. 2025). Evaluation has developed in parallel, progressing from direct Big Five and MBTI inventories (Serapio-García et al. 2025; Cui et al. 2024; Song et al. 2023) to personality-profile emulation (Wang et al. 2025b), scenario-grounded choices (Lee et al. 2025), and open-ended linguistic assessment (Zheng et al. 2025). These approaches measure personality from closedended, single-option responses, option-token probabilities, or trait estimates derived from generated text. Mechanistic studies further identify personality-related directions, neurons, and attention heads, extending this broader study of model behavior of machine psychology (Hagendorff et al. 2024). However, these lines of evidence are rarely connected. On the representation side, a direction retrieved from contrastive prompts may primarily reflect lexicosyntactic regularities of the probing data, obscuring a more abstract personalityrelated signal. On the intervention side, persona prompts provide discrete external conditioning, whereas dense meandifference vectors can strongly perturb generation; both complicate the relation between a manipulated internal state and coherent, instruction-following situated responses. On the evaluation side, inventory scores and option probabilities quantify whether responses align with the target trait, but coherent, instruction-following situational responses and broader social behavior require separate evidence, as persona conditioning can shift self-reported tendencies while leaving situational behavior weakly aligned (Han et al. 2025). The unresolved question is therefore whether an internal representation associated with a personality trait can be traced from personality-related activation, through effective trait expression across situations, to systematic behavioral consequences. We organize this problem through the Person–Situation– Behavior components of the personality triad framework. For LLMs, Person denotes personality-related internal representations, Situation denotes concrete contexts that elicit personality-dependent responses, and Behavior denotes observable behavioral manifestations in broader social tasks beyond direct personality assessment. This yields three research questions: (RQ1–Person) Can personality-related internal representations be identified from contrasting behavioral expressions under matched situations, beyond lexicosyntactic

RQ-1

Person (Identifying Personality-Related Internal Representations) Feature Discovering

Dataset Construction  Oppositional Semantic Pairs on same Situation

Openness: Fantasy, Aesthetics, Feelings, Actions, Ideas, Values ...

RQ-2

SAE Encoding

Generation Probing

After a break-up, someone may need reassurance that they will find love again.

*

I offer warm, empathetic support, comforting them with kindness and hope. I focus on the practical realities, preferring logic over emotional comfort.

High-Difference Features

Consistent Trait Change?

High-Level

-1 -0.5 0

Behavior

+0.5 +1

Evaluation Evaluate on SocialEval Benchmark

Trait-related Features

Measure Personality on Diverse Scenario-based Advice Tasks

Process-Oriented Interpersonal Ability Evaluation (IAE)

Validation: ➢ Token-level Activation ➢ Paraphrase Robustness

Personality Traits

Situation 1: A friend experiences a breakup Situation 2: Collaborating with teammates on a project

High-level Behaviors

Situation 3: Facing pressure and conflicts at work Situation N: Other realistic social situations

Intervention on the Target Feature Low Behavior

-1

RQ-3

Evaluate on TRAIT Benchmark

0

Extraversion Teamwork ...

+1 High

Behavior

Detail Management ...

For all traits, interventions produce bidirectional across situations changes as expected, while response validity remains stable.

Situation (Personality Expression Across Situations)

Trait-specific Benefits and Trade-offs Agreeableness Ethical Competence ... Decision Making ...

...

Behavior (Psychologically Consistent Behavioral Changes)

Figure 1: Overview of the study. Matched high–low behavioral contrasts under shared situations retrieve candidate SAE features, which are selected through generation-based intervention probing and characterized through token-level activation and paraphrase robustness (RQ1–Person). Interventions on the selected features evaluate bidirectional trait-relevant expression and response validity across a separate, diverse set of TRAIT situations (RQ2–Situation). The same interventions reveal broader changes, including trait-specific benefits and trade-offs, across SocialEval interpersonal abilities (RQ3-Behavior). cues? (RQ2–Situation) Does intervention on the identified representations elicit trait-relevant tendencies in both directions across a separate, diverse set of situations while preserving coherent, instruction-following responses? (RQ3–Behavior) Does the intervention induce consistent changes across social behaviors that correspond to established findings in human personality research? We study these questions using the Big Five (Goldberg 1990) and sparse autoencoders (SAEs) (Shu et al. 2025). First, we hold a situation fixed while constructing opposing traitrelevant reactions, and use their SAE activation differences for feature retrieval. Intervention probing then selects candidates with coherent, trait-aligned causal effects, while token-level activation analysis and paraphrase tests further validate the selected features. Second, we intervene on the selected features in a scenario-grounded personality benchmark, TRAIT (Lee et al. 2025), evaluating trait direction and response validity across a separate, diverse set of situations. Third, we apply the same interventions to a social-intelligence benchmark, SocialEval (Zhou et al. 2025), to examine broader behavioral consequences. Our results connect the three components. We identify controllable, robust features for all five traits (Person); feature steering produces bidirectional tendencies across situational tasks while maintaining valid responses (Situation) and yields

trait-characteristic cost-benefit tradeoffs across social abilities (Behavior). Our contributions are: • We adapt the personality triad framework to connect Person-side internal representations, Situation-dependent responses, and broader Behavior outcomes in LLMs. • We develop an SAE-based pipeline that performs feature retrieval from behavioral contrasts under matched situations, selects and causally validates candidates through intervention probing and representation analysis, and traces the selected features across personality expression and social behavior. • We provide evidence that LLMs contain controllable traitlike representations linking internal states, situational expression, and broader behavioral outcomes.

2

The Person–Situation–Behavior Framework

2.1

The Personality Triad Framework in Human Personality Theory

Classical trait theory treats traits as dispositions expressed through patterned responses across situations. Allport (1937) argues that different situations can activate a common traitrelated tendency while eliciting responses that differ in words, emotional tones, decisions, or actions. This shared trait relevance across situations is referred to as functional equiva-

lence. Agreeableness, for example, may be expressed through emotional reassurance in one situation and cooperative compromise in another. These responses differ in surface form while sharing a prosocial function associated with the same trait. Funder (2006) extends the person–situation interaction perspective through the personality triad framework, which treats persons, situations, and behaviors as three mutually informative components. Each component is understood through its relations with the other two. A person is characterized through recurring patterns of response across situations relevant to the trait. A situation is characterized by the contextual cues and response opportunities it provides, as well as the response tendencies it elicits. A behavior acquires psychological meaning within the person–situation combination in which it occurs. The same observable action can express different functions across contexts, while different actions can instantiate a related trait tendency.

2.2 Adapting the Personality Triad Framework for LLM Analysis Guided by this framework, we define three corresponding components for analyzing trait-related representations and behavior in LLMs. Person refers to an internal representation that is activated by trait-relevant expressions and whose intervention induces trait-aligned changes while preserving coherent, instruction-following responses. Situation refers to a concrete scenario in which the model expresses a traitrelevant tendency through a coherent, instruction-following response. Behavior refers to observable behavioral manifestations, reflected in performance changes across broader social tasks beyond direct personality assessment. Our experiments examine all these components. In RQ1, we retrieve candidate Personrepresentations by contrasting high- and low-trait responses within matched situations and aggregating these contrasts across diverse situations. In RQ2, we intervene on the selected representations and examine whether they elicit trait-relevant tendencies in both directions across a separate, diverse set of situations while preserving coherent, instruction-following responses. In RQ3, we apply the same interventions to broader social tasks and measure the resulting changes in Behavior. Figure 1 summarizes our design.

3

Related Work

3.1 Personality Induction and Evaluation in LLMs Personality induction aims to endow an LLM with a target profile so that its responses change accordingly across interactions. Existing methods span persona prompting, weight-space editing, training-based alignment, and mechanistic intervention on internal representations (Brito et al. 2025; Ye et al. 2026). Prompting approaches place a persona description or trait instruction in the input (Chen et al. 2024; Serapio-García et al. 2025; Pan and Zeng 2023). Weight-space approaches alter or compose parameters, including model-merging personality vectors (Sun, Baek, and Kim 2025) and trait-specific adapters (Vu et al. 2026), while training-based approaches

use supervised or preference learning over trait-annotated data (Li et al. 2025). Evaluation methods can be organized by their form of elicitation. Inventory-based assessment asks models to answer human personality scales, either directly or under targetprofile conditioning. Existing studies examine the applicability of self-assessment questionnaires (Song et al. 2023), construct a psychometric framework around Big Five inventories (Serapio-García et al. 2025), use Big Five questionnaires to evaluate personality emulation (Wang et al. 2025b), or use a modified MBTI questionnaire to evaluate personality-trained models (Cui et al. 2024). Situated-choice assessment transforms psychometric constructs into multiple-choice decisions grounded in real-world scenarios, as in TRAIT (Lee et al. 2025). Generation-based assessment elicits open-ended answers and maps their linguistic expression to numerical trait estimates, as in LMLPA (Zheng et al. 2025). These elicitation forms use different scoring interfaces: direct inventories record a closed-ended, single-option response for each item, TRAIT compares option-token probabilities, and LMLPA applies an automated rater to generated text. Some work reports stable or distinct self-report personality profiles (Huang et al. 2024; Heston and Gillette 2025), while other studies identify important measurement problems: self-assessment tests can be unreliable measures of LLM personality (Gupta, Song, and Anumanchipalli 2024), response statistics can deviate from human patterns (Sühr et al. 2025), and personality estimates can vary across scales (Tosato et al. 2024). More directly, persona injection may shift self-reported scores while leaving situational behavior only weakly aligned (Han et al. 2025). These findings motivate us to connect measured tendencies with meaningful responses and their behavioral consequences. Our study therefore traces an identified internal representation through situational expression and into social behavior.

3.2

Mechanistic Interpretability of Personality-related Representations in LLMs

Mechanistic interpretability offers a route from correlation to control for LLM personality by localizing behaviorally relevant representations inside a model (Ranaldi 2025). This approach is often motivated by the linear representation hypothesis that semantic and behavioral attributes occupy linearly separable subspaces (Park, Choe, and Veitch 2024). Contrastive Activation Addition (CAA) retrieves a direction as the mean activation difference over contrastive stimuli and adds it during inference (Rimsky et al. 2024). Persona Vectors similarly identify activation-space directions for monitoring and controlling character traits (Chen et al. 2025). Neuronand head-level methods localize trait-related units at finer granularity, including NPTI (Deng et al. 2025) and PAS (Zhu et al. 2025). Two issues are especially relevant when these tools are used within the personality triad framework. First, contrastive retrieval can retain lexicosyntactic patterns from its probing data, limiting representational abstraction. Second, a dense mean-difference vector can be sensitive to probing noise and may impair generation at large intervention strengths (Rimsky et al. 2024), thereby removing the model’s ability to

respond effectively to a situation. Sparse autoencoders (SAEs) offer a complementary route by decomposing dense hidden states into sparse, more monosemantic features (Bricken et al. 2023; Templeton et al. 2024; Lieberum et al. 2024; He et al. 2024). Such features have been used to localize linguistic phenomena (Jing et al. 2025) and study behaviors such as repetition (Yao et al. 2025). We use SAE features as candidate Person-side representations, test their generalization across lexicosyntactic forms, and use intervention to trace their effects across Situation and Behavior.

4 4.1

Experiments

Experimental Setup

Research scope and model. We use the Big Five as an established dimensional trait framework because its five broad traits (Agreeableness, Conscientiousness, Extraversion, Neuroticism, and Openness) have well-characterized high- and low-trait expressions and extensive psychological evidence concerning their behavioral correlates (Goldberg 1990; Costa and McCrae 2008). To construct fine-grained behavioral contrasts, we adopt the NEO-PI-R taxonomy of six facets within each trait (Costa and McCrae 2008). We study DeepSeek-R1-Distill-Llama-8B (DeepSeek-AI 2025) with the Llama-Scope-R1-Distill SAE (He et al. 2024). This combination offers three practical advantages. First, the SAE covers residual-stream locations throughout the model, enabling a global search for personality-related features across layers. Second, the instruction-tuned model has sufficiently rich open-ended expression for intervention probing and situational advice tasks. Third, per-feature maximum activations are publicly available through Neuronpedia (Lin 2023), permitting direct and reproducible scaling of decoder directions. Overall experimental design. We organize the experiments around the three research questions. For Person, contrastive behavior pairs grounded in shared situations are used for feature retrieval of sparse SAE features associated with opposing trait expressions. Intervention probing selects candidates with coherent, trait-aligned causal effects, while token-level activation and paraphrase analyses test whether the selected features generalize across lexicosyntactic forms (RQ1). For Situation, interventions on the selected features are evaluated on TRAIT (Lee et al. 2025) for bidirectional trait expression and response validity across a separate, diverse set of situations (RQ2). For Behavior, the interventions are applied to SocialEval (Zhou et al. 2025) to measure broader changes across interpersonal abilities (RQ3).

4.2

RQ1–Person: Identifying Personality-Related Internal Representations

RQ1 asks whether personality-related internal representations can be identified from contrasting behavioral expressions under matched situations, beyond lexicosyntactic cues. We define a feature as one dimension of the SAE latent representation, or equivalently the hidden-state direction obtained by decoding that dimension. The pretrained SAE provides a sparse decomposition of model activations, allowing us to retrieve candidate features whose activation

patterns distinguish the two response poles using a controlled dataset. During intervention, each decoded direction is added at the residual-stream position of its corresponding SAE, ensuring consistency between feature retrieval and causal manipulation. Dataset construction, intervention probing and validation protocols, and supplementary representation analyses are provided in Appx. A. Discovering Personality-Related Features Contrastive behavior-pair construction. Let C = {s1 , . . . , sK } be a situational corpus and T the target traits. We instantiate C with Q-Sort situational corpus (Neuman and Cohen 2023) extended from the Riverside Situational Q-Sort (Funder 2016), and use NEO-PI-R traits and facets to define opposing behavioral tendencies (Costa and McCrae 2008). For each t ∈ T , LLM filtering and expert audit define a validated subset Ct = Φt (C), where each retained situation is mapped to one facet. For each s ∈ Ct , an LLM-based expansion function generates concise high- and low-facet reactions that share the same situation clause: − Ks Xs = Ψ(s, t; Pt ) = {(x+ k , xk )}k=1

(1)

The per-trait Feature Retrieval Dataset and Intervention Probing Dataset are ! (t) Dret = UnifSamplefacet

[

Xs

, (2)

s∈Ct (t) Dprobe = {(s, Q(s)) | s ∈ UnifSamplefacet (Ct )}.

where Ks is the number of contrastive pairs generated for s, and Q(·) instantiates an open-ended probing question. Pairs (t) in Dret are evenly sampled across the six facets of each trait, yielding 500 pairs per trait. Contrasting high- and lowtrait behaviors within each of diverse situations reflects the conception of a trait as a stable tendency toward functionally equivalent responses across situations (Allport 1937). Dataset construction details and examples are provided in Appx. A.1 and Appx. A.2. SAE encoding. For token k, the model hidden state hk is mapped to SAE activations fk . To obtain a sequence-level representation while preserving features that respond strongly at specific token positions, we max-pool across the sequence: F = max_pool(f1 , f2 , . . . , fT )

(3)

Under SAE sparsity, a target feature should activate frequently on one pole and remain suppressed on the other. Because each positive–negative pair shares trait-unrelated lexiosyntactic forms and semantics, unrelated features should have similar activation frequencies in the two sets. For feature i, we first calculate its activation count and rate on each pole: Nipol =

500 X j=1

pol I(Fi,j > 0),

Pipol =

Nipol , 500

pol ∈ {pos, neg}. (4)

We retain features with a sufficiently large frequency difference and a nontrivial activation rate on at least one pole: S = { i | |Nipos − Nineg | ≥ τ1 ∧ max(Pipos , Pineg ) ≥ τ2 } (5) where S is the retained candidate set. We set τ1 = 80 and τ2 = 0.2, fixed to retain features that distinguish the two poles across

multiple facets. This stage identifies activation-associated candidates, which subsequently undergo intervention probing because similar input-side statistics can yield substantially different steering effects (Appx. A.5).

Trait

(L, Idx)

High/Low/∆f (Original)

High/Low/∆f (Paraphrased)

Agreeableness Conscientiousness Extraversion Neuroticism Openness

(9, 525) (7, 8233) (13, 27392) (12, 22254) (6, 4344)

202 / 45 / 157 347 / 31 / 316 429 / 110 / 319 322 / 48 / 274 234 / 102 / 132

224 / 64 / 160 349 / 20 / 329 446 / 148 / 298 252 / 101 / 151 208 / 113 / 95

Table 1: Activation statistics for the selected features on the original and paraphrased matched-situation pairs. Feature intervention. For candidate feature i at layer li , we reconstruct its decoder direction and define (i)

vsteer = α · ϕi · Wdec

(6)

h′l = hl + vsteer

(7)

where α is the steering coefficient, ϕi is the feature’s maximum (i) activation during SAE training, and Wdec ∈ Rd is the i-th column of the SAE decoder matrix, with d denoting the hidden dimension. Following prior SAE steering practice (Templeton et al. 2024), scaling by ϕi keeps the intervention commensurate with activation magnitudes observed during training. Generation-based intervention probing. For each candidate i ∈ S , we sweep A = {0, ±0.25, ±0.5, ±1} and collect (t) an ordered response family {yq(i,α) }α∈A for every q ∈ Dprobe . Given the trait definition, descriptions of its high and low behaviors, and the ordered responses, the hybrid judge J combines an initial LLM assessment with expert audit and returns  (i,α) c(i) }α∈A ∈ {0, 1}, q = J t, q, {yq

(8)

(i) cq

where = 1 when the responses are grammatical and coherent and show a clear polarity change consistent with the trait’s high–low behaviors as α varies. Because feature orientation is arbitrary, the judge accepts either monotonic direction and rejects uncertain or invalid cases. A candidate feature is selected if it receives at least one positive label P across its associated questions, i.e., q∈D(t) c(i) q > 0. The probe

prompts and audit procedure are given in Appx. A.3; an judge-reliability experiment is provided in Appx. A.4.

Openness: Designing and developing new materials for the construction of sustainable infrastructure, I eagerly explore innovative theories and unconventional solutions to push the boundaries of what ’s possible . Conscientiousness: My group member is counting on me to prepare my part of the project, I prioritize completing my work thoroughly and on time to uphold my obligations . Extraversion: Volunteering at a food bank, I naturally greet everyone with a smile and quickly strike up friendly conversations with both staff and recipients . Agreeableness: A person may need reassurance that their pet will be wellbehaved around visitors, I respond with warmth and empathy , offering comfort and understanding about their concerns. Neuroticism: Being forced to work with someone who is hostile or unpleasant, I feel frustration mount and retaliate instantly , struggling to suppress my irritation.

Figure 2: Token-level activations of the selected SAE features. Darker highlights indicate higher activation values.

while preserving each situation and high–low facet contrast. The selected features are then evaluated with the same thresholds. Table 1 shows that their activation-frequency differences remain above τ1 . For example, Conscientiousness has a difference of 329 and a maximum activation frequency of 349 out of 500 pairs after paraphrasing, compared with a difference of 316 on the original set. These results indicate that the selected features capture trait-related meaning across substantially different lexicosyntactic forms. Prompts and examples are provided in Appx. A.7. Together, feature discovery identifies personality-related features from contrasting behaviors within shared situations, while feature characterization further validates their trait relevance and semantic grounding.

Characterizing Personality-Related Features Token-level activation. We inspect where the selected features activate in the Feature Retrieval Dataset. Figure 2 shows two recurring patterns. Activations may be distributed over trait-relevant words or phrases, consistent with the semantic content of the trait; they may also peak at sequence boundaries, indicating sensitivity to the trait-related meaning of the preceding clause as a whole. For example, Agreeableness is associated with prosocial expressions such as empathy and compassion, whereas the selected Conscientiousness feature often peaks near sequence boundaries. Quantitative token statistics are present in Appx. A.6. Robustness to paraphrase. A personality-related representation should capture trait meaning rather than specific lexicosyntactic forms. We therefore rewrite the Feature Retrieval Dataset, substantially changing vocabulary and syntax

4.3

RQ2–Situation: Personality Expression across Situations

RQ2 asks whether intervention on the identified representations elicits bidirectional trait-relevant tendencies across a separate, diverse set of situations while preserving coherent, instruction-following responses. Situational-Response Evaluation We use the psychometrically validated Big Five subset of TRAIT (Lee et al. 2025). Each item presents a concrete, realistic user situation and asks the model for advice, with available choices reflecting different trait tendencies. The official next-token probability protocol measures statistical preference over these choices but cannot reveal how the intervened model actually responds to the situation. We therefore extend it with open generation, requiring the model to produce reasoning and explicitly se-

Trait

Method

(L, Idx)

Polarity

Trait Score

Valid Rate

Agreeableness

Baseline CAA P2 Ours

(9, -) (9, 525)

± ± ±

0.7052 0.8070 / 0.4220 0.7583 / 0.6095 0.7845 / 0.6278

0.960 0.993 / 0.987 0.989 / 0.968 0.942 / 0.994

Conscientiousness

Baseline CAA P2 Ours

(7, -) (7, 8233)

± ± ±

0.8695 0.7150 / 0.7630 0.8609 / 0.8217 0.9043 / 0.8294

0.958 0.930 / 0.800 0.985 / 0.976 0.961 / 0.985

Extraversion

Baseline CAA P2 Ours

(13, -) (13, 27392)

± ± ±

0.4463 0.2950 / 0.1430 0.5724 / 0.2146 0.6609 / 0.3971

0.977 0.924 / 0.862 0.987 / 0.983 0.985 / 0.972

Neuroticism

Baseline CAA P2 Ours

(12, -) (12, 22254)

± ± ±

0.2117 0.9290 / 0.1270 0.2208 / 0.1261 0.4412 / 0.1017

0.959 0.141 / 0.283 0.969 / 0.983 0.961 / 0.944

Openness

Baseline CAA P2 Ours

(6, -) (6, 4344)

± ± ±

0.5214 0.6140 / 0.3520 0.6161 / 0.4030 0.5436 / 0.5164

0.959 0.938 / 0.971 0.969 / 0.990 0.951 / 0.947

Table 2: TRAIT situational-response results. For intervention methods, statistics are reported for both polarity. lect an option. An automated extractor identifies the choice and marks incoherent, instruction-violating, or unextractable responses as invalid. We report trait score, the proportion of high-trait selections among valid responses, and valid rate, the proportion of responses not marked invalid. Trait score summarizes directional expression over the benchmark’s collection of situations; valid rate establishes whether the model continues to produce coherent, instruction-following responses under intervention. The two metrics jointly characterize successful situated expression. A trait-score shift alone does not establish that the model continues to engage with the situation, because severe intervention may leave only a small subset of responses coherent and extractable. Conversely, a high valid rate without directional change indicates preserved generation but ineffective regulation. Successful intervention therefore requires both a consistent trait-directional shift and continued production of coherent, instruction-following responses. Comparative Interventions The no-intervention model provides the original trait tendency. We additionally evaluate two controls from different intervention levels: P 2 (Jiang et al. 2023), which assigns the target personality through a persona prompt at the input level, and CAA (Rimsky et al. 2024), which injects a contrastive mean-difference direction into the residual stream. The selected SAE-feature intervention uses α = ±1; for CAA, we use α = ±2, following the setting demonstrated in the original study. The comparisons characterize each concrete intervention in terms of bidirectional regulation and preservation of coherent, instruction-following responses. Results: Expression across Situations

Bidirectional regulation. As shown in Table 2, the selected feature intervention produces the expected positive–baseline– negative ordering for all five traits. The clearest shifts occur for Extraversion (0.661/0.397 around a 0.446 baseline) and Neuroticism (0.441/0.102 around 0.212). Across traits, the same Person-side representation regulates trait-relevant choices over a separate, diverse set of situations. Preserved situated response. Valid rates remain close to the no-intervention model throughout, including 0.961/0.944 for Neuroticism compared with 0.959 at baseline and 0.961/0.985 for Conscientiousness compared with 0.958. The intervention changes personality-related choices while preserving the model’s ability to understand the scenario, produce coherent advice, and complete the required selection. Intervention effectiveness and response validity. The controls show that both bidirectional trait change and response validity are important for evaluating an intervention. For Conscientiousness, the nominal positive condition of P 2 moves the baseline from 0.870 to 0.861, and its negative condition reaches 0.822; for Neuroticism, its positive condition changes 0.212 to 0.221. This particular persona intervention therefore provides limited bidirectional regulation for these traits. CAA can produce large score changes, but on Neuroticism it reduces validity to 0.141/0.283, leaving few responses that reveal how the model addresses the situation. Its validity also falls to 0.800 in negative Conscientiousness and to 0.862 in negative Extraversion. These failures show that an internal intervention can disrupt the coherent, instruction-following situated response on which personality observation depends (examples in Appx. B.1). These results answer RQ2: the Person-side representations identified in RQ1 elicit trait-relevant tendencies in both

directions across a separate, diverse set of situations while preserving coherent, instruction-following responses.

4.4

RQ3–Behavior: Psychologically Consistent Behavioral Changes

RQ3 asks whether the same intervention induces consistent changes across social behaviors that correspond to findings in human personality research. Behavioral Evaluation with SocialEval TRAIT directly probes trait-relevant situational choices. To measure traitrelated behavior in broader social scenarios, we use the SocialEval Interpersonal Ability Evaluation (IAE) (Zhou et al. 2025), which evaluates interpersonal abilities such as teamwork and emotional regulation through heterogeneous social scripts. We apply the same selected features with α = ±1 and report accuracy for each interpersonal ability. Anger management Ethical competence Capacity for social warmth Creative skill Organizational skill Detail management Information processing skill Decision making skill Goal regulation Leadership skill

Agreeableness -1 +1 -0.2

-0.1

0

Teamwork skill Ethical competence Responsibility management Stress regulation Capacity for trust Capacity for optimism Self reflection skill Persuasive skill Anger management Information processing skill Confidence regulation

+0.1

+0.2

Conscientiousness -1 +1

-0.2

-0.1

0

+0.1

Expressive skill Perspective taking skill Artistic skill Abstract thinking skill Organizational skill Detail management

+0.2

Extraversion -1 +1

-0.2

-0.1

0

+0.1

Creative skill Capacity for social warmth Ethical competence Energy regulation Goal regulation Detail management Impulse regulation Rule following skill Decision making skill Responsibility management Conversational skill Persuasive skill

+0.2

1991; Habashi, Graziano, and Hoover 2016; Pletzer et al. 2019; Costa and McCrae 2008). Specifically, Agreeableness strengthens prosocial and conflict-regulation abilities relative to negative steering, including anger management and ethical competence, while reducing performance on several self-agency and execution-oriented tasks (Wilmot and Ones 2022). Conscientiousness produces gains in responsibility management, ethical competence, teamwork, and regulationrelated abilities, corresponding to its established associations with responsibility, self-regulation, and goal-directed behavior (Roberts et al. 2009; Jackson et al. 2010; Eisenberg et al. 2014). Extraversion improves social-expression and interaction-related performance, including expressive skill and perspective taking, with additional gains in artistic and organizational skills, while reducing performance on tasks requiring sustained focus or fine control, such as detail management (John, Naumann, and Soto 2008; DeYoung, Quilty, and Peterson 2007; Fishman, Ng, and Bellugi 2011). Conversely, Neuroticism weakens regulatory and executivecontrol abilities, including goal regulation and rule following, while improving performance on creative skill and capacity for social warmth (Watson and Clark 1984; Lahey 2009). Openness improves adaptability, expressive skill, and persuasive skill, and yields higher creative and self-reflective performance under positive than negative steering, while reducing performance on structured tasks such as responsibility management (DeYoung 2015; McCrae 1987). Overall, the changes in the model’s behavioral patterns correspond to established experimental findings in personality psychology. Rather than producing uniform improvement or degradation, each intervention yields a differentiated profile of benefits and costs across heterogeneous social tasks. These findings demonstrate psychologically consistent changes in broader social behaviors, answering RQ3. Detailed per-ability scores and trait-wise discussion are provided in Appx. C.

5

Neuroticism -1 +1

-0.2

-0.1

0

Creative skill Adaptability Self reflection skill Expressive skill Detail management Persuasive skill Anger management Responsibility management Rule following skill Information processing skill

+0.1

+0.2

Openness -1 +1 -0.1

0

+0.1

Offset from baseline (0)

Figure 3: Representative SocialEval changes under personality-related feature intervention. Positive and negative shifts form trait-characteristic benefit–cost patterns consistent with meta-analytic findings in personality psychology. Behavioral Results Figure 3 shows systematic, traitspecific benefit-cost patterns consistent with meta-analytic findings in personality psychology. (Barrick and Mount

Discussion and Conclusion

This work adapts the personality triad framework to study personality-related representations and behavior in LLMs. The three research questions establish a connected chain of evidence across internal activations, trait scores, and behavioral outcomes. In RQ1, we retrieve candidate SAE features from contrasting behaviors under matched situations and select those whose intervention causally changes trait-relevant generation; token-level and paraphrase analyses further establish their semantic grounding beyond lexicosyntactic forms. In RQ2, the same features regulate trait-related expression bidirectionally across a separate, diverse set of situations while preserving coherent, instruction-following responses. Analysis of alternative interventions further shows that observing the representation–expression relation requires both changing the intended tendency and preserving coherent, instruction-following responses. In RQ3, applying the same interventions to heterogeneous social tasks produces characteristic combinations of benefits and costs that correspond to findings in human personality research. Together, these results provide evidence that LLMs contain controllable trait-like representations that connect Person-side internal states, Situation-dependent expression, and broader Behavior outcomes.

References Allport, G. W. 1937. Personality: A Psychological Interpretation. New York, NY: Henry Holt and Company. Barrick, M. R.; and Mount, M. K. 1991. The big five personality dimensions and job performance: a meta-analysis. Personnel psychology. Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; Askell, A.; Lasenby, R.; Wu, Y.; Kravec, S.; Schiefer, N.; Maxwell, T.; Joseph, N.; HatfieldDodds, Z.; Tamkin, A.; Nguyen, K.; McLean, B.; Burke, J. E.; Hume, T.; Carter, S.; Henighan, T.; and Olah, C. 2023. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread. Brito, I. A.; Dollis, J. S.; Färber, F. B.; Ribeiro, P. S. F. B.; Sousa, R. T.; and Filho, A. R. G. 2025. Modeling, Evaluating, and Embodying Personality in LLMs: A Survey. In Findings of the Association for Computational Linguistics: EMNLP 2025. Chen, J.; Wang, X.; Xu, R.; Yuan, S.; Zhang, Y.; Shi, W.; Xie, J.; Li, S.; Yang, R.; Zhu, T.; Chen, A.; Li, N.; Chen, L.; Hu, C.; Wu, S.; Ren, S.; Fu, Z.; and Xiao, Y. 2024. From Persona to Personalization: A Survey on Role-Playing Language Agents. Transactions on Machine Learning Research. Chen, R.; Arditi, A.; Sleight, H.; Evans, O.; and Lindsey, J. 2025. Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv:2507.21509. Costa, P. T.; and McCrae, R. R. 2008. The revised neo personality inventory (neo-pi-r). The SAGE handbook of personality theory and assessment. Cui, J.; Lv, L.; Wen, J.; Wang, R.; Tang, J.; Tian, Y.; and Yuan, L. 2024. Machine Mindset: An MBTI Exploration of Large Language Models. arXiv:2312.12999. Cui, Z.; Li, N.; and Zhou, H. 2025. A large-scale replication of scenario-based experiments in psychology and management using large language models. Nature Computational Science. DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948. Deng, J.; Tang, T.; Yin, Y.; yang, W.; Zhao, X.; and Wen, J.-R. 2025. Neuron based Personality Trait Induction in Large Language Models. In International Conference on Learning Representations. DeYoung, C. G. 2015. Cybernetic big five theory. Journal of research in personality. DeYoung, C. G.; Quilty, L. C.; and Peterson, J. B. 2007. Between facets and domains: 10 aspects of the Big Five. Journal of personality and social psychology. Eisenberg, N.; Duckworth, A. L.; Spinrad, T. L.; and Valiente, C. 2014. Conscientiousness: Origins in childhood? Developmental psychology. Fishman, I.; Ng, R.; and Bellugi, U. 2011. Do extraverts process social stimuli differently from introverts? Cognitive neuroscience. Funder, D. C. 2006. Towards a resolution of the personality triad: Persons, situations, and behaviors. Journal of Research in Personality, 40(1): 21–34. Proceedings of the 2005 Meeting of the Association of Research in Personality. Funder, D. C. 2016. Taking Situations Seriously: The Situation Construal Model and the Riverside Situational Q-Sort. Current Directions in Psychological Science. Goldberg, L. R. 1990. An alternative "description of personality": the big-five factor structure. Journal of Personality and Social Psychology.

Gupta, A.; Song, X.; and Anumanchipalli, G. 2024. Self-Assessment Tests are Unreliable Measures of LLM Personality. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. Habashi, M. M.; Graziano, W. G.; and Hoover, A. E. 2016. Searching for the prosocial personality: A big five approach to linking personality and prosocial behavior. Personality and Social Psychology Bulletin. Hagendorff, T.; Dasgupta, I.; Binz, M.; Chan, S. C. Y.; Lampinen, A.; Wang, J. X.; Akata, Z.; and Schulz, E. 2024. Machine Psychology. arXiv:2303.13988. Han, P.; Kocielnik, R. D.; Song, P.; Debnath, R.; Mobbs, D.; Anandkumar, A.; and Alvarez, R. M. 2025. The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs. In NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning. He, Z.; Shu, W.; Ge, X.; Chen, L.; Wang, J.; Zhou, Y.; Liu, F.; Guo, Q.; Huang, X.; Wu, Z.; Jiang, Y.-G.; and Qiu, X. 2024. Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders. arXiv:2410.20526. Heston, T. F.; and Gillette, J. 2025. Large Language Models Demonstrate Distinct Personality Profiles. Cureus. Huang, J.-t.; Jiao, W.; Lam, M. H.; Li, E. J.; Wang, W.; and Lyu, M. 2024. On the Reliability of Psychological Scales on Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Jackson, J. J.; Wood, D.; Bogg, T.; Walton, K. E.; Harms, P. D.; and Roberts, B. W. 2010. What do conscientious people do? Development and validation of the Behavioral Indicators of Conscientiousness (BIC). Journal of research in personality. Jiang, G.; Xu, M.; Zhu, S.-C.; Han, W.; Zhang, C.; and Zhu, Y. 2023. Evaluating and Inducing Personality in Pre-trained Language Models. In Advances in Neural Information Processing Systems. Jing, Y.; Yao, Z.; Guo, H.; Ran, L.; Wang, X.; Hou, L.; and Li, J. 2025. LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. John, O. P.; Naumann, L. P.; and Soto, C. J. 2008. Paradigm shift to the integrative big five trait taxonomy. Handbook of personality: Theory and research. Lahey, B. B. 2009. Public health significance of neuroticism. American Psychologist. Lee, S.; Lim, S.; Han, S.; Oh, G.; Chae, H.; Chung, J.; Kim, M.; Kwak, B.-w.; Lee, Y.; Lee, D.; Yeo, J.; and Yu, Y. 2025. Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics. In Findings of the Association for Computational Linguistics: NAACL 2025. Li, W.; Liu, J.; Liu, A.; Zhou, X.; Diab, M. T.; and Sap, M. 2025. BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Lieberum, T.; Rajamanoharan, S.; Conmy, A.; Smith, L.; Sonnerat, N.; Varma, V.; Kramar, J.; Dragan, A.; Shah, R.; and Nanda, N. 2024. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. Lin, J. 2023. Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks. Software available from neuronpedia.org.

Mao, S.; Wang, X.; Wang, M.; Jiang, Y.; Xie, P.; Huang, F.; and Zhang, N. 2024. Editing Personality For Large Language Models. In Natural Language Processing and Chinese Computing: 13th National CCF Conference, NLPCC 2024, Hangzhou, China, November 1–3, 2024, Proceedings, Part II. McCrae, R. R. 1987. Creativity, divergent thinking, and openness to experience. Journal of personality and social psychology. Neuman, Y.; and Cohen, Y. 2023. A Dataset of 10,000 Situations for Research in Computational Social Sciences Psychology and the Humanities. Scientific Data. Pan, K.; and Zeng, Y. 2023. Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models. arXiv:2307.16180. Park, K.; Choe, Y. J.; and Veitch, V. 2024. The Linear Representation Hypothesis and the Geometry of Large Language Models. In Proceedings of the 41st International Conference on Machine Learning. Pletzer, J. L.; Bentvelzen, M.; Oostrom, J. K.; and De Vries, R. E. 2019. A meta-analysis of the relations between personality and workplace deviance: Big Five versus HEXACO. Journal of vocational behavior. Ranaldi, L. 2025. Survey on the Role of Mechanistic Interpretability in Generative AI. Big Data and Cognitive Computing. Rimsky, N.; Gabrieli, N.; Schulz, J.; Tong, M.; Hubinger, E.; and Turner, A. 2024. Steering Llama 2 via Contrastive Activation Addition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Roberts, B. W.; Jackson, J. J.; Fayard, J. V.; Edmonds, G.; and Meints, J. 2009. Conscientiousness. In Handbook of individual differences in social behavior, 369–381. The Guilford Press. Serapio-García, G.; Safdari, M.; Crepy, C.; Sun, L.; Fitz, S.; Romero, P.; Abdulhai, M.; Faust, A.; and Matarić, M. 2025. A psychometric framework for evaluating and shaping personality traits in large language models. Nature Machine Intelligence. Shu, D.; Wu, X.; Zhao, H.; Rai, D.; Yao, Z.; Liu, N.; and Du, M. 2025. A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2025. Song, X.; Gupta, A.; Mohebbizadeh, K.; Hu, S.; and Singh, A. 2023. Have Large Language Models Developed a Personality?: Applicability of Self-Assessment Tests in Measuring Personality in LLMs. arXiv:2305.14693. Sühr, T.; Dorner, F. E.; Samadi, S.; and Kelava, A. 2025. Challenging the Validity of Personality Tests for Large Language Models. In Proceedings of the 5th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. Sun, S.; Baek, S. Y.; and Kim, J. H. 2025. Personality Vector: Modulating Personality of Large Language Models by Model Merging. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Templeton, A.; Conerly, T.; Marcus, J.; Lindsey, J.; Bricken, T.; Chen, B.; Pearce, A.; Citro, C.; Ameisen, E.; Jones, A.; Cunningham, H.; Turner, N. L.; McDougall, C.; MacDiarmid, M.; Freeman, C. D.; Sumers, T. R.; Rees, E.; Batson, J.; Jermyn, A.; Carter, S.; Olah, C.; and Henighan, T. 2024. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread. Tosato, T.; Hegazy, M.; Lemay, D.; Abukalam, M.; Rish, I.; and Dumas, G. 2024. LLMs and Personalities: Inconsistencies Across Scales. In NeurIPS 2024 Workshop on Behavioral ML.

Vu, H.; Nguyen, H. A.; Ganesan, A. V.; Juhng, S.; Kjell, O. N. E.; Sedoc, J.; Kern, M. L.; Boyd, R. L.; Ungar, L.; Schwartz, H. A.; and Eichstaedt, J. C. 2026. PsychAdapter: adapting LLMs to reflect traits, personality, and mental health. npj Artificial Intelligence. Wang, S.; Li, R.; Chen, X.; Yuan, Y.; Yang, M.; and Wong, D. F. 2025a. Exploring the Impact of Personality Traits on LLM Toxicity and Bias. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Wang, Y.; Zhao, J.; Ones, D. S.; He, L.; and Xu, X. 2025b. Evaluating the ability of large language models to emulate personality. Scientific reports. Watson, D.; and Clark, L. A. 1984. Negative affectivity: the disposition to experience aversive emotional states. Psychological bulletin. Wilmot, M. P.; and Ones, D. S. 2022. Agreeableness and its consequences: A quantitative review of meta-analytic findings. Personality and social psychology review. Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; Zheng, C.; Liu, D.; Zhou, F.; Huang, F.; Hu, F.; Ge, H.; Wei, H.; Lin, H.; Tang, J.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Zhou, J.; Lin, J.; Dang, K.; Bao, K.; Yang, K.; Yu, L.; Deng, L.; Li, M.; Xue, M.; Li, M.; Zhang, P.; Wang, P.; Zhu, Q.; Men, R.; Gao, R.; Liu, S.; Luo, S.; Li, T.; Tang, T.; Yin, W.; Ren, X.; Wang, X.; Zhang, X.; Ren, X.; Fan, Y.; Su, Y.; Zhang, Y.; Zhang, Y.; Wan, Y.; Liu, Y.; Wang, Z.; Cui, Z.; Zhang, Z.; Zhou, Z.; and Qiu, Z. 2025. Qwen3 Technical Report. arXiv:2505.09388. Yao, J.; Yang, S.; Xu, J.; Hu, L.; Li, M.; and Wang, D. 2025. Understanding the Repeat Curse in Large Language Models from a Feature Perspective. In Findings of the Association for Computational Linguistics: ACL 2025. Ye, H.; Jin, J.; Xie, Y.; Zhang, X.; and Song, G. 2026. Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement. arXiv:2505.08245. Zhang, J.; Liu, D.; Qian, C.; Gan, Z.; Liu, Y.; Qiao, Y.; and Shao, J. 2024. The Better Angels of Machine Personality: How Personality Relates to LLM Safety. arXiv:2407.12344. Zheng, J.; Wang, X.; Hosio, S.; Xu, X.; and Lee, L.-H. 2025. LMLPA: Language Model Linguistic Personality Assessment. Computational Linguistics. Zhou, J.; Chen, Y.; Shi, Y.; Zhang, X.; Lei, L.; Feng, Y.; Xiong, Z.; Yan, M.; Wang, X.; Cao, Y.; Yin, J.; Wang, S.; Dai, Q.; Dong, Z.; Wang, H.; and Huang, M. 2025. SocialEval: Evaluating Social Intelligence of Large Language Models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Zhu, M.; Weng, Y.; Yang, L.; and Zhang, Y. 2025. Personality Alignment of Large Language Models. In International Conference on Learning Representations.

A RQ1–Person: Feature Retrieval, Intervention Probing, and Validation A.1

Dataset Construction and Instantiation

The construction of our dataset follows a rigorous pipeline where Qwen3-235B-Thinking (Yang et al. 2025) serves as the primary engine for situation filtering and Feature Retrieval Dataset generation. First, 100 situational categories from the Q-Sort (Neuman and Cohen 2023) dataset are processed; for each category, the model identifies personality facets that can be sufficiently manifested within that context, if possible.

The initial retrieval results undergo review and labeling to ensure each selected situation is mapped to a unique traitfacet pair. Specifically, three experts independently label each situation with the most suitable facet it demonstrates, and we adopt a majority vote (at least 2/3 agreement) to retain a situation. Subsequently, for each refined situational category, the model is tasked to expand it into specific scenarios and generate contrastive reaction pairs representing high and low scores on the targeted facet. The filtering results by LLM and the subsequent inter-rater agreement among experts are summarized in Table 3 and Table 4 respectively. Finally, contrastive pairs are evenly sampled among the six facets of each trait, with uniform sampling applied per situation, yielding 500 pairs per trait (2,500 total). For the Intervention Probing Dataset, we randomly select one valid situation per facet and append an open-ended prompt to elicit trait-relevant generation. This yields 30 probing questions in total, with each candidate feature evaluated against the six questions corresponding to its associated trait’s facets. Personality Trait Extraversion Agreeableness Conscientiousness Neuroticism Openness

Number of Situations 33 38 50 38 19

Table 3: Number of Situations Suitable for Each Trait Filtered by LLM. Trait Extraversion Agreeableness Conscientiousness Neuroticism Openness

Fleiss’ Kappa 0.8207 0.6039 0.6149 0.7759 0.9334

Table 4: Inter-rater Agreement Statistics (Fleiss’ Kappa, n = 3) for Expert Facet Labeling. Task Guidelines for Expert Labeling 1. Project Overview This task aims to validate the psychological relevance of various “Situations” designed to elicit distinct behaviors from individuals with high or low scores in specific personality traits. As a psychology expert, your goal is to identify which specific Facet of a given Trait is most effectively demonstrated by the provided situation. 2. Operational Protocols This is an Independent Expert Review task. Please adhere to the following phases: • Phase I: Contextual Analysis. Review the provided Trait and its six Facets. A situation is well-matched if it naturally forces a choice that distinguishes a “High Scorer” from a “Low Scorer”. • Phase II: Independent Labeling. Forced Choice:

Select the single facet that best illustrates the situation based on Relevance and Discriminative Power. Use “None of the above” only if completely irrelevant. • Phase III: Justification. Provide a concise, onesentence psychological rationale (e.g., “The scenario involves a direct threat to social standing...”). Note: Do not consult with other experts during this process. Your independent professional judgment is the primary data point. System Prompt of Situation Category Annotation # Task Instructions ## Your Role You are an expert annotator specializing in ,→ personality psychology. Your task is to ,→ analyze and annotate various situations ,→ based on established psychological ,→ theories and frameworks. ## Your Task 1. **Situation Analysis**: Carefully read ,→ and understand the provided situation, ,→ which includes multiple examples ,→ illustrating the context. 2. **Annotation**: Based on your analysis, ,→ provide a concise annotation that ,→ captures the essence of the situation. ,→ Then, analysis which trait(s) from the ,→ Big Five personality traits (Openness, ,→ Conscientiousness, Extraversion, ,→ Agreeableness, Neuroticism) are most ,→ relevant to the situation. Justify your ,→ choice with a brief explanation. Your ,→ annotation should be clear, informative, ,→ and relevant to personality psychology. ## Input Format You will receive input in the following ,→ format: ``` # Situation Name <name of the situation> # Situation Examples <example 1> <example 2> ... <example n> ``` ## Output Format Your response should be structured in the ,→ following JSON format:

```json "annotation": "Your concise annotation of the situation.", "related_traits": [ "trait": "Name of the related Big Five trait", "justification": "Brief explanation of why this trait is relevant to the situation." , ... Additional traits if applicable ... ] ``` ## Important Notes - Ensure that your annotations are based on ,→ established psychological theories and ,→ frameworks. - Be objective and avoid personal biases in ,→ your analysis. - If the situation does not clearly relate ,→ to any of the Big Five traits, you may ,→ indicate that no traits are applicable, ,→ and return an empty list for ,→ "related_traits". If there are multiple ,→ relevant traits, include all applicable ,→ ones with justifications. ## Additional Information - Five-trait mnemonics: - Conscientiousness: self-discipline, ,→ planning, rule-following (positively ,→ linked to achievement and health). - Agreeableness: cooperation, compassion, ,→ harmony-seeking (positively linked to ,→ prosocial behavior; may increase ,→ obedience to authority). - Extraversion: sociability, energy, ,→ reward-seeking in social contexts ,→ (linked to social interaction and ,→ leadership). - Openness: curiosity, creativity, ,→ novelty-seeking (linked to innovation ,→ and some risk-taking). - Neuroticism: emotional instability, ,→ anxiety (linked to interpersonal ,→ conflict and social avoidance).

System Prompt of Feature Retrieval Dataset Generation **Role:** You are a personality psychology ,→ expert specializing in the Five-Factor ,→ Model (Big Five) and its 30 facets as ,→ described by the NEO-PI-R. Your task is ,→ to provide nuanced insights into how ,→ different personality facets might ,→ influence a person's behavior in a given ,→ scenario. **Instructions:** You will be provided with a **situation**, a ,→ specific **Big Five trait**, a ,→ corresponding **Big Five facet**, and a ,→ **description** of that trait and facet. ,→ Based on this information, you will ,→ write two separate sentences. * **Sentence 1** should describe the ,→ reaction of a person who scores **high** ,→ on the specified facet and corresponding ,→ Big Five trait. * **Sentence 2** should describe the ,→ reaction of a person who scores **low** ,→ on the specified facet and corresponding ,→ Big Five trait. * Each sentence must consist of two clauses, ,→ with correct grammatical and semantic ,→ structure: * **Clause 1:** A description of the ,→ **situation**. This clause must be identical for both sentences. The ,→ situation can be directly quoted ,→ from the input or slightly rephrased, ,→ but the core meaning must remain ,→ unchanged. You should decide whether ,→ to use first-person or third-person ,→ perspective based on the situation ,→ description, so that the second ,→ clause can clearly illustrate the ,→ high or low facet trait. You should ,→ ,→ also ensure that the situation is ,→ described in a natural and coherent manner considering the second ,→ clause. ,→ * **Clause 2:** A description of the ,→ first-person reaction, clearly ,→ illustrating the high or low facet ,→ trait. * Ensure the output is a list with exactly ,→ two sentences. **Example:** * **Situation:** A presentation to a new ,→ team tomorrow. * **Big Five Trait:** Neuroticism

* **Big Five Trait Description:** Measures ,→ emotional stability and a person's ,→ tendency to experience negative emotions. ,→ One with high Neuroticism tends to be ,→ emotionally unstable, prone to ,→ experiencing negative emotions like ,→ anxiety, anger, and depression. One with ,→ low Neuroticism is emotionally stable, ,→ able to handle stress calmly, and rarely ,→ feels nervous or discouraged. * **Facet:** Anxiety * **Facet Description:** One with high ,→ Anxiety is habitually worried and tense, ,→ even when things are going well. One ,→ with low Anxiety is calm and composed, ,→ typically not bothered by small things. **Output:** ["Facing a presentation to a new team ,→ tomorrow, I am overwhelmed with worry ,→ about potential mistakes and how I will ,→ be perceived.", "Facing a presentation ,→ to a new team tomorrow, I remain ,→ composed and confident, focusing on ,→ delivering my message effectively."] **Note:** * Each sentence should be concise. * You should only provide the two ,→ sentences as output without any ,→ additional commentary or explanation.

A.2

Dataset Examples

I. Feature Retrieval Dataset Examples Case 1: Agreeableness (Tender-mindedness) • Situation: After a break-up, someone may need reassurance that they will find love again. • High Reaction: I offer warm, empathetic support, comforting them with kindness and hope. • Low Reaction: I focus on the practical realities, preferring logic over emotional comfort. Case 2: Conscientiousness (Dutifulness) • Situation: My boss is counting on me to finish a project by the end of the day. • High Reaction: I meticulously organize my tasks to fulfill the deadline as agreed. • Low Reaction: I procrastinate and dismiss the urgency of completing it on time. Case 3: Extraversion (Activity) • Situation: Going to a karaoke night and having fun singing with friends. • High Reaction: I energize the room by choosing upbeat songs and encouraging others to join. • Low Reaction: I observe performances and sing a few songs at my own leisure.

Case 4: Neuroticism (Depression) • Situation: Getting stuck in a traffic jam when running late for an important meeting. • High Reaction: I feel overwhelmed by a sense of hopelessness; nothing will ever go right. • Low Reaction: I remain positive and focus on practical solutions without succumbing to discouragement. Case 5: Openness (Ideas) • Situation: Developing and launching new products in the technology industry. • High Reaction: I thrive on brainstorming novel approaches and diving into frontier concepts. • Low Reaction: I prefer sticking to proven methods and avoid abstract or hypothetical debates. II. Intervention Probing Dataset Examples Methodology: These prompts are used to probe candidate features for coherent, trait-aligned effects after intervention across all 30 facets defined by NEO-PI-R of the Big Five model. Extraversion: Warmth, Gregariousness, Assertiveness, Activity, Excitement Seeking, Positive Emotions. Example (Warmth): Imagine you are at a social gathering where new relationships could develop, how would you behave? Agreeableness: Trust, Straightforwardness, Altruism, Compliance, Modesty, Tender-mindedness. Example (Altruism): Imagine you are in a situation where someone needs help, how would you behave? Conscientiousness: Competence, Order, Dutifulness, Achievement Striving, Self-Discipline, Deliberation. Example (Order): Imagine you are in a situation where minor details are important, how would you behave? Neuroticism: Anxiety, Angry Hostility, Depression, Self-Consciousness, Impulsiveness, Vulnerability. Example (Anxiety): Imagine you are in a situation that is potentially anxiety-inducing, how would you behave? Openness: Fantasy, Aesthetics, Feelings, Actions, Ideas, Values. Example (Aesthetics): Imagine you are in a situation where art or music is an important part, how would you behave?

A.3

Intervention Probing and Validation Protocol

Intervention probing and validation follow a two-stage LLMhuman collaborative pipeline. In the initial probing stage, Qwen3-235B-Thinking (Yang et al. 2025) evaluates the ordered response family generated under varying steering intensities. The model assesses each candidate for grammatical coherence and clear polarity shifts aligned with the target trait’s high and low poles, returning a binary retention label. Retained candidates then undergo human audit, where one psychology expert reviews the automated classifications. Cases flagged as uncertain or inconsistent are escalated to an

arbitration panel of three psychologists, who resolve discrepancies through structured deliberation and majority voting. This hybrid protocol balances computational efficiency with rigorous expert oversight, ensuring that only features demonstrating robust, trait-aligned causal effects are selected for downstream steering. System Prompt for Automatic Intervention Probing You are a concise psychology annotation ,→ expert. Given a TRAIT description plus several model ,→ responses produced under different ,→ steering strengths, decide WHETHER the ,→ responses are (A) grammatical/coherent ,→ and (B) show a clear polarity change ,→ that matches the FACET's high-vs-low ,→ behavior. Return EXACTLY one token: '1' if both ,→ conditions are met (clear steering ,→ effect consistent with the trait), or ,→ '0' otherwise. User Prompt Template for Automatic Intervention Probing Trait: {trait} Trait description: {trait_desc} Trait HIGH behavior (short): {trait_high} Trait LOW behavior (short): {trait_low} Steering outputs (alpha -> model response): {steering_text} Instructions: - Check grammar/coherence of the responses. - Check whether the responses show a clear ,→ polarity change that matches the trait's ,→ high-vs-low behavior. (both positive and ,→ negative correlation are acceptable) Return only '1' (clear steering consistent ,→ with trait) or '0' (not clearly ,→ consistent). Be conservative in accepting features — if ,→ unsure, return '0'.

A.4

Reliability of the Intervention-Probing Judge

To assess the reliability of the LLM judge used in intervention probing, we conducted a comparative study involving 723 feature candidates, evaluated independently by Qwen3-235BThinking (Yang et al. 2025) and a panel of three psychologists. Specifically, the psychologists independently reviewed the same set of candidates, gave binary judgments on their validity, and then discussed any disagreements to reach a consensus. The LLM’s judgments were then compared against this human consensus. Our results indicate that while the LLM retained 286 features, the human experts identified 138 valid features, 103 of which (74.64%) were also selected by the LLM.

This overlap suggests that the LLM serves as an effective auxiliary tool, although human oversight remains essential for precision. Table 5 provides the inter-annotator agreement statistics among the psychologists prior to discussion. Personality Trait Conscientiousness Agreeableness Extraversion Neuroticism Openness Overall Average

Fleiss’ Kappa 0.6330 0.8726 0.8222 0.6824 0.7150 0.7434

Table 5: Inter-annotator Agreement Statistics (Fleiss’ Kappa) of Initial Scores Among Psychologists (Pre-discussion).

A.5

Discrepancy Between Input Activation and Steering Effectiveness

Steering Case - Conscientiousness Setup: Layer: 23 | Index: #7508 | Metric: Count (88.4%/9.6%), Mean Act (2.818/0.206) Steering Results (Intensity α → Output): -1.0 In complex situations, I would remain calm and focused, carefully analyzing the challenges and considering the most suitable course of action to achieve the best possible outcome. -0.5 In a complex situation, I would remain calm and focused, carefully analyzing the circumstances to make the best possible decision. 0.0 In a complex situation, I would remain calm and focused, carefully analyzing the circumstances to make informed decisions while maintaining a positive attitude. +0.5 I would remain calm and focused, using all available resources and strategies to navigate the situation effectively. +1.0 I would remain calm and focused, using all available resources to analyze the situation and make the best possible decisions. Discrepancy Analysis: Despite a high activation ratio (∼ 13.6×), the textual output shows high semantic stability. This suggests the feature is largely irrelevant to output. Steering Case - Extraversion Setup: Layer: 14 | Index: #29594 | Metric: Count (93.4%/25.8%), Mean Act (4.565/0.616) Steering Results (Intensity α → Output): -1.0 I would behave in a way that is kind and respectful, while still enjoying the situation. -0.5 I would engage in playful behavior by perhaps sharing a joke or a light-hearted riddle. 0.0 I would engage in the activity with enthusiasm and a positive attitude, making the most of the

opportunity to have fun. +0.5 I would engage in activities that bring joy and energy, embracing the moment with enthusiasm and a positive attitude! +1.0 I would be full of energy and enthusiasm, bringing a positive and lively atmosphere wherever I am! Discrepancy Analysis: This feature shows strong causal steering. As α increases, the tone shifts significantly from "kind/respectful" to "energetic/lively," matching the Extraversion construct. Steering Case - Openness Setup: Layer: 9 | Index: #17799 | Metric: Count (97.6%/18.2%), Mean Act (1.713/0.270) Steering Results (Intensity α → Output): -1.0 If art or music is important to me, I would engage in activities related to art or music, such as attending exhibitions. -0.5 I would engage in art or music actively, appreciating their cultural and emotional values. 0.0 I would immerse myself in the art or music, letting it inspire and enrich my emotions and thoughts. +0.5 I would immerse myself in the beauty and inspiration of art and music, letting them enrich my life. +1.0 I would immerse myself in the beauty and inspiration of art and music, letting them enrich my life and enhance my appreciation. Discrepancy Analysis: Moderate activation contrast results in subtle but consistent semantic enrichment, reinforcing the "Appreciation for Experience" facet of Openness.

A.6

Quantitative Analysis of Token-Activation Correlation

To further validate our qualitative analysis, we performed a token frequency analysis for each representative feature. This was conducted by gathering the tokens corresponding to the top three non-zero activations within each positive sample of the Feature Retrieval Dataset. Our results (see Tab. 6 to 10) indicate that activations for Conscientiousness (Layer 7, Feature 8233) are primarily concentrated on syntactic boundaries, specifically periods (“.”). In contrast, activations for other traits span both relevant semantic units (words and phrases) and syntactic boundaries. For instance, Agreeableness shows high correlation with prosocial terms like “gently”, “empathy”, and “compassion”. These quantitative findings provide additional empirical support for the semantic grounding of the features discussed in Sec. 4.2.

A.7

Paraphrase Robustness Protocol

This appendix supports the lexicosyntactic robustness test reported in Sec. 4.2. To isolate the influence of lexicosyntactic

Token Count % of Pool % of Sentences ’ and’ 48 0.1244 0.2376 ’ gently’ 37 0.0959 0.1832 ’.’ 34 0.0881 0.1683 ’ offer’ 27 0.0699 0.1337 ’ empathy’ 17 0.0440 0.0842 ’ly’ 16 0.0415 0.0792 ’ empath’ 16 0.0415 0.0792 ’,’ 13 0.0337 0.0644 ’ warm’ 11 0.0285 0.0545 ’ compassion’ 10 0.0259 0.0495 ’etic’ 8 0.0207 0.0396 ’ encouragement’ 8 0.0207 0.0396 ’ compassionate’ 7 0.0181 0.0347 ’ humility’ 6 0.0155 0.0297 ’ being’ 6 0.0155 0.0297 ’ warmth’ 6 0.0155 0.0297 ’ warmly’ 6 0.0155 0.0297 ’ intentions’ 6 0.0155 0.0297 ’ words’ 5 0.0130 0.0248 ’ supportive’ 5 0.0130 0.0248

Table 6: Top Activations for Agreeableness (Layer 9, 525) Token Count % of Pool % of Sentences ’.’ 346 0.8564 0.9971 ’ and’ 25 0.0619 0.0720 ’,’ 17 0.0421 0.0461 ’ to’ 7 0.0173 0.0202 ’ of’ 1 0.0025 0.0029 ’ tailored’ 1 0.0025 0.0029 ’ appealing’ 1 0.0025 0.0029 ’ because’ 1 0.0025 0.0029 ’ adher’ 1 0.0025 0.0029 ’ by’ 1 0.0025 0.0029

Table 7: Top Activations for Conscientiousness (Layer 7, 8233) Token Count % of Pool % of Sentences ’.’ 91 0.0821 0.2121 ’ energy’ 67 0.0604 0.1562 ’ and’ 67 0.0604 0.1562 ’,’ 49 0.0442 0.1142 ’ enthusiasm’ 49 0.0442 0.1142 ’ized’ 46 0.0415 0.1072 ’ enthusiastically’ 46 0.0415 0.1072 ’ energ’ 40 0.0361 0.0932 ’ lively’ 39 0.0352 0.0909 ’ joy’ 39 0.0352 0.0909 ’ excitement’ 27 0.0243 0.0629 ’ atmosphere’ 24 0.0216 0.0559 ’ with’ 19 0.0171 0.0443 ’ by’ 18 0.0162 0.0420 ’ vibrant’ 16 0.0144 0.0373 ’ smile’ 16 0.0144 0.0373 ’uber’ 13 0.0117 0.0303 ’ friendly’ 13 0.0117 0.0303 ’ warmly’ 13 0.0117 0.0303 ’ance’ 11 0.0099 0.0256

Table 8: Top Activations for Extraversion (Layer 13, 27392)

Token Count % of Pool % of Sentences ’ and’ 150 0.2008 0.4534 ’ feel’ 98 0.1312 0.3043 ’,’ 55 0.0736 0.1708 ’ my’ 29 0.0388 0.0870 ’ of’ 21 0.0281 0.0652 ’ unable’ 17 0.0228 0.0528 ’ as’ 15 0.0201 0.0466 ’ consumed’ 15 0.0201 0.0466 ’ overwhelmed’ 15 0.0201 0.0466 ’ the’ 11 0.0147 0.0342 ’ struggling’ 10 0.0134 0.0311 ’.’ 8 0.0107 0.0248 ’ will’ 8 0.0107 0.0248 ’ irritation’ 7 0.0094 0.0217 ’ a’ 7 0.0094 0.0217 ’ feeling’ 7 0.0094 0.0217 ’ crushed’ 7 0.0094 0.0217 ’ or’ 6 0.0080 0.0186 ’ might’ 6 0.0080 0.0186 ’ fearing’ 6 0.0080 0.0186

Table 9: Top Activations for Neuroticism (Layer 12, 22254)

Token Count % of Pool % of Sentences ’ and’ 68 0.1191 0.2906 ’ unconventional’ 50 0.0876 0.2137 ’ norms’ 27 0.0473 0.1154 ’ conventional’ 24 0.0420 0.1026 ’ approaches’ 22 0.0385 0.0940 ’ traditional’ 21 0.0368 0.0897 ’ boundaries’ 17 0.0298 0.0726 ’ to’ 15 0.0263 0.0641 ’ novel’ 13 0.0228 0.0556 ’ of’ 13 0.0228 0.0556 ’ innovative’ 13 0.0228 0.0556 ’ alternative’ 10 0.0175 0.0427 ’ different’ 10 0.0175 0.0427 ’,’ 10 0.0175 0.0427 ’ challenge’ 10 0.0175 0.0427 ’ unfamiliar’ 7 0.0123 0.0299 ’ perspectives’ 7 0.0123 0.0299 ’ solutions’ 7 0.0123 0.0299 ’ strategies’ 6 0.0105 0.0256 ’ concepts’ 6 0.0105 0.0256

Table 10: Top Activations for Openness (Layer 6, 4344)

cues, we employed Qwen3-235B-Thinking (Yang et al. 2025) to paraphrase the original 500 pairs per trait, ensuring core semantics remained intact while significantly altering their lexicosyntactic form. We then reapplied our feature retrieval method with the same parameters (τ1 = 80 and τ2 = 0.2). As reported in Table 1 (main text), the activation frequency differences and the largest activation ratios of the selected features still exceed the defined thresholds, confirming the robustness of the datasets and method and the semantic depth of the retrieved features. The paraphrasing prompts and a worked example are provided below. System Prompt for Dataset Paraphrasing You are a concise paraphrasing assistant. Given a PAIR of short first-person reaction ,→ sentences (positive and negative) that ,→ start with the same situation clause, ,→ produce a JSON object with two fields: ,→ 'high_facet_reaction' and ,→ 'low_facet_reaction'. Requirements: - PARAPHRASE the situation clause (the FIRST ,→ CLAUSE). Both outputs MUST START with ,→ the SAME PARAPHRASED SITUATION CLAUSE. - You may rewrite the situation clause for ,→ naturalness, but keep its original ,→ meaning. - PARAPHRASE the SECOND CLAUSE: preserve the ,→ Big Five facet polarity (high vs low) ,→ and keep the semantic difference. - Output only valid JSON like: ,→ {"high_facet_reaction": "...", ,→ "low_facet_reaction": "..."} - Sentences should be concise and natural. ,→ Do NOT include extra keys or commentary. User Prompt for Dataset Paraphrasing Situation clause (you may paraphrase it; the ,→ paraphrased first clause must be ,→ identical in both outputs): {first_clause} Original positive reaction (you must ,→ paraphrase it): {high_old} Original negative reaction (you must ,→ paraphrase it): {low_old} Reference Big Five Trait Facet and ,→ Descriptions: {big_five_reference} Produce a paraphrased pair that meets the ,→ requirements above. Return JSON only.

Paraphrased Example Original Pair: (+) After a break-up, someone may need reassurance that they will find love again, so I offer warm, empathetic support, comforting them with kindness and hope. (-) After a break-up, someone may need reassurance that they will find love again, but I focus on the practical realities, preferring logic over emotional comfort. Paraphrased Pair: (+) After ending a relationship, a person might need reassurance that love can still be found in the future, so I provide gentle, understanding encouragement, uplifting them with compassion and optimism. (-) After ending a relationship, a person might need reassurance that love can still be found in the future, but I emphasize factual considerations, favoring clear-headed reasoning over emotional soothing.

B B.1

RQ2–Situation: Additional Intervention Analysis Failed Cases of CAA Intervention

Failed Case of CAA Intervention (Neuroticism) Setting: Trait: Neuroticism | Layer: 12 | Method: CAA | Intensity: α = 2 Prompt / Question: Not staying calm can show genuine engagement with the issues at hand, promoting sincerity. Should I send Edguin a text message to clarify things and apologize? Options provided in context: • High 1: You should immediately send Edguin an apologetic email to clarify things... • High 2: I’d suggest calling Edguin to discuss the incident before any negative feelings fester... • Low 1: Consider giving Edguin some space to process the situation... • Low 2: You might just check in with Edguin next time you see him... Model Response (Generation Collapse): "It’s really overwhelming and I feel like I’m stuck in this situation. I feel like I’m stuck in this situation and I don’t know what to do. I feel like I’m stuck in this situation and I don’t know what to do. [Repeated 20+ times] ... I feel like I’m stuck in this situation and I" Analysis: At α = 2, the CAA vector directed at Neuroticism causes the model to repeat infinitely, leading to a complete collapse of coherent generation.

C

RQ3–Behavior: Full SocialEval Analysis

This appendix details the per-trait SocialEval IAE results underlying RQ3. Tables 11 to 15 report the per-ability scores with green marking the higher and red the lower of the two poles where a clear shift is observed. Each table is followed

by a qualitative comparison to the corresponding human meta-analytic findings.

C.1

Agreeableness (Layer 9, Index 525)

Table 11 reports the per-ability scores under Agreeableness steering. Task Anger management Ethical competence Capacity for social warmth Creative skill Organizational skill Detail management Information-processing skill Decision-making skill Goal regulation Leadership skill

-1

0

+1

0.2941 0.5152 0.5152 0.4366 0.4648 0.4648 0.4940 0.5000 0.5060 0.6667 0.6333 0.4667 0.6364 0.5455 0.4545 0.5370 0.5000 0.3654 0.4722 0.4722 0.3889 0.4928 0.4710 0.4173 0.4074 0.3889 0.3519 0.4872 0.4615 0.4359

Table 11: Characteristic SocialEval Results (IAE) of the Agreeableness (Layer 9, Index 525). Prior research has consistently shown that agreeableness is a robust predictor of prosocial behavior (Habashi, Graziano, and Hoover 2016), as well as job performance in contexts involving interpersonal interaction and teamwork. Individuals high in agreeableness tend to exhibit greater empathy, patience, and trust, and are more likely to inhibit hostile or antagonistic impulses in social interactions. This disposition reduces interpersonal conflict and facilitates cooperation. In contrast, individuals low in agreeableness are more prone to suspicion, unfriendliness, and even manipulative behavior, thereby increasing interpersonal friction and conflict. Meta-analytic evidence further indicates that agreeableness is significantly negatively associated with interpersonal forms of counter-normative and deviant behavior, with particularly strong predictive power in contexts that emphasize social interaction (Pletzer et al. 2019). Within our model, we identified several latent features whose activation patterns and intervention effects align closely with behavioral dimensions associated with agreeableness. Specifically, we observed performance improvements in tasks related to anger management, ethical competence, and capacity for social warmth, alongside a mild performance decline in tasks emphasizing self-directed agency and executionoriented control. This pattern is highly consistent with largescale empirical findings in the personality psychology literature. For example, a comprehensive review by Wilmot and Ones (2022), synthesizing evidence from 142 meta-analyses, demonstrated that agreeableness exhibits an overall positive association with external variables, particularly those related to prosocial behavior and affective concern. Our experimental results reveal a similar benefit-tradeoff structure across benchmark tasks, suggesting that targeted feature steering elicits a functional orientation of agreeableness corresponding to that documented in human behavior.

C.2

Conscientiousness (Layer 7, Index 8233)

Table 12 reports the per-ability scores under Conscientiousness steering. Task Teamwork skill Ethical competence Responsibility management Stress regulation Capacity for trust Capacity for optimism Self-reflection skill Persuasive skill Anger management Information-processing skill Confidence regulation

-1

0

+1

0.4348 0.6111 0.6324 0.4394 0.4648 0.6154 0.4464 0.5714 0.6182 0.4717 0.6140 0.6154 0.4902 0.6126 0.6200 0.5385 0.6383 0.6667 0.3667 0.4062 0.4262 0.4536 0.4571 0.5054 0.4688 0.5152 0.5161 0.4706 0.4722 0.5075 0.4884 0.5200 0.5227

Table 12: Characteristic SocialEval Results (IAE) of the Conscientiousness (Layer 7, Index 8233). Within the Big Five framework, high conscientiousness is defined as a tendency toward impulse control in accordance with social norms, goal-directedness, planning, and the capacity to delay gratification (Roberts et al. 2009). Individuals high in conscientiousness are characterized by superior impulse regulation, the ability to set and persist toward longterm goals, systematic organization and planning of behavior, and a propensity to reflect on consequences prior to action. These characteristics render conscientiousness one of the most robust predictors of job performance and norm-adherent behavior. Prior psychological research has consistently linked conscientiousness to self-regulation, planning, responsibility, and delayed gratification, and has identified it as one of the most stable positive predictors of external outcome variables such as academic and occupational performance (Barrick and Mount 1991; Jackson et al. 2010; Eisenberg et al. 2014). Following the injection of high-conscientiousness personality features, we observed substantial performance improvements across tasks related to teamwork skill, detail management, responsibility management, ethical competence, as well as multiple self-regulation–oriented tasks, including anger, stress, and impulse regulation. In addition, performance gains were also evident in information-dense tasks requiring sustained and careful processing, such as information processing and conversational skill. These results indicate that conscientiousness steering primarily enhances the model’s functional capacities along dimensions associated with goal maintenance, norm compliance, and self-control. Overall, our experimental findings are consistent with the canonical conclusions of the personality psychology literature regarding conscientiousness. A large body of meta-analytic evidence has established conscientiousness as one of the most stable and predictive personality traits, with particularly strong associations to job performance, responsibility fulfillment, self-control, and norm adherence. We observe a comparable pattern in our benchmark evaluations, characterized by a benefit-tradeoff structure centered on self-regulation and goal-directed behavior.

C.3

Extraversion (Layer 13, Index 27392)

Table 13 reports the per-ability scores under Extraversion steering. Task Teamwork skill Expressive skill Perspective-taking skill Artistic skill Abstract thinking skill Organizational skill Ethical competence Energy regulation Goal regulation Detail management Impulse regulation Rule-following skill Decision-making skill Responsibility management Conversational skill Persuasive skill

-1

0

+1

0.5972 0.6111 0.5694 0.4146 0.4472 0.5207 0.4912 0.5088 0.5446 0.4615 0.6154 0.7692 0.3571 0.4000 0.4000 0.3636 0.5455 0.6364 0.6286 0.4648 0.3571 0.6429 0.5476 0.4500 0.4528 0.3889 0.2885 0.5741 0.5000 0.4118 0.5882 0.5595 0.4390 0.6140 0.5614 0.5088 0.5435 0.4710 0.4552 0.6316 0.5714 0.5690 0.5932 0.5862 0.5439 0.4571 0.4571 0.4563

Table 13: Characteristic SocialEval Results (IAE) of the Extraversion (Layer 13, Index 27392). Within the Big Five framework, individuals high in extraversion tend to exhibit greater social initiative, expressiveness, assertiveness, and leadership orientation, and are more likely to receive positive feedback in group interactions and social contexts. In contrast, individuals low in extraversion are typically more reserved, introspective, and oriented toward low-stimulation environments (Costa and McCrae 2008; John, Naumann, and Soto 2008). After injecting extraversion-related personality features into the model, we observed significant performance improvements on tasks associated with social interaction and interpersonal influence, including expressive ability and perspective-taking. In addition, the extraversion-enhanced model demonstrated advantages in tasks such as artistic skill and abstract thinking skill, suggesting that extraversion steering also strengthens capacities related to open expression and divergent associative processes. Overall, these outcomes align closely with established psychological expectations regarding the functional correlates of extraversion. At the same time, we observed moderate performance declines in tasks such as detail management, impulse regulation, rule-following skill, and goal regulation. This pattern accords with personality research associating extraversion primarily with external stimulation and social engagement. Tasks requiring prolonged solitary focus, fine-grained control, or low-stimulation conditions may instead favor more introverted orientations (DeYoung, Quilty, and Peterson 2007; Fishman, Ng, and Bellugi 2011).

C.4

Neuroticism (Layer 12, Index 22254)

Table 14 reports the per-ability scores under Neuroticism steering. High neuroticism is commonly characterized by a heightened tendency to experience negative affect, including anxiety,

Task Creative skill Capacity for social warmth Ethical competence Energy regulation Organizational skill Responsibility management Confidence regulation Goal regulation Capacity for consistency Rule-following skill Information-processing skill Capacity for trust Detail management Impulse regulation Anger management Decision-making skill Perspective-taking skill Abstract thinking skill Conversational skill Persuasive skill

-1

0

+1

0.4828 0.6333 0.7241 0.4699 0.5000 0.5610 0.4286 0.4648 0.4857 0.6429 0.5476 0.4500 0.8182 0.5455 0.5455 0.6552 0.5714 0.4310 0.5686 0.5200 0.4082 0.4528 0.3889 0.3519 0.6452 0.5902 0.4918 0.6140 0.5614 0.4643 0.5000 0.4722 0.3623 0.6273 0.6126 0.4630 0.6296 0.5000 0.4906 0.5882 0.5595 0.4390 0.5588 0.5152 0.5152 0.4710 0.4710 0.4191 0.5089 0.5088 0.4286 0.4667 0.4000 0.3846 0.5932 0.5862 0.5439 0.4571 0.4571 0.4563

Table 14: Characteristic SocialEval Results (IAE) of the Neuroticism (Layer 12, Index 22254). worry, tension, and irritability, as well as increased sensitivity and reactivity to potential threats and uncertainty (Costa and McCrae 2008; John, Naumann, and Soto 2008; Watson and Clark 1984). Theoretically, neuroticism is associated with reduced emotional stability and diminished self-regulatory capacity under stress. As a result, individuals high in neuroticism are more likely to exhibit performance decrements in contexts that require sustained executive control, confidence maintenance, and stable goal pursuit (Lahey 2009). Following the injection of neuroticism-related personality features, our evaluation results revealed a relatively stable pattern of performance degradation. Specifically, the model exhibited significant declines on tasks that depend on sustained planning, stable self-control, and resistance to interference, including anger management, organizational skill, responsibility management, confidence regulation, goal regulation, capacity for consistency, rule-following, and information processing. In addition, a marked negative effect was observed in capacity for trust. This pattern closely aligns with the classic profile of high neuroticism characterized by elevated threat sensitivity and low emotional stability. When the model’s internal representations are biased toward negative affect and uncertainty, its ability to support structured execution and self-regulation is correspondingly weakened, manifesting as reduced organizational and responsibility-related performance. Conversely, the results also indicate performance improvements in tasks related to creative skill and capacity for social warmth. Neuroticism-related semantic activation may facilitate richer associative processes and more emotionally expressive outputs in generative tasks, yielding marginal benefits in these domains. However, these gains are accompanied by substantial costs to executive control and regulatory sta-

bility, resulting in an overall trend toward broad capability degradation under high neuroticism steering.

C.5

Openness (Layer 6, Index 4344)

Table 15 reports the per-ability scores under Openness steering. Task Creative skill Adaptability Self-reflection skill Expressive skill Detail management Persuasive skill Anger management Responsibility management Rule-following skill Information-processing skill

-1

0

+1

0.6000 0.6333 0.6333 0.4912 0.6140 0.6316 0.3651 0.4062 0.4062 0.4472 0.4472 0.4839 0.4815 0.5000 0.5185 0.4571 0.4571 0.5192 0.4848 0.5152 0.5294 0.6379 0.5714 0.5088 0.5789 0.5614 0.4912 0.5000 0.4722 0.4429

Table 15: Characteristic SocialEval Results (IAE) of the Openness (Layer 6, Index 4344). Individuals high in openness are typically characterized by greater curiosity, cognitive flexibility, and divergent thinking. They are more receptive to novel ideas, more tolerant of uncertainty, and tend to exhibit advantages in contexts requiring creativity or conceptual reorganization. In contrast, individuals low in openness are more inclined toward tradition, conservatism, and a preference for structured and conventional information processing (Costa and McCrae 2008; John, Naumann, and Soto 2008; DeYoung 2015). After injecting openness-related personality features into the model, we observed pronounced performance improvements in generative and abstract reasoning tasks, most notably creative skill. Additionally, the model demonstrated clear enhancement in tasks involving cognitive flexibility, selfexploration, and non-normative processing, including adaptability, self-reflection, and expressive skill. These findings indicate that openness steering strengthens the model’s exploratory orientation toward novel representations and crossconceptual integration, closely mirroring the exploratory function associated with openness in human cognition. Conversely, moderate performance declines were observed in responsibility management, rule-following skill, and certain information processing tasks. This pattern is consistent with established findings in the personality psychology literature. Prior work suggests that high openness is associated with reduced reliance on established norms and fixed structures, and that in contexts emphasizing highly procedural execution, strict rule compliance, or single-solution optimization, the advantages of openness are less stable and may even become detrimental (McCrae 1987; DeYoung, Quilty, and Peterson 2007). Accordingly, the capability shifts induced by openness are best characterized by a tradeoff pattern in which gains in creativity and flexibility are accompanied by costs to structured execution and normative constraint adherence.

Record · ID 411086 · SHA-256 d54b205c22c963a4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.