The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
YUAN YUAN 1 [email protected]
arXiv:2607.02368v1 [stat.ML] 2 Jul 2026
Abstract
paradigm, consistently revealing systematic response patterns and biases (Safdari et al., 2023; Jiang & Zou, 2024; Argyle et al., 2023). Furthermore, the geometric factor structure of personality is a psychological cornerstone, and this structure reliably emerges in aggregate LLM response data (Liu et al., 2023). However, a critical gap persists: prevailing research operates almost exclusively on aggregate feature averages (e.g., mean dimension scores), collapsing the within-instance correlation structure that defines individual differences. This renders the recovered “persona” a sample-level artifact and obscures a fundamental question.
Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation structure. We test whether this geometric structure is intrinsic or frame-dependent. Constructing withininstance correlation matrices from IPIP-50 responses, we analyze geometry on SPD manifolds under manipulated question orderings in GPT-4o simulating American and Chinese-American personas. We find that persona expression comprises two dissociable components: aggregated features (Big Five scores) degrade under randomization (21% drop) but are frame-robust; geometric features (SPD manifold) collapse under frame misalignment (42% drop) but recover substantially (to 84%) under shared frames, surpassing aggregated features (76%). This collapse-recovery pattern reveals that persona geometry is not intrinsic but a frame-dependent coordination pattern encoding information invisible to aggregation.
This omission is especially pressing given the autoregressive nature of LLMs (Radford et al., 2019; Vaswani et al., 2017). The standard practice of using a fixed question order conflates two potential sources of the observed geometry: is it a stable, intrinsic trait of the model, or is it merely an epiphenomenon of a specific, shared temporal frame during measurement? Consequently, the pivotal inquiry is not whether a geometric structure exists—it is mathematically given—but what its nature is, and why its frame-dependence has been systematically neglected.
Our findings establish a dual-nature framework for LLM personas—frame-dependent geometry versus frame-robust aggregates—necessitating frame-aware evaluation and challenging static trait conceptions.
1.1. Research Question and Hypotheses We hypothesize that the apparent stability of geometric bias structure is an artifact of fixed frames. To test this, we propose three competing hypotheses that capture distinct possibilities in the literature: 1. Intrinsic Structure (H1): Geometric features capture stable model properties, as assumed in trait-based personality assessment (McCrae & John, 1992).
1. Introduction The use of psychometric questionnaires (e.g., IPIP, Big Five) to investigate LLM personas is a well-established
2. Measurement Artifact (H2): The apparent structure is spurious and destroyed by perturbation, analogous to order effects in survey methodology (Schuman & Presser, 1996).
Yuan Yuan 1 Independent Researcher. Correspondence to: <[email protected]>. This paper was submitted to ICML 2026 but has been withdrawn by the authors and is published on Arxiv as an independent preprint.>.
3. Frame-Dependent Coordination (H3): Geometric features encode relational patterns that require temporal alignment—a novel hypothesis motivated by LLMs’ autoregressive nature (Radford et al., 2019).
rd
Proceedings of the 43 International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).
1
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
Visual predictions. Figure 1 plots the predicted clustering accuracy (y-axis) across our three analytical conditions—Fixed Order (FO), Random Order Native Frame (RO), and Random Order Bootstrap Shared Frame (ROBTSP)—for each hypothesis. H1 predicts a high, flat line; H2 predicts high accuracy only in FO with collapse in both randomized conditions; and H3 uniquely predicts a V-shaped pattern: collapse in RO (frame misalignment) followed by recovery in RO-BTSP (shared frame realignment).
This clean dissociation reveals that bias in LLM output space comprises two dissociable components: one tied to how dimensions coordinate during sequence processing (geometry), and the other reflecting what values are typically generated (aggregation). This discovery challenges the prevailing monolithic view of bias, establishing a dual nature framework for understanding LLM bias: as simultaneously a frame-dependent coordination pattern and a frame-robust aggregate tendency. Our Contributions:
We systematically tests these hypotheses through controlled order manipulation and geometric analysis.
• Empirical: We demonstrate—via a novel itemdimension matrix and bootstrap protocol—that the geometric structure of LLM bias is not intrinsic but frame-dependent: it collapses under misalignment but recovers under shared frames, revealing a dual nature distinct from aggregated tendencies.
H1: Intrinsic Structure (Failed) H2: Measurement Artifact (Failed) H3: Frame-Dependence (Supported)
Predicted Clustering Accuracy (%)
100 90 80 70
Collapse
Recovery
Random Order (Frame Misaligned)
RO-Bootstrap (Frame Realigned)
• Theoretical: We establish a dual-nature framework for LLM bias, distinguishing between frame-dependent coordination geometry and frame-robust aggregated tendencies.
60 50 40
Fixed Order (Frame-Aligned)
Experimental Conditions
• Methodological: We propose a new standard for LLM evaluation that emphasizes frame-aligned analysis and explicit decomposition of order vs. frame effects. Beyond substantive findings, we introduce: (1) a matrix approach based on item-dimensions within each LLM instance allowing geometric analysis of individual LLM responses, (2) systematic manipulation of temporal frames to dissociate intrinsic structure from coordination artifacts, and (3) validation across sample sizes (N ≈ 100 and N = 2000, see Appendix A) demonstrating robustness and clarifying optimal sample selection for geometry analysis bias.
Figure 1. Differential predictions for geometric bias features. Only frame-dependence (H3) predicts the observed collapse-recovery V-shape pattern across conditions (FO → RO → RO-BTSP).
1.2. Methodological Innovation and Dual-Nature Discovery To test these hypotheses and dissociate content from temporal structure effects, we introduce a novel methodological framework. We develop the Item-Dimension Matrix method, constructing within-instance correlation matrices from questionnaire responses, enabling geometric analysis on the manifold of symmetric positive definite (SPD) matrices. Critically, we systematically manipulate question ordering to test whether geometric features exhibit the invariance expected of intrinsic structures.
2. Related Work 2.1. LLM Personality Assessment Recent work has systematically applied personality inventories to LLMs (Safdari et al., 2023; Jiang & Zou, 2024). These studies typically adopt human psychometric assumptions without questioning their applicability to autoregressive architectures. Our work directly tests these assumptions through sequence manipulation.
Our investigation reveals a fundamental dissociation: all geometry-based features (SPD manifold, eigenvalues, eigenvectors) catastrophically collapse under frame misalignment, but substantially recover under shared frames. This reversible collapse indicates that what appears as “bias geometry” is not a static property but a frame-dependent coordination pattern that can only be measurable in temporal alignment. In stark contrast, aggregation-based features (Big Five scores) show the opposite sensitivity: they are robust to frame misalignment but degrade under content randomization.
2.2. Order Effects in Measurement Human assessment shows modest order effects (typically < 10%) attributed to cognitive consistency mechanisms (Schuman & Presser, 1996). LLMs likely operate differently through context accumulation rather than self-consistency maintenance. 2
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
2.3. Temporal Effects in LLMs
details). After filtering, the final counts were FO (US=96, CA=97, total=193), RO (US=92, CA=95, total=187). This yields a balanced design with sufficient power for our hypothesis tests.
While recent work has documented position-dependent biases in LLMs due to attention mechanisms (Vaswani et al., 2017), these studies typically focus on local context effects (e.g., recency bias) rather than systematic geometric structures emerging from sequential coordination. Our work extends this line by asking whether the global correlational geometry observed in persona assessments is itself framedependent, a question that has not been addressed in the prior literature on temporal effects.
Data Collection For each LLM call, we first simulated American or Chinese-American personas through cultural prompts (see the Appendix C.1). Following cultural induction, we administer the 50-item International Personality Item Pool (IPIP-50) (Goldberg, 1992). Items were adapted to first-person statements for LLM comprehension (complete list in Appendix C.2).
2.4. Geometric Methods and Methodological Rationale We draw methodological inspiration from functional connectivity (FC) analysis in neuroscience, where correlation matrices capture how brain regions coordinate over time (Barachant et al., 2013). This framework better captures dynamic systems than static trait models. Similarly, we hypothesize LLM bias involves inter-dimensional coordination during autoregressive generation—how personality dimensions covary in sequence—invisible to dimensionwise aggregation.
Cultural Persona Induction and Justification We selected cultural bias as an experimental platform because it provides a well-documented and robust signal for testing geometric representations. Previous work establishes that LLMs exhibit distinct response patterns when simulating American versus Chinese-American perspectives (Jiang & Zou, 2024; Santurkar et al., 2023). These established differences in bias content provide a strong signal for testing whether geometric representations vary independently of our core manipulation: temporal structure (question ordering). Although cultural identity is multifaceted, this binary classification serves as a controlled testbed for frame-dependence mechanisms, not as exhaustive cultural representation.
Correlation matrices naturally lie on the manifold of symmetric positive definite (SPD) matrices, where Riemannian metrics (e.g., log-Euclidean (Arsigny et al., 2007)) provide principled distances respecting manifold geometry. SPD methods have proven effective in brain-computer interfaces (Barachant et al., 2013) and computer vision (Huang & Van Gool, 2017), but assume the geometric structure is intrinsic. Our work tests this assumption for LLMs.
3.2. Item-Dimension Matrix and Correlation Construction
Critically, recent studies document that LLMs exhibit position-dependent biases due to attention mechanisms and autoregressive processing (Vaswani et al., 2017). Unlike human cognitive consistency effects, these architectural properties may create frame-dependent structures. By systematically manipulating temporal frames, we test whether bias geometry is intrinsic or an artifact of measurement alignment—a question that has not been addressed in prior geometric analysis.
Methodological Rationale Our research question, whether geometric bias structures are intrinsic properties or frame-dependent artifacts, requires an analytic framework that captures within-instance correlation patterns while enabling systematic temporal frame manipulation. Traditional aggregate methods (Big Five means) collapse within-instance structure, while factor analysis on pooled data (Liu et al., 2023) recovers only sample level patterns. We develop the Item-Dimension Matrix approach to construct instance-specific correlation matrices C(π) that encode how dimensions covary during sequential generation under ordering π. This conceptually parallels functional connectivity analysis in neuroscience (Barachant et al., 2013), where correlation matrices capture temporal coordination. Crucially, different orderings π produce different C(π) responses from identical responses, directly operationalizing the frame variation while maintaining the content constant.
3. Method 3.1. Experimental Design and Data Generation Model and Instrument We used the OpenAI API (gpt-4o-2024-05-13, temperature=0.7) to collect responses, targeting 100 LLM calls per cell1 ; we retained only complete, well-formed answers (see the Appendix C.3 for 1
This sample size was selected to balance discriminable cultural signals with sufficient data for stable correlation estimation, while avoiding over-aggregation effects that dilute group differences (see Appendix A for large-sample validation at N = 2000 demonstrating robustness of findings and rationale for this choice).
Since correlation matrices reside on the SPD manifold, we map them to the tangent space at identity via log(C) (Arsigny et al., 2007), enabling standard Euclidean operations while respecting the manifold structure. 3
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
Constructing item-dimension matrix using Multivariate Time Series from Questionnaire Responses We conceptualize the sequential response process as a multivariate (5-channel) time series. Each channel corresponds to one of the Big Five personality dimensions. Time progresses with the presentation of each question in the order π. (i)
• RO-BTSP (Random Order, Bootstrap Shared Frame): For each bootstrap iteration b, a random πb is drawn and used to recompute C(πb ) for all RO in(i) stances, imposing a shared frame: CBTSP,b = C(πb ) . This methodology isolates the effect of temporal coordination from the response content.
50
For a given instance with a response vector r ∈ R and a specific question order π, we construct a 10 × 5 ItemDimension Matrix X(π) as follows:
3.3. Instruments Selection and Methodology Validation Instrument Selection The construction of wellconditioned correlation matrices for the LLM instance C(π) requires a matrix of the full dimension of the item X(π) . Many popular personality instruments fail this requirement. The Ten-Item Personality Inventory (TIPI) (Gosling et al., 2003) provides only 2 items per dimension, insufficient for a reliable correlation estimation within an instance. The IPIP-NEO-300 (Johnson, 2014) measures 30 facets with 10 items each, producing a matrix 10 × 30 with too few items per facet. The balanced structure of IPIP-50 (10 items × 5 dimensions) provides full-rank X(π) while maintaining comparability with previous research on the LLM persona (Goldberg, 1992).
1. Each of the 50 IPIP items is pre-mapped to one of the five dimensions via a fixed function dim(·). 2. We iterate through the sequence π. At time step t (corresponding to the t-th question in π, denoted π[t]), we obtain the numerical score r(π[t]). 3. We place this score r(π[t]) in the column of X(π) that corresponds to the dimension j = dim(π[t]). The score is appended to the next available row position within that column, preserving its temporal order of occurrence in π. 4. After processing all 50 questions in π, each of the 5 columns contains exactly 10 scores—all responses for that dimension—in the exact order in which they were encountered during the sequence π.
Validity of the Item-Dimension Matrix Approach The item-dimension matrix reorganizes the responses from the validated IPIP-50 (Goldberg, 1992) to allow the analysis of correlations within the instance. This is a data organization method, not a new psychometric instrument—it preserves all information from original responses while enabling geometric analysis on SPD manifolds. Its validity rests on: (1) the established psychometric properties of IPIP-50 and (2) the full-rank structure (10×5) that ensures well-conditioned correlation matrices. The approach is designed to test framedependence hypotheses by systematically manipulating temporal order while preserving response content.
Thus, the column j of X(π) represents the temporal response sequence for dimension j as it was sampled intermittently during order π. Different orders π produce different temporal arrangements of the same 10 responses within each column. Computing Within-Instance Correlation Matrices From the matrix X(π) , we compute the 5×5 within-instance Pearson correlation matrix: C(π) = corr(X(π) ),
Advantages Over Traditional Approaches Traditional personality assessment relies on aggregate scores, discarding the within-instance correlation structure. Our itemdimension matrix enables geometric analysis at the instance level, analogously to functional connectivity analysis in neuroscience (Barachant et al., 2013). This approach is necessary because: (1) it preserves temporal coordination information lost in aggregation; (2) it allows testing of frame-dependence via order manipulation; (3) it provides a mathematically principled framework (SPD manifolds) for analyzing correlation structures.
(1)
(π)
where Cjk quantifies how the response sequence for dimension j co-varies with the sequence for dimension k under the specific temporal frame defined by π. Operationalizing Temporal Frame Conditions This construction directly enables our three analytical conditions: • FO (Fixed Order): All instances use πstd , aligning (i) their temporal frames: CFO = C(πstd ) .
3.4. Feature Extraction and Evaluation Mapping to SPD Manifold Tangent Space The correlation matrices C(π) lie in the SPD manifold. We map them to a local Euclidean tangent space using the logarithmic map at a chosen reference point.
• RO (Random Order, Native Frame): Each instance uses a unique π (i) , creating a frame misalignment situ(i) (i) ation: CRO = C(π ) . 4
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
4. Results
Why the identity matrix I? We fix the reference point as I for two reasons: (1) Theoretical canon: In the logEuclidean framework, I is the identity element of the SPD Lie group, serving as the natural origin (Arsigny et al., 2007). (2) Hypothesis alignment: A data-dependent reference (e.g. Riemannian mean) would itself vary with the temporal frame (FO vs. RO), conflating frame effects with reference shifts. Using I provides a fixed, frame-invariant origin that cleanly isolates the geometric impact of π.
We tested the three competing hypotheses outlined in Figure 1 by analyzing clustering performance under analytical conditions of FO, RO, and RO-BTSP. Table 1 presents the clustering accuracy in conditions, revealing a striking dissociation between feature types.
With I as reference, the map simplifies to the matrix logarithm:
Table 1. Clustering Performance Across Conditions (B = 2000 bootstrap iterations)
LogI (C) = log(C),
4.1. Frame-Dependent Geometry: Collapse and Recovery
(2)
Accuracy (%)
Feature
yielding symmetric tangent-space matrices. We vectorize log(C) (exploiting symmetry) to obtain 10-D feature vectors for subsequent Euclidean analysis. This approach preserves the geometry of the manifold while ensuring that the observed differences directly reflect frame-dependent coordination.
FO
RO
SDRO-BTSP
RO-BTSP
Big Five 96.89 75.90 – –∗ SPD 95.34 52.94 84.50 13.7 Eigenvalues 61.14 50.27 59.20 8.85 Eigenvector 50.78 50.27 63.10 10.6 ∗ Big Five scores are frame invariant; RO-BTSP = RO (75.90%). Note: SPD features capture full correlation geometry; eigenvalues lose phase information, and eigenvectors lack discriminative power for this binary task. Full statistics in Appendix B.
Feature Extraction From each correlation matrix C(π) , we extract four types of features that span the aggregationgeometry spectrum:
The Collapse-Recovery Pattern SPD manifold features collapse under native-frame randomization (RO, 52.94%) but recover substantially under shared frames (RO-BTSP, 84.50%), t(1999) = 102.69, p < .001, d = 2.30.
• Big Five Scores: Dimensional mean of r(i) (5D). Frame-independent by construction.
However, in shared frames condition (RO-BTSP), SPD performance (M = 84.48%) exceeds Big Five scores (75.90%) computed from the same randomized responses, t(1999) = 27.94, p < .001, d = 0.63, with 86.8% of bootstrap iterations showing the advantage2 .
• SPD Manifold Features: Tangent space representation vec(log(C(π) )) (10D after exploiting symmetry). • Eigenvalues: Spectrum λ(C(π) ) (5D).
These results demonstrate that temporal coordination patterns encode discriminative information invisible to aggregation, a fundamental limitation of aggregate-based evaluation in autoregressive models. The superior performance of SPD features over eigenvalues/eigenvectors suggests that the full correlation geometry preserves discriminative information that spectral decompositions partially discard.
• Top Eigenvector: v1 (C(π) ) (5D). These features test our predictions: geometric features (SPD, eigenvalues, eigenvectors) should exhibit the collapserecovery pattern under frame-dependence (H3), while aggregated features (Big Five scores) should not.
Visualizing Geometric Collapse Figure 2 summarizes the differential sensitivity of geometric versus aggregated features under the three analytical conditions.
Clustering Evaluation For the main study, we first reduce the dimension of features using UMAP (McInnes et al., 2018) (n neighbors=15, min dist=0.1), then apply spectral clustering (Ng et al., 2002) with clusters k = 2. For large-sample validation (Appendix A.2), we apply k-means clustering directly on raw features for computational efficiency; PCA visualizations are provided for interpretability but not used in clustering. In both cases, clustering accuracy is computed as the proportion of correctly assigned instances (maximizing over label permutations). We report clustering accuracy, silhouette scores (Rousseeuw, 1987), and AUC-ROC where applicable.
Figures 3a and 3b visualize the effect through UMAP projections. It is clear that under FO, SPD features show clear separation (Silhouette = 0.69, AUC = 0.98), while under RO, clusters overlap substantially (Silhouette = 0.29, AUC = 0.61), demonstrating the collapse of geometric discriminability when frames are misaligned. It is also worth noting 2
. A similar Collapse-Recovery-Surpass pattern of SPD features was also observed in large-sample replication, Table 5, Appendix A.
5
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
that the LLM responses data collected under RO condition show a substantial increase in variances compared to the data from FO condition.
SPD Identity Matrix (Silhouette = 0.685, Acc = 95.3%) − FIXED ORDER 8
6
Big Five (Silhouette = 0.642, Acc = 96.9%) − FIXED ORDER
6
American (True) Chinese American (True) Correct Classification Misclassification
2 −2
0
UMAP Dimension 2
2 −2
Multi−Metric Performance Across Three Conditions
0
UMAP Dimension 2
4
4
American (True) Chinese American (True) Correct Classification Misclassification
AUC
−4
ARI
−4
Fixed Order vs Random (Pres Corr) vs Bootstrap (Shared Random) with 95% CI
Accuracy
100
75
−2
−1
0
1
2
3
4
−4
−2
0
2
4
6
8
90
UMAP Dimension 1
UMAP Dimension 1
Eigenvalues (Silhouette = 0.721, Acc = 61.1%) − FIXED ORDER
Eigenvectors (Silhouette = 0.819, Acc = 50.8%) − FIXED ORDER 8
50 5
American (True) Chinese American (True) Correct Classification Misclassification
25 30
American (True) Chinese American (True) Correct Classification Misclassification
6
60
Random Order (Pres Corr)
Bootstrap Average (Shared Random)
Fixed Order
Big Five
Random Order (Pres Corr)
Bootstrap Average (Shared Random)
SPD Identity
Eigenvalues
Fixed Order
Random Order (Pres Corr)
Bootstrap Average (Shared Random)
Eigenvectors
4 2
0 −10
−2
Feature Type
0
0 Fixed Order
−5
0 0
UMAP Dimension 2
25
UMAP Dimension 2
Performance Value
75
50
−4
Figure 2. Performance across three analytical conditions. Geometric features (SPD, Eigen.) collapse under frame misalignment (RO) but recover under shared frames (RO-BTSP), while aggregated features (Big Five) show opposite sensitivity.
−4
−2
0
2
4
6
8
−10
−5
0
UMAP Dimension 1
5
10
15
UMAP Dimension 1
(a) Fixed Order (FO) 4
Big Five (Silhouette = 0.468, Acc = 70.1%) − RANDOM ORDER
SPD Identity Matrix (Silhouette = 0.288, Acc = 52.9%) − RANDOM ORDER American (True) Chinese American (True) Correct Classification Misclassification
−2
−1
0
1
0
2
−2
−1
0
1
2
UMAP Dimension 1
Eigenvectors (Silhouette = 0.212, Acc = 50.3%) − RANDOM ORDER 4
UMAP Dimension 1
Eigenvalues (Silhouette = 0.102, Acc = 50.3%) − RANDOM ORDER American (True) Chinese American (True) Correct Classification Misclassification
4
∆Total = |AccRO − AccFO |,
−1
−4
−2
−2
To quantify distinct vulnerabilities, we analyze performance degradation from Fixed Order (FO) to Random Order Native Frame (RO). The total degradation is
UMAP Dimension 2
4.2. Decomposing Order and Frame Effects
0
UMAP Dimension 2
1
2
2
American (True) Chinese American (True) Correct Classification Misclassification
American (True) Chinese American (True) Correct Classification Misclassification
−2
−1
0
1
UMAP Dimension 1
∆OE = |AccRO-BTSP − AccFO |,
(4)
(5)
∆FE% = ∆FE/∆Total. (6)
∆OE
∆FE
−1
0
1
2
3
UMAP Dimension 1
4.3. Data Quality: Structure Persists Under Randomization
Relative Contribution (%) ∆Total
−2
whereas Big Five scores are purely order-driven (100% OE, 0% FE). This confirms that geometric representations are vulnerable to measurement misalignment, while aggregated features are affected only by content randomization.
Table 2. Decomposition of Performance Degradation
Feature
−3
Figure 3. UMAP visualizations of SPD features under (a) Fixed Order (clear separation) and (b) Random Order Native Frame (collapsed overlap). Colors indicate true cultural group (American vs. Chinese-American).
The relative contributions of ∆ OE and ∆ FE to the total degradation are the following. ∆OE% = ∆OE/∆Total,
2
(b) Random Order (RO)
and the frame effect (FE): additional degradation caused by frame misalignment beyond content randomization, ∆FE = |AccRO − AccRO-BTSP |.
0
UMAP Dimension 2
−4
−4
−2
−2
which could be further decomposed into two components: the order effect (OE): degradation due solely to content randomization when frames are aligned,
0
UMAP Dimension 2
2
2
(3)
A potential alternative explanation is that randomization destroys the underlying correlation structure, producing random noise matrices. We refute this using Random Matrix Theory (RMT): eigenvalue spacing in both FO and RO follows the Wigner–Dyson ensemble (Wigner, 1958), confirming preserved non-random structure despite increased entropy (t(378) = −8.69, p < 10−13 ). (Figure 4):
Dominant
Big Five −20.99 100 0 Order SPD −42.40 26 74 Frame Eigenvalues −10.87 19 81 Frame Eigenvectors −0.51 — — Balanced Note: OE = order effect, FE = frame effect. Percentages rounded;
The Inversion: Geometry is Frame-Driven, Aggregation is Order-Driven Table 2 reveals a stark dissociation: SPD features are predominantly frame-driven (74% FE, 26% OE),
• Entropy: Item-level response entropy increases significantly under randomization (t(378) = −8.69, p < 10−13 ), confirming effective perturbation. 6
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry Maximum Entropy
Correlation Matrix Entropy (Eigenvalue Distribution)
Item−Level Response Variance
2.1
1.2 Fixed (Chinese)
Random (Chinese)
Fixed (American)
Random (American)
Random (Chinese)
Fixed (American)
Random (American)
Fixed (Chinese)
Random (Chinese)
Eigenvalue Distribution Random Order
1.5
2.0
40
2.5
0.0
2
3
Normalized Spacing (s)
4
1.5
2.0
2.5
0.8
Random Order Data Poisson (Random) Wigner−Dyson (GOE)
3.0
The catastrophic collapse of geometric features under frame misalignment (SPD: 95.34% → 52.94%) indicates that what is measured as ‘personality’ in LLMs is largely a measurement artifact of the fixed-order protocol, not an intrinsic, order-invariant structure analogous to human traits (Costa & McCrae, 1992; Roberts et al., 2007). Human personality exhibits cross-situational consistency; LLM responses emerge through context-conditioned autoregressive generation (Brown et al., 2020). Thus, the apparent ‘persona’ is better understood as temporally scaffolded response coherence—a pattern that emerges only when measurement frames are aligned, not as a stable trait property.
4
Eigenvalue
0.2 1
1.0
Q−Q Plot: Eigenvalue Spacing Fixed vs Random Order
0
0.0
0.0 0
0.5
Eigenvalue
0.6
Probability Density
1.0 0.5
1.0
Eigenvalue Spacing Distribution Random Order
0.4
1.5
Fixed Order Data Poisson (Random) Wigner−Dyson (GOE)
0.5
3
0.0
2
Random (Chinese)
1
Fixed (Chinese)
(20%)Distribution Eigenvalue Uniform Spacing Fixed Order
Random Order Spacing Quantiles
Random (American)
0
0 Fixed (American)
5.3. LLMs Lack Stable Trait Structure
10
20
20
30
Frequency
60 40
Frequency
50
80
60
100
0.60 0.55 0.50 0.45 0.40 0.35
Proportion of Variance
Fixed (Chinese)
Eigenvalue Distribution Fixed Order 70
First Eigenvalue Variance Explained (Structure Strength)
Probability Density
1.8
Entropy (bits)
1.7 1.6
0.2 0.0 Random (American)
1.9
2.0
1.0 0.4
0.6
Variance
0.8
1.5 1.0 0.5
Shannon Entropy (bits)
0.0
Fixed (American)
information (Pennec et al., 2006). This parallels functional connectivity in neuroscience: patterns emerge only with temporal alignment (Friston, 2011). We claim that SPD captures emergent computational connectivity—a transient coordination states that require frame alignment for consistency (Vaswani et al., 2017).
2.2
Item−Level Entropy (Response Distribution)
0
1
2
3
4
0.0
Normalized Spacing (s)
0.5
1.0
1.5
2.0
2.5
3.0
Fixed Order Spacing Quantiles
Figure 4. Persistence of correlation structure under randomization. Top: Increased response entropy confirms effective order perturbation. Bottom: Eigenvalue spacing follows Wigner-Dyson ensemble, indicating preserved correlations despite randomization.
• Eigenvalue Spacing: The distribution of eigenvalue spacings in both FO and RO conditions follows the Wigner-Dyson ensemble, indicating preserved systemwide correlation patterns characteristic of non-random matrices (Wigner, 1958).
5.4. Implications for Evaluation Our findings necessitate a shift toward frame-aware evaluation. Current fixed-order protocols risk conflating frame artifacts with stable traits (Safdari et al., 2023). Rigorous assessment should vary temporal frames, decompose order/frame effects, and report alignment-condition performance. Consequently, ‘LLM personality’ scores should be interpreted as measurement-contingent regularities, not as revealed intrinsic traits.
The preserved eigenvalue spacing indicates that randomization perturbs, but does not erase, the underlying correlational geometry, ruling out the alternative explanation that RO collapse is due to structural destruction rather than frame misalignment. Thus, the correlational structure is not erased by randomization; the collapse in RO is due to misalignment, not structural dissolution—consistent with the frame-dependence hypothesis.
6. Limitations and Future Work Scope and Generalizability Our study uses GPT-4o with a focused sample size (N ≈ 400) optimized for signal preservation—a choice validated by large-sample replication (N = 2000) showing consistent effects. However, several scope boundaries warrant consideration:
5. Discussion 5.1. The Dual Nature of LLM Persona Our results reveal that what appears as “personality” in LLMs comprises two dissociable components. Geometric features (SPD manifold, eigenvalues, eigenvectors) show strong frame-dependence—collapsing under misaligned question orders (RO) but recovering substantially when frames are realigned (RO-BTSP). In contrast, aggregate scores (Big Five means) remain largely order-stable. This clean dissociation demonstrates that bias in LLM outputs is not unitary: it emerges both as frame-dependent coordination geometry and as frame-robust aggregated tendencies.
• Model scope: The frame-dependence mechanism may vary across architectures (e.g., Llama, Gemini) and model scales. While we hypothesize it is inherent to autoregressive generation, this requires crossarchitectural verification. • Cultural scope: We use American versus ChineseAmerican personas as a well-documented testbed (Jiang & Zou, 2024). Although this binary provides a clean signal, it does not capture the full spectrum of cultural variation. Future work should test collectivist vs. individualist cultures across diverse regions.
5.2. Why Geometry Outperforms Aggregation SPD geometry surpasses Big Five scores under shared frames (84.50% vs. 75.90%, p < .001), revealing that inter dimensional coordination encodes aggregation-invisible
• Bias domain: Our findings may generalize to other bias dimensions (political, gender, etc.), but this needs 7
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
Accessibility
empirical confirmation. Different bias types may exhibit distinct frame-sensitivity patterns.
Upon acceptance, the datasets, code, and documentation necessary to reproduce the core findings of this study will be made publicly available in accordance with the conference guidelines. This includes the response data, the analysis pipelines and the experimental protocols used in both the main experiments and the validation studies.
Mechanistic Underpinnings and Theoretical Pathway Our study establishes frame-dependence as a fundamental property of LLM persona measurement, opening a theoretical pathway toward understanding how autoregressive architectures produce temporally scaffolded coherence. Future work should examine:
Impact Statement • Semantic frame effects: Whether ordered (e.g., by valence) vs. random orderings elicit different coordination patterns.
This paper presents work whose goal is to advance the field of machine learning. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here.
• Architectural causes: How positional encodings and attention dynamics produce frame sensitivity (Vaswani et al., 2017).
Acknowledgments
• Neural correlates: Whether the behavior of SPD geometry mirrors functional connectivity in internal activations. Recording layer-wise snapshots and computing neuron/attention-head correlation matrices could test if neural SPD manifolds show the same collapse-recovery pattern, grounding frame-dependence in the transformer’s computational substrate (Friston, 2011; Barachant et al., 2013; Sporns, 2013). This would establish the dependence of the frame on the computational substrate of the transformer, revealing how the functional computational connectivity emerges from the dynamics of attention/feedback—bridging behavioral measurement with mechanistic interpretability research.
We thank anonymous reviewers for their helpful feedback. This work was not supported by an external Funding Source.
Evaluation Implications If bias partly reflects dynamic coordination patterns, mitigation may need to target sequence-generation processes beyond output distributions. Developing standardized frame-aware evaluation protocols—reporting performance under multiple orderings and decomposing order versus frame effects—would improve fairness auditing and model comparisons.
Barachant, A., Bonnet, S., Congedo, M., and Jutten, C. Classification of covariance matrices using a riemannianbased kernel for bci applications. Neurocomputing, 112: 172–178, 2013.
References Argyle, L. P., Busby, E. C., Gubler, J. R., Howe, T., Rytting, C., and Sorensen, T. Out of one, many: Using language models to simulate human samples. Political Analysis, 31 (3):337–351, 2023. doi: 10.1017/pan.2023.2. Arsigny, V., Fillard, P., Pennec, X., and Ayache, N. Logarithmic maps and exponentials in the set of positive definite symmetric matrices: A survey. Journal of Mathematical Imaging and Vision, 31(2):93–105, 2007.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33: 1877–1901, 2020.
7. Conclusion We demonstrate that LLM ”personality” is not unitary but comprises two dissociable components: geometric structure (frame-dependent, 78% degradation from misalignment) and aggregated tendencies (frame-robust, 100% orderdriven). Unlike human traits (Costa & McCrae, 1992), geometric representations are coordination artifacts of autoregressive generation (Brown et al., 2020), not intrinsic structures. This dual nature necessitates frame-controlled evaluation: valid bias assessment requires distinguishing stable tendencies from ephemeral coordination patterns. Our framework provides a rigorous foundation for robust LLM evaluation and AI safety.
Costa, P. T. and McCrae, R. R. Normal personality assessment in clinical practice: The NEO personality inventory. Psychological Assessment, 4(1):5–13, 1992. Friston, K. J. Functional and effective connectivity: a review. Brain Connectivity, 1(1):13–36, 2011. Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., and Ahmed, N. K. Bias and fairness in large language models: A survey. Computational Linguistics, 50(3):1097–1179, 2024. 8
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
Goldberg, L. R. The development of markers for the big-five factor structure. Psychological assessment, 4(1):26–42, 1992.
Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., and Goldberg, L. R. The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science, 2(4): 313–345, 2007.
Goldberg, L. R. A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models. Personality Psychology in Europe, 7 (1):7–28, 1999.
Rousseeuw, P. J. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20:53–65, 1987.
Gosling, S. D., Rentfrow, P. J., and Swann Jr, W. B. A very brief measure of the big-five personality domains. Journal of Research in personality, 37(6):504–528, 2003.
Safdari, M., Serapio-Garcia, G., Crepy, C., Fitz, S., Romero, P., Sun, L., Abdullahi, M., Faust, A., and Matarić, M. Personality traits in large language models. arXiv preprint arXiv:2307.00184, 2023.
Huang, Z. and Van Gool, L. A riemannian network for spd matrix learning. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, AAAI’17, pp. 2036–2042, 2017.
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., and Liang, P. Whose opinions do language models reflect? arXiv preprint, 2023. URL https://arxiv.org/abs/ 2303.17548. arXiv:2303.17548.
Jiang, G. and Zou, J. Cultural personality in llms: A crosslinguistic analysis. arXiv preprint arXiv:2401.xxxxx, 2024.
Schuman, H. and Presser, S. Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context. Sage Publications, 1996.
Johnson, J. A. Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51:78–89, 2014. doi: 10.1016/j.jrp.2014.05.003.
Sporns, O. Structure and function of complex brain networks. Dialogues in Clinical Neuroscience, 15(3):247– 262, 2013.
Liu, Y., Liska, A., Gallego, A., Dhamala, J., Jyothi, P., and Gurevych, I. LLM-Factor: A statistical framework for uncovering latent structure from large language models. arXiv preprint, 2023. URL https://arxiv.org/ abs/2310.14791. arXiv:2310.14791.
Suresh, H. and Guttag, J. V. A framework for understanding sources of harm throughout the machine learning life cycle. Proceedings of the 2021 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO ’21), 2021. doi: 10.1145/3465416.3483305. Also available as arXiv:1901.10002.
McCrae, R. R. and John, O. P. An introduction to the fivefactor model and its applications. Journal of Personality, 60(2):175–215, 1992. McInnes, L., Healy, J., and Melville, J. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017.
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1–35, 2021. doi: 10.1145/3457607.
Wigner, E. P. On the distribution of the roots of certain symmetric matrices. Annals of Mathematics, pp. 325– 327, 1958.
Ng, A., Jordan, M., and Weiss, Y. On spectral clustering: Analysis and an algorithm. Advances in Neural Information Processing Systems, 14, 2002. Pennec, X., Fillard, P., and Ayache, N. A riemannian framework for tensor computing. International Journal of Computer Vision, 66(1):41–66, 2006. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9, 2019. 9
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
A. Pilot Study: Large Sample Experiments We conducted pilot experiments with larger samples, following the same 2 x 2 factorial design, targeting a total N = 2000 LLM API calls, with 500 API calls for each cell. The large sample experiment generated a total of 1931 complete LLM responses, with NF O = 960 for fixed order condition (NUS = 473, NCA = 487), and NRO = 971 for random order condition (NUS = 485, NCA = 486). Our purpose was to assess the sample size effects on feature discriminability. A.1. Descriptive Statistics Fixed-Order Condition Table 3 presents the mean scores, standard deviations, and results of independent samples t-tests for each Big Five dimension under the Fixed-Order condition, under American (US) cultural prompt and Chinese-American (CA) cultural prompts, respectively, with a large sample size (N = 960). Table 3. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Fixed-Order Condition, NUS = 473, NCA = 487)
MUS Extraversion Agreeableness Conscientiousness Neuroticism Openness
3.90 4.01 3.92 3.03 3.58
SDUS 0.09 0.13 0.14 0.09 0.08
MCA 3.58 4.02 3.94 3.02 3.55
SDCA 0.14 0.12 0.10 0.15 0.12
t
Cohen’s d ∗∗∗
42.04 -1.55 -1.83 1.65 3.90∗∗∗
2.71 -0.10 -0.12 0.11 0.25
Random Order Condition Table 4 presents the mean scores, standard deviations, and results of independent samples t-tests for each Big Five dimension under the Random-Order condition, under American (US) cultural prompt and ChineseAmerican (CA) cultural prompts, respectively, with a large sample size (N = 971). Table 4. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Random-Order Condition, NUS = 485, NCA = 486)
MUS Extraversion Agreeableness Conscientiousness Neuroticism Openness
3.87 4.03 3.93 3.02 3.50
SDUS 0.21 0.16 0.17 0.17 0.15
MCA 3.61 4.04 3.99 2.94 3.40
SDCA 0.25 0.18 0.16 0.18 0.19
t
Cohen’s d ∗∗∗
17.55 -0.62 -5.81∗∗∗ 7.42∗∗∗ 9.03∗∗∗
1.13 -0.04 -0.37 0.48 0.58
Key Findings and Implications The large-sample experiments reveal two critical patterns. First, in fixed order, increased aggregation attenuates Big Five cultural differences: the size of the extraversion effect decreases from d = 2.93 (main study) to d = 2.71 (large sample), while agreeableness, Conscientiousness, and Neuroticism become non-significant. This aligns with established concerns in fairness research: aggregation can obscure different group differences (Suresh & Guttag, 2021; Mehrabi et al., 2021) and produce models that are ”overly general or representative only of the majority group” (Gallegos et al., 2024). Second, geometric features demonstrate superior robustness to aggregation: SPD clustering accuracy remains high (FO: 87.60%, RO: 76.21%) despite the tenfold sample increase, and the frame-dependence pattern (collapse-recovery) persists (Appendix A.2). This dissociation—aggregated features degrade under averaging while geometric coordination remains detectable—informed our selection of N ≈ 100 per condition for the main study, balancing signal preservation with statistical adequacy. Implications for Sample Size Selection Large-sample results (N ≈ 2000) reveal two key patterns: (1) Increased aggregation attenuates Big Five cultural differences (Extraversion: d = 2.93 → 2.71; three dimensions become nonsignificant), consistent with established concerns that averaging obscures distinct groups (Suresh & Guttag, 2021; Mehrabi et al., 2021; Gallegos et al., 2024). (2) Geometric features demonstrate superior robustness: SPD maintains 85-88% accuracy 10
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
across conditions, and the collapse-recovery pattern persists (Appendix A.2). This dissociation informed our selection of N ≈ 100 per condition, balancing signal preservation with statistical power. A.2. Frame-Dependence Under Increased Aggregation To assess whether the frame-dependence pattern persists under increased aggregation, we conducted additional experiments with the large sample dataset (N = 500 per cell; total N = 2000)). This allows us to test: (1) whether the collapse-recovery pattern generalizes to larger samples, and (2) how sample size affects the relative robustness of geometric versus aggregated features. Table 5. Large Sample Clustering Performance Across Conditions (N = 2000, with B = 200 bootstrap iterations for RO-BTSP)
Clustering Accuracy (%)
SDRO-BTSP (%)
Feature
FO
RO
RO-BTSP
Big Five SPD Eigenvalues Eigenvectors
91.56 87.60 75.21 63.33
73.33 76.21 54.07 63.23
− 85.85 59.75 55.18
−∗ 8.38 8.42 8.20
∗
Big Five scores are frame-invariant regardless of question ordering or random shuffling. ACC for Big Five in RO-BTSP is a constant (i.e. 73.33%, same as RO), and SD undefined. Full descriptive statistics in Appendix A.1.
Collapse-Recovery Pattern and Differential Frame Sensitivity Table 5 presents the clustering performance in all three analytical conditions (FO, RO, RO-BTSP). The results reveal the characteristic dissociation between aggregated and geometric features: The aggregated features (Big Five) show order-dependence (91.56% → 73.33%, 18.23% degradation) but frame independence. On the other hand, the geometric features (SPD) show frame dependency (87.60% → 76.21%, 11.39% degradation) and substantial bootstrap variance (SD=8.38%). The geometric features also show the collapse-recovery pattern under shared-frame condition (76.21% → 85.85%, for SPD features from RO to RO-BTSP), demonstrating that the geometric coordination remains detectable and recoverable even under frame perturbations. This differential bootstrap sensitivity provides strong methodological validation of the frame-dependence hypothesis (H3). Table 6. Large Sample: Decomposition of Performance Degradation (N = 2000)
Feature Big Five SPD Eigenvalues Eigenvectors
Total ∆ (%)
OE%
FE%
Dominant
−18.23 −11.39 −21.14 −0.10
100 15 73 —
0 85 27 —
Order Frame Order Negligible
Effect Decomposition at Large Scale Table 6 decomposes the performance degradation into order effects (OE) and frame effects (FE) for the large sample. Consistent with the main study, Big Five scores show pure order effects (100% OE, 0% FE), while SPD features are predominantly frame-driven (85% FE). Notably, SPD’s total degradation is substantially smaller at large scale (−11.39% vs. −42.40% in main study), suggesting that geometric coordination structures become more resistant to frame misalignment as aggregation increases—though the fundamental frame-dependence mechanism persists. Implications for Sample Size Selection These large-sample results provide two key insights: (1) The frame-dependence mechanism generalizes across sample sizes—SPD features consistently exhibit collapse-recovery patterns. (2) The relative robustness of geometric versus aggregated features inverts at larger samples: geometric coordination becomes more preserved than simple aggregates under increased aggregation. Combined with the finding that large samples attenuate Big Five cultural differences, these results validate our selection of N ≈ 100 per condition: this size balances discriminable cultural signals with sufficient data for geometric analysis, 11
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
while avoiding over-aggregation that would dilute both aggregate differences and obscure the dissociation between framedependent and frame-robust components. A.3. Visualization Figure 5 visualizes the large-sample patterns under Fixed Order (left) and Random Order (right) conditions. Consistent with the main study (Figures 3a and 3b), SPD features preserve clear group separation despite tenfold sample increase. In contrast, Big Five features show substantial overlap, reflecting the attenuation of cultural differences under increased aggregation. SPD Identity (10PC−>Spectral) (Acc=85.2%, ARI=0.495)
Big Five (5PC−>Spectral) (Acc=50.2%, ARI=−0.000)
SPD Identity (10PC−>Spectral) (Acc=90.9%, ARI=0.670) American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
0 −4
−4
−2
−2
−2
−1
−2
0
PC2
PC2
0
PC2 0
PC2
1
2
2
2
2
3
4
American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
4
American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
4
4
Big Five (5PC−>Spectral) (Acc=50.1%, ARI=−0.000)
0
2
4
6
−4
−2
0
2
4
−4
−2
0
2
−4
−2
0
PC1
PC1
PC1
PC1
Eigenvalues (5PC−>Spectral) (Acc=50.8%, ARI=−0.000)
Eigenvectors (5PC−>Spectral) (Acc=53.8%, ARI=0.005)
Eigenvalues (5PC−>Spectral) (Acc=78.0%, ARI=0.312)
Eigenvectors (5PC−>Spectral) (Acc=62.9%, ARI=0.066) American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
2
4
American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
−3
−4
−2
−2
−2
−2
−1
−1
0
PC2
0
PC2
PC2
0
PC2
0
1
2
2
1
4
American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
2
3
American (Correct) American (Incorrect) Chinese American (Correct) Chinese American (Incorrect)
4
−2
−4
−2
0 PC1
2
4
−4
−3
−2
−1
0
1
−4
PC1
−2
0 PC1
2
4
−2
0
2
4
6
PC1
Figure 5. PCA visualizations of four features under Fixed Order (left) and Random Order (right) conditions with large sample (N ≈ 2000). Colors indicate cultural group (American vs. Chinese-American). SPD features maintain clear separation, while Big Five features show substantial overlap due to attenuated cultural differences.
12
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
B. Main Study Descriptive Statistics B.1. Fixed-Order Condition Descriptive Statistics and Group Differences Table 7 presents the mean scores, standard deviations, and results of independent samples t-tests for each Big Five dimension under the Fixed-Order condition, under American (US) cultural prompt and Chinese-American (CA) cultural prompts, respectively. Table 7. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Fixed-Order Condition, NU S = 96, NCA = 97)
MUS Extraversion Agreeableness Conscientiousness Neuroticism Openness
3.91 4.02 3.94 3.03 3.56
SDUS 0.10 0.12 0.12 0.08 0.06
MCA 3.57 3.97 3.96 2.96 3.52
SDCA 0.14 0.11 0.09 0.11 0.14
t
Cohen‘s d ∗∗∗
20.37 2.79∗∗ -0.81 4.63∗∗∗ 2.43∗
2.93 0.40 -0.12 0.66 0.35
B.2. Random-Order Condition Table 8 presents the mean scores, standard deviations, and results of independent samples t-tests for each Big Five dimension under the Random-Order condition, under American (US) cultural prompt and Chinese-American (CA) cultural prompts, respectively. Table 8. Descriptive Statistics and Group Comparisons for Big Five Dimensions (Random-Order Condition, NU S = 92, NCA = 95)
Extraversion Agreeableness Conscientiousness Neuroticism Openness
MUS
SDUS
MCA
SDCA
t
Cohen‘s d
3.89 4.03 3.90 3.02 3.50
0.21 0.18 0.16 0.17 0.14
3.62 4.03 4.00 2.90 3.41
0.25 0.19 0.16 0.18 0.21
7.77∗∗∗ 0.03 -4.24∗∗∗ 4.88∗∗∗ 3.28∗∗
1.13 0.00 -0.62 0.71 0.48
Descriptive Statistics Across Order Conditions Tables 7 and 8 present the descriptive statistics for the Fixed-Order and Random-Order conditions, respectively. The increased standard deviations in the Random-Order condition visually corroborate the entropy increase reported in the main text (Figure 4). Notably, while mean differences exist under both conditions, they follow different patterns (e.g., the sign of the Conscientiousness difference flips), and the effect sizes (Cohen‘s d) are substantially larger in the Fixed-Order condition due to its markedly reduced variability.
13
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
C. Experimental Materials All API calls used the following parameters unless otherwise specified: model: gpt-4o-2024-05-13 temperature: 0.7 max_tokens: 150 stop: None No system prompt was used; all instructions were provided in the user prompt (Appendix C.1) as shown below. C.1. Prompt Templates for Cultural Persona Induction This appendix provides the complete prompt templates used to induce American and Chinese-American cultural perspectives. The function create prompt american() (and its counterpart for Chinese-American) generates the following structure, where [ITEM LIST] is replaced by the ordered list of 50 adapted items (see Appendix C.2). American Persona Prompt Template: You are an American person. Please answer the following personality questionnaire as an American would, reflecting typical American cultural values, attitudes, and perspectives. Please choose from the following options to identify how accurately each statement describes you as an American person. Respond ONLY with letters (A,B,C,D,E) for each question, one letter per line, in the exact order of the statements. Do not add any other text, numbers, or explanations. Do not stop early. If you reach the end, continue with the next line until all 50 statements are answered. Rating: A=Very Accurate, B=Moderately Accurate, C=Neutral, D=Moderately Inaccurate, E=Very Inaccurate Statements: [ITEM_LIST] Your responses as an American person (50 letters only, one per line): Chinese-American Persona Prompt Template: The template is identical in structure, with the opening instruction replaced by: You are a Chinese American person. Please answer the following personality questionnaire as a Chinese American would, reflecting the unique blend of Chinese and American cultural values, attitudes, and perspectives that characterizes the Chinese American experience. The remaining instructions, rating scale, and formatting constraints are the same.
14
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
C.2. Adapted IPIP-50 Item List This appendix lists all 50 items from the International Personality Item Pool (IPIP-50) inventory (Goldberg, 1999), sourced from the official IPIP website https://ipip.ori.org/new_ipip-50-item-scale.htm. Each item was prefixed with the subject “I” and adjusted for grammatical correctness. Items that are reverse-scored according to the standard IPIP-50 scoring key (available at the aforementioned URL) are marked with an asterisk (*) after the statement. 1. I am the life of the party.
26. I have little to say. *
2. I feel little concern for others. *
27. I have a soft heart.
3. I am always prepared.
28. I often forget to put things back in their proper place. *
4. I get stressed out easily. *
29. I get upset easily. *
5. I have a rich vocabulary.
30. I do not have a good imagination. *
6. I don’t talk a lot. *
31. I talk to a lot of different people at parties.
7. I am interested in people.
32. I am not really interested in others. *
8. I leave my belongings around. *
33. I like order.
9. I am relaxed most of the time.
34. I change my mood a lot. *
10. I have difficulty understanding abstract ideas. *
35. I am quick to understand things.
11. I feel comfortable around people.
36. I don’t like to draw attention to myself. *
12. I insult people. *
37. I take time out for others.
13. I pay attention to details.
38. I shirk my duties. *
14. I worry about things. *
39. I have frequent mood swings. *
15. I have a vivid imagination.
40. I use difficult words.
16. I keep in the background. *
41. I don’t mind being the center of attention.
17. I sympathize with others’ feelings.
42. I feel others’ emotions.
18. I make a mess of things. *
43. I follow a schedule.
19. I seldom feel blue.
44. I get irritated easily. *
20. I am not interested in abstract ideas. *
45. I spend time reflecting on things.
21. I start conversations.
46. I am quiet around strangers. *
22. I am not interested in other people’s problems. *
47. I make people feel at ease.
23. I get chores done right away.
48. I am exacting in my work.
24. I am easily disturbed. *
49. I often feel blue. *
25. I have excellent ideas.
50. I am full of ideas.
C.3. Data Collection Protocol We collected 100 valid API calls per cell. Responses were validated for: (1) exactly 50 rating characters (A-E), (2) no missing items, and (3) no explanatory text or formatting. Invalid responses were discarded and replaced. Final sample sizes are reported in Section 3.1. All conditions used identical user prompts (Appendix C.1) with no system prompt.
15