Linguistic Monoculture in LLM-Assisted Language Use Suhas Thejaswi§ , Juhi Kulshreshta§ , and Lutz Oettershagen: §
Aalto University, Finland [email protected] :
University of Liverpool, United Kingdom [email protected]
arXiv:2607.27134v1 [cs.AI] 29 Jul 2026
Abstract Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and co-evolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author–model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.
Keywords:
1
Algorithmic monoculture, Linguistic monoculture, Price of Monoculture
Introduction
Large language models (LLMs) are increasingly used to draft, revise, and polish text. Although such assistance can improve clarity, reduce errors, and help authors meet institutional expectations, widespread reliance on shared models may pull users toward similar model-mediated linguistic patterns. Moreover, LLMgenerated or LLM-assisted text may enter future training and personalization pipelines, creating a coupled population-level system in which authors adapt to model outputs and models may in turn adapt to language already influenced by earlier models. Repeated interaction with shared generative systems may therefore reduce variation in language through which people develop, communicate and distinguish ideas [32, 36]. We call this reduction in population-level variation across lexical choices, syntactic constructions, discourse markers, and other stylistic patterns as linguistic monoculture. Recent empirical work provides evidence that LLM assistance can homogenize human text, ideas, expressions, and stylistic choices across users [3, 13, 23, 31, 32, 37]. Large-scale studies have also documented shifts in word frequencies and LLMassociated stylistic markers in scientific articles and abstracts [17, 18, 26, 28]. This concern is especially 1
consequential in academia, where scientific writing helps make concepts legible, establish accepted methods, and shape what intellectual communities recognize as a contribution [4, 22, 27, 38]. We therefore ask: Under what forms of repeated author–model interaction does LLM assistance drive a population toward a shared linguistic norm, and when can heterogeneous preferences or personalization preserve diversity? Our analysis focuses on diversity in linguistic-feature distributions rather than convergence in semantic content, reasoning, or intellectual perspective. Linguistic variation is nevertheless a distinct and measurable dimension of expression through which authors establish voice, structure arguments and make distinctions legible to others [14, 15, 20, 30]. At the same time, not all convergence is undesirable: shared conventions can improve clarity, reduce errors, and make writing easier to understand and evaluate [7, 22]. The relevant question is therefore when shared AI assistance produces useful standardization and when it causes an excessive loss of population-level linguistic diversity. To answer this question, we develop a reduced-form framework representing authors and prompt-conditioned LLM output distributions over linguistic features. It isolates how sharing, deployment-level feedback, and personalization affect population-level diversity, measured by average pairwise Jensen–Shannon (JS) divergence between author distributions. We examine three author–LLM interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model updated recursively from author outputs, and personalized models updated through author-specific and population-level feedback. These mechanisms isolate whether the model is shared, whether it is recursively updated, and whether its feedback is personalized. We then endogenize conformity to a shared linguistic norm as a strategic choice and compare individually optimal conformity with the socially optimal level and quantify the loss of linguistic diversity through the price of monoculture. Our analysis of interaction dynamics initially treats conformity as exogenous. In practice, however, adopting a shared linguistic norm may be a strategic choice: standardized language can improve clarity, legibility and perceived fluency, and may be rewarded by reviewers, readers, or institutions [5, 19]. Yet such private benefits may come at the cost of an author’s distinctiveness preferences and reduce population-level linguistic diversity. We therefore compare individually optimal conformity with the socially optimal level. In detail, our contributions are as follows:1 1. A reduced-form framework for LLM-mediated dynamics. We formalize authors and promptconditioned LLM output distributions over linguistic features under three deployment mechanisms: a fixed shared system, a recursively updated shared system, and recursively updated personalized systems (Section 2). 2. Linguistic convergence and determinants of diversity. We establish convergence and convergencerate bounds for all three mechanisms and characterize how conformity pressure, author-specific preferences, recursive feedback, and personalization determine long-run population-level diversity. Shared models can drive diversity to low levels, whereas author-specific preferences and personalization can preserve diversity (Section 3). 3. Strategic conformity and the price of monoculture. We show that individually rational conformity can exceed the social optimum because authors may not internalize the value their distinctiveness provides to others; we bound the price of monoculture in a symmetric regime and identify when it diverges (Section 4). 4. Quantitative comparison of interaction mechanisms. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes across controlled parameter regimes. Paired runs on identical initializations enable controlled comparison of the interaction mechanisms (Section 5 and Appendix F). 1 All proofs and omitted details are in the Appendix.
2
Related work. Our work connects research on LLM-induced homogenization, opinion dynamics, recursive model training, algorithmic monoculture, and communication accommodation. Building on classical averaging models, we couple author adaptation to shared or personalized models updated from author outputs, and analyze linguistic-diversity equilibria and strategic welfare. LLM-induced homogenization. Empirical studies find that LLM assistance can homogenize written expression, ideas, authorial voice, and cultural style across users [1–3, 31, 32]. Analyses of academic writing and presentations similarly document increased use of LLM-associated linguistic markers [17, 18, 26, 28]. These studies motivate our problem; we complement them by formalizing how homogenization can emerge through repeated author–LLM interaction. Opinion dynamics. Our work builds on classical averaging models [12, 16, 21], using them as a reduced-form representation of an algorithmic intermediary whose output distribution may be updated from the population it influences. This introduces two deployment choices absent from human-only models: whether the intermediary is shared or personalized, and whether author outputs affect its future behavior. These choices determine whether the system converges to one shared norm or to a family of author-specific equilibria. Model collapse, knowledge collapse, and recursive feedback. Recursive training on model-generated data can degrade generative models and reduce coverage of the tails of the original data distribution [35]. Related work studies how AI-mediated information environments may narrow the range of available knowledge and loss of tails [34, 39]. This literature centers on the quality or epistemic diversity of model outputs, whereas we study the linguistic diversity of human authors who adapt to and may shape a shared model. Algorithmic monoculture. Kleinberg and Raghavan [24] show that reliance on common algorithmic advice can be individually beneficial while reducing social welfare. Recent work by Kleinberg et al. [25] quantifies this loss in matching markets, proving a tight price-of-anarchy bound of factor 2. Our work studies a linguistic analogue: shared LLM assistance may improve individual legibility while reducing population-level linguistic diversity. Communication accommodation and personalization. Communication accommodation theory studies how speakers adjust their language toward or away from interlocutors [5, 19], including in human–LLM interaction [9]. Related language-game models examine how shared conventions emerge through peer-to-peer interaction [10, 11]. In contrast, we study an LLM as a shared or personalized algorithmic interlocutor and characterize when personalization can preserve diversity across authors.
2
A Mathematical Framework for LLM-Assisted Language Use
řm We use rjs to denote t1, . . . , ju for j P N, and ∆m´1 :“ tx P Rm ě0 : k“1 xk “ 1u to denote an pm ´ 1qdimensional probability simplex. Let F “ rms be a common set of linguistic features encoding lexical choices, syntactic constructions, discourse markers or other stylistic patterns. We represent linguistic style as a probability distribution over F . For each author i P rns and time t P t0, . . . , T u, let pti P ∆m´1 denote author i’s linguistic-style distribution, with initial distribution p0i . Let q t P ∆m´1 denote the prompt-conditioned linguistic-style distribution induced by an LLM system at time t under a fixed class (or distribution) of writing prompts; if personalized to author i, we write qit . This is an output-level abstraction rather than a model-parameter representation. For notational uniformity, in shared-model settings we write qit :“ q t for every author i. At each time step t P t0, . . . , T ´ 1u, author i prompts the LLM, observes LLM-generated linguistic suggestions and updates their own style as pt`1 “ looooomooooon p1 ´ αi q pti ` i Style retention
t t α i Ai pqi , zi q loooooomoooooon
,
(1)
Model- and context-driven update
where αi P r0, 1s measures the rate at which author i adapts their linguistic style. The operator Ai : ∆m´1 ˆ Z Ñ ∆m´1 describes how the author incorporates the model’s output given an author-specific adaptation input zit P Z, which may encode external incentives, preferred style, or other contextual information. 3
The induced output distribution may itself evolve when author-produced text is incorporated into future training data, a recursive feedback mechanism studied in the model-collapse literature [35]. We model this in reduced form as t t`1 qit`1 “ looβq p1 ´ βq Bppt`1 (2) 1 , . . . , p n q, moion ` looooooooooooooomooooooooooooooon Model retention
Feedback from author outputs
where β P r0, 1s is the retention parameter: larger β places more weight on the past model distribution. The operator B maps the updated author distributions to the distribution of training data update řnused tot`1 t`1 t`1 the model; we instantiate this to be the weighted population average Bpp , . . . , p q “ w p with i n 1 i i“1 ř wi ě 0, i wi “ 1. Our main object of interest is population-level linguistic diversity, measured by the average pairwise Jensen–Shannon divergence [29] among author distributions: Dt “
ÿ 1 JSppti , ptj q, npn ´ 1q i‰j
where JSpp, qq “ 12 KLpp}xq` 21 KLpq}xq with x “ p`q 2 ; smaller values indicate convergence toward a common linguistic norm. When the model distribution evolves, we also track the average author–model divergence n
Mt “
1 ÿ JSppti , qit q, n i“1
which measures the alignment between each author and the model they interact with; in shared-model settings, M t Ñ 0 corresponds to a shared human-LLM linguistic equilibrium. Finally, when each author interacts with a personalized model, we track diversity among the personalized models, Qt “
ÿ 1 JSpqit , qjt q, npn ´ 1q i‰j
where Qt Ñ 0 means all personalized model distributions become indistinguishable. In the personalized setting, Dt measures diversity among authors, Qt diversity among personalized models and M t alignment between each author and their model. Since Jensen-Shannon divergence is symmetric and finite for probability distributions, it is a natural choice for all three measures. Remark A.1 in Appendix A justifies our choice to represent linguistic style as a distribution rather than a point in a high-dimensional feature space.
2.1
Author-LLM interaction mechanisms
We formalize three interaction mechanisms (IMs) that introduce increasing degrees of author–model coupling. Interaction Mechanism 1 (Shared model with fixed distribution). All authors interact with the same model, whose linguistic-style distribution remains fixed over time, i.e., q t “ q 0 for all t P t0, . . . , T u. The linguistic-style distribution of author i evolves as pt`1 “ p1 ´ αi q pti ` αi Ai pq 0 , zit q. i Interaction Mechanism 2 (Shared model with recursive updates). All authors interact with the same model, but the model distribution evolves over time. The coupled author–model dynamics evolve as pt`1 “ p1 ´ αi q pti ` αi Ai pq t , zit q, i t`1 q t`1 “ β q t ` p1 ´ βq Bppt`1 1 , . . . , pn q, t`1 with Bppt`1 1 , . . . , pn q “
řn
řn t`1 as the weighted-average instantiation, where wj ě 0 and j“1 wj “ 1. j“1 wj pj 4
Interaction Mechanism 3 (Personalized models with recursive updates). Each author i interacts with an author-specific personalized model with linguistic-style distribution qit , initialized with a common base distribution qi0 “ q 0 for all i P rns. The coupled author–model dynamics evolve as pt`1 “ p1 ´ αi q pti ` αi Ai pqit , zit q, i t`1 qit`1 “ β qit ` p1 ´ βq Bi ppt`1 1 , . . . , pn q,
where the feedback operator Bi mixes author i’s own updated distribution with the population-level update: t`1 t`1 Bi ppt`1 ` p1 ´ ρq 1 , . . . , pn q “ ρ p i
n ÿ
wj pt`1 j ,
j“1
with ρ P r0, 1s, wj ě 0,
řn
j“1 wj “ 1. Equivalently,
qit`1 “ β qit ` γ pt`1 `δ i
n ÿ
wj pt`1 j ,
j“1
where γ “ p1 ´ βqρ, δ “ p1 ´ βqp1 ´ ρq, and β ` γ ` δ “ 1. The parameter ρ determines the degree of personalization: when ρ “ 1, the model update depends only on author i’s own distribution; when ρ “ 0, all personalized models receive the same population-level feedback. In Appendix E, we investigate Interaction Mechanism 4, which partitions a single population into three subpopulations, each following the update rule of one of IMs 1-3.
3
Convergence of Linguistic Distributions
We analyze the three interaction mechanisms in increasing order of coupling, establishing convergence as well as convergence-rate bounds.
3.1
A shared model with fixed distribution
In Mechanism 1 the model distribution is fixed, so q t “ q 0 for all t and author i’s linguistic-style distribution evolves as pt`1 “ p1 ´ αi q pti ` αi Ai pq 0 , zit q. First, we consider the model-induced adaptation to be i homogeneous across authors and time, Ai pq 0 , zit`1 q “ Ai pq 0 , zit q “ Apq 0 , zq for all i and t, so all authors are influenced by the same fixed model distribution and have the same adaptation input. Proposition 3.1. In Mechanism 1, suppose the model-induced adaptation is homogeneous across authors and time, i.e., Ai pq 0 , zit q “ Apq 0 , zq for every author i and time t. Let αmin “ miniPrns αi ą 0 be the minimum to guarantee Dt ď ε. adaptation rate. Then, for every ε ą 0, it is sufficient to take t ě logp2{εq αmin The homogeneity assumption is strong: all authors are pulled toward the same distribution. Thus Proposition 3.1 gives a baseline showing that common adaptation and rewards cause diversity to decay exponentially, at a rate set by the slowest-adapting author. Our next result relaxes the homogeneity assumption with author-specific adaptation. Proposition 3.2. Consider Mechanism 1 with fixed model distribution q t “ q 0 . Suppose that the adaptation operator is time-invariant but author-specific, i.e., for each author i, Ai pq 0 , zit q “ ai for all t, where ai P ∆m´1 , and let αmin “ miniPrns αi ą 0. Then each author’s linguistic-style distribution converges to its author-specific adapted distribution, pti Ñ ai , exponentially fast at a rate controlled by αmin . Consequently, Dt Ñ D8 :“
ÿ 1 JSpai , aj q, npn ´ 1q i‰j
and Dt Ñ 0 if and only if ai “ aj for all i, j. 5
When adaptation is author-specific but fixed over time, authors do not collapse to a single linguistic-style distribution: each converges to their own adaptation target, and population-level diversity converges to the diversity among these targets—author-specific incentives preserve diversity even under repeated exposure to the same fixed model. If D8 ą 0, the process stabilizes at a positive level of diversity. Interplay between diversity and conformity. To interpolate between these regimes, let each author balance a conformity incentive pulling toward the shared model-induced norm against a diversity incentive preserving author-specific style. For the remainder of this section, let ri P ∆m´1 denote author i’s fixed preferred linguistic-style distribution, and set zit “ ri for all t. We model the adaptation as Ai pq 0 , ri q “ p1 ´ λi qri ` λi q 0 ,
λi P r0, 1s,
where λi measures the strength of the conformity incentive: λi “ 0 preserves the preferred style ri , while λi “ 1 fully adopts the shared model distribution q 0 . Proposition 3.3. Consider Mechanism 1 with fixed model distribution q t “ q 0 . Suppose each author has an author-specific preferred distribution ri P ∆m´1 , and the adaptation operator is Ai pq 0 , ri q “ p1 ´ λi qri ` λi q 0 ,
λi P r0, 1s.
Then each author distribution converges to ai pλi q :“ p1 ´ λi q ri ` λi q 0 . Precisely, if αmin :“ miniPrns αi ą 0 to guarantee that maxiPrns }pti ´ ai pλi q}1 ď ε. then for every ε ą 0, it is sufficient to take t ě logp2{εq αmin Consequently, ÿ 1 Dt Ñ D8 pλq :“ JS pai pλi q, aj pλj qq . npn ´ 1q i‰j In particular, Dt Ñ 0 if and only if ai pλi q “ aj pλj q for all i, j. A sufficient condition is λi “ 1 for all i, in which case ai pλi q “ q 0 for every author i. The proof establishes that for a common conformity level λi “ λ, the limiting diversity satisfies D8 pλq “ Op1 ´ λq. The parameter λ represents institutional or reward-based pressure to conform to a shared modelinduced linguistic norm: when λ is small, author-specific preferences dominate and diversity persists; when λ is large, the shared norm dominates and diversity collapses. Even with author-specific adaptation, monoculture can emerge if the reward structure places sufficiently high weight on conformity.
3.2
A shared model with recursive updates
t`1 We allow the shared model to be retrained on author outputs. Writing P t`1 :“ Bppt`1 1 , . . . , pn q “ řn next t`1 , the coupled dynamics of Mechanism 2 are i“1 wi pi
pt`1 “ p1 ´ αi q pti ` αi Ai pq t , zit q, i q t`1 “ βq t ` p1 ´ βq P t`1 . If the adaptation target does not depend on the evolving model distribution (Ai pq t , zit q “ a or ai ), the author dynamics are unchanged from the fixed-model setting and Propositions 3.1 and 3.2 apply verbatim: recursive feedback alone does not force monoculture when authors are pulled toward fixed targets. Recursive updates matter only when the adaptation operator depends non-trivially on q t . We therefore consider the conformity-mixture adaptation Ai pq t , ri q “ p1 ´ λi qri ` λi q t ,
λi P r0, 1s,
where the author’s preferred distribution ri is fixed over time and λi now measures conformity to the current model distribution.
6
t Proposition 3.4. Consider Mechanism 2 with Ai pq t ,řri q “ p1 ´ λi qri ` λi qř and λi P r0, 1q. Suppose that n n t`1 t`1 αi P p0, 1s for every author i, β P r0, 1q, and P “ i“1 wi pi , wi ě 0, i“1 wi “ 1. Then the coupled author–model dynamics converge to an equilibrium pp˚1 , . . . , p˚n , q ˚ q satisfying
p˚i “ p1 ´ λi qri ` λi q ˚ řn wi p1 ´ λi qri ˚ for every author i, and q “ i“1 řn . 1 ´ i“1 wi λi ÿ 1 Consequently, Dt Ñ D˚ :“ JSpp˚i , p˚j q. npn ´ 1q i‰j In particular, the limiting diversity is the diversity among the equilibrium author distributions p˚i , rather than the diversity among the fixed-model targets p1 ´ λi q ri ` λi q 0 . logp2{εq , The proof also shows that convergence is exponential, with a worst-case time bound of t ě p1´βq min iPrns αi p1´λi q t which is weaker than the fixed-model bound, controlled by αmin alone: the moving target q can slow the guaranteed rate, though the bound need not be tight in typical instances (see Section 5). The condition λi ă 1 ensures that each author retains some pull toward their preferred distribution ri ; if λi “ 1 for all authors, the fixed-point formula for q ˚ becomes singular and that case must be treated separately. The weaker rate bound should not be read as recursive feedback dampening homogenization: recursion changes the equilibrium, anchoring authors at the endogenous q ˚ rather than the exogenous q 0 . Proposition 3.6 makes the resulting comparison precise.
3.3
Personalized models with recursive updates
Personalization changes what the system converges to: because each model update includes an author-specific feedback component, the limiting model distributions qi˚ can remain distinct. The system may therefore converge to a family of author-specific equilibria rather than a single shared norm. The parameter ρ controls the extent of personalization. Proposition 3.5. Consider Mechanism 3 with author and model updates pt`1 “ p1 ´ αi q pti ` αi Ai pqit , ri q, i n ÿ qit`1 “ βqit ` γpt`1 ` δ wj pt`1 i j , j“1
řn where αi P p0, 1s, Ai pqit , ri q “ p1 ´ λi qri ` λi qit with λi P r0, 1q, wj ě 0, j“1 wj “ 1, and β ` γ ` δ “ 1. γ Suppose that β P r0, 1q, and let ρ “ 1´β P r0, 1s. Then the coupled personalized author–model dynamics ˚ ˚ ˚ converge to an equilibrium pp1 , . . . , pn , q1 , . . . , qn˚ q with p˚i “ ηi ri ` p1 ´ ηi qP ˚ , 1´λi and P ˚ “ where ηi “ 1´ρλ i
qi˚ “ ρηi ri ` p1 ´ ρηi qP ˚ ,
řn wi ηi ri ři“1 . Consequently, n i“1 wi ηi
Dt Ñ D˚ :“
ÿ 1 JSpp˚i , p˚j q, npn ´ 1q i‰j
Qt Ñ Q˚ :“
ÿ 1 JSpqi˚ , qj˚ q. npn ´ 1q i‰j
7
When ρ ą 0, each equilibrium model distribution qi˚ retains an author-specific component proportional to ρηi , so both author diversity D˚ and personalized-model diversity Q˚ can remain bounded away from zero; when ρ “ 0, all personalized models collapse to the same population-level distribution. Personalization thus does not prevent convergence—it changes its object, from a single shared norm to a family of distinct author–model equilibria. Because JS depends on both pairwise separation and the cloud’s location in the simplex (Remark A.3), we isolate pairwise geometry using the translation-invariant quadratic diversity p 1 , . . . , pn q :“ Dpp
ÿ 1 }pi ´ pj }22 , 2npn ´ 1q i‰j
p 8 for its value at equilibrium. and write D 1´λ . Then, for every pair i ‰ j, Proposition 3.6. Suppose λi “ λ P r0, 1q for every author i, and let η “ 1´ρλ
ai pλq ´ aj pλq “ p1 ´ λqpri ´ rj q p˚i ´ p˚j “ p1 ´ λqpri ´ rj q p˚i ´ p˚j “ ηpri ´ rj q
under IM 1,
under IM 2,
under IM 3.
Consequently, IMs 1 and 2 have identical limiting quadratic diversity, while p p 8 pIM 1q, p 8 pIM 3q “ D8 pIM 1q ě D D p1 ´ ρλq2 with equality if and only if ρλ “ 0. Under common conformity, recursion relocates the limiting author cloud from the exogenous q 0 to the endogenous q ˚ without changing its pairwise geometry. Hence any IM 1–IM 2 gap in limiting JS diversity arises only from JS’s position dependence. Personalization instead expands pairwise geometry relative to IMs 1–2 by 1{p1 ´ ρλq; heterogeneous λi can additionally create a genuine IM 1–IM 2 geometric gap (Appendix F). Summary. These results treat the conformity parameters λi as exogenous: they characterize what happens when authors have a fixed tendency to adopt the model-induced norm, but not why authors would choose to have that tendency.
4
Strategic Conformity and the Price of Monoculture
In Section 3, the conformity parameter λi quantifies how much author i is influenced by the model-induced linguistic norm, but not why. We now endogenize this parameter in the setting of Proposition 3.3: each author chooses λi to trade off individual rewards of conforming—such as legibility, perceived polish and alignment with reviewer expectations—against the loss of distinctiveness that conformity entails. This turns the dynamics into a strategic game, and we ask the linguistic analogue of the question raised by algorithmic monoculture [24]: can individually rational linguistic adaptation produce collectively inefficient homogenization? Under the fixed-model quadratic game with orthogonal signatures analyzed below, the answer is yes: the unique Nash equilibrium conformity level weakly exceeds the socially optimal level for every author, and strictly exceeds it for any author who conforms when others value distinctiveness. This gap is the value of an author’s distinctiveness to others: a payoff externality no individual internalizes and one that harms even authors who optimally choose not to conform (Lemma 4.1). In a symmetric regime, the resulting price of monoculture can be arbitrarily large across instances.
8
4.1
Conformity as a choice
We work in the setting of Mechanism 1, Proposition 3.3: the model distribution is fixed at q 0 , each author i has a preferred distribution ri P ∆m´1 and the adaptation operator is Ai pq 0 , ri q “ p1 ´ λi q ri ` λi q 0 . For any conformity profile λ “ pλ1 , . . . , λn q P r0, 1sn , Proposition 3.3 shows that the author dynamics converge exponentially to ai pλi q “ p1 ´ λi q ri ` λi q 0 . We use this convergence result to define a strategic game over long-run writing policies. Each author first chooses a fixed conformity level λi , which determines how strongly they rely on the model-induced norm. The interaction dynamics in eq. (1) then unfold and payoffs are evaluated at the limiting profile pa1 pλ1 q, . . . , an pλn qq. Thus the game is played over long-run writing styles and not over individual time-step updates. This timescale separation is justified by exponential convergence: when the evaluation horizon is long relative to the convergence time, authors’ distributions spend most of the horizon close to their limiting profile, so payoffs evaluated at the limiting profile approximate long-run average payoffs. For brevity, we write ui :“ ri ´ q 0 for author i’s signature vector —the direction and magnitude of their idiosyncrasy relative to the model’s linguistic norm—and σi :“ 1´λi P r0, 1s for their retained distinctiveness, so that ai pλi q ´ q 0 “ σi ui , ai pλi q ´ ri “ ´λi ui , ai pλi q ´ aj pλj q “ σi ui ´ σj uj . Payoffs. Given a conformity profile λ, author i’s utility is ›2 ›2 ci › bi › Ui pλi , λ´i q “ ´ ›ai pλi q ´ q 0 ›2 ´ ›ai pλi q ´ ri ›2 2 2 looooooooooomooooooooooon loooooooooomoooooooooon legibility reward
authenticity cost
ÿ› › θi ›ai pλi q ´ aj pλj q›2 , ` 2 2pn ´ 1q j‰i looooooooooooooooooomooooooooooooooooooon
(3)
distinctiveness payoff
where bi , θi ě 0 and ci ą 0. The three terms represent, respectively, the benefit of proximity to the modelinduced norm q 0 , the private cost of deviating from the author’s preferred style ri and the value of being distinguishable from other authors. The authenticity cost makes full conformity costly, while the distinctiveness payoff depends on the realized styles of others and couples with the author’s choices. Payoffs in eq. (3) use squared Euclidean distance rather than the Jensen–Shannon divergence used for Dt ; JS is locally quadratic (Remark A.3). We therefore evaluate long-run diversity using the quadratic measure p 8 pλq :“ Dpa p 1 pλ1 q, . . . , an pλn qq, which matches the payoff geometry and yields introduced in Section 3, D closed-form equilibria. We further make the following structural assumption. Definition 1. A population has orthogonal signatures if xui , uj y “ 0 for all i ‰ j, with d2i :“ }ui }22 ą 0 for all i. Orthogonality is a tractable benchmark yielding closed-form welfare results. Signatures lie in the m ´ 1dimensional tangent space of the simplex, so Definition 1 requires n ď m ´ 1: the benchmark describes populations no larger than the feature dimension. This is not restrictive when F is a rich inventory of lexical and syntactic markers, but it does couple population size to feature-set size. Appendix D relaxes orthogonality. Under orthogonality, }ai ´ aj }22 “ σi2 d2i ` σj2 d2j , and substituting in eq. (3) gives the separable form Ui pλq “
ı ÿ d2i ” θi pθi ´bi q σi2 ´ci p1 ´ σi q2 ` σ 2 d2 , 2 2pn ´ 1q j‰i j j
(4)
and the long-run diversity simplifies to n ÿ p 8 pλq “ 1 D σ 2 d2 . n i“1 i i
9
(5)
The coupling in eq. (4) is purely through the second term: author i’s choice does not change their own best response, but it does change everyone else’s payoff. Lemma 4.1. Under orthogonal signatures, for every i ‰ j, θj BUj “ ´ σi d2i ď 0, Bλi n´1 with strict inequality whenever θj ą 0 and λi ă 1. In particular, an increase in any author’s conformity strictly reduces the payoff of every other author who places positive value on distinctiveness, including authors who choose not to conform. Pairwise contrast is a shared resource: the distance }ai ´ aj }2 enters both Ui and Uj , so when author i moves toward the norm they consume contrast that author j was also drawing on. Lemma 4.1 also resolves the strategic question for non-adapters: an author for whom resisting is optimal still bears the cost of others’ adaptation, because the pool of styles against which their distinctiveness is measured collapses toward q 0 . There is no insulation in holding out. Equilibrium exceeds optimal conformity. Next we compare the conformity levels chosen by selfinterested authors at Nash equilibrium with those chosen by a social planner that maximizes total welfare. The next theorem shows that authors over-conform relative to the social optimum, which leads to weakly lower long-run diversity and weakly lower welfare at equilibrium. Theorem 4.2. Consider řthe game with payoffs in Equation (3) under orthogonal signatures, with ci ą 0 for 1 all i and let θ̄´i :“ n´1 j‰i θj denote the average distinctiveness value of the other authors. Then: (i) Every author has a strictly dominant strategy, so the game has a unique Nash equilibrium λNE , given by pbi ´ θi q` , where pxq` :“ maxt0, xu. λNE “ i pbi ´ θi q` ` ci řn (ii) Utilitarian welfare W pλq :“ i“1 Ui pλq is maximized at the unique profile λSO given by ` ˘ bi ´ θi ´ θ̄´i ` SO ˘ . λi “ ` bi ´ θi ´ θ̄´i ` ` ci NE (iii) For every author, λNE ě λSO ą 0 and θ̄´i ą 0. That is, every i i , with strict inequality if and only if λi p 8 pλNE q ď author who conforms at all over-conforms relative to the social optimum, and consequently D SO NE SO p D8 pλ q and W pλ q ď W pλ q.
Author i’s privately optimal conformity treats their distinctiveness as worth θi , while the planner values it at θi ` θ̄´i , because every other author also derives contrast from author i’s signature. Each author’s over-conformity increases with how much the rest of the community values contrast with them. Because the wedge depends on the average θ̄´i , it need not vanish in large populations. Notably, since the equilibrium is in dominant strategies, the inefficiency is driven purely by the payoff externality of Lemma 4.1—exactly as in algorithmic monoculture [24], where individually optimal reliance on a shared algorithm reduces aggregate welfare without any agent erring.
4.2
The price of monoculture
We quantify the diversity loss caused by strategic conformity via the price of monoculture (PoM), which compares the long-run diversity of the socially optimal conformity profile with that of Nash equilibrium. We evaluate PoM in a symmetric setting, where all authors have the same conformity reward, authenticity cost, value for distinctiveness and distance from the model norm. In this case, the PoM has a closed-form expression and reveals three regimes: no inefficiency, pure deadweight conformity and interior over-conformity with an arbitrarily large diversity loss. 10
Definition 2. The price of monoculture of an instance is the ratio of socially optimal to equilibrium long-run diversity, ř SO 2 2 p 8 pλSO q pσi q di D ě 1. “ ř i NE PoM :“ 2 2 NE p D8 pλ q i pσi q di Corollary 4.3. Suppose bi “ b, ci “ c, θi “ θ, and d2i “ d2 for all i. Then: (i) If b ď θ: λNE “ λSO “ 0 and PoM “ 1. Distinctiveness is at least as valuable as conformity privately, so no inefficiency arises. b´θ (ii) If θ ă b ď 2θ: λNE “ b`c´θ ą 0 while λSO “ 0, and
ˆ PoM “
b`c´θ c
˙2 .
Every author rationally conforms, yet the planner would have no one conform: equilibrium conformity is pure deadweight. b´θ b´2θ (iii) If b ą 2θ: both levels are interior, λNE “ b`c´θ ą λSO “ b`c´2θ , and
ˆ PoM “
b`c´θ b ` c ´ 2θ
˙2
ˆ “
θ 1` b ` c ´ 2θ
˙2 ,
which is increasing in θ, equals 1 at θ “ 0, and can diverge along sequences with b ` c ´ 2θ Ó 0. The per-author welfare loss at equilibrium is ¯ d2 1´ c2 θ2 W pλSO q ´ W pλNE q “ ¨ , n 2 pb ` c ´ 2θqpb ` c ´ θq2 which is strictly positive whenever θ ą 0. Three observations are worth highlighting. First, the welfare loss vanishes at θ “ 0 and is quadratic to leading order near zero; when distinctiveness has no value, conformity is simply beneficial standardization. Thus, θ distinguishes benign standardization from harmful monoculture. Second, in regime (ii), conformity is individually rational but socially wasteful: every author conforms although the planner prefers none. Third, since b ą 2θ implies b ` c ´ 2θ ą c, every regime satisfies PoM ď p1 ` θ{cq2 . The inefficiency is finite for each fixed instance but not uniformly bounded: it can diverge whenever θ{c Ñ 8, including as c Ó 0 with θ fixed. Appendix A discusses externalities on readers (Remark A.4) and strategic conformity under recursive model updates (Remark A.5).
5
Quantitative Comparison and Illustrations
We simulate n “ 100 authors over m “ 10 abstract linguistic features for 100 independent runs of T “ 200 steps (see Appendix F for details). Figure 1 reports the mean and standard deviation of population-level linguistic diversity across runs. Panel (a) compares the three interaction mechanisms using β “ 0.25 and, for IM 3, ρ “ 2{3 (γ “ 0.5 and δ “ 0.25). Across all mechanisms, LLM assistance initially reduces diversity, after which Dt stabilizes at a mechanism-specific level. Under the baseline parameterization, recursive updates of the shared model (IM 2) produce a faster decline in Dt and a lower limiting diversity than IM 1. By Proposition 3.6, under common conformity the mechanisms have identical limiting pairwise geometry, so any JS gap is purely positional. In our baseline, heterogeneous λi additionally produces the quadratic gap isolated in Appendix F. Panel (b) examines personalization in IM 3. Holding β “ 0.25 fixed, we vary ρ and set γ “ p1 ´ βqρ and δ “ p1 ´ βqp1 ´ ρq, thereby reallocating update weight from population-level to author-specific feedback. Diversity at T “ 200 increases monotonically with ρ, from approximately 0.057 at ρ “ 0 to 0.171 at ρ “ 1, showing that stronger personalization can preserve population-level linguistic diversity. Additional details and experiments are reported in Appendix F. 11
0.20
(b) D 200
(a)
0.16
Dt
0.12 IM3 IM1 IM2
0.08 0.04 0 1
5
20
Time t [log(1 + t)]
200
0
.5
Personalization ρ
1
Figure 1: Population-level linguistic diversity under LLM assistance. Panel (a) shows the evolution of Dt under IM 1–3, with time displayed on a logp1 ` tq scale and tick labels reporting the original time steps. Panel (b) shows IM 3 diversity at T “ 200 as personalization ρ increases. Lines report means over 100 runs and shading denotes ˘1 standard deviation; larger values indicate greater diversity.
6
Conclusions, Limitations and Future Work
We formalized linguistic monoculture as a population-level dynamical process over author and model distributions. Under our mechanisms, shared models can pull authors toward a common norm, while author-specific preferences and personalization can preserve diversity. Within our utility model, equilibrium conformity can exceed the social optimum, yielding a potentially unbounded price of monoculture in some regimes. When distinctiveness has little value, convergence can instead be efficient. Limitations and future work. Our work introduces a mathematical framework for formalizing linguistic monoculture in LLM-assisted language use. The framework is stylized by design: it focuses on a limited set of interaction mechanisms that make it possible to analyze how linguistic diversity evolves under different forms of author–model interaction. As with many theoretical models, this tractability relies on simplifying assumptions, which allow us to obtain explicit convergence, equilibrium and welfare guarantees. These assumptions provide a foundation for identifying and formally analyzing mechanisms of linguistic convergence, while leaving richer behavioral and empirical refinements for future work. We therefore discuss limitations of the framework and outline directions for connecting it more closely to empirical settings and broadening its applicability. Heterogeneous non-LLM language exposure. Our framework isolates LLM-mediated adaptation from other influences, such as collaborators, reading research articles, reviewers, disciplinary norms, broader media environments and other human factors that shape language use. Thus, our results characterize convergence under controlled settings rather than a complete model of linguistic change under real-world conditions. Future work should incorporate heterogeneous non-LLM exposures, social networks, recursive norm-shaping incentives, and empirical calibration of conformity and personalization parameters from longitudinal corpora of LLM-assisted writing. Fixed conformity, quadratic payoff and orthogonal signatures In Section 4, we make three simplifying assumptions to obtain closed-form expressions for equilibrium, welfare, and the price of monoculture. First, each author chooses a fixed conformity level λi , which determines their writing policy and remains unchanged over time. In practice, authors may update their conformity levels in response to experience, feedback, or changes in model behavior. Second, the analysis uses quadratic payoffs. Finally, the main-text closed-form results assume orthogonal author signatures, permitting at most m ´ 1 nonzero signatures in ∆m´1 . These assumptions isolate the conformity externality, but real signatures may be correlated and distinctiveness need not be valued quadratically. Extending the framework to dynamic conformity choices, richer similarity structures, and alternative payoff models is an important direction for future work. Empirical validation Our theoretical analysis is supported by simulations intended to illustrate the qualitative
12
behavior predicted by our framework, rather than to provide empirically calibrated forecasts. The parameter choices allow us to compare interaction mechanisms under controlled conditions and should not be interpreted as estimates of linguistic convergence. A natural next step is to verify these findings in controlled longitudinal studies of LLM-assisted writing, measuring how authors’ linguistic-style distributions evolve under fixed, recursively updated and personalized assistance. In this sense, our framework provides a foundation for future empirical work by identifying the interaction mechanisms, parameters, and diversity measures that such studies could estimate or experimentally vary. Acknowledgments and Funding. Suhas Thejaswi acknowledges support from the Technology Industries of Finland Centennial Foundation through a grant awarded to Aalto University. The authors declare that the funding source does not create any conflict of interest in relation to this work. GenAI usage statement. LLMs were used to edit and polish author-written text, and to assist with implementation and code refinement.
References [1] Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Leibo, Max Kleiman-Weiner, and Natasha Jaques. How LLMs distort our written language. arXiv preprint arXiv:2603.18161, 2026. [2] Dhruv Agarwal, Mor Naaman, and Aditya Vashistha. AI suggestions homogenize writing toward western styles and diminish cultural nuances. In Proceedings of the CHI conference on human factors in computing systems, pages 1–21, 2025. [3] Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. Homogenization effects of large language models on human creative ideation. In Proceedings of the 16th conference on creativity & cognition, pages 413–425, 2024. [4] Charles Bazerman. Shaping Written Knowledge: The Genre and Activity of the Experimental Article in Science. University of Wisconsin Press, Madison, WI, 1988. ISBN 978-0299116941. [5] Allan Bell. Language style as audience design. Language in Society, 13(2):145–204, 1984. [6] Douglas Biber and Susan Conrad. Register, genre, and style. Cambridge University Press, 2019. [7] Lucas M Bietti and Adrian Bangerter. Will the widespread use of large language models in scientific writing undermine scientists’ critical thinking? PLoS biology, 24(6):e3003801, 2026. [8] Melanie Brucks and Olivier Toubia. Prompt architecture induces methodological artifacts in large language models. PLOS one, 20(4):e0319159, 2025. [9] Rong-Ching Chang and Hao-Chuan Wang. Communication Accommodation Between Large Language Models and Users Across Cultures (Student Abstract). In AAAI Conference on Artificial Intelligence, pages 29331–29333, 2025. doi: 10.1609/AAAI.V39I28.35241. URL https://mlanthology.org/aaai/2025/ chang2025aaai-communication/. [10] Guanrong Chen and Yang Lou. Naming game. Switzerland: Springer International Publishing, 2019. [11] Bart De Vylder and Karl Tuyls. How to reach linguistic consensus: A proof of convergence for the naming game. Journal of theoretical biology, 242(4):818–831, 2006. [12] Morris H DeGroot. Reaching a consensus. Journal of the American Statistical association, 69(345): 118–121, 1974. [13] Anil R Doshi and Oliver P Hauser. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28):eadn5290, 2024. 13
[14] Nick J Enfield. Linguistic relativity from reference to agency. Annual Review of Anthropology, 44: 207–224, 2015. [15] Nicholas Evans and Stephen C Levinson. The myth of language universals: Language diversity and its importance for cognitive science. Behavioral and Brain Sciences, 32(5):429–448, 2009. [16] Noah E Friedkin and Eugene C Johnsen. Social influence and opinions. Journal of Mathematical Sociology, 15(3-4):193–206, 1990. [17] Mingmeng Geng and Roberto Trotta. Human-LLM coevolution: Evidence from academic writing. In Findings of the Association for Computational Linguistics: ACL 2025, pages 12689–12696, 2025. [18] Mingmeng Geng, Caixi Chen, Yanru Wu, Yao Wan, Pan Zhou, and Dongping Chen. The impact of large language models in academia: from writing to speaking. In Findings of the Association for Computational Linguistics: ACL 2025, pages 19303–19319, 2025. [19] Howard Giles. Communication accommodation theory: Negotiating personal relationships and social identities across contexts. Cambridge University Press, 2016. [20] John J. Gumperz and Stephen C. Levinson, editors. Rethinking Linguistic Relativity, volume 17 of Studies in the Social and Cultural Foundations of Language. Cambridge University Press, Cambridge, 1996. ISBN 9780521567067. [21] Rainer Hegselmann and Ulrich Krause. Opinion dynamics and bounded confidence: models, analysis and simulation. J. Artif. Soc. Soc. Simul., 5(3), 2002. [22] Ken Hyland. Academic Discourse: English in a Global Context. Continuum, London, 2009. [23] Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. Co-writing with opinionated language models affects users’ views. In Proceedings of the CHI conference on human factors in computing systems, pages 1–15, 2023. [24] Jon Kleinberg and Manish Raghavan. Algorithmic monoculture and social welfare. Proceedings of the National Academy of Sciences, 118(22):e2018340118, 2021. [25] Robert Kleinberg, Erald Sinanaj, and Éva Tardos. Price of anarchy of algorithmic monoculture. arXiv preprint arXiv:2604.00444, 2026. [26] Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát, and Jan Lause. Delving into LLMassisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27): eadt3813, 2025. [27] Bruno Latour and Steve Woolgar. Laboratory Life: The Construction of Scientific Facts. Princeton University Press, Princeton, NJ, 2nd edition, 1986. ISBN 9780691028323. [28] Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sirui Liu, Safyr He, Yifan Cui, et al. Quantifying large language model usage in scientific papers. Nature Human Behaviour, 9(12):2599–2609, 2025. doi: 10.1038/s41562-025-02273-8. [29] Jianhua Lin. Divergence measures based on the shannon entropy. IEEE Transactions on Information theory, 37(1):145–151, 1991. [30] John A Lucy. Linguistic relativity. Annual review of anthropology, 26(1):291–312, 1997. [31] Kibum Moon, Adam E. Green, and Kostadin Kushlev. Homogenizing effect of large language models (llms) on creative diversity: An empirical comparison of human and chatgpt writing. Computers in Human Behavior: Artificial Humans, 6:100207, 2025. doi: 10.1016/j.chbah.2025.100207.
14
[32] Vishakh Padmakumar and He He. Does writing with language models reduce content diversity? In International Conference on Learning Representations, volume 2024, pages 642–669, 2024. [33] Valentina N Pescuma, Dina Serova, Julia Lukassek, Antje Sauermann, Roland Schäfer, Aria Adli, Felix Bildhauer, Markus Egg, Kristina Hülk, Aine Ito, et al. Situating language register across the ages, languages, modalities, and cultural aspects: Evidence from complementary methods. Frontiers in Psychology, 13:964658, 2023. [34] Andrew J Peterson. AI and the problem of knowledge collapse. AI & Society, 40(5):3249–3269, 2025. [35] Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. AI models collapse when trained on recursively generated data. Nature, 631(8022):755–759, 2024. [36] Zhivar Sourati, Farzan Karimi-Malekabadi, Meltem Ozcan, Colin McDaniel, Alireza Ziabari, Jackson Trager, Ala Tak, Meng Chen, Fred Morstatter, and Morteza Dehghani. The shrinking landscape of linguistic diversity in the age of large language models. arXiv preprint arXiv:2502.11266, 2025. [37] Zhivar Sourati, Alireza S Ziabari, and Morteza Dehghani. The homogenizing effect of large language models on human expression and thought. Trends in Cognitive Sciences, 2026. [38] John M. Swales. Genre Analysis: English in Academic and Research Settings. Cambridge University Press, Cambridge, UK, 1990. [39] Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christensen, Chan Young Park, and Isabelle Augenstein. Epistemic diversity and knowledge collapse in large language models. arXiv preprint arXiv:2510.04226, 2025.
15
A
Remarks and Further Clarifications
This section provides additional clarifications and remarks that are omitted from the main text due to space constraints. Remark A.1 (Distributional representation of linguistic style). We represent linguistic style as a distribution rather than a single point in feature space because neither authors nor models exhibit one fixed style across contexts. Human language varies systematically with task, genre, and audience—the linguistic notion of register [6, 33]—and the style a model induces depends on the prompt, task, and framing [8]. A distribution captures this context-dependent variation while still giving a well-defined way to measure diversity across authors. The representation is also general: to study convergence in a specific feature ( e.g., syntactic or formatting conventions) or register ( e.g., academic prose), the feature set F can be restricted accordingly while other aspects of language are treated as fixed. Remark A.2 (Reparameterizing the recursive model update). The model update in Mechanism 2 can equivalently be written as n ÿ
q t`1 “ βq t `
γj pt`1 j ,
where β ě 0, γj ě 0, β `
j“1
n ÿ
γj “ 1.
j“1
Here, γj is the contribution of author j’s output to the model update, and setting γj “ p1 ´ βqwj recovers the weighted-average aggregation. This form keeps the update a convex combination of the old model distribution and the new population data. Remark A.3 (Quadratic surrogate for JS). Payoffs in (3) use squared Euclidean distance rather than the Jensen–Shannon divergence used for Dt . The two agree to second order: for nearby distributions, JSpp, qq “
1 ÿ ppk ´ qk q2 ` op}p ´ q}2 q, 8 k xk
x“
p`q . 2
Thus, when feature probabilities are bounded below, the two metrics are equivalent up to constants. Unlike squared Euclidean distance, however, JS also depends on the midpoint x: clouds with identical pairwise difference vectors can receive different JS diversity when located in different regions of the simplex. All p 8 and the quadratic payoffs in (3); Appendix F additionally uses a translationresults in Section 4 concern D invariant quadratic diagnostic to separate changes in pairwise geometry from this positional sensitivity. Remark A.4 (Reader welfare). Theorem 4.2 counts only authors’ payoffs in welfare, so each pairwise contrast is valued by exactly its two endpoints and the wedge θ̄´i is independent of n. If population diversity also benefits agents outside the game—readers, students, downstream researchers who never choose a λi but p 8 pλq that no consume the variety the community produces—then welfare contains an additional term Θ ¨ D Θ 2 2 author internalizes even partially. Under orthogonal signatures, this adds n σi di to the social value of author i’s retained distinctiveness. Thus the planner’s effective distinctiveness value for author i increases by 2Θ{n in the first-order condition, and if the societal value Θ scales with the size of the audience, the wedge grows with the community’s reach. The inefficiency identified in the main text is therefore a lower bound: it is the loss that remains even when only the strategic participants are counted. Remark A.5 (Strategic conformity under recursive updates). We have endogenized λ in the fixed-model setting of Mechanism 1. In the recursive setting of Mechanism 2, the equilibrium norm itself depends on the conformity profile, ř wi p1 ´ λi qri , q ˚ pλq “ i ř 1 ´ i wi λi by Proposition 3.4. An author’s conformity then exerts a second externality: by feeding less of their signature back into training data, they shift the shared norm itself toward the signatures of the remaining
16
non-conformists, reweighting whose style defines “polished” for everyone. For λi ă 1, the fixed point remains well-defined and can be viewed as a weighted average of the preferred distributions ri , with weights proportional to wi p1 ´ λi q. The expression becomes singular only at the boundary where all λi “ 1; as the profile approaches that boundary, the limiting norm can depend on the relative rates at which the terms 1 ´ λi vanish. This boundary behavior is the strategic counterpart of the singularity noted after Proposition 3.4 and of model-collapse dynamics [35]. Characterizing equilibria of the two-externality game, and whether normshaping incentives dampen or amplify over-conformity, is an open direction we view as the natural next step.
B
Omitted Proofs from Section 3
Proposition 3.1. In Mechanism 1, suppose the model-induced adaptation is homogeneous across authors and time, i.e., Ai pq 0 , zit q “ Apq 0 , zq for every author i and time t. Let αmin “ miniPrns αi ą 0 be the minimum to guarantee Dt ď ε. adaptation rate. Then, for every ε ą 0, it is sufficient to take t ě logp2{εq αmin Proof. Let a :“ Apq 0 , zq. By assumption, the update rule becomes pt`1 “ p1 ´ αi q pti ` αi a. i By subtracting a on both sides followed by induction on t, we get, pt`1 ´ a “ p1 ´ αi qp pti ´ aq i pti ´ a “ p1 ´ αi qt pp0i ´ aq Taking ℓ1 -norms and relying on the fact that both p0i and a are probability distributions, so their ℓ1 distance is at most 2, we have ` ˘ }pti ´ a}1 “ p1 ´ αi qt } p0i ´ a }1 ď 2 p1 ´ αi qt . Using the standard inequality 1 ´ x ď e´x for x P r0, 1s, p1 ´ αi qt ď e´αi t , we get }pti ´ a}1 ď 2e´αi t . Now consider any pair of authors i, j. By triangle inequality and using the bound above, it follows that, }pti ´ ptj }1 ď }pti ´ a}1 ` }ptj ´ a}1 ď 2e´αi t ` 2e´αj t ď 4e´αmin t . Now, using the standard bound JSppti , ptj q ď 12 }pti ´ ptj }1 , we obtain JSppti , ptj q ď 2e´αmin t . And, averaging over all ordered pairs i ‰ j, we get Dt “
ÿ 1 JSppti , ptj q ď 2e´αmin t . npn ´ 1q i‰j
Therefore, to ensure Dt ď ε, it is sufficient that 2e´αmin t ď ε ùñ t ě
logp2{εq “O αmin
ˆ
logp1{εq αmin
˙ ,
which completes the proof. Proposition 3.2. Consider Mechanism 1 with fixed model distribution q t “ q 0 . Suppose that the adaptation operator is time-invariant but author-specific, i.e., for each author i, Ai pq 0 , zit q “ ai for all t, where ai P 17
∆m´1 , and let αmin “ miniPrns αi ą 0. Then each author’s linguistic-style distribution converges to its author-specific adapted distribution, pti Ñ ai , exponentially fast at a rate controlled by αmin . Consequently, Dt Ñ D8 :“
ÿ 1 JSpai , aj q, npn ´ 1q i‰j
and Dt Ñ 0 if and only if ai “ aj for all i, j. Proof of Proposition 3.2. For each author i, define ai :“ Ai pq 0 , zit q, which is fixed over time by assumption. The update rule becomes pt`1 “ p1 ´ αi q pti ` αi ai . i Subtracting ai from both sides, followed by induction gives pt`1 ´ ai “ p1 ´ αi q ppti ´ ai q i pti ´ ai “ p1 ´ αi qt pp0i ´ ai q. Taking ℓ1 -norms and using the fact that both p0i and ai are probability distributions whose ℓ1 distance is at most 2, gives }pti ´ ai }1 ď 2 p1 ´ αi qt ď 2 e´αi t ď 2 e´αmin t . Thus pti Ñ ai for every author i, as t Ñ 8. Since Jensen–Shannon divergence is continuous on the probability simplex, for every pair i, j, JSppti , ptj q Ñ JSpai , aj q. Averaging over all ordered pairs i ‰ j, we obtain Dt “
ÿ ÿ 1 1 JSppti , ptj q Ñ JSpai , aj q “ D8 . npn ´ 1q i‰j npn ´ 1q i‰j
Therefore Dt Ñ 0 if ai “ aj for all i, j. Conversely, since JSpai , aj q “ 0 if and only if ai “ aj , we have D8 “ 0 only when all limiting adapted distributions are identical. Convergence of linguistic diversity. The time steps necessary for convergence of authors to their own limits within ε ą 0, i.e. }pti ´ ai }1 ď ε, follows from the analysis in Proposition 3.1, and it is enough to 8 8 t require t ě logp2{εq αmin . However, if D ą 0, then for any ε ă D , there is no convergence time to D ď ε. The process stabilizes at a positive level of linguistic diversity. Proposition 3.3. Consider Mechanism 1 with fixed model distribution q t “ q 0 . Suppose each author has an author-specific preferred distribution ri P ∆m´1 , and the adaptation operator is Ai pq 0 , ri q “ p1 ´ λi qri ` λi q 0 ,
λi P r0, 1s.
Then each author distribution converges to ai pλi q :“ p1 ´ λi q ri ` λi q 0 . Precisely, if αmin :“ miniPrns αi ą 0 then for every ε ą 0, it is sufficient to take t ě logp2{εq to guarantee that maxiPrns }pti ´ ai pλi q}1 ď ε. αmin Consequently, ÿ 1 Dt Ñ D8 pλq :“ JS pai pλi q, aj pλj qq . npn ´ 1q i‰j In particular, Dt Ñ 0 if and only if ai pλi q “ aj pλj q for all i, j. A sufficient condition is λi “ 1 for all i, in which case ai pλi q “ q 0 for every author i.
18
Proof of Proposition 3.3. Using ai pλi q :“ p1 ´ λi qri ` λi q 0 , the update rule of the author’s distribution can be written as pt`1 “ p1 ´ αi qpti ` αi ai pλi q i ` ˘ pt`1 ´ ai pλi q “ p1 ´ αi q pti ´ ai pλi q i ` ˘ pti ´ ai pλi q “ p1 ´ αi qt p0i ´ ai pλi q }pti ´ ai pλi q}1 “ p1 ´ αi qt }p0i ´ ai pλi q}1 , where the second line subtracts ai pλi q on both sides, the third follows by induction, and the fourth takes ℓ1 -norms. Since both p0i and ai pλi q are probability distributions, their ℓ1 -distance is at most 2. Therefore, using the standard inequality 1 ´ x ď e´x for x P r0, 1s and αmin :“ miniPrns αi , }pti ´ ai pλi q}1 ď 2p1 ´ αi qt ď 2e´αi t ď 2e´αmin t , max }pti ´ ai pλi q}1 ď 2e´αmin t . iPrns
To guarantee maxiPrns }pti ´ ai pλi q}1 ď ε, it is sufficient that 2e´αmin t ď ε, equivalently t ě logp2{εq αmin . This proves the claimed convergence-time bound. Since pti Ñ ai pλi q for every author i, and Jensen–Shannon divergence is continuous on the probability simplex, JSppti , ptj q Ñ JSpai pλi q, aj pλj qq. Averaging over all ordered pairs i ‰ j gives Dt Ñ D8 pλq “
ÿ 1 JS pai pλi q, aj pλj qq . npn ´ 1q i‰j
Because every Jensen–Shannon term is nonnegative and JSpp, qq “ 0 if and only if p “ q, we have D8 pλq “ 0 if and only if ai pλi q “ aj pλj q for all i, j. In particular, if λi “ 1 for all i, then ai pλi q “ q 0 for every i, and therefore Dt Ñ 0. Finally, suppose that λi “ λ for all authors. For any 0 ď λ1 ă λ2 ď 1, define c :“
1 ´ λ2 P r0, 1s. 1 ´ λ1
Then, for every author i, ai pλ2 q “ c ai pλ1 q ` p1 ´ cqq 0 . By joint convexity of Jensen–Shannon divergence, ` ˘ ` ˘ JS ai pλ2 q, aj pλ2 q ď c JS ai pλ1 q, aj pλ1 q . Averaging over all ordered pairs shows that D8 pλq is nonincreasing in the common conformity level λ. Moreover, ai pλq ´ aj pλq “ p1 ´ λqpri ´ rj q,
}ai pλq ´ aj pλq}1 “ p1 ´ λq}ri ´ rj }1 .
Using JSpp, qq ď 12 }p ´ q}1 , we obtain D8 pλq ď p1 ´ λq
ÿ 1 }ri ´ rj }1 . 2npn ´ 1q i‰j
Thus D8 pλq Ñ 0 as λ Ñ 1, and D8 pλq “ Op1 ´ λq.
19
t Proposition 3.4. Consider Mechanism 2 with Ai pq t ,řri q “ p1 ´ λi qri ` λi qř and λi P r0, 1q. Suppose that n n t`1 t`1 αi P p0, 1s for every author i, β P r0, 1q, and P “ i“1 wi pi , wi ě 0, i“1 wi “ 1. Then the coupled author–model dynamics converge to an equilibrium pp˚1 , . . . , p˚n , q ˚ q satisfying
p˚i “ p1 ´ λi qri ` λi q ˚ řn wi p1 ´ λi qri ˚ for every author i, and q “ i“1 řn . 1 ´ i“1 wi λi ÿ 1 Consequently, Dt Ñ D˚ :“ JSpp˚i , p˚j q. npn ´ 1q i‰j In particular, the limiting diversity is the diversity among the equilibrium author distributions p˚i , rather than the diversity among the fixed-model targets p1 ´ λi q ri ` λi q 0 . Proof of Proposition 3.4. At equilibrium pt`1 “ pti “ p˚i and q t “ q ˚ . Since αi ą 0, the author distributions i satisfy p˚i “ p1 ´ αi qp˚i ` αi pp1 ´ λi qri ` λi q ˚ q αi p˚i “ αi pp1 ´ λi qri ` λi q ˚ q p˚i “ p1 ´ λi qri ` λi q ˚
(6)
Moreover, at equilibrium the model update satisfies q ˚ “ βq ˚ ` p1 ´ βqP ˚ , n ÿ q˚ “ P ˚ “ wi p˚i i“1
q˚ “ ˚
q “
n ÿ i“1 n ÿ
`
wi p1 ´ λi qri ` λi q ˚ ˜ wi p1 ´ λi qri `
i“1
n ÿ
˘ ¸
wi λi
q˚ ,
i“1
whereřthe third line substitutes Equation (6). Since λi ă 1 for every i, and since wi ě 0 with n have i“1 wi λi ă 1. Therefore, řn wi p1 ´ λi q ri . q ˚ “ i“1 řn 1 ´ i“1 wi λi
ř
i wi “ 1, we
(7)
Equations (6) and (7) give the claimed expressions for p˚i and q ˚ at equilibrium. It remains to show that the distributions converge to this equilibrium. Define the deviations from equilibrium by xti :“ pti ´ p˚i and y t :“ q t ´ q ˚ . Using the update rule for pt`1 and the equilibrium identity for p˚i , i xt`1 “ pt`1 ´ p˚i “ p1 ´ αi qppti ´ p˚i q ` αi λi pq t ´ q ˚ q i i “ p1 ´ αi qxti ` αi λi y t , }xt`1 }1 ď p1 ´ αi q }xti }1 ` αi λi }y t }1 . i ␣ ( Let Rt :“ max maxiPrns }xti }1 , }y t }1 and µ :“ miniPrns αi p1 ´ λi q. Since αi ą 0 and λi ă 1 for all i, we have µ ą 0. Then, for every i, }xt`1 }1 ď p1 ´ αi qRt ` αi λi Rt “ p1 ´ αi p1 ´ λi qq Rt ď p1 ´ µqRt . i
20
Next, consider the model deviation y t`1 “ q t`1 ´ q ˚ . Using the update rule q t`1 “ βq t ` p1 ´ βqP t`1 and the equilibrium identity q ˚ “ βq ˚ ` p1 ´ βqP ˚ , we get y t`1 “ βpq t ´ q ˚ q ` p1 ´ βqpP t`1 ´ P ˚ q. řn řn řn wi p˚i , we have P t`1 ´ P ˚ “ i“1 wi xt`1 . Substituting back and and P ˚ “ i“1 Using P t`1 “ i“1 wi pt`1 i i ř using }xt`1 }1 ď p1 ´ µqRt together with i wi “ 1, i y t`1 “ βy t ` p1 ´ βq
n ÿ
wi xt`1 i
i“1
}y t`1 }1 ď β}y t }1 ` p1 ´ βq
n ÿ
wi }xt`1 }1 i
i“1
`
˘ ` ˘ ď β ` p1 ´ βqp1 ´ µq Rt “ 1 ´ p1 ´ βqµ Rt . Let κ :“ 1´p1´βqµ. Since β ă 1 and µ ą 0, we have κ ă 1. Also κ ě 1´µ, because κ “ 1´p1´βqµ ě 1´µ. Therefore, combining the bounds for xt`1 and y t`1 , we get Rt`1 ď κRt , and applying this recursively yields i t t 0 R ď κ R . Since κ ă 1, it follows that Rt Ñ 0 as t Ñ 8. Hence, pti Ñ p˚i for every i, and q t Ñ q ˚ . This establishes the convergence. Since Jensen–Shannon divergence is continuous on the probability simplex, JSppti , ptj q Ñ JSpp˚i , p˚j q for every pair i, j, and averaging over all ordered pairs i ‰ j yields ÿ ÿ 1 1 JSppti , ptj q Ñ JSpp˚i , p˚j q “ D˚ . Dt “ npn ´ 1q i‰j npn ´ 1q i‰j Convergence time. The convergence is exponential. Since all initial and equilibrium quantities are probability distributions, R0 ď 2, so ! ) max max }pti ´ p˚i }1 , }q t ´ q ˚ }1 ď 2κt . i
Consequently, to guarantee maxi }pti ´ p˚i }1 ď ε and }q t ´ q ˚ }1 ď ε, it is sufficient to take tě
logp2{εq . p1 ´ βq miniPrns αi p1 ´ λi q
Thus the worst-case convergence guarantee for the recursive-update setting is weaker than for the fixedmodel setting, where the corresponding bound is controlled only by αmin ; the bound need not be tight, and empirically the recursive dynamics can homogenize faster (Section 5). This completes the proof. Proposition 3.5. Consider Mechanism 3 with author and model updates pt`1 “ p1 ´ αi q pti ` αi Ai pqit , ri q, i n ÿ qit`1 “ βqit ` γpt`1 ` δ wj pt`1 i j , j“1
řn where αi P p0, 1s, Ai pqit , ri q “ p1 ´ λi qri ` λi qit with λi P r0, 1q, wj ě 0, j“1 wj “ 1, and β ` γ ` δ “ 1. γ Suppose that β P r0, 1q, and let ρ “ 1´β P r0, 1s. Then the coupled personalized author–model dynamics converge to an equilibrium pp˚1 , . . . , p˚n , q1˚ , . . . , qn˚ q with p˚i “ ηi ri ` p1 ´ ηi qP ˚ , 1´λi where ηi “ 1´ρλ i
qi˚ “ ρηi ri ` p1 ´ ρηi qP ˚ ,
řn wi ηi ri and P ˚ “ ři“1 . Consequently, n i“1 wi ηi
Dt Ñ D˚ :“
ÿ 1 JSpp˚i , p˚j q, npn ´ 1q i‰j 21
Qt Ñ Q˚ :“
ÿ 1 JSpqi˚ , qj˚ q. npn ´ 1q i‰j
Proof of Proposition 3.5. First, we characterize the equilibrium, and then show that the dynamics converge to it. Characterizing the equilibrium. At equilibrium, for every author i, we have pt`1 “ pti “ p˚i and qit`1 “ i ˚ t qi “ qi . The author update gives ` ˘ p˚i “ p1 ´ αi qp˚i ` αi p1 ´ λi qri ` λi qi˚ p˚i “ p1 ´ λi qri ` λi qi˚
since αi ą 0.
Next, the personalized model update gives, with P ˚ :“
řn
˚ j“1 wj pj ,
qi˚ “ βqi˚ ` γp˚i ` δP ˚ p1 ´ βqqi˚ “ γp˚i ` δP ˚ δ γ p˚ ` P ˚. qi˚ “ 1´β i 1´β γ Using γ ` δ “ 1 ´ β and ρ :“ 1´β , this becomes
qi˚ “ ρp˚i ` p1 ´ ρqP ˚ . We now solve for p˚i . Substituting the expression for qi˚ into the equilibrium equation for p˚i , ` ˘ p˚i “ p1 ´ λi qri ` λi ρp˚i ` p1 ´ ρqP ˚ p1 ´ ρλi qp˚i “ p1 ´ λi qri ` λi p1 ´ ρqP ˚ p˚i “
λi p1 ´ ρq ˚ 1 ´ λi ri ` P , 1 ´ ρλi 1 ´ ρλi
1´λi , we have where the division is valid since λi ă 1 and ρ P r0, 1s imply 1 ´ ρλi ą 0. Defining ηi :“ 1´ρλ i i p1´ρq 1 ´ ηi “ λ1´ρλ , and therefore i
p˚i “ ηi ri ` p1 ´ ηi qP ˚ . řn It remains to identify P ˚ . By definition, P ˚ “ i“1 wi p˚i . Substituting the expression for p˚i and using ř i wi “ 1, ˜ ¸ n n ÿ ÿ ˚ P “ wi ηi ri ` 1 ´ wi ηi P ˚ i“1
˜
n ÿ
i“1
¸ w i ηi
P˚ “
i“1
n ÿ
wi ηi ri i“1 řn wi ηi ri P ˚ “ ři“1 , n i“1 wi ηi
where the division is valid since λi ă 1 implies ηi ą 0, hence ηi ri ` p1 ´ ηi qP ˚ into qi˚ “ ρp˚i ` p1 ´ ρqP ˚ ,
ř
i wi ηi
ą 0. Finally, substituting p˚i “
qi˚ “ ρηi ri ` ρp1 ´ ηi qP ˚ ` p1 ´ ρqP ˚ “ ρηi ri ` p1 ´ ρηi qP ˚ . This proves the claimed form of the equilibrium.
22
Convergence to equilibrium. Define the deviations xti :“ pti ´ p˚i and yit :“ qit ´ qi˚ . Using the author update and the equilibrium identity, xt`1 “ pt`1 ´ p˚i “ p1 ´ αi qppti ´ p˚i q ` αi λi pqit ´ qi˚ q i i “ p1 ´ αi qxti ` αi λi yit , }xt`1 }1 ď p1 ´ αi q}xti }1 ` αi λi }yit }1 . i Define Rt :“ max tmaxi }xti }1 , maxi }yit }1 u and µ :“ miniPrns αi p1 ´ λi q. Since αi ą 0 and λi ă 1 for every i, we have µ ą 0. Then ` ˘ }xt`1 }1 ď 1 ´ αi p1 ´ λi q Rt ď p1 ´ µqRt . i Next, consider the model deviations. From the model update and the equilibrium identity, yit`1 “ βyit ` γxt`1 `δ i
n ÿ
wj xt`1 j
j“1
}yit`1 }1 ď β}yit }1 ` γ}xt`1 }1 ` δ i
n ÿ
wj }xt`1 j }1
j“1
ď pβ ` pγ ` δqp1 ´ µqq Rt “ pβ ` p1 ´ βqp1 ´ µqq Rt “ p1 ´ p1 ´ βqµq Rt , ř t using }xt`1 j }1 ď p1 ´ µqR , j wj “ 1, and γ ` δ “ 1 ´ β. Let κ :“ 1 ´ p1 ´ βqµ. Since β ă 1 and µ ą 0, we have κ ă 1; moreover κ ě 1 ´ µ. Therefore, the bounds for both xt`1 and yit`1 imply Rt`1 ď κRt , and i ˚ t t 0 t t t recursively R ď κ R . Since κ ă 1, we have R Ñ 0, so pi Ñ pi and qi Ñ qi˚ for every author i, establishing convergence. Convergence time. Since all p0i , p˚i , qi0 , qi˚ are probability distributions, their ℓ1 -distances are at most 2. Hence R0 ď 2 and Rt ď 2κt . To guarantee maxi }pti ´ p˚i }1 ď ε and maxi }qit ´ qi˚ }1 ď ε, it is sufficient that 2κt ď ε, equivalently t ě logp2{εq ´ log κ . Since κ “ 1 ´ p1 ´ βqµ and ´ logp1 ´ zq ě z for z P p0, 1q, it is sufficient to take logp2{εq . tě p1 ´ βq mini αi p1 ´ λi q Finally, since Jensen–Shannon divergence is continuous on the probability simplex, for every pair i, j, JSppti , ptj q Ñ JSpp˚i , p˚j q,
JSpqit , qjt q Ñ JSpqi˚ , qj˚ q.
Averaging over all ordered pairs i ‰ j, we obtain Dt Ñ D˚ and Qt Ñ Q˚ . This completes the proof. 1´λ Proposition 3.6. Suppose λi “ λ P r0, 1q for every author i, and let η “ 1´ρλ . Then, for every pair i ‰ j,
ai pλq ´ aj pλq “ p1 ´ λqpri ´ rj q
under IM 1,
p˚i ´ p˚j “ p1 ´ λqpri ´ rj q under IM 2, p˚i ´ p˚j “ ηpri ´ rj q under IM 3. Consequently, IMs 1 and 2 have identical limiting quadratic diversity, while p p 8 pIM 3q “ D8 pIM 1q ě D p 8 pIM 1q, D p1 ´ ρλq2 with equality if and only if ρλ “ 0. Proof. For IM 1, Proposition 3.3 gives ai pλq “ p1 ´ λqri ` λq 0 , hence ai pλq ´ aj pλq “ p1 ´ λqpri ´ rj q. Under Proposition 3.4, p˚i “ p1 ´ λqri ` λq ˚ , giving the same difference for IM 2. Under Proposition 3.5, ηi “ η for ř p is homogeneous of every i, hence P ˚ “ i wi ri and p˚i “ ηri ` p1 ´ ηqP ˚ , so p˚i ´ p˚j “ ηpri ´ rj q. Since D degree two in pairwise differences and η{p1 ´ λq “ 1{p1 ´ ρλq, the diversity relation follows. 23
C
Omitted Proofs from Section 4
Lemma 4.1. Under orthogonal signatures, for every i ‰ j, θj BUj “ ´ σi d2i ď 0, Bλi n´1 with strict inequality whenever θj ą 0 and λi ă 1. In particular, an increase in any author’s conformity strictly reduces the payoff of every other author who places positive value on distinctiveness, including authors who choose not to conform. θ
j σi2 d2i . Differentiating with σi “ 1 ´ λi Proof. From (4), λi enters Uj (j ‰ i) only through the term 2pn´1q
θ
j gives BUj {Bλi “ ´ n´1 σi d2i .
Theorem 4.2. Consider řthe game with payoffs in Equation (3) under orthogonal signatures, with ci ą 0 for 1 all i and let θ̄´i :“ n´1 j‰i θj denote the average distinctiveness value of the other authors. Then: (i) Every author has a strictly dominant strategy, so the game has a unique Nash equilibrium λNE , given by pbi ´ θi q` , where pxq` :“ maxt0, xu. λNE “ i pbi ´ θi q` ` ci řn (ii) Utilitarian welfare W pλq :“ i“1 Ui pλq is maximized at the unique profile λSO given by ` ˘ bi ´ θi ´ θ̄´i ` SO ˘ λi “ ` . bi ´ θi ´ θ̄´i ` ` ci NE (iii) For every author, λNE ě λSO ą 0 and θ̄´i ą 0. That is, every i i , with strict inequality if and only if λi p 8 pλNE q ď author who conforms at all over-conforms relative to the social optimum, and consequently D SO NE SO p 8 pλ q and W pλ q ď W pλ q. D
Proof. (i) By (4), Ui is additively separable: the only term involving λi is hi pσi q :“
ı d2i ” pθi ´ bi qσi2 ´ ci p1 ´ σi q2 , 2
which is independent of λ´i . Hence the maximizer of hi over σi P r0, 1s is a dominant strategy. We have ” ı ” ı h1i pσi q “ d2i pθi ´ bi qσi ` ci p1 ´ σi q “ d2i ci ´ pbi ` ci ´ θi qσi . Case bi ą θi . Then bi ` ci ´ θi ą ci ą 0, hi is strictly concave, and the unconstrained maximizer σi˚ “ ˚ NE ci {pbi ` ci ´ “ θi q P p0, 1q is interior,‰ giving λi “ 1 ´ σi “ pbi ´ θi q{pbi ` ci ´ θi q. Case bi ď θi . Then 1 2 hi pσi q “ di ci p1 ´ σi q ` pθi ´ bi qσi ě 0 on r0, 1s, strictly positive on r0, 1q, so hi is maximized at σi “ 1, i.e. λNE “ 0. Both cases match the stated formula, and uniqueness of the maximizer in each case gives i uniqueness of the equilibrium. (ii) Summing (4) over i and collecting, for each i, the coefficient of σi2 d2i {2—namely pθi ´ bi q from Ui and θj n´1 from each Uj with j ‰ i, the latter summing to θ̄´i —we obtain W pλq “
n ÿ d2 ”` i
i“1
2
ı ˘ θi ` θ̄´i ´ bi σi2 ´ ci p1 ´ σi q2 .
This is again additively separable across authors, and each summand has exactly the form hi with θi replaced by θi ` θ̄´i . The argument of part (i) applied verbatim with this replacement yields the stated λSO and its i uniqueness. 24
(iii) Define, for x ě 0, φi pxq :“
pbi ´ θi ´ xq` , pbi ´ θi ´ xq` ` ci
so that λNE “ φi p0q and λSO “ φi pθ̄´i q. The map y ÞÑ y{py ` ci q is strictly increasing on r0, 8q and i i x ÞÑ pbi ´ θi ´ xq` is nonincreasing, strictly decreasing while positive; hence φi is nonincreasing, strictly with strictness exactly when φi p0q ą 0 and θ̄´i ą 0. The ď λNE decreasing while positive. This gives λSO i i diversity comparison follows from (5), since σiNE ď σiSO for every i; the welfare comparison holds because λSO maximizes W . Corollary 4.3. Suppose bi “ b, ci “ c, θi “ θ, and d2i “ d2 for all i. Then: (i) If b ď θ: λNE “ λSO “ 0 and PoM “ 1. Distinctiveness is at least as valuable as conformity privately, so no inefficiency arises. b´θ (ii) If θ ă b ď 2θ: λNE “ b`c´θ ą 0 while λSO “ 0, and
ˆ PoM “
b`c´θ c
˙2 .
Every author rationally conforms, yet the planner would have no one conform: equilibrium conformity is pure deadweight. b´θ b´2θ (iii) If b ą 2θ: both levels are interior, λNE “ b`c´θ ą λSO “ b`c´2θ , and
ˆ PoM “
b`c´θ b ` c ´ 2θ
˙2
ˆ “
1`
θ b ` c ´ 2θ
˙2 ,
which is increasing in θ, equals 1 at θ “ 0, and can diverge along sequences with b ` c ´ 2θ Ó 0. The per-author welfare loss at equilibrium is ¯ d2 1´ c2 θ2 W pλSO q ´ W pλNE q “ ¨ , n 2 pb ` c ´ 2θqpb ` c ´ θq2 which is strictly positive whenever θ ą 0. Proof. Parts (i) and (ii) and the equilibrium expressions in (iii) follow from Theorem 4.2 with θ̄´i “ θ; the p 8 “ σ 2 d2 and σ NE “ c , while σ SO “ 1 PoM expressions follow from (5), since in the symmetric case D b`c´θ c in regime (iii). in regime (ii) and σ SO “ b`c´2θ For the welfare loss in (iii), write A :“ b ` c ´ 2θ ą 0. From the proof of Theorem 4.2(ii), per-author 2 welfare at a symmetric profile σ is d2 wpσq with wpσq “ p2θ ´ bqσ 2 ´ cp1 ´ σq2 “ ´Aσ 2 ` 2cσ ´ c. Then wpσ SO q “ wpc{Aq “ c2 {A ´ c, while with σ NE “ c{pA ` θq, wpσ NE q “ ´
Ac2 2c2 c2 pA ` 2θq ´c“ ` ´ c. 2 pA ` θq A`θ pA ` θq2
Subtracting, wpσ SO q ´ wpσ NE q “ c2 ¨
pA ` θq2 ´ ApA ` 2θq c2 θ 2 “ , ApA ` θq2 ApA ` θq2
since pA ` θq2 ´ ApA ` 2θq “ θ2 . Substituting A “ b ` c ´ 2θ and A ` θ “ b ` c ´ θ completes the proof.
25
D
Correlated Author Signatures
We now relax the orthogonality assumption in Definition 1 and characterize the strategic-conformity game for arbitrary author signatures. Recall that ui :“ ri ´ q 0 ,
d2i :“ ∥ui ∥22 ,
σi :“ 1 ´ λi ,
so that the limiting style of author i satisfies ai pλi q ´ q 0 “ σi ui . Let gij :“ xui , uj y and let G “ pgij qi,jPrns denote the Gram matrix of the author signatures. Thus, gii “ d2i , while gij measures the alignment between the directions in which authors i and j differ from the shared model-induced norm. Orthogonal signatures correspond to gij “ 0 for every i ‰ j. Proposition D.1 (Strategic conformity with correlated signatures). Consider the game with payoffs in Equation (3), without imposing orthogonality. 1. For every author i, Ui pσq “
‰ d2i “ pθi ´ bi qσi2 ´ ci p1 ´ σi q2 2 ÿ ÿ θi θi ` d2j σj2 ´ σi gij σj . 2pn ´ 1q j‰i n ´ 1 j‰i
(8)
The corresponding quadratic long-run diversity is D8 pλq “
n ÿ 1 ÿ 2 2 2 di σi ´ gij σi σj . n i“1 npn ´ 1q iăj
(9)
2. For every pair i ‰ j, ˘ BUj θj ` σj gij ´ σi d2i . (10) “ Bλi n´1 Consequently, an increase in author i’s conformity imposes a nonpositive externality on author j if and only if σi d2i ě σj gij . (11) In particular, the externality is nonpositive at every conformity profile whenever gij ď 0. 3. Define Ai :“ bi ` ci ´ θi . If Ai ą 0, author i’s payoff is strictly concave in σi , conditional on σ´i , and their unique best response is » fi ř θi c ´ g σ — i pn ´ 1qd2i j‰i ij j ffi ffi , BRi pσ´i q “ Πr0,1s — (12) – fl Ai where Πr0,1s denotes projection onto r0, 1s. If, in addition, ÿ θi |gij | ă 1, 2 iPrns pn ´ 1qdi Ai j‰i
max
26
(13)
then the joint best-response map is a contraction and the game has a unique Nash equilibrium. If the Nash equilibrium is interior, its retained-distinctiveness vector satisfies N σ NE “ h,
(14)
where hi :“ ci d2i and Nii “ d2i pbi ` ci ´ θi q,
Nij “
θi gij , n´1
i ‰ j.
(15)
4. Let θ̄´i :“
1 ÿ θj . n ´ 1 j‰i
Utilitarian welfare can be written as 1 W pσq “ C ` hJ σ ´ σ J Sσ, 2
(16)
where C is independent of σ, hi “ ci d2i , and ` ˘ Sii “ d2i bi ` ci ´ θi ´ θ̄´i , Sij “
θi ` θ j gij , n´1
i ‰ j.
(17) (18)
If S is positive definite, welfare is strictly concave and has a unique maximizer on r0, 1sn . If the maximizer is interior, it is characterized by Sσ SO “ h.
(19)
Proof. For any i ‰ j, ∥ai ´ aj ∥22 “ ∥σi ui ´ σj uj ∥22 “ σi2 d2i ` σj2 d2j ´ 2σi σj gij .
(20)
Substituting Equation (20) into Equation (3), the terms involving σi2 d2i appear once for every j ‰ i, giving ÿ θi θi σi2 d2i “ σi2 d2i . 2pn ´ 1q j‰i 2 Collecting this term with the legibility and authenticity terms gives the first line of Equation (8). The remaining squared-signature and cross terms give the second line. For quadratic diversity, summing Equation (20) over unordered pairs gives ÿ` ÿ ˘ σi2 d2i ` σj2 d2j “ pn ´ 1q σi2 d2i . iăj
i
Substitution into the definition of quadratic long-run diversity gives Equation (9). To obtain the externality formula, consider i ‰ j. The only part of Uj that depends on λi is the pairwise term involving authors i and j. Since Bσi “ ´1, Bλi
27
we have BUj θj B “ ∥σj uj ´ σi ui ∥22 Bλi 2pn ´ 1q Bλi θj “ xσj uj ´ σi ui , ui y n´1 ˘ θj ` σj gij ´ σi d2i . “ n´1
(21)
This proves Equations (10) and (11). If gij ď 0, both σj gij and ´σi d2i are nonpositive, so the externality is nonpositive. Differentiating Equation (8) with respect to σi gives θi ÿ BUi “ d2i rci ´ pbi ` ci ´ θi qσi s ´ gij σj . Bσi n ´ 1 j‰i
(22)
Moreover, B 2 Ui “ ´d2i Ai . Bσi2 Thus, Ai ą 0 implies strict concavity in σi . Solving Equation (22) and projecting the resulting unconstrained maximizer onto the feasible interval gives Equation (12). Projection onto a closed interval is nonexpansive. Hence, for any two profiles σ and σ 1 , ˇ ˇ 1 ˇBRi pσ´i q ´ BRi pσ´i qˇ ď ď
ÿ ˇ ˇ θi |gij | ˇσj ´ σj1 ˇ 2 pn ´ 1qdi Ai j‰i ÿ θi |gij |∥σ ´ σ 1 ∥8 . pn ´ 1qd2i Ai j‰i
(23)
Condition (13) therefore makes the joint best-response map a contraction in the sup norm. Banach’s fixedpoint theorem gives existence and uniqueness of the Nash equilibrium. At an interior equilibrium, Equation (22) is zero for every i, which is exactly the linear system N σ NE “ h. For welfare, each unordered pair ti, ju appears in both authors’ distinctiveness payoffs, with total coefficient θi ` θ j . 2pn ´ 1q Therefore, ‰ 1 ÿ“ 2 2 bi di σi ` ci d2i p1 ´ σi q2 2 i ÿ θi ` θ j ` ∥σi ui ´ σj uj ∥22 . 2pn ´ 1q iăj
W pσq “ ´
(24)
Expanding the squared distances and collecting coefficients gives Equations (16)–(18). If S is positive definite, the Hessian of welfare is ´S, so welfare is strictly concave. Its interior first-order condition is ∇W pσq “ h ´ Sσ “ 0, which gives Equation (19). Corollary D.2 (Over-conformity with nonpositively aligned signatures). Suppose gij ď 0
for every i ‰ j. 28
(25)
Assume that the Nash equilibrium and social optimum are interior and that θi ÿ |gij | d2i pbi ` ci ´ θi q ą n ´ 1 j‰i and ` ˘ d2i bi ` ci ´ θi ´ θ̄´i ą
1 ÿ pθi ` θj q|gij | n ´ 1 j‰i
(26)
(27)
for every author i. Then the Nash equilibrium and social optimum are unique and σiSO ě σiNE
for every i.
(28)
λNE ě λSO i i
for every i.
(29)
θ̄´i ą 0.
(30)
Equivalently, The inequality for author i is strict whenever σiNE ą 0
and
Consequently, D8 pλNE q ď D8 pλSO q.
(31)
Proof. Under Equation (25), both N and S have nonpositive off-diagonal entries. Conditions (26) and (27) give positive, strictly dominant diagonal entries. Consequently, N and S are nonsingular M -matrices and have entrywise nonnegative inverses. Furthermore, S ď N entrywise. On the diagonal, Sii “ Nii ´ d2i θ̄´i ď Nii . For i ‰ j, Sij ´ Nij “
θj gij ď 0. n´1
Because σ NE ě 0, Sσ NE ď N σ NE “ h. Multiplication by the nonnegative matrix S ´1 gives σ NE ď S ´1 h “ σ SO , which proves Equations (28) and (29). For strictness, observe that h ´ Sσ NE “ pN ´ Sqσ NE . Under Equation (25), the matrix N ´ S is entrywise nonnegative. Its ith diagonal contribution is d2i θ̄´i σiNE , which is strictly positive whenever Equation (30) holds. Because S ´1 is entrywise nonnegative and has strictly positive diagonal entries, it follows that σiSO ą σiNE , or equivalently, λNE ą λSO i i . Finally, when gij ď 0, each pairwise squared distance ∥σi ui ´ σj uj ∥22 “ σi2 d2i ` σj2 d2j ´ 2σi σj gij is nondecreasing in each of σi and σj on r0, 1s2 . The componentwise inequality σ SO ě σ NE therefore implies Equation (31). 29
Remark D.3 (Positive alignment can reverse the externality). For unrestricted correlated signatures, the conformity externality need not be negative. For example, suppose that authors i and j have identical signature vectors, ui “ uj , so that gij “ d2i “ d2j . If σj ą σi , Equation (10) gives θj d2i BUj “ pσj ´ σi q ą 0. Bλi n´1 In this case, author i already lies closer to the shared norm than author j. Additional conformity by author i moves their realized styles farther apart and therefore increases author j’s distinctiveness payoff. Thus, no theorem asserting a negative conformity externality at every profile can hold for arbitrary positively correlated signatures. The orthogonal-signature result is therefore one member of a wider class rather than an implication that holds for every signature geometry. Orthogonality eliminates strategic cross terms and yields the closed-form dominant-strategy expressions in Theorem 4.2. More generally, nonpositive signature alignments preserve the negative externality, and under the conditions of Corollary D.2 they preserve componentwise over-conformity. Positive alignments make the sign and magnitude of the externality depend on the authors’ relative retained distinctiveness.
E
A Heterogeneous Population with Multiple Interaction Mechanisms
This section introduces a fourth interaction mechanism that combines Mechanisms 1, 2 and 3 within a single population by partitioning authors into subpopulations, each following the update rule of its assigned mechanism. Interaction Mechanism 4 (A heterogeneous population with multiple interaction mechanisms). The population is partitioned into three disjoint subpopulations rns “ I1 Y I2 Y I3 , where authors in I1 , I2 and I3 follow Interaction Mechanisms 1, 2 and 3, respectively. Let q 0 denote the common base model distribution. Authors in I1 interact with a shared model whose distribution remains fixed at q0 . Their linguistic-style distributions evolve as pt`1 “ p1 ´ αi q pti ` αi Ai pq 0 , ri q, i P I1 . i Authors in I2 interact with a shared model whose linguistic style distribution q t is recursively updated using feedback from authors in I2 . Their coupled author–model dynamics evolve as pt`1 “ p1 ´ αi q pti ` αi Ai pq t , ri q, i P I2 , i ` t`1 ˘ t`1 t q “ β q ` p1 ´ βq B ppj qjPI2 , where B aggregates the updated author distributions in I2 . In particular, we use the weighted-average update ÿ ÿ ` ˘ wj pt`1 wj ě 0, wj “ 1. B ppt`1 j qjPI2 “ j , jPI2
jPI2
Each author i P I3 interacts with an author-specific personalized model qit , initialized from a common base distribution qi0 “ q 0 . Their coupled author–model dynamics are pt`1 “ p1 ´ αi q pti ` αi Ai pqit , ri q, i P I3 , i ÿ qit`1 “ β qit ` γ pt`1 `δ wj pt`1 i P I3 , i j , jPI3
30
ř where β, γ, δ ě 0, β ` γ ` δ “ 1, wj ě 0, and jPI3 wj “ 1. The three subpopulations evolve according to their respective update rules, with no additional coupling across I1 , I2 and I3 . Thus, Mechanism 4 captures a heterogeneous population in which different groups use LLM assistance through different interaction mechanisms. Remark E.1 (Convergence under Interaction Mechanism 4). Under the conformity-mixture adaptation operators and parameter conditions of Propositions 3.3 to 3.5, Interaction Mechanism 4 does not introduce additional coupling across the three subpopulations. Therefore, convergence follows by applying the corresponding result separately within each subpopulation. Authors in I1 converge to the fixed-model limits characterized in Proposition 3.3; authors in I2 , together with their shared recursively updated model, converge to the endogenous shared equilibrium characterized in Proposition 3.4; and authors in I3 , together with their personalized models, converge to the family of author-specific equilibria characterized in Proposition 3.5. Consequently, the population-level diversity Dt also converges. Its limit is the average pairwise Jensen– Shannon divergence among the limiting author distributions across the whole population. In particular, if p˚i denotes the limiting distribution of author i under the mechanism assigned to their subpopulation, then ˚ Dt ÝÑ Dmix :“
ÿ 1 JSpp˚i , p˚j q. npn ´ 1q i‰j
Equivalently, this limiting diversity decomposes into within-subpopulation terms, determined by the limiting diversity inside each of IMs 1-3, and between-subpopulation terms, determined by the distances between the equilibria reached by authors using different mechanisms. Thus, the mixed mechanism can yield an intermediate level of diversity in simulations, because the population contains authors governed by fixed, recursive, and personalized interaction rules. However, the exact limiting value depends on the sizes of the subpopulations, their initial and preferred distributions, and the parameters of each mechanism; no universal ordering relative to IMs 1-3 follows without additional assumptions.
F
Additional Experimental Results
This appendix reports robustness and diagnostic experiments complementing Figure 1. We study sensitivity to population size, Dirichlet concentration, conformity heterogeneity, and the diversity metric, together with auxiliary convergence measures, personalization dynamics, and a heterogeneous population whose subgroups follow IM 1–3. Our implementation2 is written in Python programming language using standard numerical libraries. All experiments were executed on a commodity laptop. The source code is available as open-source under a modest license agreement.3 Unless stated otherwise, we simulate n “ 100 authors over m “ 10 abstract linguistic features for T “ 200 steps. We initialize p0i , q 0 , and ri independently from Dirichletp1, . . . , 1q and use the same initialization across mechanisms within each run. We independently sample αi „ Uniformr0.05, 0.50q and λi „ Uniformr0, 1q, and use uniform author weights wi “ 1{n. The baseline parameters are β “ 0.25 and ρ “ 2{3 (equivalently, γ “ 0.5 and δ “ 0.25). All panels report means over 100 independent paired runs, with shading denoting ˘1 sample standard deviation. Time is displayed on a logp1 ` tq scale to resolve the rapid initial transient; tick labels report the original time steps. The master random seed is 123456789. The absolute plateau levels and the ordering between IM 1 and IM 2 depend on the initialization and parameter prior. The experiments therefore illustrate mechanism-specific dynamics in controlled parameter regimes rather than establish a universal ordering. Alongside the Jensen–Shannon diversity Dt , we report the translation-invariant quadratic diagnostic ÿ m r t :“ }pt ´ ptj }22 . D 8npn ´ 1q i‰j i 2 LLMs (e.g., ChatGPT Codex, CoPilot) were used to assist with implementation and code refinement. All generated code was reviewed and verified by the authors, who remain responsible for its correctness. 3 https://github.com/suhastheju/llm-monoculture-code
31
Linguistic diversity D t
a = 0.1
a=1 IM 1 IM 2 IM 3
0.5 0.4
a=5 0.04
0.15 0.03 0.10
0.3
0.02 0.2
0.05 0 1
5
20
Time t [log(1 + t)]
200
0.01 0 1
5
20
Time t [log(1 + t)]
200
0 1
5
20
200
Time t [log(1 + t)]
Figure 2: Linguistic-diversity trajectories Dt under IM 1–3 for Dirichletpa, . . . , aq initialization. The endpoint ordering IM 2 ă IM 1 ă IM 3 holds at every concentration. Lines show means over 100 runs, and shading denotes ˘1 sample standard deviation. This quantity depends only on the pairwise difference vectors pti ´ ptj , so a common translation of an author cloud leaves it unchanged. The factor m{8 matches the second-order JS scaling when pairwise midpoints lie r t only as a diagnostic; the paper’s primary diversity measure remains Dt . at the barycenter. We use D Dirichlet concentration and diversity metric. We first vary the initialization concentration a P t0.1, 1, 5u in Dirichletpa, . . . , aq. The ordering IM 2 ă IM 1 ă IM 3 holds at T “ 200 under both Dt r t for every value of a. For the baseline a “ 1, the quadratic endpoints are approximately 0.067 for and D IM 2, 0.083 for IM 1, and 0.107 for IM 3. Thus, the observed endpoint ordering is also present under the quadratic diagnostic and is not solely an artifact of the positional sensitivity of Jensen–Shannon divergence. Conformity heterogeneity. To separate positional and geometric effects, we sample λi „ Uniformr0.5 ´ w, 0.5 ` ws for w P t0, 0.25, 0.5u. At w “ 0, IM 1 and IM 2 have identical limiting pairwise difference vectors, and their quadratic gap is numerically zero (8.9ˆ10´9 ), while their JS gap is approximately 0.0071. The latter therefore reflects JS’s dependence on location within the simplex. As conformity becomes heterogeneous, the quadratic gap rises to approximately 0.0041 at w “ 0.25 and 0.0163 at w “ 0.5, showing that heterogeneous responses to different anchors create a genuine difference in pairwise geometry. Population-size robustness. Figure 4 compares the three interaction mechanisms for n “ 100 and n “ 500. The qualitative ordering is stable across the two population sizes: recursive shared feedback produces the lowest long-run diversity, while personalized feedback preserves the most diversity. Variability across runs is smaller at n “ 500, as expected since Dt averages over Opn2 q author pairs. Auxiliary convergence measures. Figure 5 reports author–model divergence M t and personalizedmodel diversity Qt . Because IM 1 and IM 2 use a single shared model, their model diversity is identically zero and is omitted from panel (b). Under IM 3, M t rapidly falls to a small positive level while Qt remains positive, indicating that personalized models become distinct while remaining closely aligned with their respective authors. The limiting model diversity is smaller than the limiting author diversity (Q‹ « 0.041 versus D‹ « 0.095), consistent with Proposition 3.5: each qi‹ retains the author-specific component ri with weight ρηi ă ηi . Personalization and transient dynamics. To complement the endpoint comparison in Figure 1, Figure 6 shows the full IM 3 trajectory for ρ P t0, 2{3, 1u. We hold β “ 0.25 fixed and set γ “ p1 ´ βqρ and
32
D 200 gap (IM 1 − IM 2)
JS diversity D
0.025
Quadratic D ̂ (transl. inv.)
0.020 0.015 0.010 0.005 0.000 0.00
0.25
0.50
Conformity half-width w
Figure 3: IM 1 minus IM 2 endpoint diversity as conformity heterogeneity increases. Under common conformity, the quadratic gap vanishes although the JS gap remains positive; with heterogeneous conformity, a genuine geometric gap emerges under both metrics. δ “ p1 ´ βqp1 ´ ρq. Increasing ρ therefore reallocates feedback weight from the shared population to the associated author. The endpoints provide a check on Proposition 3.5. At ρ “ 0 the personalized update reduces exactly to IM 2, and the observed D200 « 0.057 matches the IM 2 plateau. At ρ “ 1 we have ηi “ 1, so p‹i “ qi‹ “ ri ; since ri and p0i are drawn from the same Dirichlet, diversity should return to approximately its initial level, and the observed 0.171 is consistent with D0 « 0.175. All mechanisms exhibit a non-monotone transient: Dt undershoots its limit near t « 5 before partially recovering. Because personalized models are initialized at the common base distribution q 0 , every author is initially pulled toward the same point; the qit differentiate only later, allowing authors to drift back toward their preferred ri . The recovery scales with ρ, consistent with this account. Heterogeneous interaction mechanisms. For IM 4, we assign proportions 0.33, 0.33, and 0.34 of the population to IM 1, IM 2, and IM 3, respectively, using assignment seed 123456789. Each subgroup follows its own update rule. In particular, the population-level component received by personalized models is computed only from authors in the IM 3 subgroup, as specified in Section E; there is no additional cross-subpopulation feedback. All three quantities stabilize; the positive limiting value of Qt reflects differences among the authorfacing models used across and within the three subgroups.
33
n = 100
0.20
IM 1
IM 2
IM 3
IM 1
IM 2
IM 3
Dt
0.15
0.10
0.05
n = 500
0.20
Dt
0.15
0.10
0.05 0
1
5
20
200
Time t [log(1 + t)]
0.20
(a)
(b)
IM 1 IM 2 IM 3
0.15
Model diversity Q t
Author--model divergence M t
Figure 4: Population-size robustness of linguistic diversity Dt under IM 1–3 for n “ 100 (top) and n “ 500 (bottom). Lines show means over 100 runs and shading denotes ˘1 standard deviation.
0.10 0.05
0.04 0.03 0.02 0.01 IM 3
0.00
0.00 0
1
5
20
200
Time t [log(1 + t)]
0
1
5
20
200
Time t [log(1 + t)]
Figure 5: Auxiliary convergence measures under the baseline parameters: (a) author–model divergence M t for IM 1–3 and (b) personalized-model diversity Qt for IM 3. Lines indicate mean over 100 runs and shading denotes ˘1 standard deviation.
34
Linguistic diversity D t
0.200
ρ=0
0.175
ρ = 2/3
ρ=1
0.150 0.125 0.100 0.075 0.050 0
1
5
20
200
Time t [log(1 + t)]
(a)
(b)
Model diversity Q t
Author diversity D t
0.18
Author--model divergence M t
Figure 6: IM 3 linguistic-diversity trajectories for three personalization levels. Larger ρ preserves more authorspecific feedback and produces higher long-run Dt . Lines show mean over 100 runs and shading denotes ˘1 standard deviation.
0.15 0.12 0.09 0.06
0.06
0.04
0.02
0.00 0 1
5
20
Time t [log(1 + t)]
200
0 1
5
20
Time t [log(1 + t)]
200
0.20
(c)
0.15
0.10
0.05
0 1
5
20
200
Time t [log(1 + t)]
Figure 7: Evolution of (a) author diversity Dt , (b) diversity among author-facing model distributions Qt , and (c) author–model divergence M t under heterogeneous IM 4. Lines show mean over 100 runs and shading denotes ˘1 standard deviation.
35