Conceptio › Archive › arXiv CS
arXiv CSopen access

Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents Hafsa Akbar1 , Daniel Platnick2,3 , Marjan Alirezaie2,3 , Hossein Rahnama1,2,3 1 MIT Media Lab, Massachusetts Institute of Technology 2 Flybits Labs, Creative AI Hub 3 Toronto Metropolitan University Correspondence: [email protected]

arXiv:2609.21997v1 [cs.MA] 18 Sep 2026

Abstract LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model’s training prior. We introduce Bayesian Chronicle Agents (BCA), a minimal belief layer separating what an agent believes from how it speaks. Each stance is a probability, updated by one Bayesian step per utterance heard. A single prior-strength parameter κ encodes stubbornness, modeled after its role in Friedkin–Johnsen (FJ) opinion dynamics. We then sweep this parameter to yield three canonical regimes of opinion dynamics on demand (consensus, persistent disagreement, committed-minority influence), with persistent disagreement matching the FJ closedform fixed points at R2 = 0.93–0.99. We further show that prescribed κ remains recoverable after the language round-trip, with perfect rank-order recovery across all four models. Explicit belief also makes simulation auditable: the layer surfaces systematic per-model stance biases that end-to-end simulation would silently absorb.

1

Introduction

Large language models increasingly serve as the inhabitants of social simulations: synthetic populations that discuss, persuade, and form collective opinions (Park et al., 2023; Chuang et al., 2024). In most such systems the agent’s opinion is revised by the LLM itself, in context. Two well-documented problems follow. First, populations tend to converge toward model-inherent biases despite assigned personas, a consensus collapse that makes simulated societies unrealistically uniform and prompt-sensitive (Taubenfeld et al., 2024; Chuang et al., 2024). Second, a recent evidence-integration study finds that LLM confidence revisions violate Bayesian norms and are

poorly calibrated under conflicting evidence (Kim et al., 2025). Classical opinion dynamics offers the interpretable machinery these LLM-based simulations lack. DeGroot averaging (DeGroot, 1974) predicts consensus in connected populations; Friedkin– Johnsen (FJ) (Friedkin and Johnsen, 1990) explains lasting disagreement through per-agent stubbornness; committed-minority models study a small unyielding faction moving a flexible majority (Xie et al., 2011; Centola et al., 2018). However, these models operate on bare scalars, with no notion of language which is precisely what LLMs now offer. We ask whether that language capability can be harnessed without losing the interpretability of the classical models. Proposed architecture. To combine the mathematical control of classical opinion dynamics with the language capabilities of LLM agents, we add a small, explicit belief layer between an agent’s persona and its actions (i.e., speech generation for our set-up); we call the resulting agents Bayesian Chronicle Agents (BCA) (Fig. 1). An agent’s stance on a concept is a probability held outside the language model, the simplest instance of a concept– assertion chronicle, our structured identity representation (Platnick et al., 2025; Alirezaie et al., 2024). When the agent hears an utterance, a separate calibrated appraiser model reads it into scalar evidence, and the belief takes one Bayesian update step (§3). A single per-agent knob, the prior strength κ, sets how far each step moves the agent: a stubborn agent carries a large prior tally, so one observation barely moves it. With this architecture, what an agent believes follows a transparent probabilistic rule; how it acts/speaks in consequence is decided by the LLM. The Bayesian belief update layer contributes the following to LLM-based social simulation. (1) Recoverability. We prescribe κ, run the full language

evidence e ∈ [0, 1]

speaker belief α Beta(α, β), b = α+β

generator (LLM, T =0.9)

Expressed Opinion: “Rail would ease the daily gridlock downtown. . . ”

latent, logged

appraiser (LLM, T =0, calib. τ )

listener update b′ =(1−η)b + ηe + ρ(b0 −b)

only signal listeners see

weights η, ρ set by κ, γ

κ κ→0 DeGroot: consensus

0<κ<∞ FJ: persistent disagreement

κ→∞ committed anchor

Figure 1: One round through the belief layer. A speaker’s latent belief is rendered to text by the generator LLM; an appraiser LLM maps the text back to scalar evidence e; each listener takes one Bayesian update step modeled after the FJ update (Eq. 2). κ places the agent on the classical stubbornness axis (bottom).

round-trip, and recover its rank order from the resulting belief dynamics (Spearman 1.0 on every model tested; §5.1). (2) Regime fidelity. One peragent knob (at a global fixed γ, §3) sweeps the classical spectrum: DeGroot consensus (κ → 0), FJ persistent disagreement (finite κ), committed anchors (κ → ∞), each validated against the corresponding closed-form reference on four LLMs (§5.2). (3) Auditability. Because every belief change is a logged event, we can compare what a speaker believed with what listeners were told, i.e., how faithfully the language channel transmits stance, surfacing a systematic distortion in every model we test and which regimes it corrupts. We release our code, prompts, and logged run data to support reproducibility.1

2

Related Work

We situate our work at the intersection of three lines of research: Classical Opinion Dynamics. DeGroot averaging (DeGroot, 1974), Friedkin–Johnsen (Friedkin and Johnsen, 1990), and committed-minority models (Xie et al., 2011; Centola et al., 2018), introduced in §1, offer interpretable, mathematically tractable accounts of consensus, disagreement, and minority influence. A related line, bounded confidence (Hegselmann and Krause, 2002), instead lets agents ignore opinions too far from their own. These models are transparent by construction, but operate on bare scalars, with no notion of language which is exactly the gap an LLM agent can fill. LLM-Based Opinion Simulation. LLM agents converge toward model-inherent biases (Chuang et al., 2024) and only partially toward assigned personas (Taubenfeld et al., 2024), yet can also reproduce classical signatures such as minority tipping (Flint Ashery et al., 2025). At the individual 1

https://github.com/hafsa-akbar/ bayesian-chronicle-agents

level, LLM belief revision itself is found to violate Bayesian norms under conflicting evidence (Kim et al., 2025). Those dynamics emerge from the model’s prior; ours are prescribable per agent and, as we show, recoverable. Explicit Belief State for LLM Agents. Generative-agent memory architectures (Park et al., 2023) and structured identity models (Platnick et al., 2025; Alirezaie et al., 2024) persist in content but leave belief revision to the LLM entirely or keep it static. The closest work to fill that gap is the concurrent Belief Engine (Yang et al., 2026): an auditable log-odds accumulator with two empirical controls (evidence uptake, prior anchoring), evaluated through controlled singleand two-agent debates and human-trajectory replay. Our approach is complementary and differs on three axes: our single knob has an exact classical correspondence rather than an empirical role; we test parameter recoverability (prescribe, then recover through the language round-trip), unevaluated in Belief Engine; and we validate at population scale against closed-form regime references.

3

The Belief Layer

The belief layer has three parts: a structured representation of what an agent believes (§3.1), a pair of LLM components that translate between belief and language (§3.2), and a Bayesian rule that updates belief from what the agent hears (§3.3). 3.1

Identity Representation.

We represent agent identity as a graph of beliefholding concept nodes (propositions the agent holds opinions about) held outside the language model, a simplified variation of the identity chronicle i.e., knowledge graph learned from a person’s digital footprint and shareable as a “borrowable identity” (Alirezaie et al., 2024). Each concept carries an assertion set, the competing natural-

language stances on it (here an opposing pair {a+ , a− }), and a credence: a probability distribution over those assertions, whose dominant entry is the stance surfaced in agent behavior. 3.2

From Belief to Action.

Two LLM components connect the numeric belief state to natural language (Fig. 1). The generator is the agent’s voice: when the agent speaks, its current credence b is rendered as a plain-language stance descriptor in the prompt (see App. D for the exact mapping), conditioning the generated action (here, the expressed opinion). The appraiser is the agent’s ears: when the agent hears another speaker, the appraiser reads the utterance and returns a single number e ∈ [0, 1], its judgment of how strongly the text supports a+ over a− for the concept under discussion (the architecture supports appraising multiple concepts per utterance, but our single-concept experiments below exercise only one). Only this appraised evidence reaches the belief: the LLM controls what is said and how it is understood, while every change to what the agent believes passes through the Bayesian update, keeping the dynamics controllable and auditable. 3.3

The Belief Update

The layer has two parameters: prior strength κ = α0 + β0 (α0 , β0 are the prior pseudo-counts the agent starts with), which encodes the agent’s stubbornness, and a forgetting factor γ ∈ (0, 1] that sets how quickly old evidence fades relative to the founding prior. We hold γ fixed at 0.7 for every agent (chosen by a no-LLM sweep, App. A), so κ is the single per-agent knob we move; exact Bayes (γ=1) provably forces every population to consensus (App. E.3) and is kept as the no-forgetting ablation (§5.2). Each heard utterance contributes one unit of evidence, split between the two stances by the appraised value e and weighted by a fixed evidence weight w. We fix w = 1 in all our experiments: α ← α0 + γ(α − α0 ) + we, β ← β0 + γ(β − β0 ) + w(1 − e).

(1)

Intuitively, the total pseudo-count n = α + β is how much the agent already “knows”. Written in terms of the credence, the core of one step of Eq. 1 is the convex blend b′ = (1 − η) b + η e,

(2)

with susceptibility η = w/n′ , where n′ is the postdiscount count (App. E.2): a stubborn agent (large κ, hence large n) barely moves, a pliable one moves a lot. With forgetting on (γ < 1), the exact step also adds a pull ρ (b0 − b) back toward the agent’s initial opinion—the defining ingredient to letting population dynamics like FJ persist. Closed forms and derivations are in App. E.

4

Experimental Setup

All runs use a single synthetic policy question with two opposing stances, so no model brings a strong real-world prior to it (concept, stances, and prompts in App. D). We run the experiment with N =20 agents on a complete graph, 5 seeds × 20 roundrobin turns per condition (a fixed κ assignment across agents, defined in § 5). Each round every agent speaks once (generator temp. 0.9); each comment is read once by the appraiser (temp. 0), which turns it into the scalar evidence e of Eq. 1, consumed by every listener. Finally, an independent per-round 0–100 self-report on that particular concept checks that agents faithfully express what they actually believe (probe prompt in App. D; audit in App. C). We repeated our experiments with gpt-5.4-mini, gpt-5.4 (OpenAI), Llama-4-Scout-17B-16E (open weights, Meta), and claude-sonnet-4-6 (Anthropic), under the same sampling parameters across all four models. We found LLM appraisers tend to state overconfident probabilities, so we fit one temperature τ per model on 100 labelled utterances, reducing held-out calibration error (App. B).

5

Results

5.1

Prescribed κ is recoverable through the language channel

Is the prescribed stubbornness recoverable from behaviour after the full language round-trip? For each model we sweep κ ∈ {0.5, 1, 2, 4, 8, 16, 32} (14,000 utterances, 266,000 listener events per model) and invert the exact one-step identity (App. E) per event: a listener at belief b who hears a speaker with latent belief s and moves to b′ yields κ̂ =

w (s − b) − (b′ − b) (γm + w) , (b′ − b) − (1 − γ) (b0 − b)

(3)

with m the listener’s discounted evidence count and b0 its initial belief (notation, derivation, and numerical-exactness check in App. E). Using the

Figure 2: One knob, three classical regimes (gpt-5.4-mini; one line per agent, averaged over seeds; all four models replicate these regimes, Tab. 2). (a) When everyone is pliable, opinions merge into a single consensus (the late drift is mild channel bias, App. C). (b) Add stubborn agents and disagreement persists, settling onto the dotted theory-predicted levels. (c) Turn forgetting off and the same population collapses to consensus anyway. (d) A minority that never budges pulls the majority smoothly, with no tipping point.

speaker’s latent belief s instead of the appraised evidence e (which would recover κ by construction) means recovery succeeds only if the generated text carried that belief and the appraiser read it back out, so any remaining error is attributable to the language channel. We discard ill-conditioned events where the listener already nearly agrees with the speaker (|s − b| ≤ 0.05): such utterances barely move the listener regardless of κ, so they carry no information about stubbornness (retention statistics in App. A). On all four models the channel is faithful (appraised evidence tracks latent speaker belief, Pearson 0.94–0.97) and κ is recovered perfectly in rank order (Spearman 1.0 on every model; Fig. 5 and per-model numbers in App. C). Recovered magnitudes are uniformly attenuated (e.g., κ=32 → 25– 30; per-condition medians in Tab. 3): the channel inflates the evidence–belief gap (App. C), so agents appear somewhat more pliable than prescribed. Thus, within BCA, prescribed κ remains behaviorally recoverable after passing through the language channel. 5.2

One knob, three classical regimes

We dial only κ and ask whether the population produces each canonical regime, judged against the corresponding closed-form reference (Fig. 2; full cross-model numbers in Tab. 2, App. C). Consensus (pliable limit κ → 0; here κ=0.1). When no one is stubborn, DeGroot theory predicts consensus near the average initial opinion. Spread indeed collapses by three orders of magnitude on all four models, but where consensus lands exposes the channel: each channel misreads stances in a characteristic direction, and pliable agents— believing whatever they hear—accumulate the misreading. Against the true initial mean of

0.50, consensus lands at 0.50 (gpt-5.4) and 0.44 (gpt-5.4-mini), whose symmetric misreadings largely cancel, but at 0.15 (claude-sonnet-4-6) and 0.01 (Llama-4-Scout), whose one-sided misreadings compound. An end-to-end simulation would report these displaced consensuses as findings; the belief layer instead detects and measures the bias (App. C). Persistent disagreement (FJ; stubborn camps κ=16 at the extremes, pliable κ=1 agents between). Once some agents are stubborn, opinions should stop merging: camps hold their ground and pliable agents settle between them. We observe this stable spread, with per-agent final beliefs matching the FJ fixed point at R2 = 0.93–0.99 across models. At γ=1 the same population collapses back to near-consensus and κ-recovery degrades sharply, showing forgetting is the enabling ingredient (§3; full ablation on gpt-5.4-mini: App. A, Tab. 1). Minority influence (committed κ → ∞ minority vs. free κ=4 majority). The affine update provably cannot produce a tipping point (App. E.5), and none appears: sweeping the committed fraction from 5% to 40%, the majority’s mean rises smoothly on every model (Fig. 2d), in contrast to critical-mass experiments (Xie et al., 2011; Centola et al., 2018).

6

Conclusion

A Bayesian belief layer makes LLM social simulation controllable in the currency the opiniondynamics field already trusts: prescribed κ remains recoverable through the full language round-trip in exact rank order (Spearman 1.0 on four models), it controls the classical regime spectrum against closed-form references, and its logged events make simulation auditable, surfacing per-model channel bias. Multi-concept identities, directional (not just

confidence) channel correction, and heterogeneous topologies are natural next steps.

Limitations Our study isolates the mechanism at the cost of scope: one synthetic topic, a complete graph, N =20 agents, binary stances, and validation against classical theory rather than human trajectories, so these simulations are not predictions of human opinion change. Our engine is also sequential (one utterance at a time), whereas the theoretical reference is synchronous mean-field; nevertheless, terminal beliefs match the corresponding FJ fixed points at R2 = 0.93–0.99 across models, suggesting that this mismatch is small in the setting studied (App. E.3). The appraiser is itself an LLM: calibration corrects the appraiser’s confidence, but directional distortion of the channel (whether introduced in rendering or reading) passes through and still shapes dynamics in the fully pliable regime. On the data side, all conversation in our runs is model-generated by construction; the only humanlabelled data is the small, single-topic calibration set of App. B, so the appraiser’s calibration is only as strong as those 100 judgments. And because the topic is deliberately fictional, chosen so that models bring no strong pre-trained stance, our results do not speak to debates where they do. Finally, our pipeline requires two LLM calls per utterance across four models, which may limit exact reproducibility without comparable API access.

References Marjan Alirezaie, Hossein Rahnama, and Alex Pentland. 2024. Structural learning in the design of perspectiveaware AI systems using knowledge graphs. In AAAI 2024 International Workshop on AI for Digital Human. Damon Centola, Joshua Becker, Devon Brackbill, and Andrea Baronchelli. 2018. Experimental evidence for tipping points in social convention. Science, 360(6393):1116–1119. Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T. Rogers. 2024. Simulating opinion dynamics with networks of llmbased agents. In Findings of the Association for Computational Linguistics: NAACL 2024. Morris H. DeGroot. 1974. Reaching a consensus. Journal of the American Statistical Association, 69(345):118–121.

Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli. 2025. Emergent social conventions and collective bias in llm populations. Science Advances, 11(20):eadu9368. Noah E. Friedkin and Eugene C. Johnsen. 1990. Social influence and opinions. The Journal of Mathematical Sociology, 15(3–4):193–206. Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, pages 1321–1330. Rainer Hegselmann and Ulrich Krause. 2002. Opinion dynamics and bounded confidence: Models, analysis, and simulation. Journal of Artificial Societies and Social Simulation, 5(3). Minsu Kim, Sangryul Kim, and James Thorne. 2025. From evidence to belief: A bayesian epistemology approach to language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies. Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Daniel Platnick, Mohamed E. Bengueddache, Marjan Alirezaie, Dava J. Newman, Alex Pentland, and Hossein Rahnama. 2025. ID-RAG: Identity retrievalaugmented generation for long-horizon persona coherence in generative agents. In LLAIS 2025: Workshop on LLM-Based Agents for Intelligent Systems, at ECAI 2025. Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Goldstein. 2024. Systematic biases in llm simulations of debates. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 251–267. Jierui Xie, Sameet Sreenivasan, Gyorgy Korniss, Weituo Zhang, Chjan Lim, and Boleslaw K. Szymanski. 2011. Social consensus through the influence of committed minorities. Physical Review E, 84(1):011130. Joshua C. Yang, Maurice Flechtner, Damian Dailisan, and Michiel A. Bakker. 2026. Belief engine: Configurable and inspectable stance dynamics in multi-agent llm deliberation. arXiv preprint arXiv:2605.15343.

A

Choosing γ: the oracle sweep and full ablation

A mechanism-only sweep (oracle evidence, no LLM calls) over γ ∈ {1.0, 0.95, 0.9, 0.8, 0.7, 0.6, 0.5} shows final FJ cross-agent variance rising from 9×10−5 at γ=1 to 0.056 at the γ=0.7 operating point (0.066 by γ=0.5), while DeGroot remains at consensus for every γ (Fig. 3). We fix γ=0.7 before any paid LLM run and hold it constant across all models and regimes. Table 1 reports the full ablation through the language channel.

reads in deployment. Each calibration utterance carries a soft label: a graded gold target y ∈ [0, 1] for its stance strength, assigned by hand by a single human annotator (e.g., 0.8 for a clear but not absolute lean toward a+ , 0.5 for balanced); each label carries a one-line rationale in the released data. The appraiser states p+ ; calibration applies temperature scaling (Guo et al., 2017), pcal = σ(logit(p+ )/τ ), with one τ per model fit by minimising the cross-entropy −[y log p + (1 − y) log(1−p)] on an 80-example split and evaluated on 20 held-out utterances. Validation ECE improves 0.101 → 0.052 (gpt-5.4-mini), 0.103 → 0.046 (gpt-5.4), 0.040 → 0.034 (Llama-4-Scout), 0.081 → 0.033 (claude-sonnet-4-6). Calibration is independent of κ and γ; each fit is frozen and reused by all downstream runs of that model.

C

Figure 3: Oracle γ sweep (belief mechanism only, no LLM calls): final cross-agent variance in the FJ configuration rises as soon as γ < 1 and saturates near the chosen operating point, while the DeGroot configuration reaches consensus for every γ.

Metric (language channel)

γ=0.7

γ=1

FJ final variance / persists FJ R2 vs. fixed point κ-recovery rel. error Well-conditioned events DeGroot final variance Belief–slider audit (r)

0.053 / yes 0.97 0.49 0.76 5×10−5 0.99

10−4 / no −1.1 1.41 0.12 7×10−7 0.98

Table 1: Forgetting ablation (gpt-5.4-mini, 5 seeds). Removing forgetting kills FJ persistence and substantially degrades κ-recoverability; consensus and auditability are unaffected

Retention of well-conditioned events after the §5.1 filters rises monotonically with prescribed stubbornness (gpt-5.4-mini: 0.45 at κ=0.5 to 0.92 at κ=32, closely matched by the other three models): low-κ conditions are noisier, not selectively discarded.

B

Appraiser calibration per model

The 100 utterances were LLM-generated to span all stance bands, matching the medium the appraiser

Auditing the language channel

How each model transmits stances. For every utterance the layer logs two numbers: the belief the speaker actually held (s) and the evidence the appraiser reported to listeners (e). Averaging e against s draws each model’s channel transfer curve (Fig. 4a); a perfectly faithful channel would sit on the diagonal. All models show some deviation. The GPT channels exaggerate: moderate stances arrive as more extreme, on both sides of neutral. Llama-4-Scout transmits every stance as leaning somewhat more toward a− , its whole curve sitting below the diagonal. The Claude channel exaggerates only the a− side (e.g., a mild 0.35 stance arrives as 0.24) while staying roughly faithful on the a+ side. These are properties of how each model transmits stance through the generate– appraise round-trip, not opinions about the topic itself. The curve measures the generator and appraiser jointly; separating rendering from reading (e.g., by cross-appraising one model’s utterances with another model’s appraiser) is left to future work. Why the shape of the misreading matters. A symmetric exaggeration pushes some readings up and others down, so over a balanced population the errors cancel and the consensus stays put— which is why the GPT populations land near 0.50 (Fig. 4b). A shifted or one-sided curve instead injects a small push in the same direction every round. Pliable agents have nothing to resist it with, so the pushes accumulate into the large drifts of Llama

and Claude; stubborn agents are re-anchored by their priors at every step, so the same push cannot accumulate, and the FJ regime lands near theory on all four models (Fig. 4c). Why appraiser calibration addresses a different failure mode. Temperature scaling fixes overconfidence: an appraiser that tags evidence e = 0.9 on an utterance that supports the assertion only at 0.7 is pulled back toward the label. But it rescales confidence symmetrically about the neutral point, so it cannot raise, lower, or bend one side of the transfer curve: directional bias of the language channel passes through. That bias is an artefact of the channel, not of the belief-update layer, and, because every event is logged, it is measurable in our architecture rather than silently absorbed. Does the hidden belief govern behaviour? A separate check: each round every agent also gives an independent 0–100 self-report of its stance (App. D), which never enters any belief update. These self-reports track the hidden belief of an agent closely (r ≈ 0.98–0.99 over 18,000 reports per model), so the Beta state is not internal bookkeeping: it is what the agent expresses in a personacoherent manner which is verified through a channel separate from the generator–appraiser loop. Table 2 collects the headline metrics for all four models.

D

Prompts and protocol

Concept. The simulated debate concerns one municipal policy question for the fictional town of Aldenvale (transit_priority), with the opposing assertion pair a+ : “expand rail” and a− : “keep roads”. The town and the question are invented so that no model brings a strong pre-trained prior to either side. Generator (system): “You are a resident of the fictional town of Aldenvale taking part in a community discussion about local transportation policy. You speak naturally, like a real person at a town meeting.” (user): the two stances, a naturallanguage descriptor of the agent’s private stance strength, and an instruction to write a 1–2 sentence comment without mentioning numbers, probabilities, or internal variables. Appraiser (system): “You are an impartial stance classifier . . . judging the evidence in the text rather than your own opinion.” (user): both stances, the statement, and a request for {"p_plus": <0..1>}. Probe: an inde-

pendent request for a single 0–100 integer stance report. Full templates, seeds, and the round-robin event loop are in the released code. The stance descriptor is a fixed nine-bin qualitative mapping from b to phrases, from “completely opposed to rail expansion and fully committed to roads” (b < 0.10) through “genuinely torn and balanced” (0.45 < b ≤ 0.55) to “completely committed to rail expansion” (b ≥ 0.90); no numbers appear in any prompt.

E

Derivations and proof sketches

Notation. An agent’s state on the concept is the pair of pseudo-counts (α, β): accumulated evidence for a+ and a− respectively. The total count is n = α + β and the credence is b = α/n. The prior counts are (α0 , β0 ) with prior strength κ = α0 +β0 and prior mean b0 = α0 /κ (the agent’s initial stance). Each heard utterance delivers appraised evidence e ∈ [0, 1] with weight w; c counts how many utterances the agent has absorbed; γ ∈ (0, 1] is the forgetting factor. Subscripts index update steps: nc is the count after c updates. E.1

The count is deterministic

Summing the two lines of Eq. 1 makes the evidence e cancel: nc+1 = κ + γ (nc − κ) + w. Unrolling from n0 = κ gives ( c κ + w 1−γ 1−γ nc = κ + wc

γ < 1, γ = 1.

So the count, that is how much the agent “already knows”, is a known function of (κ, γ, c), independent of what was said. This is the fact the κrecovery estimator exploits (§E.4). For γ < 1 the count saturates at n⋆ = κ + w/(1 − γ); for γ = 1 it grows forever. E.2

One step is a Friedkin–Johnsen update

Substituting α = b n into the α-line of Eq. 1 and dividing by nc+1 yields the exact per-step identity b′ = b + ηc (e − b) + ρc (b0 − b), with susceptibility ηc = w/nc+1 and anchor weight ρc = κ(1 − γ)/nc+1 . Each utterance therefore moves the credence toward the evidence, plus (when forgetting is on) a pull back toward the

Figure 4: The belief layer as a measurement instrument. (a) Every model’s channel transfer function deviates from identity, each in its own way (GPTs: symmetric expansion; Llama: uniform downward shift; Claude: asymmetric negative-side amplification). (b) Pliable DeGroot populations compound asymmetric distortion into a confidently wrong consensus (final values labelled; reference 0.50). (c) Displacement from the theoretical reference: large for the pliable regime, small for the prior-anchored FJ regime on the same models—stubbornness absorbs channel bias. gpt-5.4-mini

gpt-5.4

Llama-4-Scout claude-sonnet-4-6

Weights / provider closed / OpenAI closed / OpenAI Appraiser temperature τ 1.89 2.15 Channel alignment (Pearson) 0.96 0.97 DeGroot consensus (ref. 0.50) 0.44 0.50 FJ persists / R2 yes / 0.97 yes / 0.98 Committed: smooth, no tipping yes yes κ-recovery (Spearman) 1.00 1.00 κ-recovery (rel. error) 0.49 0.46 Auditability, belief–slider (r) 0.99 0.99

open / Meta 1.51 0.94 0.01 yes / 0.93 yes 1.00 0.33 0.98

closed / Anthropic 1.69 0.96 0.15 yes / 0.99 yes 1.00 0.27 0.99

Table 2: Full pipeline across four models (γ=0.7, 5 seeds, 20 rounds, per-model τ refit). Structural results replicate on all four; κ rank-recovery is Spearman 1.0 for every individual seed, not only the pooled per-condition medians of Fig. 5.

prescribed κ

0.5

1

2

4

8

16

32

gpt-5.4-mini 0.0 0.3 1.0 2.5 5.6 12.1 25.3 gpt-5.4 0.0 0.3 1.0 2.6 6.0 12.6 26.2 Llama-4-Scout −0.1 0.4 1.6 3.8 7.6 14.8 30.0 claude-sonnet-4-6 0.3 0.5 1.3 3.4 7.0 14.0 28.7

Table 3: Recovered stubbornness (oracle-aligned; percondition medians pooled over seeds) for every prescribed κ. Rank order is exact in every row; magnitudes are uniformly attenuated, increasingly so at low κ. Figure 5: Prescribe-then-recover: κ is recoverable through the language channel. Prescribed stubbornness (x) vs. value recovered from listeners’ belief movements (Eq. 3). Rank order is exact on all four models; magnitudes are uniformly attenuated by the channel (Tab. 3).

agent’s initial stance—exactly the two defining ingredients of the FJ update: susceptibility to the social signal and anchoring to one’s founding opinion (Friedkin and Johnsen, 1990). Stubbornness enters as claimed: ηc is strictly decreasing in κ. At γ=1 the anchor vanishes (ρc = 0) and the step reduces exactly to the two-term blend of Eq. 2 with η = w/(nc + w).

E.3

Exact Bayes forces consensus; forgetting prevents it

At γ=1 the update accumulates evidence without discount, so after t updates  b(t) = ρ(t) b0 + 1 − ρ(t) ē(t) ,

where ρ(t) = κ/(κ + wt) is the residual prior weight and ē(t) the running average of the evidence heard. The prior weight ρ(t) → 0: every agent’s stance converges onto the shared evidence stream, so on a connected graph disagreement is transient— no assignment of κ can sustain it. For γ < 1 the count saturates (§E.1), so η stays bounded away from zero and setting b′ = b in the identity of §E.2 gives the stationary point b⋆ = λ0 b0 + (1 − λ0 ) ē,

κ(1−γ) λ0 = κ(1−γ)+w ,

E.5

No tipping point

Consider the complete graph with a fraction f of agents committed (frozen) at b=1 and the free (0) agents sharing κ and a common starting stance bF . By symmetry all free agents hold a common belief (t) bF , and the average evidence a free agent hears in (t) round t is f · 1 + (1 − f ) bF (exactly f N/(N −1) once a free agent’s own belief is excluded from its neighbours; with N =20 the two are indistinguishable). Substituting into the blend gives (t+1)

the classical FJ steady state: the initial stance keeps a permanent weight that grows with stubbornness, so heterogeneous κ sustains heterogeneous opinions forever. (λ0 is the equilibrium mixture weight ρ∞ /(ρ∞ + η∞ ), not the limit of the per-step ρc of §E.2.) For the population reference of Fig. 2b and the R2 of §5.2: with N agents on the complete graph, zero-diagonal uniform weights Wij = 1/(N −1), and per-agent susceptibility λi = w/(κi (1 − γ) + w), the theory lines are the synchronous (meanfield) FJ fixed point b⋆ = (I − ΛW )−1 (I − Λ) b0 with Λ = diag(λi ), i.e. the stationary belief vector if every agent updated from one simultaneous snapshot of its neighbours, rather than our engine’s sequential, one-utterance-at-a-time schedule. R2 compares per-agent terminal beliefs (5seed averages) against b⋆ directly, without refitting (1 − SSres /SStot , no intercept); that the fit is 0.93– 0.99 despite this mismatch indicates the sequential dynamics track the mean-field reference closely. E.4

Recovering κ from behavior

This derives Eq. 3. Take the exact one-step identity of §E.2, b′ = b+(w/n′ )(e−b)+(κ(1−γ)/n′ )(b0 − b) with n′ = κ + γm + w, where m = w(1 − γ c )/(1 − γ) is known from the listener’s event history (§E.1) and b0 is its logged initial belief. Substituting the speaker’s latent belief s for the evidence e (the channel test of §5.1) and solving the resulting linear equation for κ gives Eq. 3. Because the identity is exact at every γ, so is the inversion: applied to the logged events with the consumed evidence e it returns the prescribed κ to numerical precision on all four models, and at γ=1 it reduces to the familiar κ̂ = wη̂ −1 − m − w with η̂ = (b′ − b)/(s − b). All deviation that remains when s replaces e is therefore attributable to the language channel.

(t)  1 − bF , Y  (T ) (0)  bF = 1 − 1 − bF 1 − ηt f .

1 − bF

= 1 − ηt f



t<T

As a function of f this is continuous, smooth, and strictly increasing: no discontinuity, no bistability, no critical mass. (With forgetting on, the one-step (0) (t) coefficients gain the anchor term ρt (bF − bF ) of §E.2 and remain affine in f ; composing T such (T ) steps makes bF polynomial, hence still smooth, in f .) Tipping requires a nonlinearity (e.g., acceptance thresholds or majority rules) that an affine update does not contain; the experiments confirm the predicted smooth dose–response. The dashed curve in Fig. 2d is a stylised full-following reference, not a proven bound: it instantiates one aggregated unitweight update per round, an idealisation of the sequential engine (which applies N −1 per-utterance updates per round); with forgetting on, the anchor term of §E.2 additionally pulls free agents toward their starting stance. Both effects place the measured dose–response below this reference curve in our runs.

Record · ID 1006884 · SHA-256 71cb460a7cdb2df4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.