Conceptio › Archive › arXiv CS
arXiv CSopen access

Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

arXiv:2605.10931v1 [math.AP] 11 May 2026

Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

Albert Alcalde Department of Mathematics Friedrich-Alexander-University Erlangen-Nürnberg [email protected]

Leon Bungert Institute of Mathematics CAIDAS University of Würzburg [email protected]

Konstantin Riedl Mathematical Institute University of Oxford Reuben College [email protected]

Tim Roith CIT School Technical University of Munich Munich Center for Machine Learning [email protected]

Abstract Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the evolution of tokens in deep encoder-only transformers at inference time which is described in the large-token limit by a mean-field continuity equation. Leveraging ideas from the convergence analysis of interacting multiparticle systems, with particles corresponding to tokens, we prove that the token distribution rapidly concentrates onto the push-forward of the initial distribution under a projection map induced by the key, query, and value matrices, and remains metastable for moderate times. Specifically, we show that the Wasserstein distance p of the two distributions scales like log(β+1)/β exp(Ct) + exp(−ct) in terms of the temperature parameter β −1 → 0 and inference time t ≥ 0. For the proof, we establish Lyapunov-type estimates for the zero-temperature equation, identify its limit as t → ∞, and employ a stability estimate in Wasserstein space together with a quantitative Laplace principle to couple the two equations. Our result implies that for time scales of order log β the token distribution concentrates at the identified limiting distribution. Numerical experiments confirm this and, beyond that, complement our theory by showing that for finite β and large t the dynamics enter a different terminal phase, dominated by the spectrum of the value matrix.

1

Introduction

The transformer architecture has driven in recent years the unprecedented successes of modern machine learning, establishing itself as the fundamental backbone for large language models [7, 12, 23, 30], and foundation models [11]. Central to its design is the self-attention mechanism [8, 45], which enables the model to capture intricate, long-range contextual dependencies within the input data. A rigorous characterization of how these internal representations evolve through the network, however, remains a significant analytical challenge. To mathematically study how tokens {xi }ni=1 ⊂ Rd evolve through the self-attention modules of a deep encoder-only transformer, we follow [14, 16, 29, 41] and model the forward pass of a Preprint.

transformer with single-head attention and without fully-connected MLPs, in the infinite-depth limit, as a continuous-time dynamical interacting particle system of the form   n X exp (β⟨Q(t)x (t), K(t)x (t)⟩) i j Pn ẋi (t) = Pxi (t)  V (t)xj (t). (1) exp (β⟨Q(t)x (t), K(t)x (t)⟩) i k k=1 j=1 System (1) governs the evolution of tokens {xi }ni=1 , interpreted as particles, through the self-attention layers of the transformer at inference. Here, Q(t), K(t) ∈ Rdint ×d and V (t) ∈ Rd×d denote the query, key, and value weight matrices of the attention modules, while β > 0 is an inverse-temperature parameter [14, 18, 29]. The weight matrices have been obtained during training and are fixed for the inference dynamics, which are the focus of the present work. Here, Px (y) := y − ⟨x, y⟩ x for y ∈ Rd is the orthogonal projection onto the tangent space Tx Sd−1 of the sphere Sd−1 at a point x ∈ Sd−1 , and models layer normalization. This is motivated by the simplified root mean square normalization x 7→ g ⊙ x/∥x∥ with gain g ∈ Rd [47], which is used, e.g., in Llama models [44]. We set g = 1 and refer to [31] and [16, Section 1(b)] for a recent review of modeling choices for normalizations. Mathematical setting. Let us now introduce the large-token regime of (1), which serves as the analytical framework for this paper. As done in the largest part of the current literature [14, 16, 29, 41] and justified by memory-efficient variants of transformers such as Universal Transformers [21] and ALBERT [34], we assume in our analysis that the weight matrices Q, K, and V are identical throughout the layers, i.e., that Q(t) ≡ Q, K(t) ≡ K, and V (t) ≡ V for all t ≥ 0. For notational simplicity and without loss of generality, we henceforth write B := Q⊤ K ∈ Rd×d . After positional encoding [45, Section 3.5], the transformer dynamics (1) are invariant under permutaPn tions allowing the empirical measure ρβ,n = n1 i=1 δxi (t) to fully characterize the dynamics, which t is described by a continuity equation of the form (2) below. Our theoretical analysis will take place in the infinite-token regime, where the temporal dynamics of the token distribution (ρβt )t≥0 ⊂ P(Sd−1 ) with ρβ0 = ρ0 ∈ P(Sd−1 ) is captured by the transformer continuity equation !! Z exp (β⟨x, By⟩) β β β ∂t ρt = − div ρt Px V y dρt (y) , (2) R β Sd−1 Sd−1 exp (β⟨x, Bz⟩) dρt (z) which describes a flow in the space P(Sd−1 ) of probability measures over the sphere Sd−1 . Here, div and ∇ denote the manifold divergence and gradient on the sphere Sd−1 , respectively. It is well-known that the empirical measure ρβ,n defined above is a weak solution to (2), however, this equation allows t for more complex solutions if the initial condition is not supported on a finite set. As shown in the seminal work [14] under invertibility assumptions on B and V , and Assumption 3 on the initial datum ρ0 , solutions of (2) converge in the zero-temperature limit β → ∞ to those of    V B⊤x ∂t ρt = − div ρt Px . (3) ∥B ⊤ x∥

Moreover, for large times, solutions of (3) concentrate on the dominant eigenspace of the matrix V B ⊤ , which we denote by E. Correspondingly, for such a linear subspace E of Rd , we denote by PE : Rd → E the orthogonal projection onto E and define Π : Sd−1 \ E ⊥ → Sd−1 as Π(x) :=

PE x . ∥PE x∥

(4)

Main Contributions. In this work, we study the evolution of tokens in deep encoder-only transformers at inference time by analyzing the mean-field token distribution ρβt of the transformer continuity equation (2). We quantify its concentration at the push-forward Π♯ ρ0 of the initial datum under the projection map (4), which is determined by the model weights, in terms of the temperature parameter β −1 and inference time t. Our proof leverages tools from the analysis of interacting particle systems, in particular, stability estimates in Wasserstein space involving a quantitative Laplace principle, and Lyapunov techniques to quantify convergence speeds. Our main contributions are as follows. • Identification of Π♯ ρ0 as the stationary distribution of the zero-temperature continuity equation (3) for symmetric matrices V B ⊤ , and proof of exponential convergence of the zerotemperature solution ρt to the dominant eigenspace of V B ⊤ in the Wasserstein distance. 2

• Derivation of a quantitative Wasserstein stability estimate between solutions ρβt of the transformer continuity equation (2) and the zero-temperature limit ρt .

• Quantification of the approximation error between ρβt and Π♯ ρ0 in Wasserstein distance as p

log(β+1)/β exp(Ct) + exp(−ct), indicating an interplay between temperature β −1 and time t.

• Characterization of the non-asymptotic concentration regime of the mean-field transformer •

dynamics up to times t = O(log β), corresponding to the early-phase behavior most relevant for practical finite-depth models. Experimental identification of the terminal phase after times t = O(log β), where concentration shifts from the dominant eigenspace of V B ⊤ to the one of V .

Organization. Following the presentation and discussion of our theoretical contributions in Section 2, Section 3 puts these findings into context within the current literature. Thereafter, in Section 4, we present proof details for our statements, with technical results deferred to Appendices A and B. Section 5 then illustrates and complements our theoretical findings with numerical experiments, with additional experiments provided in Appendix C. We conclude the paper with Section 6.

2

Main results

This section is dedicated to discussing the main theoretical contributions of our paper. We first state the assumptions needed for proving our main results, before elaborating on them in more detail. Assumption 1. The weight matrices B and V are such that A1 the matrix B = Q⊤ K is invertible, A2 the matrix V B ⊤ ∈ Rd×d is symmetric, with largest eigenvalue µ1 having multiplicity k and second largest eigenvalue µ2 satisfying γ := µ1 − µ2 > 0 (with the convention that γ = ∞ if k = d). We denote by E the eigenspace of V B ⊤ with respect to µ1 and by PE : Rd → E the orthogonal projection onto E. R 2p E )x∥ Assumption 2. The initial datum ρ0 ∈ P(Sd−1 ) is such that Vp (ρ0 ) := Sd−1 ∥(Id−P dρ0 (x) ∥PE x∥2p is finite for some p ∈ (0, 1]. Assumption 3. The initial datum ρ0 ∈ P(Sd−1 ) has a Lipschitz continuous density f0 with respect to the surface measure Hd−1 Sd−1 that satisfies ℓ0 := minx∈Sd−1 f0 (x) > 0.

Under Assumptions 1–3 on the weight matrices B and V , and the initial datum ρ0 , we derive in the subsequent statement quantitative concentration bounds for the transformer continuity equation (2). We denote the smallest and largest singular values of a matrix by σmin (•) and σmax (•). Theorem 1. Assume that the matrices B and V satisfy Assumption 1 and that the initial distribution ρ0 ∈ P(Sd−1 ) satisfies Assumption 2 for some p ∈ (0, 1] as well as Assumption 3. Let (ρβt )t∈[0,T ] denote the weak solution to the transformer continuity equation (2). Then it holds q    pγ C1 t C0 t W2 (ρβt , Π♯ ρ0 ) ≤ 2 log(β+1) e − e + V (ρ ) exp − t , p 0 β σmax (B) for constants C1 = C1 (d, σmax (B), σmax (V )) > C0 = C0 (d, σmax (B), σmax (V ), σmin (B), ℓ0 ) > 0. Before discussing Theorem 1 and its assumptions in the remainder of this section, let us provide as a corollary a quantitative convergence estimate for the zero-temperature transformer continuity equation (3). It follows directly from the proof of Theorem 1 (see equation (10)) in Section 4. Corollary 2. Assume that the matrices B and V satisfy Assumption 1. Moreover, let the initial distribution ρ0 be such that Assumption 2 holds for some p ∈ (0, 1]. Let (ρt )t∈[0,T ] denote the weak solution to the zero-temperature transformer continuity equation (3). Then it holds   pγ W2 (ρt , Π♯ ρ0 ) ≤ Vp (ρ0 ) exp − σmax t . (B) To the best of our knowledge, Corollary 2 is the first result which identifies the stationary distribution of the zero-temperature transformer continuity equation (3). Note, however, that as observed in [14, Remark 3.6] there exist weight matrices B and V (which do not satisfy A2 in Assumption 1) such that only the collapse onto E but no convergence can be expected. 3

(a) t = 0

(b) t = 2.5

(c) t = 5

Figure 1: Illustration of Theorem 1. We observe concentration of ρβ,n at E ∩ Sd−1 (orange circle) t ⊤ when running (1) with n = 2, 000 tokens to simulate (2) for B = R diag(5, 5, 1)R and V = Id, where R is a rotation by π/8. We use β = 30 and a von Mises–Fisher mixture as initial density ρ0 . Discussion of the assumptions. Demanding invertibility of the product B = Q⊤ K of the key and query matrices in A1 of Assumption 1 is relatively common in the field, see, e.g., [14, 17, 28, 29, 40]. It is especially important when working with the zero-temperature continuity equation (3) which is not well-defined unless B is invertible. The symmetry assumption on V B ⊤ from A2 might appear to be relatively strong. It is important to note, however, that without an assumption of this form one cannot hope to be able to characterize the stationary distribution Π♯ ρ0 for (3) as done p in Corollary 2. Still, the quantitative stability result from Proposition 6, stating that W2 (ρβt , ρt ) ≤ 2 log(β+1)/β (eC1 t − eC0 t ), is true without A2. As a consequence, whenever one can analyze the convergence of the zerotemperature solution ρt under more relaxed assumptions than A2, such results extend to ρβt for large β by the triangle inequality. See also Appendix C.2 for related numerical experiments. Moving on to Assumption 2, we note that this assumption implies that ρ0 (E ⊥ ∩ Sd−1 ) = 0 since otherwise the denominator would be zero on a set of positive measure. This is necessary: Since E ⊥ ∩ Sd−1 is invariant for (3) when V B ⊤ is symmetric, any mass initially supported on E ⊥ remains there for all times and cannot flow toward E. However, Assumption 2 is stronger requiring that not too much mass accumulates close to E ⊥ ∩ Sd−1 either. This is necessary to prove Corollary 2 with Lyapunov techniques. We suspect that it might be possible to weaken this assumption or even replace it by ρ0 (E ⊥ ∩ Sd−1 ) = 0 if one is willing to worsen or sacrifice having a convergence rate altogether. Assumption 3 is only needed for the quantitative Laplace principle from Lemma 5. Loosely speaking, for token distributions ρβt with zero mass on open subsets of the sphere, the vector field in (2) is a poor approximation of the one in (3) even for large β. The authors of [14] also used Assumption 3 to show that ρβt has a positive density for all finite times t > 0, and similar assumptions were used in [13, 19, 40]. One possible way to remove this assumption is to add noise, leading to noisy transformer models as recently considered in [10, 40, 43]. This is also exploited in the analysis of the conceptually similar consensus-based optimization model [26]. Note, however, that the exponential convergence of the solution to the zero-temperature continuity equation (3) from Corollary 2 holds without Assumption 3 and is, in particular, valid for singular initial distributions concentrated on finitely many points as long as they do not lie on the lower-dimensional sphere E ⊥ ∩ Sd−1 . Let us remark that at first glance it might seem like Assumptions 2 and 3 contradict each other. After all, the former prevents mass from accumulating close to E ⊥ ∩ Sd−1 , while the latter demands a strictly positive density. However, this is no contradiction if the parameter p and the dimension of the eigenspace E are in a correct relationship to each other, as demonstrated by the following example. Example 3. Since Assumption 3 implies that ρ0 ≥ ℓ0 Hd−1 Sd−1 , it is necessary for AssumpR 2p −2p tion 2 to verify whether Sd−1 ∥(Id − PE )x∥ ∥PE x∥ dHd−1 < ∞. In fact, it suffices to check R −2p whether Sd−1 ∥PE x∥ dHd−1 (x) < ∞. Note that, for p > 0, the singular set of the integrand is ⊥ d−1 S := E ∩ S and equals a (d − k − 1)-dimensional sphere. In a neighborhood of S the quantity ∥PE x∥ scales like the distance of x to S.R The integral over a tubularR ε-neighborhood SRε of S in spherε ε −2p ical coordinates therefore scales like Sε ∥PE x∥ dHd−1 ∼ 0 r12p rk−1 dr = 0 rk−1−2p dr, which is finite if and only if k > 2p. While for p < 12 this condition is automatically satisfied, it 4

can be a restriction for larger p. For instance, if p = 1 it requires the dimension of the dominant eigenspace E of V B ⊤ to be at least three. Interestingly, the condition dim E ≥ 3 also appeared in [19, Theorem 3.1] for analyzing clustering phenomena of mean-field transformer models. Discussion of the main results. Our main result, Theorem 1, encapsulates two different effects that are at play for solutions to the transformer continuity equation (2). The first one is a temperature effect stating that for large values of β (corresponding to low temperatures), solutions ρβt of (2) are close to solutions ρt of the zero-temperature continuity equation p (3) in the Wasserstein distance. In fact, Proposition 6 in Section 4.1 states that W2 (ρβt , ρt ) ≤ 2 log(β+1)/β (eC1 t − eC0 t ), which quantifies [14, Theorem 3.2]. This estimate is not uniform in time due to the exponential factor, and we do not expect it to be either, as demonstrated by our numerical experiments in Section 5. The second effect is the exponential convergence of the zero-temperature limit (3) to a stationary distribution ρ∞ which we show to depend explicitly on the initial datum and to have the form ρ∞ := Π♯ ρ0 . Taking into account the definition of Π in (4), ρ∞ can be thought of as a projection of ρ0 onto E ∩ Sd−1 . The time scales at which these two effects are present can be derived directly from Theorem 1 by investigating which of the two terms dominates. Until times of order O(log β), the exponential convergence prevails. Thereafter, the approximation error due to positive temperature takes over, see also Figure 4 for a numerical illustration of this behavior.    pγ  exp − σmax t , 0 ≤ t ≲ log β, (B) W2 (ρβt , Π♯ ρ0 ) ≲ q  eC1 t log(β+1) , t ≳ log β. β Hence, the most relevant regime for our result is that of time scales of at most O(log β). On such scale, the following corollary of Theorem 1 states that for any given target accuracy ε > 0, there exists β > 0 large enough such that after a time of order log 1ε , the Wasserstein distance of ρβt to Π♯ ρ0 is below ε in a proper time interval. See Appendix A.3 for the proof of the following result. Corollary 4. For any ε > 0 there exists β ≳ ε14 log 1ε large enough such that it holds W2 (ρβt , Π♯ ρ0 ) ≤ ε for all t ∈ [t1 , t2 ] with t1 < t2 , where q     2Vp (ρ0 ) (B) β ε 1 := log 1 + t1 := σmax log and t 2 pγ ε C1 4 log(β+1) . Conjectured long-time behavior. We close this section by emphasizing that there is no reason to believe that after a time scale of order O(log β) the token distribution ρβt is still close to Π♯ ρ0 . In fact, we have numerical evidence suggesting convergence to the dominant eigenspace of V (either in magnitude or not, denoted as F abs and F , respectively) for large times in certain cases (see Figure 2). This observation is supported by the asymptotic stability results from [4], and it neither contradicts our result nor the three-phase analysis of [14]. Phases 2 and 3 described therein apply only once the distribution ρβt has fully collapsed onto E ∩ Sd−1 , and they are derived under the assumption that V = Id. Under this assumption, the following conjecture becomes trivial, since F = Rd and hence F ∩ Sd−1 = Sd−1 . We provide heuristics and numerical experiments supporting the conjecture in Appendices B and C.4, respectively. Conjecture 1. There exist weight matrices B and V ̸= Id satisfying Assumption 1, and a probability measure ρ∞ ∈ P(Sd−1 ) supported on F ∩ Sd−1 or F abs ∩ Sd−1 such that limt→∞ W2 (ρβt , ρ∞ ) = 0.

3

Related work

Let us now relate our results to previous works in the literature on the mathematical analysis of transformer models, and outline their connection to a class of interacting particle systems. Mathematical analysis of transformers. A central body of work models transformer layers through the lens of continuous dynamical systems and partial differential equations (PDEs). This perspective was pioneered in [41] by linking attention to transport-type PDEs, and expanded in [29] by incorporating key architectural components such as layer normalization within a mean-field framework. More recent contributions such as [17] provide a unified treatment of these approaches. Together, 5

(a) t = 0

(b) t = 2.5

(c) t = 4

(d) t = 10

Figure 2: Illustration of Conjecture 1. Using the setup of Figure 1, with the same matrix V B ⊤ , but with V = diag(1, 1, 2) and B adapted accordingly, we observe initial concentration of ρβ,n at t E ∩ Sd−1 (orange circle) before the terminal clustering in F ∩ Sd−1 (red crosses). these works yield a concise mathematical description of transformer dynamics that serves also as our modeling foundation. The authors of [14] identify the zero-temperature limit (3) and establish convergence towards it as β −1 → 0. However, their result is not quantitative, i.e., no rate in β nor its interplay with inference time t is provided. The analysis further identifies additional phases of the dynamics once the solution concentrates on the dominant eigenspace E of V B ⊤ . In contrast, our work takes low but non-zero temperature effects into account and provides quantitative convergence guarantees. This perspective is motivated by the numerical evidence in Section 5 which suggests that β −1 > 0 may prevent full collapse onto the eigenspace E. Relatedly, [16] also focuses on equation (2) with β −1 > 0 and characterize the stationary states of (2) the case B = ±V , where RRin ⟨x,By⟩ the system can be interpreted as a gradient flow of the energy E(ϱ) = e dϱ(x) dϱ(y), either maximizing it if B = V or minimizing it if B = −V . We discuss this case in Appendix C.3 and only emphasize here that our results complement their theory in two ways: we go beyond the gradient flow setting by allowing B ̸= V and, in particular, make concrete statements about the asymptotic behavior of ρβt . Besides mean-field descriptions, a complementary line of research investigates transformer dynamics at the finite-particle level, with an emphasis on qualitative and asymptotic behavior. In particular, clustering and concentration phenomena are analyzed in [19, 29, 40], with extensions to causal attention settings in [1, 32]. The long-time behavior of transformer dynamics has also been studied, including the emergence of metastability [3, 27], multistability with several equilibria related to the spectrum of the value matrix [4], and phase transitions [20], as well as propagation-of-chaos results in large-token limits [13]. Our work differs by considering token distributions with positive densities and providing quantitative estimates of concentration phenomena in an initial non-asymptotic phase of the dynamics where the attention matrix B plays a crucial role. Interacting particle systems. Transformers share striking similarities with classical interacting particle systems (e.g., the Kuramoto model [37]), and optimization methods (e.g., gradient descent [46] or the Frank–Wolfe algorithm [3]). In particular, they are connected (as noted in [29]) to consensus-based optimization (CBO) [36, 39] which we formalize and exploit in what follows. In line with our setting, we consider CBO constrained to the sphere, as in [25]. The aim is to locate a global minimizer of some objective function J : Sd−1 → R by evolving a particle ensemble {xi }ni=1 that follows a coupled system of stochastic differential equations of the form  ẋi (t) = Pxi (t) mβ,ρβ,n + σ · noise . (5) t

The so-called consensus point mβ,ϱ , defined for any probability measure ϱ ∈ P(Sd−1 ), is a weighted mean to which particles with smaller values of J contribute more. In fact, particle weights can also be determined through a pairwise interaction via polarization, as proposed in [15]. Mathematically, this amounts to kernelizing the consensus point as Z κ(x, y) exp (−βJ(y)) R mκβ,ϱ (x) = y dϱ(y), (6) Sd−1 Sd−1 κ(x, z) exp (−βJ(z)) dϱ(z) p where κ : Rd × Rd → R is some kernel. With ∥•∥B := ⟨•, B •⟩ being the energy-norm, we choose   2 2 J(x) = − 12 ∥x∥B , and κ(x, y) = exp − β2 ∥x − y∥B . 6

In the noiseless case σ = 0, equation (5) with consensus point as in (6) leads exactly to the transformer model (1) when V ≡ Id and B ≡ Q⊤ K is time-independent. This interpretation motivated us to leverage techniques from the mean-field analysis of interacting particle systems for optimization— most prominently, the quantitative Laplace principle and Lyapunov techniques—to analyze the transformer model (2). Apart from this, by embedding the self-attention dynamics (1) into the rich family of CBO models [15, 26, 38], one can hope to derive more expressive transformers by changing the kernel, the objective function, or by adding additional drift terms and noise (the latter of which leads to noisy transformers as previously investigated in, e.g., [2, 10, 24, 33, 35, 40, 42, 43]).

4

Proof details for the main results

The proof of Theorem 1 relies on two key ingredients: quantitative stability estimates for continuity equations in Wasserstein space which leverage a quantitative Laplace principle, and Lyapunov-type estimates for the Wasserstein distance to control the convergence of the zero-temperature transformer continuity equation (3). While the former allow to control the Wasserstein distance between the flows ρβt and ρt (see Proposition 6), the latter enable us to quantify the convergence rate of ρt to Π♯ ρ0 (see Proposition 8 in combination with Lemmas 7 and 9). Before presenting these key ingredients in Sections 4.1 and 4.2, we give the proof of Theorem 1, which demonstrates how those tools are used. Proof of Theorem 1. Denoting by (ρβt )t∈[0,T ] the solution of (2), and letting (ρt )t∈[0,T ] denote the solution of (3) with coinciding initial datum ρ0 = ρβ0 , we can bound by the triangle inequality W2 (ρβt , Π♯ ρ0 ) ≤ W2 (ρβt , ρt ) + W2 (ρt , Π♯ ρ0 ).

(7)

The first term on the right-hand side of (7) can be estimated by employing the quantitative stability estimate in Wasserstein space from Proposition 6 for the continuity equations (2) and (3), yielding q  W2 (ρβt , ρt ) ≤ 2 log(β+1) eC1 t − eC0 t (8) β for all t ∈ [0, T ]. For the second term on the right-hand side of (7), we first note that thanks to Proposition 8 it holds for the functional Vp defined in (14) that   2pγ t . (9) Vp (ρt ) ≤ Vp (ρ0 ) exp − σmax (B) With Vp (ρ0 ) < ∞ as of Assumption 2, we have that Vp (ρt ) < ∞ for all t ∈ [0, T ]. This implies that ρt (Sd−1 ∩ E ⊥ ) = 0 for all t ∈ [0, T ], since otherwise the denominator in the integrand of Vp would be zero on a set of positive measure. Since according to Lemma 9 it furthermore holds that Π♯ ρt = Π♯ ρ0 for all t ∈ [0, T ], an application of Lemma 7 together with the bound (9) shows that q q   pγ W2 (ρt , Π♯ ρ0 ) = W2 (ρt , Π♯ ρt ) ≤ 2Vp (ρt ) ≤ 2Vp (ρ0 ) exp − σmax t , (10) (B) p verifying that 2Vp (ρt ) is a suitable Lyapunov functional for W2 (ρt , Π♯ ρ0 ). Inserting (8) and (10) into (7) concludes the proof. 4.1

Quantitative estimates for the Wasserstein distance between the flows ρβt and ρt

To estimate the Wasserstein distance W2 (ρβt , ρt ) of the solutions ρβt and ρt of the transformer continuity equations (2) and (3) in terms of the temperature parameter β, we develop a stability estimate in Wasserstein space, which is based on a quantitative version [26] of the Laplace principle [22]. Denoting by ϱ ∈ P(Sd−1 ) a fully supported probability measure, for any fixed x ∈ Sd−1 the vector  Z exp β⟨B ⊤ x, y⟩ R := mβ,ϱ (x) y dϱ(y) (11) ⊤ Sd−1 Sd−1 exp (β⟨B x, z⟩) dϱ(z) converges as β → ∞ to a maximizer of the function Sd−1 ∋ y 7→ ⟨B ⊤ x, y⟩, which is given by y ∗ (x) :=

B⊤x , ∥B ⊤ x∥

x ∈ Sd−1 ,

if B is invertible. More rigorously, the next lemma holds; see Appendix A.1.1 for the proof. 7

(12)

Lemma 5. Assumethat the matrix B satisfies A1. Let ϱ ∈ P(Sd−1 ), r > 0, and define the spherical cap Br (y ∗ (x)) := y ∈ Sd−1 : ∥y − y ∗ (x)∥ ≤ r . Then, for every β > 0 and every q > 0, it holds q 2q 2e−βq mβ,ϱ (x) − y ∗ (x) ≤ r2 + σmin (13) (B) + ϱ(Br (y ∗ (x))) . Using this result, we quantify the convergence ρβt → ρt as β → ∞, which has been shown earlier in [14, Theorem 3.2] under invertibility assumptions on B and V as well as Assumption 3 on the initial datum ρ0 . We show the following; see Appendix A.1.2 for the proof. Proposition 6. Assume that the matrix B satisfies A1. Let (ρβt )t∈[0,T ] and (ρt )t∈[0,T ] denote weak solutions to the continuity equations (2) and (3) with coinciding initial datum ρβ0 = ρ0 , which is assumed to satisfy Assumption 3. Then, for all t ∈ [0, T ], it holds q  eC1 t − eC0 t . W2 (ρβt , ρt ) ≤ 2 log(β+1) β 4.2

Lyapunov-type estimates in Wasserstein space controlling the convergence of ρt to Π♯ ρ0

To measure how well the token distribution ρt of (3) concentrates near the dominant eigenspace E of V B ⊤ , let us define for a measure ϱ ∈ P(Sd−1 ) and for any p ∈ (0, 1], the energy (Lyapunov) functional Z 2p ∥(Id − PE )x∥ Vp (ϱ) := Rp (x) dϱ(x) with Rp (x) := . (14) 2p ∥PE x∥ Sd−1 Here, Rp is equal to zero on E and infinity on E ⊥ . The following auxiliary result shows that Rp can be regarded as a distance from E; see Appendix A.2.1 for the proof. Lemma 7. Let E ⊂ Rd be a linear subspace and let p ∈ (0, 1]. Then, for all x ∈ Sd−1 \ E ⊥ , it holds 2 that ∥x − Π(x)∥ ≤ 2Rp (x) and for any probability measure ϱ ∈ P(Sd−1 ) with ϱ(E ⊥ ∩ Sd−1 ) = 0 we have Z W22 (ϱ, Π♯ ϱ) ≤ 2 Rp (x) dϱ(x) = 2Vp (ϱ). Sd−1

Along the flow ρt of the zero-temperature equation (3), we can quantify the decay of Vp (ρt ) as demonstrated in the subsequent result; see Appendix A.2.2 for the proof. Proposition 8. Assume that the matrices B and V satisfy Assumption 1. Moreover, let the initial distribution ρ0 be such that Assumption 2 holds for some p ∈ (0, 1]. Let (ρt )t∈[0,T ] denote the weak solution of (3) with initial datum ρ0 . Then, for all t ∈ [0, T ], it holds   2pγ Vp (ρt ) ≤ Vp (ρ0 ) exp − σmax (B) t . Noticing that the projection Π defined in (4) is invariant in time for a solution ρt of (3), as made rigorous by the next lemma, Lemma 7 justifies that Vp is indeed a suitable Lyapunov functional for the flow ρt since W22 (ρt , Π♯ ρ0 ) = W22 (ρt , Π♯ ρt ) ≤ 2Vp (ρt ) with Vp (ρt ) decaying exponentially fast thanks to Proposition 8. The proof of Lemma 9 is given in Appendix A.2.3. Lemma 9. Assume that the matrices B and V satisfy Assumption 1. Moreover, let the initial distribution ρ0 be such that Assumption 2 holds for some p ∈ (0, 1]. Let (ρt )t∈[0,T ] denote the weak solution of (3) with initial datum ρ0 . Then it holds Π♯ ρt = Π♯ ρ0 for all t ∈ [0, T ].

5

Numerical experiments

We provide numerical experiments that validate Theorem 1, and investigate the long-time dynamics beyond the non-asymptotic concentration regime. In particular, we observe the predicted concentration near the dominant eigenspace E of V B ⊤ up to time scales of order log β, followed by a termial phase, in which tokens concentrate near a dominant eigenspace of the value matrix V , consistent with Conjecture 1. For our simulations, we use an explicit Euler discretization of (1) with timestep ∆t = 0.01.1 The remaining parameters are specified individually for each experiment. We first fix β = 100, d = 3, 1We employ the CBX library [9] to run our normalized self-attention dynamics with suitable objective and kernel function.

8

(a) t ∈ [0, 4)

(b) t ∈ [4, 9)

(c) t ∈ [9, 20)

Figure 3: Dynamics of n = 200 tokens evolving according to (1) with β = 100. We show the dominant eigenspace of V B ⊤ (orange crosses) and the dominant eigenspace F of V (red crosses). n = 200 tokens, and parameter matrices V = diag(−1, 1, −2) and B = diag(−1, −1, 1) such that V B ⊤ = diag(1, −1, −2). In Figure 3, we depict the dynamics until a time horizon T ≈ 20. Initially, we observe concentration on the e1 -direction which corresponds to the dominant eigenspace of V B ⊤ . On a second time scale, tokens align with the e2 -direction, corresponding to F , the dominant eigenspace of V , in line with Conjecture 1. Figure 4 quantitatively compares the mean-field dynamics for moderate, large, and infinite inverse temperature β. For moderate β, the solution first approaches the projection Π♯ ρ0 associated with the dominant eigenspace E of V B ⊤ , as indicated by the decrease of the Wasserstein distance and the transient growth of the alignment with E, as predicted by Theorem 1. After a time of order log β, this metastable behavior gives way to the second time scale, on which the solution aligns instead with the dominant eigenspace F of V . As shown in panel (b) of Figure 4, increasing β delays this transition, while in the limiting case β = ∞, the second phase is absent by Corollary 2 and the dynamics remain aligned with E. We point to Appendix C for parameter choices of B and V that instead lead to concentration on F abs , the in magnitude dominant eigenspace of V .

Alignment with E

Alignment with F

W2 (ρtβ, n , Π ] ρ0 )

1.0

1.0

1.0

0.5

0.5

0.5

0.00

5

t

10

15

(a) β = 10

0.00

5

t

10

(b) β = 103

15 0.00

5

t

10

15

(c) β = ∞

Figure 4: Alignment of ρβ,n with the dominant eigenspaces E of V B ⊤ , and F of V for d = 10 t (maximum alignment corresponds to a value of 1). The blue curve shows the Wasserstein distance quantified in Theorem 1, and the vertical line marks t = log β. The diagonal matrices V , B are fixed across trials, with normally distributed entries. Curves show the mean over 20 runs with n = 500 tokens sampled from the uniform distribution ρ0 ; shaded regions show empirical 0.10–0.90 quantile.

6

Concluding remarks

We provided a quantitative analysis of token concentration phenomena in mean-field transformers, showing that their zero-temperature limit β −1 → 0 is accurate up to time scales of order log β and identifying the corresponding limiting distribution as a projection of the initial distribution ρ0 onto the dominant eigenspace of the matrix V B ⊤ together with convergence rates thereto. In addition, we provided experimental and theoretical evidence for a terminal phase in which the dynamics align with a dominant eigenspace of the value matrix V . Extensions and limitations. There are a few limitations of our work which constitute important directions to address in future work. While the symmetry assumption on the matrix V B ⊤ enables us to characterize the stationary distribution of the zero-temperature continuity equation, it restricts 9

the dynamics compared to the general case where no stationary distribution might exist, cf. [14]. Analyzing the regime of finitely many layers and tokens, i.e., time and space discretizations of (1), investigating multi-head attention, where an average over multiple attention heads precedes the projection in (1), and enriching the dynamics by adding noise [10, 40, 43] or a multi-layer perceptron [5], are further important generalizations of our work. Besides this, a systematic understanding of Conjecture 1 and the quantification of the convergence in the terminal phase observed numerically, as well as its possible connections to the Oja flow of the value matrix V , see [4], are also of interest.

Acknowledgments and Disclosure of Funding AA was funded by the European Union’s Horizon Europe MSCA project ModConFlex (grant number 101073558). LB acknowledges funding by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – project number 544579844 (GeoMAR) within DFG-SPP 2298 “Theoretical Foundations of Deep Learning” and by the COST Action InterCoML, CA24136, supported by COST (European Cooperation in Science and Technology). LB and TR acknowledge funding by the German Ministry of Research, Technology and Space (BMFTR) under grant agreement No. 01IS24072A (COMFORT). TR acknowledges the support of the Munich Center for Machine Learning and the ERC Advanced Grant NEITALG, grant agreement No. 101198055. KR was supported by the grant “DMS-EPSRC: Asymptotic Analysis of Online Training Algorithms in Machine Learning: Recurrent, Graphical, and Deep Neural Networks” (NSF DMS-2311500). For the purpose of Open Access, the authors have applied a CC BY public copyright license to any Author Accepted Manuscript (AAM) version arising from this submission.

Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. This project has received funding from the European Research Council (ERC) under the European Union’s Horizon Europe research and innovation programme (grant agreement No. 101198055, project acronym NEITALG).

References [1] Á. R. Abella, J. P. Silvestre, and P. Tabuada. The asymptotic behavior of attention in transformers. arXiv preprint arXiv:2412.02682, 2024. [2] A. Agazzi, G. Bruno, E. M. García, S. Saviozzi, and M. Romito. Stochastic scaling limits and synchronization by noise in deep transformer models. arXiv preprint arXiv:2604.26898, 2026. [3] A. Alcalde, B. Geshkovski, and D. Ruiz-Balet. Attention’s forward pass and Frank-Wolfe. arXiv preprint arXiv:2508.09628, 2025. [4] C. Altafini. Multistability of self-attention dynamics in transformers. IEEE Transactions on Automatic Control, 2026. [5] A. Álvarez-López, B. Geshkovski, and D. Ruiz-Balet. Perceptrons and localization of attention’s mean-field landscape. arXiv preprint arXiv:2601.21366, 2026. [6] L. Ambrosio, N. Gigli, and G. Savaré. Gradient flows in metric spaces and in the space of probability measures. Lectures in Mathematics ETH Zürich. Birkhäuser Verlag, Basel, second edition, 2008. 10

[7] R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. [8] D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. In Y. Bengio and Y. LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. [9] R. Bailo, A. Barbaro, S. Gomes, K. Riedl, T. Roith, C. Totzeck, and U. Vaes. CBX: Python and Julia packages for consensus-based interacting particle methods. J. Open Source Softw., 9(98):6611, 2024. [10] K. Balasubramanian, S. Banerjee, and P. Rigollet. On the structure of stationary solutions to McKean-Vlasov equations with applications to noisy transformers. arXiv preprint arXiv:2510.20094, 2025. [11] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021. [12] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. [13] G. Bruno, F. Pasqualotto, and A. Agazzi. Emergence of meta-stable clustering in mean-field transformer models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [14] G. Bruno, F. Pasqualotto, and A. Agazzi. A multiscale analysis of mean-field transformers in the moderate interaction regime. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [15] L. Bungert, T. Roith, and P. Wacker. Polarized consensus-based dynamics for optimization and sampling. Math. Program., 211(1-2):125–155, 2025. [16] M. Burger, S. Kabri, Y. Korolev, T. Roith, and L. Weigand. Analysis of mean-field models arising from self-attention dynamics in transformer architectures with layer normalization. Philos. Trans. Roy. Soc. A, 383(2298):Paper No. 20240233, 48, 2025. [17] V. Castin, P. Ablin, J. A. Carrillo, and G. Peyré. A unified perspective on the dynamics of deep transformers. arXiv preprint arXiv:2501.18322, 2025. [18] S. Chen, Z. Lin, Y. Polyanskiy, and P. Rigollet. Critical attention scaling in long-context transformers. arXiv preprint arXiv:2510.05554, 2025. [19] S. Chen, Z. Lin, Y. Polyanskiy, and P. Rigollet. Quantitative clustering in mean-field transformer models. arXiv preprint arXiv:2504.14697, 2025. [20] A. Cowsik, T. Nebabu, X. Qi, and S. Ganguli. Geometric dynamics of signal propagation predict trainability of transformers. Phys. Rev. E, 112(5):Paper No. 055301, 13, 2025. [21] M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser. Universal transformers. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, 2019. [22] A. Dembo and O. Zeitouni. Large deviations techniques and applications, volume 38 of Applications of Mathematics (New York). Springer-Verlag, New York, second edition, 1998. 11

[23] J. Devlin, M. Chang, K. Lee, and K. Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, and T. Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics, 2019. [24] M. Engel and A. Shalova. Random quadratic form on a sphere: Synchronization by common noise. arXiv preprint arXiv:2603.06187, 2026. [25] M. Fornasier, H. Huang, L. Pareschi, and P. Sünnen. Consensus-based optimization on hypersurfaces: Well-posedness and mean-field limit. Math. Models Methods Appl. Sci., 30(14):2725– 2751, 2020. [26] M. Fornasier, T. Klock, and K. Riedl. Consensus-Based Optimization Methods Converge Globally. SIAM J. Optim., 34(3):2973–3004, 2024. [27] B. Geshkovski, H. Koubbi, Y. Polyanskiy, and P. Rigollet. Dynamic metastability in the self-attention model. arXiv preprint arXiv:2410.06833, 2024. [28] B. Geshkovski, C. Letrouit, Y. Polyanskiy, and P. Rigollet. The emergence of clusters in selfattention dynamics. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 16, 2023, 2023. [29] B. Geshkovski, C. Letrouit, Y. Polyanskiy, and P. Rigollet. A mathematical perspective on transformers. Bull. Amer. Math. Soc. (N.S.), 62(3):427–479, 2025. [30] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783, 2024. [31] N. Karagodin, S. Ge, Y. Polyanskiy, and P. Rigollet. Normalization in attention dynamics. arXiv preprint arXiv:2510.22026, 2025. [32] N. Karagodin, Y. Polyanskiy, and P. Rigollet. Clustering in causal attention masking. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 15, 2024, 2024. [33] H. Koubbi, B. Geshkovski, and P. Rigollet. Homogenized transformers. arXiv preprint arXiv:2604.01978, 2026. [34] Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, 2020. [35] M. A. Peletier and A. Shalova. Nonlinear diffusion limit of non-local interactions on a sphere. arXiv preprint arXiv:2512.03185, 2025. [36] R. Pinnau, C. Totzeck, O. Tse, and S. Martin. A consensus-based model for global optimization and its mean-field limit. Math. Models Methods Appl. Sci., 27(01):183–204, 2017. [37] Y. Polyanskiy, P. Rigollet, and A. Yao. Synchronization of mean-field models on the circle. arXiv preprint arXiv:2507.22857, 2025. [38] K. Riedl. Leveraging memory effects and gradient information in consensus-based optimisation: onn global convergence in mean-field law. European J. Appl. Math., 35(4):483–514, 2024. [39] K. Riedl. Mathematical Foundations of Interacting Multi-Particle Systems for Optimization. PhD thesis, Technische Universität München, 2024. 12

[40] P. Rigollet. The mean-field dynamics of transformers. arXiv preprint arXiv:2512.01868, 2025. [41] M. E. Sander, P. Ablin, M. Blondel, and G. Peyré. Sinkformers: Transformers with doubly stochastic attention. In International Conference on Artificial Intelligence and Statistics, pages 3515–3530. PMLR, 2022. [42] A. Shalova. Noisy gradient flows: with applications in machine learning. PhD thesis, Eindhoven University of Technology, 2025. [43] A. Shalova and A. Schlichting. Solutions of stationary McKean–Vlasov equation on a highdimensional sphere and other Riemannian manifolds. Adv. Nonlinear Anal., 15(1):20250141, 2026. [44] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. [45] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017. [46] J. Von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov. Transformers learn in-context by gradient descent. In International Conference on Machine Learning, pages 35151–35174. PMLR, 2023. [47] B. Zhang and R. Sennrich. Root mean square layer normalization. In H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alché-Buc, E. B. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 12360–12371, 2019.

13

Supplementary material This supplementary material is organized into the following appendices. A Proof details for the main result, Theorem 1

14

A.1 Quantitative estimate for the Wasserstein distance between the flows ρβt and ρt

. .

14

A.1.1 A quantitative Laplace principle . . . . . . . . . . . . . . . . . . . . . . .

14

A.1.2 Stability estimates in Wasserstein space . . . . . . . . . . . . . . . . . . .

16

A.2 Convergence of the flow ρt to Π♯ ρ0

. . . . . . . . . . . . . . . . . . . . . . . . .

19

A.2.1 Auxiliary result for Lyapunov-type convergence estimates in Wasserstein space 19 A.2.2 Lyapunov-type convergence estimates in Wasserstein space . . . . . . . . .

20

A.2.3 Invariance of Π under the zero-temperature equation . . . . . . . . . . . .

22

A.3 Proof of Corollary 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

24

B Heuristics supporting Conjecture 1

24

C Additional numerical experiments

25

A

C.1 Details on the discretization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

25

C.2 Non-symmetric V B ⊤ effects . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

26

C.3 Gradient flow case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

26

C.4 Experiments supporting Conjecture 1 . . . . . . . . . . . . . . . . . . . . . . . .

29

Proof details for the main result, Theorem 1

In this appendix, we present auxiliary technical results for our main theoretical result, Theorem 1, together with their proof details. A.1

Quantitative estimate for the Wasserstein distance between the flows ρβt and ρt

The objective of this section is to derive a bound on W2 (ρβt , ρt ) in terms of the temperature parameter β to prove Proposition 6. Our proof is based on a stability estimate in Wasserstein space, which we detail in Appendix A.1.2. It employs the quantitative Laplace principle from Lemma 5, see also [26], to control the discrepancy between the vector fields of the continuity equations (2) and (3). The proof of the quantitative Laplace principle is provided in Appendix A.1.1. A.1.1

A quantitative Laplace principle

We first give a proof of the quantitative Laplace principle on the sphere Sd−1 , Lemma 5, which was originally introduced in [26] on the plane. Here we extend it to the sphere and specialize it to a certain linear function on the sphere. Proof of Lemma 5. The proof follows the lines of [26, Proposition 4.5], adapted to the sphere Sd−1 . To simplify notation, let us define for x ∈ Sd−1 the non-negative functions  Ex (y) := B ⊤ x, y ∗ (x) − B ⊤ x, y = B ⊤ x 1 − ⟨y ∗ (x), y⟩ , 14

y ∈ Sd−1 ,

⊤

⊤

B x β⟨B x,y⟩ where we recall from (12) that y ∗ (x) = ∥B = ⊤ x∥ . With this definition, we have the identity e ⊤

eβ∥B x∥ e−βEx (y) . Using this, we obtain R −βEx (y) y dϱ(y) d−1 e mβ,ϱ (x) = RS , −βEx (y) dϱ(y) e Sd−1 ⊤

where we canceled eβ∥B x∥ in the numerator and denominator. Using that we are on the sphere Sd−1 , where ∥y − y ∗ (x)∥2 = ∥y∥2 − 2⟨y ∗ (x), y⟩ + ∥y ∗ (x)∥2 = 2 − 2⟨y ∗ (x), y⟩, we notice that ⊤ Ex (y) = ∥B 2 x∥ ∥y − y ∗ (x)∥2 . With this we directly observe that Ex,r :=

sup y∈Br (y ∗ (x))

Ex (y) =

∥B ⊤ x∥ 2 r . 2

The quantity Zβ (x) := Sd−1 e−βEx (y) dϱ(y) can be lower bounded by Z Zβ (x) ≥ e−βEx (y) dϱ(y) ≥ inf∗ e−βEx (y) ϱ(Br (y ∗ (x))) = e−βEx,r ϱ(Br (y ∗ (x))), R

y∈Br (y (x))

Br (y ∗ (x))

(15) where the first inequality is due to the integrand being positive. Now let r̃ ≥ r > 0. We can decompose (13) into contributions from inside and outside of Br̃ (y ∗ (x)). By Jensen’s inequality we have ∥mβ,ϱ (x) − y ∗ (x)∥ Z Z e−βEx (y) e−βEx (y) ≤ ∥y − y ∗ (x)∥ dϱ(y) + ∥y − y ∗ (x)∥ dϱ(y). (16) Zβ (x) Zβ (x) Br̃ (y ∗ (x)) Br̃ (y ∗ (x))c Inside Br̃ (y ∗ (x)) we have ∥y − y ∗ (x)∥ ≤ r̃, which allows to bound the first term in (16) by r̃ since R −βEx (y) the integrand is nonnegative and Br̃ (y∗ (x)) e Zβ (x) dϱ(y) ≤ 1. For the second term in (16), we use the lower bound (15), which gives Z e−βEx (y) ∥y − y ∗ (x)∥ dϱ(y) Zβ (x) Br̃ (y ∗ (x))c Z eβEx,r ∥y − y ∗ (x)∥ e−βEx (y) dϱ(y) (17) ≤ ϱ(Br (y ∗ (x))) Br̃ (y∗ (x))c    2 ≤ exp −β inf E (u) − E . x x,r ϱ(Br (y ∗ (x))) u∈Br̃ (y ∗ (x))c To obtain the last inequality, we first note that ∥y − y ∗ (x)∥ ≤ 2 since y, y ∗ (x) ∈ Sd−1 . Secondly, inside Br̃ (y ∗ (x))c we have Ex (y) ≥ inf u∈Br̃ (y∗ (x))c Ex (u), or equivalently after taking the exponen tial e−βEx (y) ≤ e−β inf u∈Br̃ (y∗ (x))c Ex (u) . And thirdly, ϱ Br̃ (y ∗ (x))c ≤ 1. Combining (17) with the observation directly before allows us to conclude (16) with the bound    2 ∗ exp −β inf Ex (u) − Ex,r . (18) ∥mβ,ϱ (x) − y (x)∥ ≤ r̃ + ϱ(Br (y ∗ (x))) u∈Br̃ (y ∗ (x))c ⊤

⊤

Now, since inf u∈Br̃ (y∗ (x))c Ex (u) = inf u∈Br̃ (y∗ (x))c ∥B 2 x∥ ∥u − y ∗ (x)∥2 = ∥B 2 x∥ r̃2 , we choose r̃ such that inf∗

u∈Br̃ (y (x))

i.e., r̃ = bound

q

Ex (u) − Ex,r = c

∥B ⊤ x∥ 2 (r̃ − r2 ) = q, 2

−βq 2 q + r2 . With this, (17) is bounded by ϱ(B2e and we obtain for (18) the upper ∗ ∥B ⊤ x∥ r (y (x)))

∥mβ,ϱ (x) − y ∗ (x)∥ ≤ r̃ +

2e−βq . ϱ(Br (y ∗ (x)))

Noticing that ∥B ⊤ x∥ ≥ σmin (B) gives (13), which concludes the proof. 15

In order to employ Lemma 5, we require a lower bound for the mass ρβt (Br (y ∗ (x))) which is uniform for finite time intervals [0, T ]. Under Assumption 3, which guarantees that the initial distribution ρβ0 has a positive density with respect to the surface measure Hd−1 Sd−1 on Sd−1 , this is the case and the following result holds for ρβt for all t ∈ [0, T ], see [14].

Lemma 10. Assume that the matrix B satisfies A1. Let (ρβt )t∈[0,T ] denote a weak solution to the continuity equation (2). Assume that ρβ0 satisfies Assumption 3. Then, there exists a constant C0 = C0 (d, σmax (B), σmax (V )) > 0 such that the density ftβ of ρβt satisfies 1 min ftβ (x) ≥ ℓ0 e−C0 t 2 x∈Sd−1 for all t ∈ [0, T ]. Proof. This result follows from [14, Lemma A.10]. A.1.2

Stability estimates in Wasserstein space

We now have available all necessary technical tools to provide a proof of Proposition 6. Proof of Proposition 6. To keep the notation concise, let us write Z  β := := vβ [ρt ](x) Px V mβ,ρβ (x) with mβ,ρβ (x) R t t Sd−1

exp (β⟨x, By⟩)

exp (β⟨x, Bz⟩) dρβt (z) Sd−1

y dρβt (y)

as well as B⊤x ∥B ⊤ x∥ for the vector fields of the continuity equations (2) and (3), which can be written compactly with the β β β  above definitions as ∂t ρt = − div ρt vβ [ρt ] and ∂t ρt = − div (ρt v∞ ), respectively. v∞ (x) := Px (V y ∗ (x))

with

y ∗ (x) :=

Denoting by πt ∈ Π(ρβt , ρt ) the optimal coupling between ρβt and ρt for the Wasserstein distance, we differentiate the squared Wasserstein distance along the solutions to the associated continuity equations [6, Theorem 8.4.7] to obtain for almost every t ∈ [0, T ] that ZZ d d 2 β W2 (ρt , ρt ) = ∥x − z∥2 dπt (x, z) dt dt Sd−1 ×Sd−1 ZZ =2 x − z, vβ [ρβt ](x) − v∞ (z) dπt (x, z) Sd−1 ×Sd−1 sZ Z ≤ 2W2 (ρβt , ρt )

Sd−1 ×Sd−1

vβ [ρβt ](x) − v∞ (z)

2

dπt (x, z),

where we used the Cauchy–Schwarz inequality to obtain the last step. Since by the chain rule it holds d d W22 (ρβt , ρt ) = 2W2 (ρβt , ρt ) dt W2 (ρβt , ρt ), we can divide both sides by 2W2 (ρβt , ρt ) and get that dt sZ Z d 2 (19) W2 (ρβt , ρt ) ≤ vβ [ρβt ](x) − v∞ (z) dπt (x, z). dt d−1 d−1 S ×S Let us now estimate the integrand in (19). For this, we decompose this term into an approximation error due to the finite inverse temperature β, and a stability term. More precisely, with the triangle inequality we have for all x, z ∈ Sd−1 that

vβ [ρβt ](x) − v∞ (z) ≤ vβ [ρβt ](x) − v∞ (x) + v∞ (x) − v∞ (z) . (20) To estimate the approximation error, i.e., the first term on the right-hand side of (20), we note after recalling ∥Px ∥ ≤ 1, that   vβ [ρβt ](x) − v∞ (x) = Px V mβ,ρβ (x) − Px V y ∗ (x) t   (21) = Px V mβ,ρβ (x) − y ∗ (x) t

≤ σmax (V ) mβ,ρβ (x) − y ∗ (x) . t

16

We can now employ Lemma 5 to estimate for arbitrary r > 0 and q > 0 that s 2e−βq 2q ∗ mβ,ρβ (x) − y (x) ≤ r2 + + β . t σmin (B) ρt Br (y ∗ (x))

(22)

To estimate the second summand on the right-hand side, we first note that for every spherical cap Br (y) ⊂ Sd−1 with y ∈ Sd−1 it holds   Z β β β d−1 ρt (Br (y)) = ft (z) dH (z) ≥ inf ft (z) Hd−1 (Br (y)) z∈Br (y)

Br (y)

1 −C0 t d−1 ℓ0 e H (Br (y)) 2 √ with the second line being thanks to Lemma 10. To bound Hd−1 (Br (y)), we observe that for r ≥ 2 the spherical cap Br (y) includes the halfsphere Sd−1 := {x ∈ Sd−1 : ⟨x, y⟩ ≥ 0}. To see this, note y d−1 that for x ∈ Sy we have ≥

2

∥x − y∥ = 2 − 2⟨x, y⟩ ≤ 2 ≤ r2

√ and hence x ∈ Br (y). In particular, the surface measure of Br (y) with r ≥ 2 can be bounded from d/2 below by half of the measure of the sphere S√d−1 , i.e., the number Cd := π /Γ(d/2) where Γ denotes the Gamma function. Furthermore, for r < 2, there exists a dimensional constant cd > 0 such that Hd−1 (Br (y)) ≥ cd rd−1 , which can be seen by spherical integration. Using the above, the surface measure of such a spherical cap can be bounded from below as ( √ cd rd−1 , r < 2, d−1 H (Br (y)) ≥ √ Cd , r ≥ 2, and we can continue (22) by

s mβ,ρβ (x) − y ∗ (x) ≤ t

   

4e−βq

2q ℓ0 e−C0 t cd rd−1 r2 + + −βq σmin (B)   4e  , ℓ0 e−C0 t Cd

for all r > 0 and q > 0. Let us now choose d log(β + 1) q := qβ := 2β

and

r := rβ :=

√

,

r< r≥

√ √

2, (23) 2,

qβ

d

and notice that by definition e−βqβ = (β + 1)− 2 . Using these definitions in (23), we get  d √ 4(β + 1)− 2  s   , r < 2,   β  ℓ0 e−C0 t cd rβd−1 2 mβ,ρβ (x) − y ∗ (x) ≤ qβ 1 + + t  σmin (B) −d √  2   4(β + 1) , rβ ≥ 2, ℓ0 e−C0 t Cd    d−1 2  4eC0 t √ 2β s −d     2, (β + 1) r < 2,  β d log(β + 1) 2 ≤ 1+ + cd ℓ0 d log(β + 1)  C t β σmin (B) √   4e 0 (β + 1)− d2 ,  rβ ≥ 2, Cd ℓ0 s log(β + 1) C0 t ≤ C(d, σmin (B), ℓ0 ) e , β d−1  2β (β+1)−d and (β+1)−d where we used in the last step that d log(β+1) dominates both β d log(β+1) for all β > 0, and that eC0 t ≥ 1 for all t ≥ 0. Using this, we can conclude (21) by s log(β + 1) C0 t β vβ [ρt ](x) − v∞ (x) ≤ C(d, σmin (B), σmax (V ), ℓ0 ) e . β 17

(24)

For the stability term, i.e., the second term in (20), we first notice that for all x, z ∈ Sd−1 it holds by the triangle inequality and reverse triangle inequality that ∥y ∗ (x) − y ∗ (z)∥ =

∥B ⊤ x − B ⊤ z∥ 1 B⊤x B⊤z 1 ≤ − + ∥B ⊤ z∥ − ∥B ⊤ x∥ ∥B ⊤ z∥ ∥B ⊤ x∥ ∥B ⊤ x∥ ∥B ⊤ z∥

∥B ⊤ x − B ⊤ z∥ ∥B ⊤ z∥ − ∥B ⊤ x∥ ∥B ⊤ x − B ⊤ z∥ ≤ 2 + ∥B ⊤ x∥ ∥B ⊤ x∥ ∥B ⊤ x∥ σmax (B) ≤2 ∥x − z∥, σmin (B) =

where we have used that B is invertible and hence ∥B ⊤ x∥ ≥ σmin (B) for all x ∈ Sd−1 . Moreover, one has for all x, z ∈ Sd−1 that ∥Px − Pz ∥ = ∥xx⊤ − zz ⊤ ∥ ≤ ∥x(x − z)⊤ ∥ + ∥(x − z)z ⊤ ∥ ≤ 2∥x − z∥.

Using these two auxiliary estimates, and recalling that ∥Px ∥ ≤ 1, ∥y ∗ (x)∥ = 1, we obtain by the triangle inequality v∞ (x) − v∞ (z) = Px (V y ∗ (x)) − Pz (V y ∗ (z))  ≤ Px V (y ∗ (x) − y ∗ (z)) + (Px − Pz )(V y ∗ (z)) ≤ Px

V (y ∗ (x) − y ∗ (z)) + Px − Pz

V y ∗ (z)

σmax (B) ∥x − z∥ + 2σmax (V )∥x − z∥ σmin (B)   σmax (B) = 2σmax (V ) 1 + ∥x − z∥. σmin (B)

(25)

≤ 2σmax (V )

Inserting (24) and (25) into (20), we obtain vβ [ρβt ](x) − v∞ (z)

s

≤ C(d, σmin (B), σmax (B), σmax (V ), ℓ0 )

! log(β + 1) C0 t e + ∥x − z∥ . β

Exploiting this estimate in (19) and abbreviating C1 := C(d, σmin (B), σmax (B), σmax (V ), ℓ0 ) > 0 as well as C̃1 := 2C0 + C1 , we get s ! log(β + 1) C0 t d β β W2 (ρt , ρt ) ≤ C1 W2 (ρt , ρt ) + e dt β s ! log(β + 1) β C t e1 W2 (ρt , ρt ) + ≤C e 0 , (26) β where the last bound is trivial by the definition of C̃1 , and we also the RR recall that—as πt denotes β 2 β 2 optimal coupling between ρt and ρt —we have W2 (ρt , ρt ) = Sd−1 ×Sd−1 ∥x − z∥ dπt (x, z). Using the integrating factor method with integrating factor e−C1 t , we multiply both sides of (26) by the integrating factor to observe that, after a little bit of algebra, (26) is equivalent to s  d  −Ce1 t β e1 log(β + 1) e(C0 −Ce1 )t . e W2 (ρt , ρt ) ≤ C dt β e

e1 > C0 , we can integrate this inequality in time to obtain that Since C s Z log(β + 1) t (C0 −Ce1 )s e1 t β −C e e W2 (ρt , ρt ) ≤ C1 e ds β 0 s  e1 C log(β + 1)  e = 1 − e(C0 −C1 )t , e1 − C0 β C 18

where we used that ρβ0 = ρ0 . Dividing now both sides by the integrating factor and using the definition of C̃1 , we are left with s  e1 C log(β + 1)  Ce1 t β W2 (ρt , ρt ) ≤ e − eC0 t e1 − C0 β C s  2C0 + C1 log(β + 1)  (2C0 +C1 )t e − eC0 t = C0 + C 1 β s  log(β + 1)  (2C0 +C1 )t ≤2 e − e C0 t β for all t ∈ [0, T ], which concludes the proof, upon renaming the constants. A.2

Convergence of the flow ρt to Π♯ ρ0

The aim of this section is to estimate the convergence of the flow ρt to Π♯ ρ0 , i.e., to estimate W2 (ρt , Π♯ ρ0 ) by proving Proposition 8 in combination with Lemmas 7 and 9. As a first step, we derive in Proposition 8 a time-evolution inequality for the functional Vp defined in (14), as we detail in Appendix A.2.2. Using Π♯ ρt = Π♯ ρ0 , as demonstrated by Lemma 9, it holds W2 (ρt , Π♯ ρ0 ) = W2 (ρt , Π♯ ρt ), we verify in Lemma 7 that the functional Vp isp indeed a suitable Lyapunov functional for the flow ρt since W2 (ρt , Π♯ ρ0 ) = W2 (ρt , Π♯ ρt ) ≤ 2Vp (ρt ), where Vp (ρt ) decays exponentially fast thanks to Proposition 8. A.2.1

Auxiliary result for Lyapunov-type convergence estimates in Wasserstein space

We first show that Rp as defined in (14) can be regarded as a distance from E (despite not formally being a distance). Proof of Lemma 7. Fix x ∈ Sd−1 \E ⊥ . To simplify notation, let us use the orthogonal decomposition of x, x = u + v,

with

u = PE x ∈ E,

v = (Id − PE )x ∈ E ⊥ ,

2

2

∥u∥ + ∥v∥ = 1.

u We note that u ̸= 0 since x ∈ Sd−1 \ E ⊥ . Using that Π(x) = ∥u∥ , a simple computation shows 2

∥x − Π(x)∥ = 2 (1 − ∥u∥) . Denoting 2

q(x) :=

∥(Id − PE )x∥

2

∥PE x∥

2

2

2

2

=

∥v∥

2,

∥u∥

∥v∥ ∥u∥ +∥v∥ 1 1 and observing that 1 + q(x) = 1 + ∥u∥ = ∥u∥ 2 = 2 and hence ∥u∥ = √ ∥u∥2

1+q(x)

2

∥x − Π(x)∥ = 2 1 − p

1

, we have

!

1 + q(x)

.

Consider now the scalar function f (q) := 1 − √

1 , 1+q

q ≥ 0.

First, f (q) ≤ 1 for all q≥ 0. Second, we claim that f (q) ≤ q for all q ≥ 0. To verify this, define 1 g(q) := q − 1 − √1+q . It holds g(0) = 0 and g ′ (q) = 1 − 12 (1 + q)−3/2 ≥ 21 > 0 for all q ≥ 0, showing that g(q) ≥ 0, or equivalently f (q) ≤ q. Hence f (q) ≤ min{1, q}. Since p ∈ (0, 1], we have for all q ≥ 0 that min{1, q} ≤ q p . Indeed, if 0 ≤ q ≤ 1, then q ≤ q p , while if q ≥ 1, then 1 ≤ qp . 19

Combining the above bounds shows 1 − √ 1

1+q(x)

!

1

2

∥x − Π(x)∥ = 2 1 − p

≤ q(x)p and therefore

1 + q(x)

≤ 2q(x)p = 2Rp (x),

validating the first part of the claim. To prove the second part of the statement, define the coupling π := (Id, Π)♯ ϱ between ϱ and Π♯ ϱ. Then, by definition of the Wasserstein distance, ZZ 2 W22 (ϱ, Π♯ ϱ) ≤ ∥x − y∥ dπ(x, y) d−1 d−1 Z S ×S Z 2 = ∥x − Π(x)∥ dϱ(x) ≤ 2 Rp (x) dϱ(x), Sd−1

Sd−1

which concludes the proof. A.2.2

Lyapunov-type convergence estimates in Wasserstein space

We now have available all the necessary technical tools to prove Proposition 8. Proof of Proposition 8. Let us denote by ρt the solution to the continuity equation (3) in the zerotemperature limit. We first deal with the trivial case of k = dim E = d in which case V B ⊤ equals a multiple of the identity matrix and hence the continuity equation (3) simplifies to ∂t ρt = 0 (using that Px (x) = 0). Hence, trivially, ρt = ρ0 for all t > 0 and Π = Id such that Vp (ρt ) = 0 and also W2 (ρt , Π♯ ρt ) = 0 for all t ≥ 0.

Ideally, we would like to use Rp as defined in (14) as test function for the weak formulation of (3). However, since Rp is infinite on E ⊥ , this is not possible. We therefore consider the regularization Rp,K (x) = min {Rp (x), K} for K > 0, and the corresponding regularized Lyapunov functional Z Vp,K (ρ) := Rp,K (x) dρ(x). Sd−1

Since Rp,K is a Lipschitz function, it is differentiable almost everywhere by Rademacher’s theorem. Hence, by chain rule and using the weak formulation of (3), it holds for almost all t ∈ [0, T ] that  Z ∇Rp,K (x), Px V B ⊤ x d Vp,K (ρt ) = dρt (x), (27) dt ∥B ⊤ x∥ Sd−1 and it remains to estimate the integrand from above. We start by deriving an upper bound for the inner product ⟨∇Rp,K (x), Px (V B ⊤ x)⟩. To simplify notation, let us use the orthogonal decomposition of x ∈ Sd−1 , x = u + v,

with

u = PE x ∈ E,

v = (Id − PE )x ∈ E ⊥ ,

2

2

∥u∥ + ∥v∥ = 1,

2p

∥v∥ p allowing us to write Rp (x) = ∥u∥ 2p = R1 (x) . Using the idempotency of PE , we obtain for p = 1

and for all x ∈ Sd−1 \ E ⊥ that 2

∇R1 (x) =

2

2v ∥u∥ − 2u ∥v∥ 4

∥u∥

= 2R1 (x)

v

2 −

∥v∥

!

u 2

∥u∥

.

(28)

Furthermore, by the chain rule, it holds that d−1

and for all x ∈ S

∇Rp (x) = pR1 (x)p−1 ∇R1 (x),

(29)

∇Rp,K (x) = 1Rp (x)≤K ∇Rp (x),

(30)

that

20

i.e., ∇Rp,K (x) = 0 if Rp (x) > K, which in particular holds for all x ∈ Sd−1 ∩ E ⊥ . Let us now bound the quantity ⟨∇Rp,K (x), Px (V B ⊤ x)⟩. First, note that ⟨∇R1 (x), x⟩ = 0 since by (28) + * 2 2 ⟨v, u⟩ ∥v∥ v u ∥u∥ ⟨u, v⟩ , x = + − − 2 2 2 − ∥u∥2 = 0, ∥u∥2 ∥v∥2 ∥v∥ ∥v∥ ∥u∥

where we have used that ⟨u, v⟩ = 0 by orthogonality. Recalling that Px (V B ⊤ )x = V B ⊤ x − ⟨x, V B ⊤ x⟩x, we compute using (30) for all x ∈ Sd−1 with R1 (x) ≤ K that ⟨∇R1,K (x), Px (V B ⊤ x)⟩ = ⟨∇R1 (x), Px (V B ⊤ x)⟩

= ⟨∇R1 (x), V B ⊤ x⟩ − ⟨x, V B ⊤ x⟩⟨∇R1 (x), x⟩ = ⟨∇R1 (x), V B ⊤ x⟩

= ⟨∇R1 (x), V B ⊤ (u + v)⟩ = ⟨∇R1 (x), V B ⊤ u⟩ + ⟨∇R1 (x), V B ⊤ v⟩ ! ⟨v, V B ⊤ u⟩ ⟨u, V B ⊤ u⟩ ⟨v, V B ⊤ v⟩ ⟨u, V B ⊤ v⟩ = 2R1 (x) − + − 2 2 2 2 ∥v∥ ∥u∥ ∥v∥ ∥u∥ ! ⟨v, V B ⊤ v⟩ ⟨u, V B ⊤ u⟩ = 2R1 (x) − 2 2 ∥v∥ ∥u∥

≤ −2(µ1 − µ2 )R1 (x) = −2γR1 (x), where we used (28) to obtain the fourth line and for the fifth line that by orthogonality and since E is an eigenspace of V B ⊤ , ⟨v, V B ⊤ u⟩ = µ1 ⟨v, u⟩ = 0 as well as ⟨u, V B ⊤ v⟩ = ⟨V B ⊤ u, v⟩ = 2 µ1 ⟨u, v⟩ = 0. To obtain the inequality in the last line, note that ⟨v, V B ⊤ v⟩ ≤ µ2 ∥v∥ and 2 ⊤ ⊤ ⟨u, V B u⟩ = µ1 ∥u∥ , since E is the dominant eigenspace of V B . Further recall the identity γ = µ1 − µ 2 . Combining (29) with the above inequality yields the case of general p ∈ (0, 1] that ⟨∇Rp,K (x), Px (V B ⊤ x)⟩ = ⟨∇Rp (x), Px (V B ⊤ x)⟩

= pR1 (x)p−1 ⟨∇R1 (x), Px (V B ⊤ x)⟩ ≤ −2pγRp (x).

(31)

Employing (30) together with (31) in (27) and noticing that B ⊤ x ≤ σmax (B) ∥x∥ = σmax (B), we obtain the differential inequality Z d ⟨∇Rp (x), Px (V B ⊤ x)⟩ Vp,K (ρt ) = 1Rp (x)≤K dρt (x) dt ∥B ⊤ x∥ Sd−1 Z 2pγ ≤− 1R (x)≤K Rp (x) dρt (x). σmax (B) Sd−1 p Integrating this inequality in t yields by the fundamental theorem of calculus Z tZ 2pγ Vp,K (ρt ) − Vp,K (ρ0 ) ≤ − 1R (x)≤K Rp (x) dρs (x) ds. (32) σmax (B) 0 Sd−1 p Let us now take the limit K → ∞. Noticing that Rp,K (x) converges point-wise from below to Rp (x) as K → ∞, by the Beppo Levi monotone convergence theorem, Vp,K (ρt ) → Vp (ρt ) and Vp,K (ρ0 ) → Vp (ρ0 ). Analogously, since 1Rp (x)≤K Rp (x) converges point-wise from below to Rp (x) as K → ∞, by the Beppo Levi monotone convergence theorem it holds Z Z 1Rp (x)≤K Rp (x) dρs (x) → Rp (x) dρs (x) = Vp (ρs ). Sd−1

Sd−1

Since this last convergence is also monotone for every s, we can apply the Beppo Levi monotone convergence theorem once again, and from (32), after taking the limit K → ∞ on both sides of the inequality, infer Z t 2pγ Vp (ρt ) − Vp (ρ0 ) ≤ − Vp (ρs ) ds. σmax (B) 0 An application of Grönwall’s inequality yields   2pγ Vp (ρt ) ≤ Vp (ρ0 ) exp − t , σmax (B) which concludes the proof. 21

A.2.3

Invariance of Π under the zero-temperature equation

Before giving the proof of Lemma 9, let us first verify the following technical result. Lemma 11. Assume that the matrices B and V satisfy Assumption 1. Then, for x ∈ Sd−1 \ E ⊥ , the Jacobian DΠ(x) of Π as defined in (4) exists and equals ! ⊤ 1 PE x (PE x) DΠ(x) = PE − . 2 ∥PE x∥ ∥PE x∥ Moreover, it holds  DΠ(x)Px

V B⊤x ∥B ⊤ x∥

 = 0.

Proof. To obtain the first claim, we compute for x ∈ Sd−1 \ E ⊥ with quotient rule that ! ⊤ ⊤ 1 PE PE x (PE x) PE x (PE x) = DΠ(x) = − PE − . 3 2 ∥PE x∥ ∥PE x∥ ∥PE x∥ ∥PE x∥ 2

For the second claim, using that Px (y) = y − ⟨x, y⟩x and that ⟨PE x, x⟩ = ∥PE x∥ , we derive for arbitrary y ∈ Sd−1 that ! ! ⟨PE x, y⟩ PE x⟨PE x, x⟩ PE x(PE x)⊤ Px (y) = PE y − PE − 2 2 PE x − ⟨x, y⟩ PE x − 2 ∥PE x∥ ∥PE x∥ ∥PE x∥ (33) ⟨PE x, y⟩ = PE y − 2 PE x. ∥PE x∥ To simplify notation, let us use the orthogonal decomposition of x, x = u + v,

with

2

v = (Id − PE )x ∈ E ⊥ ,

u = PE x ∈ E,

2

∥u∥ + ∥v∥ = 1.

⊤

⊤

µ1 u VB x VB v ⊤ ⊥ Specifically, for ŷ = ∥B (here, we crucially ⊤ x∥ , we have ŷ = ∥B ⊤ x∥ + ∥B ⊤ x∥ . Since V B v ∈ E use the symmetry of V B ⊤ ), it holds

PE ŷ =

µ1 u µ1 PE x = ⊤ ∥B x∥ ∥B ⊤ x∥

2

and hence

⟨PE x, ŷ⟩ = µ1

Using those identities in the last step, we have with (33) that ! 1 1 PE x(PE x)⊤ Px (ŷ) = DΠ(x)Px (ŷ) = PE − 2 ∥PE x∥ ∥PE x∥ ∥PE x∥

∥PE x∥ . ∥B ⊤ x∥

PE ŷ −

⟨PE x, ŷ⟩ ∥PE x∥

2 PE x

! = 0,

which concludes the proof. With this auxiliary technical tool, we can now provide a proof of Lemma 9. Proof of Lemma 9. To keep the notation concise, let us write v∞ (x) := Px (V y ∗ (x))

with

y ∗ (x) :=

B⊤x ∥B ⊤ x∥

for the vector field of the continuity equation (3). Let φ ∈ C ∞ (Sd−1 , R) be a smooth test function and choose a smooth cut-off function η ∈ C ∞ ([0, ∞), [0, 1]) satisfying η(r) = 0

for r ≤ 1,

η(r) = 1

for r ≥ 2,

For δ > 0, define ηδ (r) := η(r/δ) and φδ (x) := ηδ (∥PE x∥) φ(Π(x)). 22

|η ′ | ≤ 2.

Clearly, φδ ∈ C ∞ (Sd−1 , R) and hence an admissible test function. Using the weak formulation of (3) for φδ , it holds Z Z d φδ (x) dρt (x) = ⟨∇φδ (x), v∞ (x)⟩ dρt (x), dt Sd−1 Sd−1 where we can compute by product and chain rule ∇φδ (x) = ηδ (∥PE x∥)∇(φ ◦ Π)(x) + ηδ′ (∥PE x∥)φ(Π(x)) ∇∥PE x∥,

x ∈ Sd−1 ,

noting that ∇ ∥PE x∥ exists on the support of ηδ′ . Since by Lemma 11, DΠ(x)v∞ (x) = 0 for all x ∈ Sd−1 \ E ⊥ , we have for all x ∈ Sd−1 \ E ⊥ that ⟨∇(φ ◦ Π)(x), v∞ (x)⟩ = DΠ(x)⊤ ∇φ(Π(x)), v∞ (x) = ⟨∇φ(Π(x)), DΠ(x)v∞ (x)⟩ = 0. PE x Furthermore, for x ∈ Sd−1 \ E ⊥ we have ∇ ∥PE x∥ = ∥P . Therefore, we obtain E x∥   Z Z d PE x ′ φδ (x) dρt (x) = , v∞ (x) dρt (x). ηδ (∥PE x∥) φ(Π(x)) dt Sd−1 ∥PE x∥ Sd−1

(34)

To simplify notation, let us use the orthogonal decomposition of x, 2

2

u = PE x ∈ E, v = (Id − PE )x ∈ E ⊥ , ∥u∥ + ∥v∥ = 1.  ⊤  ⊤ B ⊤ x⟩x VB x , we compute = V B x−⟨x,V Using this notation and that v∞ (x) = Px ∥B ⊤ x∥ ∥B ⊤ x∥ x = u + v,

PE x , v∞ (x) ∥PE x∥

with

=

⟨u, v∞ (x)⟩ µ1 ∥u∥2 − ⟨x, V B ⊤ x⟩∥u∥2 ∥u∥(µ1 − ⟨x, V B ⊤ x⟩) = = , ⊤ ∥u∥ ∥B x∥ ∥u∥ ∥B ⊤ x∥

where we used in the second step that by the invariance of E and E ⊥ , it holds PE (V B ⊤ x) = µ1 u. This allows us to bound for a constant C > 0 depending only on B and V B ⊤ that   PE x , v∞ (x) ≤ C∥u∥ = C∥PE x∥. ∥PE x∥ Since ηδ′ is supported in Aδ := {x ∈ Sd−1 : δ ≤ ∥PE x∥ ≤ 2δ}, we can bound (34) as Z Z d |ηδ′ (∥PE x∥)| |φ(Π(x))| ∥PE x∥ dρt (x) ≤ Cρt (Aδ ), φδ (x) dρt (x) ≤ C {z } | {z } dt Sd−1 Sd−1 | ≤ δ2

1Aδ

(35)

≤2δ

where C > 0 is another constant depending solely on B, V B ⊤ , and the maximum of φ ∈ C ∞ (Sd−1 , R) on the compact sphere Sd−1 . It remains to estimate ρt (Aδ ). Since on Aδ it holds ∥PE x∥2 ≤ 4δ 2 and for δ ≤ 1/4 moreover 2

2

2

∥(Id − PE )x∥ = ∥v∥ = 1 − ∥u∥ = 1 − ∥PE x∥2 ≥ 1 − 4δ 2 ≥ 3/4, we have for all x ∈ Aδ that Rp (x) =

∥(Id − PE )x∥2p ≥ cp δ −2p ∥PE x∥2p

for a suitable constant cp > 0. We can therefore estimate Z Z Z δ 2p δ 2p δ 2p ρt (Aδ ) = 1 dρt (x) ≤ Rp (x) dρt (x) ≤ Rp (x) dρt (x) = Vp (ρt ). cp Aδ cp Sd−1 cp Aδ Thanks to Proposition 8, supt∈[0,T ] Vp (ρt ) ≤ Vp (ρ0 ) < ∞, and thus, as δ → 0, we get from (35) that Z d sup φδ (x) dρt (x) ≤ Cδ 2p Vp (ρ0 ) → 0, (36) t∈[0,T ] dt Sd−1 23

where C > 0 now depends also on p. By the fundamental theorem of calculus,  Z Z Z Z t d φδ (x) dρs (x) ds. φδ (x) dρt (x) − φδ (x) dρ0 (x) = ds Sd−1 0 Sd−1 Sd−1

(37)

Recall that by definition of the cut-off function η it holds |φδ | ≤ φ ◦ Π for any δ > 0 and φδ → φ ◦ Π d−1 ⊥ pointwise \ E ⊥ as δ → 0. R on S R Since Vp (ρt ) < ∞ by Proposition 8, it holds ρt (E ) = 0. Hence, Sd−1 (φ ◦ Π)(x) dρt (x) = Sd−1 \E ⊥ (φ ◦ Π)(x) dρt (x) ≤ maxz∈Sd−1 φ(z) < ∞ since φ is a smooth function on a compact space. We can now pass to the limit δ → 0 using the dominated convergence theorem to Rinfer the convergence of the integrals on the left-hand side of (37) to R φ(Π(x)) dρt (x) − Sd−1 φ(Π(x)) dρ0 (x). With the right-hand side of (37) vanishing due to d−1 S (36) as δ → 0, we obtain that Z Z φ(x) d(Π♯ ρt )(x) = φ(Π(x)) dρt (x) d−1 Sd−1 ZS Z = φ(Π(x)) dρ0 (x) = φ(x) d(Π♯ ρ0 )(x). Sd−1

Sd−1

Since this holds for all smooth φ, we conclude that Π♯ ρt = Π♯ ρ0 , as desired. A.3

Proof of Corollary 2

Proof. It is easy to verify that for t ≥ t1 it holds exp (−pγ/σmax (B)t) ≤ 2ε , and that for t ≤ t2 it p p holds 2 log(β+1)/β (eC1 t − eC0 t ) ≤ 2 log(β+1)/β (eC1 t − 1) ≤ 2ε . By choosing β large enough (in fact, β ≳ ε14 log 1ε suffices) we have t1 < t2 and the conclusion follows from Theorem 1.

B

Heuristics supporting Conjecture 1

In this section, we provide heuristics supporting Conjecture 1. Some of them can actually be found in the recent work [4], but we include them here for completeness. One-point dynamics. First, we consider the trivial case where the token distribution is completely collapsed. This captures the situation where the initial distribution ρ0 is just supported on a single point and approximates the more interesting situation where at some positive time t > 0 the token distribution is (almost) collapsed. We can show that from there on the distribution aligns with the dominant eigenspace of V . To see this note that for ϱ = δx0 with x0 ∈ Sd−1 , we have Z exp(β⟨x, By⟩) R V y dϱ(y) = V x0 . Sd−1 Sd−1 exp(β⟨x, Bz⟩) dϱ(z) It is then easy to observe that the solution of (2) is given by ρβt = δx(t) where t 7→ x(t) solves the Oja flow ẋ(t) = Px(t) V x(t), x(0) = x0 . It follows from [14, Lemmas A.13, A.15] (or a straightforward computation using the quotient rule) y(t) that x(t) = ∥y(t)∥ where t 7→ y(t) solves the system of linear ordinary differential equations ẏ(t) = V y(t). Assuming for simplicity that V is symmetric and that its largest eigenvalue is simple, it follows from the theory of linear ordinary differential equations that y(t) = exp(V t)y(0) and hence that x(t) converges exponentially fast to an eigenvector v lying in the eigenspace of V that corresponds to its largest eigenvalue as t → ∞ provided that ⟨x0 , v⟩ ̸= 0. As a consequence, W2 (ρβt , δv ) = ∥x(t) − v∥ → 0 as t → ∞. This reasoning can be extended to more general classes of matrices V by using their Jordan normal form. 24

Some stationary distributions. The above shows, in particular, that δv is a stationary distribution of (2). Let us now investigate slightly Pn more complicated stationary distributions of this equation. 1 For this, let us assume that ϱ = i=1 δvi where vi are eigenvectors of V with eigenvalue λ > 0. n Pn Abbreviating N (x) := i=1 exp (β⟨x, Bvi ⟩), we can compute the drift vector field in (2) as Z n exp(β⟨x, By⟩) 1 X R V y dϱ(y) = exp (β⟨x, Bvi ⟩) V vi N (x) i=1 Sd−1 Sd−1 exp(β⟨x, Bz⟩) dϱ(z) n

=

λ X exp (β⟨x, Bvi ⟩) vi , N (x) i=1

∀x ∈ Sd−1 .

Hence, for any smooth ϕ ∈ C ∞ (Sd−1 , R), using that Pvi (vi ) = 0, we have   Z Z exp(β⟨x, By⟩) R V y dϱ(y) dϱ(x) ∇ϕ(x), Px Sd−1 Sd−1 Sd−1 exp(β⟨x, Bz⟩) dϱ(z)  + * n n X 1X λ = ∇ϕ(vi ), Pvi  exp (β⟨vi , Bvj ⟩) vj  n i=1 N (vi ) j=1  + * n X λX 1 ∇ϕ(vi ), Pvi  exp (β⟨vi , Bvj ⟩) vj  . = n i=1 N (vi )

(38)

j̸=i

If n = 1, meaning that ϱ = δv for some v ∈ Sd−1 with V v = λv, then (38) equals zero. This shows that ϱ = δv is a stationary distribution of (2). A slightly more complex example is n = 2 and ϱ = 12 (δv + δ−v ), for which we have Pv (−v) = −v − ⟨v, −v⟩v = 0 and P−v (v) = v − ⟨−v, v⟩(−v) = 0. Hence, (38) equals zero and ϱ is a stationary distribution of (2). Such situation is depicted in Figure 2. These two classes of stationary distributions concentrated on eigenvectors of V were already identified in [4, Lemma 7] and were referred to as consensus and bipartite consensus therein. In [4, Theorem 2], it was also shown that consensus on an eigenvector with respect to the largest eigenvalue of V is a stable equilibrium and that bipartite consensus can be stable under certain conditions.

C

Additional numerical experiments

This appendix contains implementation details and supplementary numerical experiments. We begin by detailing the time discretizations used in our simulations, together with the initialization procedure for the particle simulations. We then present several additional experiments: examples showing effects that can occur when V B ⊤ is not symmetric, simulations in the gradient flow case B = ±V , and numerical evidence supporting Conjecture 1. C.1

Details on the discretization

We provide pseudocode for the employed Euler discretization of (1), and the zero-temperature limit (3). Let us highlight here that Algorithm 1 is not a faithful discretization of the zero-temperature equation. Namely, formally setting β = ∞ in (1) amounts to replacing the softmax in lines 2–3 of Algorithm 1 by a hardmax function. In the finite-particle case, however, this does not recover Algorithm 2. We note that whenever we run a numerical experiment in what follows with β = ∞, this means we employ Algorithm 2. Initialization. For the visualization in Figure 1, we employ a mixture of von Mises–Fisher distributions, with densities, that up to normalization are given as p(x|µi , κi ) ∼ exp(κi µ⊤ i x)

with parameters κi ∈ [0, ∞) and µi ∈ Rd for i = 1, . . . , m. For weights wi with mixture then fulfills m X p(x|{(µi , κi )}m ) ∼ p(x|µi , κi ). i=1 i=1

25

Pm

i=1 wi = 1, the

Algorithm 1 Discretization of (1) Require: Tokens {xi }ni=1 ⊂ Sd−1 , matrices B, V ∈ Rd×d , β, ∆t > 0. 1: for i = 1, . . . , n do exp (β⟨xi , Bxj ⟩) 2: wij ← Pn , j = 1, . . . , n Pnk=1 exp (β⟨xi , Bxk ⟩) 3: mi ← j=1 wij xj ▷ softmax-weighted consensus 4: xi ← xi + ∆t Pxi V mi 5: xi ← xi / ∥xi ∥ ▷ retract to Sd−1 6: end for Algorithm 2 Discretization of the zero-temperature limit (3) Require: Tokens {xi }ni=1 ⊂ Sd−1 , matrices B, V ∈ Rd×d , ∆t > 0 1: for i = 1, . . . , n do 2: mi ← B ⊤ xi / ∥B ⊤ xi ∥ 3: xi ← xi + ∆t Pxi V mi 4: xi ← xi / ∥xi ∥ 5: end for

▷ retract to Sd−1

For our examples, we choose µ1 = (1, −0.3, −0.2), µ2 = (0, 1, −0.3), µ3 = (−1, 1, 1), κ1 = 2, κ2 = 10, κ3 = 5, w1 = w2 = w3 = 1/3. C.2

Non-symmetric V B ⊤ effects

We illustrate the role of the symmetry assumption on V B ⊤ through two examples. To do so, we consider the following matrices ! −1 1 0 V1 = −2 1 0 , B1 = diag(−1, −1, 1), (39) 0 0 −2 ! −1 1 0 V2 = diag(−1, 1, −2), B2 = −2 1 0 . (40) 0 0 1 In this case, note that the eigenvalues of Vi Bi⊤ and Vi are not necessarily real. In fact, the eigenvalues of V1 B1⊤ and V1 equal ±i and √ −2, whereas the eigenvalues of V2 B2⊤ , despite being a non-symmetric matrix, are real and equal 1 ± 2 and −2. Those of V2 are equal to ±1 and −2. The trajectories for these two settings are displayed in Figure 5. In the top row corresponding to V1 and B1 , the complex eigenvalues lead to a visible rotation of the trajectories, and no clear alignment is observed. In the bottom row, however, V2 B2⊤ is still non-symmetric but has real eigenvalues, and the same two-step behavior as in the symmetric case is recovered. C.3

Gradient flow case

In this section, we study the alignment behavior in the gradient flow case, i.e., B = ±V . In particular, as highlighted in [16], in this case the self-attention dynamics try to either maximize or minimize the energy Z Z EB (ϱ) = e⟨x,By⟩ dϱ(x) dϱ(y). Sd−1

Sd−1

These experiments complement our theory by testing the predicted alignment behavior in the more specific gradient flow setting, where the dynamics admit an additional variational interpretation. In the following, we perform numerical experiments for four cases: maximization and minimization, each with positive and negative definite matrices. Throughout, vmax (B) and vmin (B) denote a unit eigenvector of B associated with its largest and smallest eigenvalues, respectively. 26

(a1) t ∈ [0, 4)

(b1) t ∈ [4, 9)

(c1) t ∈ [9, 20)

(a2) t ∈ [0, 1.5)

(b2) t ∈ [1.5, 6)

(c2) t ∈ [6, 20)

Figure 5: Dynamics of tokens evolving according to (1) with V B ⊤ not symmetric. In the top row, corresponding to the choice V1 , B1 in (39), we observe that trajectories rotate around the sphere and do not display clear convergence or clustering. In the bottom row, corresponding to V2 , B2 in (39), we can see exactly the same behavior as in Figure 2: initial alignment with E, followed by a second phase of alignment with F . Maximization case. For B = V , the self-attention dynamics try to maximize the energy EB and the optimal stationary state is a single Dirac at ±vmax (B), i.e., ρopt = δvmax (B) or ρopt = δ−vmax (B) . However, the example in [16, Figure 2] suggests that, for large β, the stationary state supported fully on the dominant eigenspace of B is more likely to occur. This is consistent with Theorem 1, since in this case V B ⊤ = B 2 . We display this behavior in Figure 6, where the purple line displays the opt evolution of EB (ρβ,n t )/EB (ρ ), i.e., the energy of the ensemble normalized by the theoretical best energy. For small values of β, the iteration indeed converges towards an energetically optimal state, which is a single Dirac. However, since the initial configuration is sampled uniformly on the sphere, Π♯ ρn0 is likely clustered both at vmax (B) and −vmax (B), which causes the Wasserstein distance in Figure 6 to be large. For larger values of β, one instead observes convergence towards energetically suboptimal states, which are closer to Π♯ ρn0 . On the other hand, if B is negative semi-definite, the optimal state is ρopt = 21 (δvmin (B) + δ−vmin (B) ), where in this case vmin (B) = vmax (B 2 ). In Figure 7, we observe that even for small values of β, we indeed obtain convergence of ρβ,n towards Π♯ ρn0 . Moreover, we also obtain convergence towards t energetically optimal states for large values of β. Minimization case. We now consider the case when B = −V , in which case the self-attention dynamics try to minimize the energy EB . Thus, the purple line in the following figures will now display 1−EB (ρopt )/EB (ρβ,n t ). We first consider the case where B is negative definite. Then, the optimal state is given as ρopt = δvmin (B) and here vmin (B) does not coincide with vmax (V B ⊤ ) = vmax (−B 2 ), but instead vmax (V ) = vmax (−B) = vmin (B). This can be observed in Figure 9, where for small β, we obtain convergence towards an energetically optimal state and alignment with F . For intermediate β, we obtain the effect highlighted in Figure 2, where we first obtain a phase of alignment with E and then align to F , which here is also a phase where EB is minimized. For β = ∞, as expected, we only obtain alignment towards E and the energy of the ensemble stays constant. Since B = −V , this case suggests that the second phase emerges from a disagreement between the dominant eigenspaces of B and V . Finally, we consider the minimization case with B being symmetric positive definite. 27

W2 (ρβ,n t , Π] ρ 0 )

Alignment with E

(a) β = 1

EB

(c) β = ∞

(b) β = 100

Figure 6: Alignment in the gradient flow maximization case, i.e., B = V , with B being a symmetric positive definite matrix. We run the experiment for 10 different random configurations. We consider n = 100 particles initialized independently from ρ0 , the uniform distribution on Sd−1 .

W2 (ρβ,n t , Π] ρ 0 )

Alignment with E

(a) β = 1

(b) β = 100

EB

(c) β = ∞

Figure 7: Alignment in the gradient flow maximization case, i.e., B = V , with B being a symmetric negative definite matrix. The remaining setup is the same as in Figure 6.

Alignment with E

(a) β = 1

W2 (ρβ,n t , Π] ρ 0 )

(b) β = 100

Alignment with F

EB

(c) β = ∞

Figure 8: Alignment in the gradient flow minimization case, i.e., B = −V , with B being a symmetric negative definite matrix. The remaining setup is the same as in Figure 6.

28

W2 (ρβ,n t , Π] ρ 0 )

Alignment with E

(a) β = 1

(c) β = ∞

(b) β = 100

Figure 9: Alignment in the gradient flow minimization case, i.e., B = −V , with B being a symmetric positive definite matrix. The remaining setup is the same as in Figure 6.

Here, the optimal states are not known analytically and we thus do not display the purple line in Figure 8. However, the experiments in [16] suggest that the optimal state has two modes at vmin (B) and −vmin (B). These modes are sharper the closer the smallest eigenvalue of B is to zero, and in fact the optimal state is 12 (δvmin (B) + δ−vmin (B) ) when it is exactly zero. Since in this case vmin (B) = vmax (V B ⊤ ), this observation is in line with Theorem 1. While in [16] this effect was observed for λmin (B) → 0, we similarly observe it in Figure 8 for β → ∞. We see that only for larger values of β the alignment and Wasserstein convergence can be observed. C.4

Experiments supporting Conjecture 1

We provide additional numerical evidence for Conjecture 1. For generic token configurations, we conjecture that the long-time behavior of (1) is governed by a dominant eigenspace of V : either F , associated with the largest eigenvalue of V , or F abs , associated with the eigenvalue of largest magnitude. The simulations below illustrate both possibilities.

Alignment with E

1.0 0.5 0.00

1.0

Alignment with F

0.5 10

t

20

30 0.00

(a) β = 10

1.0 0.5

10

t

20

30 0.00

(b) β = 10

1.0

1.0

0.5

0.5

0.5

10

t

20

(d) β = 103

30 0.00

10

t

20

(e) β = 103

10

t

20

30

(c) β = 10

1.0

0.00

Alignment with F abs

30 0.00

10

t

20

30

(f) β = 103

Figure 10: Alignment of ρβ,n with the dominant eigenspace E of V B ⊤ and with the eigenspaces t abs F and F of V , for d = 10 (maximum alignment equals 1). The vertical line marks t = log β. Each column uses a different pair of random diagonal matrices V and B, with independent normally distributed diagonal entries. The initial condition consists of n = 500 tokens sampled from the uniform distribution ρ0 on Sd−1 . 29

Figure 10 shows three representative random diagonal instances of V and B. For β = 10, panel (a) exhibits asymptotic alignment with F . In panel (b), F = F abs , and the dynamics align with this common eigenspace. In panel (c), E = F abs , and the limiting alignment is with this eigenspace. For β = 103 , panels (e) and (f) show the same qualitative behavior as panels (b) and (c), respectively, up to the expected delay on the time scale t ≈ log β. In panel (d), however, the limiting alignment changes from F to F abs . These simulations therefore support the prediction that the dynamics align with a dominant eigenspace of V , while indicating that the selected eigenspace can be either F or F abs , depending on the matrix realization, the initial configuration, and the temperature regime.

30

Record · ID 175240 · SHA-256 520fe86985f9aeb9
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.