ConceptioArchivearXiv CS
arXiv CSopen access

Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere Hao Ye Xi’an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences University of Chinese Academy of Sciences [email protected]

arXiv:2607.24502v1 [math.DS] 27 Jul 2026

July 2026

Abstract Rotary position embeddings (RoPE) modify attention scores through position-dependent rotations, but their effect on normalized token dynamics is not captured by the vanilla spherical self-attention model. We study the continuous-time dynamics obtained when queries and keys are rotated while values remain on the unit sphere. The resulting attention kernel is reversible and admits a sharp uniform softmax floor, yet the natural RoPE interaction energy has derivatives of both signs within one fixed nontrivial system. Every consensus state remains an equilibrium, and its transverse linearization is a reversible Markov operator whose kernel depends on the consensus point through its energy across RoPE planes. On a resonant single-frequency ring we derive an exact Bessel-aliasing spectrum, including non-coprime frequencies and the correct fixed-ring large-β asymptotics. Globally, closed hemispheres are invariant, while pairwise non-obtuse configurations and strict open semicircles contract with explicit half-angle and single-point tail bounds. These regional estimates instantiate a kernel-generic positivity principle with the sharp RoPE softmax floor. RoPE also selects an explicit score-flattening twisted branch; the generic resonant family is non-hyperbolic and linearly unstable, whereas an odd antipodal family becomes a hyperbolic saddle after quotienting global rotation. In multiple dimensions, the local consensus gap can depend non-monotonically on the allocation of energy across frequency planes, so no universal ordering by frequency is valid. Independent matrix, finite-difference, and nonlinear-flow computations cross-check the theorem boundaries and the reported constants.

1

Introduction

Self-attention is the interaction mechanism at the core of the Transformer [22]. In a deep residual stack, the layer index can be idealized as time, turning token representations into an interacting particle system. This viewpoint has exposed clustering, metastability, and synchronization phenomena that are difficult to see from a layerwise algebraic description alone [6–8, 19]. Normalization is especially important in this limit because it confines the particles to a sphere, changes contraction speeds, and can expose gradient or spectral structure [3, 12, 13]. Rotary position embedding encodes position by rotating query and key blocks, thereby converting absolute rotations into relative phase differences inside the attention score [20]. This geometric modification is not an innocuous relabeling of the spherical dynamics: the rotated score controls the weights, while the unrotated value vectors determine the motion. Related analyses show that masking, normalization, and positional encoding can each alter asymptotic token behavior

1

[1, 11, 18, 23], but the normalized query/key-only RoPE flow requires its own equilibrium, spectrum, and convergence analysis. The central observation is that RoPE preserves a symmetric positive kernel at consensus even though it precludes an energy-sign law uniform over the RoPE model class away from consensus. This separation leads to two complementary conclusions. Locally, consensus is governed by the spectral gap of a reversible Markov kernel and can be much slower than vanilla attention. Globally, convergence remains provable in contractive geometric regions without invoking a Lyapunov function. The same phase geometry also produces twisted and antipodal equilibria that are absent from the simplest consensus narrative. The paper makes the following contributions. 1. Starting from a normalized residual attention step, we formulate query/key-only RoPE selfattention on (Sd−1 )n , prove global well-posedness, reversibility, and the sharp score-range attention floor, and give exact same-parameter witnesses showing that the natural RoPE interaction energy has no uniform monotonic sign. 2. We derive the full tangent linearization at consensus. On a resonant single-frequency ring, we obtain an exact Bessel-aliasing spectrum that handles gcd(m, n) > 1, and we separate the fixed-ring exponential large-β regime from an unproved dense-phase joint limit. 3. We instantiate a kernel-generic positivity argument with the sharp RoPE floor to obtain forward invariance of fixed closed hemispheres and explicit convergence rates in pairwise non-obtuse regions and strict open semicircles, including a geodesic tail bound to the limiting consensus point. 4. We characterize a RoPE-locked, score-flattening twisted branch and its linear spectrum, including the odd antipodal exception, and prove that a multi-frequency consensus kernel is positive semidefinite with a potentially non-monotone dependence on the consensus-plane energies. Independent numerical cross-checks complement the proofs for the most convention-sensitive formulas, including Bessel aliasing, twisted-state Jacobians and drift, Dini inequalities, positive semidefiniteness, frequency aliasing, and nonlinear decay rates. The next section compares these claims at theorem level with normalized attention dynamics, positional-encoding analyses, and oscillator interpretations. Our results do not assert synchronization from arbitrary initial data or claim that twisted states, Toeplitz score structure, or the RoPE– oscillator dictionary are new in themselves. The contribution is the exact asymptotic theory of the normalized query/key-only flow: its RoPE-selected equilibrium branch, row-normalized consensus spectrum, and regional convergence rates.

2

Related work and contribution boundary

Normalized and spherical attention dynamics. Continuous-depth self-attention has been studied as a finite-particle or mean-field system, with results on clustering, metastability, normalization, and gradient-flow structure [3, 6, 7, 12]. Recent finite-particle work classifies consensus, bipartite, clustered, and polygonal equilibria [2], analyzes spectral selection through the learned value matrix under symmetric weights [13], and extends energy arguments to multiple heads [17]. These results establish that polygonal equilibria and spherical spectral analysis are not specific to RoPE. Our identity-value model instead breaks permutation symmetry through fixed query/key

2

Table 1: Closest-work comparison for standard RoPE (rotated queries/keys and unrotated values). Work

State/model

Main overlap

Boundary relative to this paper

Karagodin et al. [12]; Kuehn and Yoon [13]

Unit-sphere attention without positional encoding Spherical self-attention with learned value matrix Unnormalized Euclidean dynamics with positional encoding Static logits and signal-level RoPE theory Modified causal torus architecture

Normalized flow, rates, equilibria, spectral structure Polygonal equilibria and instability

No RoPE position kernel or Bessel-aliasing gap

Standard query/key-only RoPE on (Sd−1 )n

Equilibria, local spectrum, and convergence

Altafini [2]

Pham et al. [18]

Gu et al. [10]; Liu [14] Nunley [16]

This paper

Continuous-time RoPE token dynamics

Multiplicative/Toeplitz structure, phase modulation, aliasing RoPE phase drift and Kuramoto coupling

No RoPE-locked flat-score branch or its resonance spectrum Norm collapse/divergence rather than fixed-norm equilibrium and consensus rates No normalized state flow or tangent Markov generator No asymptotic theory for the standard spherical RoPE flow Exact non-coprime Bessel spectrum, regional rates, and multi-plane non-monotonicity

rotations. Its spectrum is that of a row-normalized position kernel at consensus, and its twisted branch is selected by the condition that RoPE-adjusted scores become exactly flat. Projected consensus and synchronization on spheres also predate attention [4, 21]. Accordingly, the regional results in section 6 are not presented as the first hemispherical synchronization theorem. Their role here is to give a kernel-generic contraction principle with explicit constants after inserting the sharp, state-uniform RoPE softmax floor. Positional encoding and RoPE. Pham et al. [18] directly study positional encoding in continuous-time token dynamics, including RoPE, but in an unnormalized Euclidean model with general learned matrices and norm collapse or divergence as the main asymptotic alternatives. Our states have fixed unit norm, values are unrotated, and the questions are equilibria, tangent rates, and regional consensus. Static analyses of multiplicative positional information and RoPE phase modulation already expose Toeplitz logit structure, frequency periodicity, and aliasing [10, 14]. We therefore do not claim novelty for a Toeplitz or Fourier representation by itself. The object analyzed here is the row-normalized consensus Markov kernel and, on a resonant ring, its exact Bessel-aliasing spectrum, including unreachable modes for non-coprime frequencies. The observation that standard RoPE leaves the value path position-blind, recently emphasized by the RoVE architecture [5], is also the precise modeling boundary used here. Oscillator interpretations and twisted states. Kuramoto Attention makes the phase-coupling interpretation explicit in a causal, discrete torus architecture [16]; its frustrated variant modifies the value law using learned harmonic and delayed couplings [15]. Thus the dictionary between rotary phases and oscillator coupling is an existing modeling connection, not a claim of this paper. Twisted states likewise have a long history in oscillator networks [9]. What is specific here is an asymptotic 3

result for the standard query/key-only spherical flow: an if-and-only-if resonance condition for the score-flattening branch, its exact Jacobian spectrum, and the separate odd antipodal family. To our knowledge, the combined theorem package in the final row is not contained in the works above. This statement is deliberately narrower than claiming the first RoPE dynamics, the first oscillator interpretation, or the first polygonal equilibrium.

3

RoPE spherical self-attention

Let n ≥ 2 and let d = 2r ≥ 2. Token i has a fixed position pi ∈ R and a state xi (t) ∈ Sd−1 . Given finite frequencies Ω = (ω1 , . . . , ωr ), define the block rotation Rp =

r M cos(ωℓ p) − sin(ωℓ p)

sin(ωℓ p)

ℓ=1

!

cos(ωℓ p)

.

(1)

For finite β ≥ 0, the score, symmetric kernel, degree, and attention weights are D

E

Wij (x) = eβsij (x) ,

sij (x) = Rpi xi , Rpj xj , Zi (x) =

n X

Wik (x),

Aij (x) =

k=1

Wij (x) . Zi (x)

(2)

RoPE acts only on queries and keys in (2); the value vectors remain unrotated. This is the standard RoPE value pathway rather than the rotated-value modification of García-Castellanos et al. [5]. The continuous model is the first-order limit of a normalized residual layer. Let yiℓ =

n X

xℓ+1 = i

Aij (xℓ )xℓj ,

j=1

xℓi + hyiℓ , xℓi + hyiℓ

(3)

where h > 0 is the residual step. Since xℓi = 1, D E xℓ+1 − xℓi i (4) = yiℓ − xℓi , yiℓ xℓi + O(h) = Px⊥ℓ yiℓ + O(h). i h The remainder is uniform on the compact product sphere. Thus, under the controlled specialization Q = K = V = I apart from the prescribed query/key RoPE rotations, the depth-time limit of (3) is

ẋi = Px⊥i

n X j=1

Aij (x)xj =

n X



Aij (x) xj − ⟨xi , xj ⟩ xi ,

Px⊥ = I − xx⊤ .

(5)

j=1

after absorbing the fixed score scaling into β. Related normalized residual limits are used in spherical attention dynamics [12, 13]; here (4) fixes the precise RoPE and value-path conventions used below. Theorem 3.1 (Well-posedness, reversibility, and sharp floor). For every initial state in (Sd−1 )n , (5) has a unique global solution. At every state, W = W ⊤ , A1 = 1, and Zi πi = P . k Zk

(6)

1 e−2β = . 1 + (n − 1)e2β n − 1 + e−2β

(7)

πi Aij = πj Aji , Moreover, Aij (x) ≥ aβ,n :=

The constant is sharp over the model class. At β = 0 it equals 1/n; for β > 0 it is strictly stronger than the coarser e−2β /n bound. 4

Theorem 3.1 is proved in section A. Compactness of the product sphere supplies global existence, while (7) is the optimal softmax bound for scores in [−1, 1]. Detailed balance does not make A symmetric unless the degrees Zi are equal.

3.1

Failure of the natural RoPE energy law

On S1 , write xi = (cos θi , sin θi ) and consider one frequency ω. With δij = θi − θj + ω(pi − pj ),

n X

E+ (θ) =

eβ cos δij ,

(8)

i,j=1

the phase equation induced by (5) is θ̇i =

X

Aij (θ) sin(θj − θi ).

(9)

j

The energy in (8) is “natural” because it recovers the standard positive-sign interaction law when the rotary phase vanishes. Indeed, with ui =

X

Wij sin δij ,

vi =

X

j

Wij sin(θi − θj ),

j

the exact identity

X ui vi

Ė+ = 2β

(10)

Zi

i

holds. At ω = 0, ui = vi and hence Ė+ = 2β

X u2 i

i

Zi

≥ 0.

(11)

RoPE separates ui from vi , so the argument no longer supplies a sign. The following proposition shows that this loss is genuine rather than merely a gap in that calculation. Proposition 3.2 (Same-system energy no-go). Fix n = 3, (p1 , p2 , p3 ) = (0, 1, 2), β = 1, and ω = π. Then d d E+ > 0, E+ < 0. (12) dt dt θ=(0,π/6,π/3) θ=(0,π/6,0) Consequently, for this fixed nontrivial RoPE system, neither E+ nor −E+ is a Lyapunov function. In particular, this energy form has no monotonicity law of either sign that is uniform over the whole RoPE model class. This proposition is energy-specific: it neither excludes a different Lyapunov function nor claims sign-indefiniteness for every frequency. The exact closed forms behind (12) appear in section A.

4

Consensus, twisted equilibria, and local linearization

4.1

Consensus and its position kernel

Let x⋆ ∈ Sd−1 and decompose it into its r rotation planes, x⋆ = (x⋆(1) , . . . , x⋆(r) ). Define aℓ =

x⋆(ℓ)

2

,

aℓ ≥ 0,

r X ℓ=1

5

aℓ = 1.

(13)

Proposition 4.1 (Consensus kernel). The consensus manifold C = {(x⋆ , . . . , x⋆ ) : x⋆ ∈ Sd−1 } consists of equilibria. At such a point, Wij⋆ = exp βKa (pi − pj ) , 

Ka (h) =

r X

aℓ cos(ωℓ h).

(14)

ℓ=1

Thus the consensus kernel depends on x⋆ only through the plane-energy vector a, not through the within-plane phases.

4.2

An explicit twisted family on the circle

Assume d = 2 and consecutive integer positions pj = j for j = 0, . . . , n−1. The twisted configuration is θj = c − ωj (mod 2π). (15) Its RoPE-adjusted scores are all equal, so Aij = 1/n. Proposition 4.2 (Exact twisted equilibria). The configuration (15) is an equilibrium if and only if ω ∈ πZ

or

nω ∈ 2πZ.

(16)

Odd multiples of π give a bipodal configuration; even multiples give consensus. Outside (16), for every fixed ω ∈ / 2πZ, max θ̇i ≤ i

|sin(nω/2)| 1 ≤ = Oω (n−1 ). n |sin(ω/2)| n |sin(ω/2)|

(17)

The constant is not uniform when the frequency varies with n toward 2πZ.

4.3

Consensus linearization

Theorem 4.3 (Tangent linearization). Let xi = x⋆ be a consensus equilibrium and let ξi ∈ Tx⋆ Sd−1 . The tangent linearization is ξ˙i =

n X

A⋆ij ξj − ξi .

(18)

j=1

In a common orthonormal tangent basis, JC = (A⋆ − In ) ⊗ Id−1 .

(19)

For the fixed linear system, the stationary-weighted tangent mean ξ¯π =

X

Z⋆ πi⋆ = P i ⋆ , k Zk

πi⋆ ξi ,

i

(20)

is conserved. The conservation in (20) is only a linearized identity. The nonlinear quantity asserted to be conserved.

6

P

i Zi (x)xi is not

5

Resonant-ring spectra and consensus slowdown

Consider d = 2, pj = j, and the exact resonance with m ∈ Z, ω=

2πm , n

g = gcd(m, n),

L=

n . g

(21)

Here L is the number of distinct sampled phases. Exact resonance makes the consensus kernel circulant; without it, a finite chain must not be replaced by a periodic distance. Theorem 5.1 (Bessel-aliasing spectrum). Under (21), A⋆ is symmetric, doubly stochastic, and circulant. Its Fourier eigenvalues are X

λq =

Ik (β)

k∈Z mk≡q (mod n)

X

Ik (β)

,

q = 0, . . . , n − 1,

(22)

k∈Z mk≡0 (mod n)

where Ik is the modified Bessel function of the first kind. If g ∤ q, then λq = 0. For β > 0, exactly L eigenvalues are positive and n − L vanish. At β = 0, the spectrum is (1, 0, . . . , 0). The tangent consensus rate is γ = 1 − λ2 (A⋆ ),

(23)

where eigenvalues are ordered non-increasingly. Formula (22) handles non-coprime m; dropping the congruence reachability condition gives the wrong zero-eigenvalue multiplicity. Theorem 5.2 (Positive spectrum and strict slowdown). For the general multi-frequency consensus kernel (14), spec(A⋆ ) ⊂ [0, 1], 0 < γ ≤ 1. (24) If β = 0, then γ = 1. If β > 0, equality γ = 1 holds if and only if every active plane is invisible on the sampled positions: aℓ > 0

=⇒

ωℓ (pi − pj ) ∈ 2πZ

for all i, j.

(25)

Whenever at least one active plane is visible, 0 < γ < 1. The proof writes Ka as a Gram matrix and applies the Schur product theorem to its entrywise exponential. The row-normalized matrix is similar to D−1/2 W ⋆ D−1/2 ⪰ 0. Condition (25), rather than “nonzero frequency,” is the exact strictness criterion. Theorem 5.3 (Fixed-effective-period large-β asymptotics). Fix the effective period L in (21). (i) If L = 1, the kernel is constant and γ(β) = 1. (ii) If L = 2, then γ(β) = 1 − tanh β =

2e−2β ∼ 2e−2β . 1 + e−2β

(26)

(iii) If L ≥ 3, let ∆L = 1 − cos(2π/L). As β → ∞, γ(β) ∼ 2∆L e−β∆L . 7

(27)

a

Non-coprime spectrum: n = 16, m = 4, β = 2 12 unreachable modes max error = 5.6e − 16

Bessel alias formula direct circulant FFT structural zero

0.8

consensus-kernel eigenvalue λq(A ⋆ )

Fixed-ring slowdown 10

local gap γ

1.0

b

10 10 10 10

−2

−5

−8

−11

L=2 L=4

L=8 L=16

2

4

L=32 continuum reference

−14

0 0.6

6

8

10

12

14

16

inverse temperature β

c

0.4

Fixed-period asymptotic prefactor 1.0

γ/[2ΔLe −x]

0.9 0.2

0.8 0.7 0.6

L=3 L=4

0.0 0.5 0

4

8

12

15

2.5

Fourier mode q

5.0

7.5

10.0

12.5

L=8 L=16

15.0

L=32 limit 1

17.5

20.0

scaled inverse temperature x = βΔL

Figure 1: Exact resonant spectrum and fixed-ring slowdown. a, Complete Fourier-labelled spectrum for n = 16, m = 4, and β = 2. Blue stems are the congruence-filtered Bessel formula, open teal circles are the direct FFT of the exact circulant attention row, and grey crosses mark the 12 modes excluded by the exact gcd(m, n) reachability mask. b, Cancellation-free finite-ring gaps for the displayed effective periods L; the dashed continuum expression 1 − I1 (β)/I0 (β) is a reference only and is not a proved joint-limit approximation here. c, Fixed-L ratios γ/[2∆L exp(−β∆L )] against β∆L for L ≥ 3, with the unit asymptotic limit dashed. Each curve is an exact finite-ring computation and does not assert a bound uniform in L. All quantities are deterministic, so no error bars are applicable. Figure 1 separates three facts that are easy to conflate in a gap-only calculation: the exact Fourier-mode reachability pattern, the absolute fixed-ring slowdown, and the leading asymptotic prefactor. All plotted finite-ring gaps use the nonnegative-sum Fourier formula, so the smallest values are not formed by subtracting two nearby eigenvalues. The continuum expression 1 − I√1 (β)/I0 (β) ∼ 1/(2β) is not the fixed-L asymptotic. Recovering it in a joint limit β, L → ∞ with L/ β → ∞ requires a uniform aliasing error bound. That estimate is not established here, so the joint-limit claim is not stated as a theorem.

6

Global contraction in invariant geometric regions

Throughout this section, positions and frequencies are arbitrary and fixed, and aβ,n denotes (7). The conclusions are independent of resonance. The proofs use only row stochasticity and the uniform lower bound Aij ≥ aβ,n ; they are therefore a kernel-generic contraction principle instantiated with the sharp RoPE floor. Earlier sphere-consensus theory gives broader qualitative synchronization 8

results under other graph and interaction assumptions [4, 21]. The narrower regions below supply explicit state-uniform constants and tail bounds for the present attention kernel. Proposition 6.1 (Fixed closed hemispheres are invariant). Fix w ∈ Sd−1 and define hi (t) = ⟨xi (t), w⟩ ,

mw (t) = min hi (t). i

(28)

If mw (0) ≥ 0, then mw (t) ≥ mw (0) ≥ 0 for all t ≥ 0. Thus both the fixed closed hemisphere and its interior are forward invariant. This conclusion alone does not imply consensus. Theorem 6.2 (Pairwise non-obtuse contraction). Let c(t) = min ⟨xi (t), xj (t)⟩ .

(29)

D+ c(t) ≥ 2aβ,n 1 − c(t)2 . 

(30)

1 − c0 , 1 + c0

(31)

i<j

If c(0) = c0 ≥ 0, then c(t) ≥ 0 and

With ρ(t) =

1 − c(t) , 1 + c(t)

ρ0 =

one has ρ(t) ≤ ρ0 e−4aβ,n t .

(32)

Equivalently, for the pairwise geodesic diameter DS (t) = arccos c(t), DS (t) DS (0) −2aβ,n t ≤ tan e . 2 2 If c0 = 0, then c(t) ≥ tanh(2aβ,n t) > 0 for every t > 0. tan

(33)

Theorem 6.3 (Strict open-semicircle contraction). Assume d = 2 and choose continuous lifts θi (t) ∈ R of the phases. If D0 = max θi (0) − min θi (0) < π, (34) i

i

then the lifted diameter remains below π and satisfies D+ D(t) ≤ −2aβ,n sin D(t),

tan

D(t) D0 −2aβ,n t ≤ tan e . 2 2

(35)

Proposition 6.4 (Bipodal boundary counterexample). Let u ∈ Sd−1 and choose signs σi ∈ {−1, 1} with both signs present. Then xi = σi u is a stationary non-consensus configuration for every finite β, every position set, and every frequency spectrum. In d = 2 this has lifted diameter π, so the strict condition in theorem 6.3 cannot be weakened to D0 ≤ π. The normalized comparisons in figure 2 use the observable appearing in the proofs rather than fitting an unconstrained exponential. They show both the conservatism of the state-uniform floor and the sharp failure at the antipodal boundary. Corollary 6.5 (Single-point limit and explicit tail). Under the assumptions of either theorem 6.2 or theorem 6.3, all tokens converge to a common point x∞ ∈ Sd−1 . Let D(t) be the corresponding maximal geodesic diameter and set q0 = tan

D(0) , 2

Then dS (xi (t), x∞ ) ≤

1 aβ,n

Q(t) = q0 e−2aβ,n t . arsinh Q(t) ≤ 9

q0 −2aβ,n t e . aβ,n

(36)

a

Strict open-semicircle contraction

q(t)/q0, q = tan(D/2)

1.0

n = 2, β = 0 (exact)

n = 6, β = 1, ω = 0.7

n = 2, β = 2, D0 = π − 0.05

proved envelope e −τ

0.8

0.6

0.4

0.2

0.0 0

1

2

3

4

5

floor scaled time τ = 2aβ, n t

b

Boundary sharpness: n = 2, β = 2

1.0

10 unordered pairs active at t = 0

0.8

diameter D(t)/π

0.8 0.6

i<j

c(t) = min⟨xi , xj⟩

c

Sphere boundary: n = 5, d = 6, c0 = 0 1.0

0.4 plane-aligned orientation rotated orientation lower bound tanh τ

0.2 0.0 0.0

0.2

0.4

0.6

0.8

1.0

0.6 0.4 D0 = π (exact stationary)

0.2

D0 = π − 0.05 strict-region bound

0.0 1.2

0.0

scaled time τ

0.5

1.0

1.5

2.0

2.5

scaled time τ

Figure 2: Regional contraction and its sharp boundary. a, Deterministic circle trajectories plotted as q(t)/q0 , where q = tan(D/2), against τ = 2aβ,n t. The dashed curve is the proved envelope exp(−τ ); the β = 0, n = 2 case is exact. Numerical curves are stopped after q(t)/q0 < 10−10 to avoid displaying the floating-point noise floor. b, Minimum pairwise inner product for two fixed orientations of five initially orthogonal tokens in d = 6, with β = 1 and the standard three-plane RoPE frequencies. All ten unordered pairs are active at c0 = 0; both step-halved trajectories lie above the dashed lower bound tanh τ . c, The exact bipodal state with D0 = π remains stationary, whereas the strict case D0 = π − 0.05 contracts below its comparison solution for n = 2, β = 2, and ω = 0. These are deterministic ODE computations, not statistical samples; no error bars are applicable. The proof in section C treats the nonsmooth minima with finite active-set Dini calculus. In particular, negative-part Grönwall estimates close the boundary cases mw (0) = 0 and c0 = 0 rather than relying on a formal tangency statement alone. p

Rate bookkeeping. Let Y = 1 − c and let L = 2(1 − c) be the associated chordal diameter. Then Y = L2 /2 identically. The half-angle bounds imply D(t) ≤ 2q0 e−2aβ,n t ,

L(t) ≤ 2q0 e−2aβ,n t ,

Y (t) ≤ 2q02 e−4aβ,n t .

(37)

Thus the logarithmic rate of the quadratic deficit 1 − c is twice that of the chordal or angular diameter. This factor is kept explicit in all numerical rate comparisons.

10

Local versus uniform rates. The consensus gap γ in (23) is the exact slowest infinitesimal transverse exponent at a specified consensus point. By contrast, 2aβ,n is a state- and position-uniform nonlinear diameter exponent. The latter is necessarily conservative. Indeed, for β > 0, A⋆ = aβ,n 11⊤ + (1 − naβ,n )B,

B=

A⋆ − aβ,n 11⊤ , 1 − naβ,n

where B is row stochastic. On the quotient by span{1}, the rank-one term vanishes, so every non-Perron eigenvalue of A⋆ has modulus at most 1 − naβ,n . Together with the positive spectrum from theorem 5.2, this gives γ ≥ naβ,n ≥ 2aβ,n . (38) At β = 0 the same inequality holds with equality γ = naβ,n = 1. On a fixed resonant ring, the asymptotics make the separation quantitative: γ ∼ n − 1, L = 2, 2aβ,n γ ∼ (n − 1)∆L eβ(2−∆L ) , L ≥ 3. (39) 2aβ,n Thus the exact local rate and the uniform global guarantee answer different questions and should not be numerically identified.

7

Twisted-state stability and multi-frequency geometry

7.1

Linear spectrum of twisted equilibria

Proposition 7.1 (Nontrivial resonant twisted states). Let pj = j, let m ∈ Z, let ω = 2πm/n, and let m̄ be m modulo n. Assume m̄ ̸= 0. At the twisted equilibrium θj = c − ωj, the Jacobian is 1 J = C, n

2πm(i − j) Cij = cos . n

n

1 2 (mult. 2), 0 (mult. n − 2)





(40)

If 2m̄ ̸≡ 0 (mod n), then o

.

(41)

spec(J) = {1 (mult. 1), 0 (mult. n − 1)} .

(42)

spec(J) = If n is even and m̄ = n/2, then

These equilibria have growing linear modes and center modes, but no linearly stable modes. They are therefore non-hyperbolic and linearly unstable, not saddles in the usual hyperbolic sense. The trivial mode m̄ = 0 is ordinary consensus and has spectrum {0, −1 (mult. n − 1)}; it is excluded from theorem 7.1. Proposition 7.2 (Odd antipodal exception). Let d = 2, pj = j for j = 0, . . . , n − 1, n = 2r + 1, and ω ∈ (2Z + 1)π. The alternating twisted circle is an equilibrium circle, and its Jacobian has spectrum n o spec(J) = − n1 (mult. r), 0 (mult. 1), n1 (mult. r − 1), 1 (mult. 1) . (43) The unique zero mode is global phase rotation. After quotienting this symmetry, the equilibrium is a hyperbolic saddle with r stable and r unstable dimensions. The two blocks of figure 3 visualize independent consequences of the same rotary-score/unrotatedvalue structure. The energy-sign failure does not cause the twisted branch, and the figure uses no causal implication between them. 11

Twisted branch (independent consequence)

b

Twisted-state spectra ×7

Natural interaction energy

a

Natural interaction energy (θ0 = 0)

π

×2

generic resonant

3 ×4

2

×1

×1

×3

odd antipodal

0

0 −1

exact Ė +

θ2

1

0.0

0.2

0.4

0.6

0.8

1.0

twisted-Jacobian eigenvalue μ(J)

−2

Ė + > 0

Nonresonant defect: fixed-ω Oω(n −1)

c

−3

Ė + < 0

−π 0

i

θ1

π

twisted defect max|θ̇ i |

10 −π

Deterministic exact witnesses (π/6, π/3) : Ė + = + 1.666364… (π/6, 0) : Ė + = − 0.129613…

10

10

10

Thus neither E + nor −E + has a universal sign. No area or probability claim is made.

direct RHS sharp bound coarse 1/n envelope

0

−1

−2

−3

ω = 0.37 ω = 1.41421 ω = 2.17

direct/formula error ≤ 1.4e − 15

10

1

10

2

token count n

Figure 3: Natural-energy sign failure and twisted branches. a, Exact field of Ė+ for n = 3, p = (0, 1, 2), β = 1, and ω = π, after fixing the global phase by θ0 = 0. Black curves are the zero level; the two triangular markers are the deterministic witnesses (θ1 , θ2 ) = (π/6, π/3) and (π/6, 0). The field rules out a universal Lyapunov sign for E+ and −E+ only; it does not rule out other Lyapunov functions and carries no area or probability interpretation. b, Exact twisted-Jacobian spectra for the generic resonant case (n, m) = (9, 2) and the odd antipodal case (n, ω) = (9, π). Blue, grey, and red markers denote stable, center, and growing modes, with multiplicities shown. The generic branch is non-hyperbolic and linearly unstable, whereas the odd antipodal branch is a saddle after quotienting global rotation. c, Direct nonresonant drift (open circles), its sharp geometric-sum bound (solid), and the fixed-frequency 1/[n| sin(ω/2)|] envelope (dotted). The bound is Oω (n−1 ), not uniform as ω approaches 2πZ. All quantities are deterministic; no error bars are applicable.

7.2

Multi-frequency consensus spectra

Theorem 7.3 (Multi-frequency kernel and local rate). Let x⋆ be a consensus point with plane energies a as in (13), and let W , D = diag(W 1), and A = D−1 W be defined by (14). Then W ⪰ 0 and A is similar to S = D−1/2 W D−1/2 ⪰ 0. Consequently, 1 = λ1 (A) > λ2 (A) ≥ · · · ≥ λn (A) ≥ 0,

γ(a) = 1 − λ2 (A) ∈ (0, 1].

(44)

The tangent Jacobian is (A − In ) ⊗ Id−1 . If β > 0, equality γ(a) = 1 holds exactly under the active-plane invisibility condition (25).

12

For consecutive integer positions, replacing any frequency ωℓ by ±ωℓ + 2πk leaves the kernel unchanged. Thus finite sampling alone prevents any universal ordering by the numerical magnitude of a frequency. Proposition 7.4 (Analytic non-monotonicity in plane energy). Take n = 3,

p = (0, 1, 2),

3 β = log 2, 4

and let a ∈ [0, 1] be the energy in the first plane. Then √ 35 2 − 32 γ(0) = , 31 3 γ(2/3) = , 4 u(1 + 2u) , γ(1) = 1 + u + u2



(ω1 , ω2 ) =

π ,π , 2

u = 2−3/4 .



(45)

In particular, γ(2/3) > max{γ(0), γ(1)}, so a 7→ γ(a) is not generally monotone. The calculation is elementary. After scaling the diagonal kernel entries to one, let x and y be the lag-one and lag-two weights. The two non-Perron eigenvalues are λ− =

1−y , 1+x+y

λ+ =

1+y 1 + − 1. 1 + x + y 1 + 2x

(46)

At a = 2/3, x = y = 1/2, so both equal 1/4. Proposition 7.5 (Analytic frequency-order reversal). For β > 0, two tokens at positions (0, ∆), and one active frequency ω, γω (∆) =

2q , 1+q



q = exp β[cos(ω∆) − 1] .

(47)

Take ωlow = π/2 and ωhigh = π. At ∆ = 1, the high frequency is slower; at ∆ = 2, it is completely aliased and has γhigh = 1, while the low-frequency gap is strictly smaller than one. Hence no position-independent ordering by frequency magnitude exists. For the standard two-plane frequencies (1, 0.01) at β = 4, direct finite-context spectra exhibit the same reversal: n = 256 :

γlow = 0.240298,

γhigh = 0.134730,

n = 384 :

γlow = 0.112756,

γhigh = 0.135707.

(48)

At n = 384, the equal-energy mixture has gap 0.263105, larger than both endpoints. These are deterministic observations on the displayed finite contexts, not a scaling theorem or a prerequisite for the analytic counterexamples above.

13

a

b

Exact three-token branch crossing

n=128

1 − λ−

1.0

Finite contexts: (ω1, ω2) = (1, 0.01), β = 4 n=256

n=384

0.6

1 − λ+

local gap γ(a1)

γ(a) = min{1 − λ − , 1 − λ + }

0.9

0.5 0.4 0.3 0.2

n = 256: high-ω endpoint slower n = 384: high-ω endpoint faster

0.0

0.2

0.4

0.6

0.8

1.0

energy in the ω = 1 plane, a1 0.8 a = 2/3

c

1.0

a=1

0.6 a=0 γ(2/3) = 3/4 strictly exceeds both endpoints

0.0

0.2

0.4

0.6

0.8

exact two-token gap γω(Δ)

0.7

1.0

low ω = π/2 high ω = π

0.8

0.6

0.4

Δ=2

Rate parity max rel. error = 2.80e − 06 0

0.8

0.6

0.25

0.4

0.2 Δ=1

1.0

d

Exact frequency reversal

fitted nonlinear rate

local consensus gap γ(a)

0.1

0.5 0.75 1

0.5

1.0

theory gap γ

energy in the ω1 = π/2 plane, a

Figure 4: No universal monotone ordering of multi-plane consensus rates. a, For the exact three-token construction in theorem 7.4, the two thin curves are the branch-limited rates 1 − λ− and 1 − λ+ , and the thick curve is their minimum γ(a). The exact interior value γ(2/3) = 3/4 strictly exceeds both endpoints; no claim of a global maximum is needed. b, Deterministic contextual observations for (ω1 , ω2 ) = (1, 0.01) and β = 4. The plane-energy curves for n = 256 and n = 384 have opposite endpoint orderings, while the n = 128 curve is shown as a muted reference. This panel is an observation on the displayed finite contexts, not a scaling theorem. c, The exact two-token formula in theorem 7.5 gives opposite low/high-frequency orderings at position separations ∆ = 1 and ∆ = 2. d, Five deterministic slow-mode integrations at n = 16 and β = 4 compare the predicted tangent gap with the fitted nonlinear decay rate; labels give the high-frequency plane energy. No error bars are applicable.

8

Numerical validation

The computations below cross-check convention-sensitive formulas and quantify finite-precision effects; no theorem uses a numerical estimate as a proof step. Finite-chain position differences are always literal, and periodic distances are introduced only under exact resonance. Scaled Bessel functions and globally scaled kernels avoid overflow without changing normalized eigenvalues. Appendix D gives the parameter grids, error definitions, integration tolerances, and rate-fitting windows. For the resonant calculation, structural zero modes are determined by the exact gcd(m, n) reachability mask rather than by counting all floating-point zeros. This distinction matters at large 14

Table 2: Independent numerical cross-checks. Coverage counts parameter configurations, not statistical replicates. The last column reports the largest discrepancy unless a one-sided inequality residual is stated. Mathematical object

Independent computations

Coverage and discrepancy

Resonant spectrum and local rate

Direct matrix spectrum, congruence-filtered Bessel spectrum, stable Fourier gap, and nonlinear phase flow

144 spectra and 12 rate fits; full-spectrum error 1.554 × 10−15 , Bessel–Fourier gap error 6.051 × 10−15 , and relative rate error 1.059 × 10−7

Geometric contraction bounds

Direct sphere/phase vector fields against the Dini inequalities and integrated trajectories against the comparison solutions

7057 derivative, boundary, and trajectory evaluations; minimum inequality residual −3.886 × 10−16 and maximum trajectory excess 2.961 × 10−10

Twisted states and drift

Direct right-hand side, closed drift formula, analytic Jacobian, and centered finite-difference Jacobian

70 equilibrium and 50 nonresonant configurations; drift-formula error 3.764 × 10−16 and Jacobian error 3.070 × 10−9

Multi-frequency gap

Symmetric similarity transform, normalized Laplacian, generalized eigenproblem, and nonlinear sphere flow

Nine three-way gap comparisons, 320 positive-semidefiniteness stress configurations, and five rate fits; gap disagreement 1.527 × 10−15 and relative rate error 2.805 × 10−6

β, where truncating a Bessel tail can numerically erase an analytically positive reachable mode. The twelve nonlinear fits span invisible, coprime, non-coprime, and effective-period-two cases, and their fitted decay rates agree with the tangent spectral gaps at the accuracy reported in table 2. The geometric checks include orthogonal initial data, exact bipodal stationary states, allpairs-active equiangular configurations, two orientations relative to the RoPE planes, and inverse temperatures up to β = 16. The small negative Dini residual in the table is at the scale of double-precision roundoff, and the apparent trajectory excess is well below the 2 × 10−8 integrationcomparison tolerance. The twisted-state calculation separately treats generic resonance, even resonant anti-alignment, and the odd antipodal family; nonresonant rows compare the direct drift with (A-5) and its frequency-dependent Oω (n−1 ) bound. For the two-plane kernel, the three spectral formulations agree to near machine precision, while trajectories initialized along the slow tangent mode recover the predicted gap. Halving the sphereintegrator step from 2 × 10−2 to 10−2 changes the fitted rate by 1.173 × 10−8 when normalized by the predicted gap. On an independent circle problem, stage-wise retracted RK4 agrees with a high-accuracy phase-coordinate DOP853 solution, preserves unit norms to roundoff, and exhibits fourth-order step-halving behavior. Finally, the natural-energy derivative is evaluated both from its analytic directional-derivative identity and by centered differences. The two exact same-system states in theorem 3.2 reproduce opposite signs independently of any parameter scan. Similarly, the normalized residual update (3) converges at first order to the sphere vector field (5), providing a direct numerical check of the discrete-to-continuous convention.

9

Limitations and open problems

The model isolates one mechanism: RoPE rotates queries and keys, the value matrix is the identity, and normalization constrains every token to a sphere. It omits learned query/key/value maps, feed-forward layers, multiple heads, causal masking, stochastic depth, and finite-step or finite-depth 15

effects beyond the first-order normalized-residual limit (4). The theorems therefore explain a controlled dynamical limit rather than the full behavior of a trained Transformer. Architectures that rotate values or replace their coupling law, such as RoVE and oscillator-inspired attention [5, 16], lie outside this model. The global results are kernel-generic but regional. A fixed closed hemisphere is invariant but does not itself imply consensus, and the explicit contraction theorems require pairwise non-obtuse data or, on the circle, a strict open semicircle. Bipodal stationary states show that this boundary cannot be erased by a continuity argument. Two asymptotic questions remain outside the theorem set. First, the dense-phase expression 1−

I1 (β) 1 ∼ I0 (β) 2β

√ is expected when β, L → ∞ and L/ β → ∞, but the required uniform aliasing error bound is not yet proved. Second, a quantitative nonresonant finite-chain limit requires uniform boundary and Diophantine control. These statements remain a partial claim and a conjecture, respectively, and are not used by any closed theorem. Finally, the generic resonant twisted family has a large center subspace. Linear analysis proves instability but does not determine the center dynamics or justify a claim that these equilibria organize nonlinear transients. Such a claim would require a center-manifold analysis or dedicated trajectory evidence.

10

Conclusion

RoPE separates the score geometry from the value dynamics in spherical self-attention. It can break the natural vanilla energy law, but leaves enough reversible kernel structure to compute local consensus rates exactly and enough uniform positivity to prove contraction in explicit geometric regions. The resulting picture is not a universal “RoPE prevents collapse” statement: consensus remains locally stable and regionally attracting, while its rate can be exponentially small on a fixed resonant ring. RoPE selects a score-flattening twisted branch and an antipodal exception within the broader landscape of polygonal attention equilibria, while multiple frequencies make the local rate depend on finite sampling and consensus-plane energies in a way that can be non-monotone. Explicit theorem boundaries and independent numerical cross-checks keep these conclusions separate from the unresolved dense-phase and nonresonant-chain limits.

References [1] Álvaro Rodríguez Abella, João Pedro Silvestre, and Paulo Tabuada. The asymptotic behavior of attention in transformers. arXiv preprint arXiv:2412.02682, 2024. URL https://arxiv. org/abs/2412.02682. [2] Claudio Altafini. Multistability of self-attention dynamics in transformers. arXiv preprint arXiv:2511.11553, 2025. URL https://arxiv.org/abs/2511.11553. [3] Martin Burger, Samira Kabri, Yury Korolev, Tim Roith, and Lukas Weigand. Analysis of mean-field models arising from self-attention dynamics in transformer architectures with layer normalization. arXiv preprint arXiv:2501.03096, 2025. URL https://arxiv.org/abs/2501. 03096. 16

[4] Christopher Criscitiello, Quentin Rebjock, Andrew D. McRae, and Nicolas Boumal. Synchronization on circles and spheres with nonlinear interactions. arXiv preprint arXiv:2405.18273, 2024. URL https://arxiv.org/abs/2405.18273. [5] Alejandro García-Castellanos, Maurice Weiler, and Erik J. Bekkers. RoVE: Rotary value embeddings attention for relative position-dependent value pathways. arXiv preprint arXiv:2606.11275, 2026. URL https://arxiv.org/abs/2606.11275. [6] Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. The emergence of clusters in self-attention dynamics. arXiv preprint arXiv:2305.05465, 2023. URL https: //arxiv.org/abs/2305.05465. [7] Borjan Geshkovski, Hugo Koubbi, Yury Polyanskiy, and Philippe Rigollet. Dynamic metastability in the self-attention model. arXiv preprint arXiv:2410.06833, 2024. URL https://arxiv.org/abs/2410.06833. [8] Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. A mathematical perspective on transformers. Bulletin of the American Mathematical Society, 62:427–479, 2025. URL https://arxiv.org/abs/2312.10794. [9] Monica Goebel, Matthew S. Mizuhara, and Sofia Stepanoff. Stability of twisted states on lattices of kuramoto oscillators. Chaos, 2021. doi: 10.1063/5.0060095. URL https://arxiv. org/abs/2106.07119. [10] Zihan Gu, Ruoyu Chen, Han Zhang, Hua Zhang, and Yue Hu. Deconstructing positional information: From attention logits to training biases. International Conference on Learning Representations, 2026. URL https://arxiv.org/abs/2505.13027. [11] Nikita Karagodin, Yury Polyanskiy, and Philippe Rigollet. Clustering in causal attention masking. arXiv preprint arXiv:2411.04990, 2024. URL https://arxiv.org/abs/2411.04990. [12] Nikita Karagodin, Shu Ge, Yury Polyanskiy, and Philippe Rigollet. Normalization in attention dynamics. arXiv preprint arXiv:2510.22026, 2025. URL https://arxiv.org/abs/2510.22026. [13] Christian Kuehn and Jaeyoung Yoon. Spectral selection in symmetric self-attention dynamics. arXiv preprint arXiv:2604.26085, 2026. URL https://arxiv.org/abs/2604.26085. [14] Feilong Liu. Rotary positional embeddings as phase modulation: Theoretical bounds on the RoPE base for long-context transformers. arXiv preprint arXiv:2602.10959, 2026. URL https://arxiv.org/abs/2602.10959. [15] Joshua Nunley. Attention as frustrated synchronization. arXiv preprint arXiv:2606.18694, 2026. URL https://arxiv.org/abs/2606.18694. [16] Joshua Nunley. Kuramoto attention: Synchronizing self-attention on the torus. arXiv preprint arXiv:2606.11585, 2026. URL https://arxiv.org/abs/2606.11585. [17] Ayan Pendharkar. Gradient flow structure and quantitative dynamics of multi-head selfattention. arXiv preprint arXiv:2605.04279, 2026. URL https://arxiv.org/abs/2605.04279. [18] Duy-Tung Pham, An The Nguyen, Viet-Hoang Tran, Nhan-Phu Chung, Xin T. Tong, Tan M. Nguyen, and Thieu N. Vo. Dynamical properties of tokens in self-attention and effects of positional encoding. arXiv preprint arXiv:2512.03058, 2025. URL https://arxiv.org/abs/ 2512.03058. 17

[19] Philippe Rigollet. The mean-field dynamics of transformers. arXiv preprint arXiv:2512.01868, 2025. URL https://arxiv.org/abs/2512.01868. [20] Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. RoFormer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021. URL https://arxiv.org/abs/2104.09864. [21] Johan Thunberg, Johan Markdahl, Florian Bernard, and Jorge Goncalves. A lifting method for analyzing distributed synchronization on the unit sphere. arXiv preprint arXiv:1805.02528, 2018. URL https://arxiv.org/abs/1805.02528. [22] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017. URL https://arxiv.org/abs/1706.03762. [23] Xinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka, and Ali Jadbabaie. On the role of attention masks and LayerNorm in transformers. arXiv preprint arXiv:2405.18781, 2024. URL https://arxiv.org/abs/2405.18781.

A

Structural proofs

A.1

Well-posedness, detailed balance, and the sharp floor

Proof of theorem 3.1. The right-hand side of (5) is smooth on a neighborhood of (Sd−1 )n : all scores and exponentials are smooth, and every denominator Zi is strictly positive. Moreover, ⟨xi , ẋi ⟩ = 0, so the product sphere is invariant. Local existence and uniqueness therefore follow from the standard ODE theorem, and compactness of (Sd−1 )n rules out finite-time escape, giving a unique global solution. Orthogonality of the rotations gives D

E

sji = Rpj xj , Rpi xi = sij , hence W = W ⊤ . Row stochasticity is immediate from the definition of A. If ZΣ = πi Aij =

P

k Zk , then

Zi Wij Wij Wji = = = πj Aji , ZΣ Zi ZΣ ZΣ

which proves (6). Every score lies in [−1, 1]. For fixed i, j, Aij =

1+

1 1  ≥ , 1 + (n − 1)e2β k̸=j exp β(sik − sij )

P

because sik − sij ≤ 2. To see sharpness over the model class, fix a unit vector u and distinct indices i, j, and choose xk = Rp⊤k u (k ̸= j), xj = −Rp⊤j u. Then row i has sij = −1 and sik = 1 for every k ̸= j, so equality holds in (7). The comparison with e−2β /n follows by direct algebra, with equality only at β = 0.

18

A.2

The natural energy has no uniform sign law

For the circle model, put Wij = eβ cos δij ,

ui =

X

Wij sin δij ,

vi =

X

j

Wij sin(θi − θj ).

j

Since the ordered-pair energy counts both (i, j) and (j, i), ∂E+ = −2βui , ∂θi

θ̇i = −

vi , Zi

Ė+ = 2β

X ui vi

Zi

i

.

(A-1)

Proof of theorem 3.2. Substitution of (p1 , p2 , p3 ) = (0, 1, 2), β = 1, and ω = π into (A-1) gives, at the first state, √

Ė+

=

(0,π/6,π/3)

−1 + 3e1+ 3 √

e 3/2 + e1/2+ 3 + e1+ 3

≈ 1.666364304459315 > 0.

(A-2)

At the second state, the same identity yields √

Ė+

5e1+ 3/2 + 4

(0,π/6,0)

√ √ =− √ ≈ −0.1296131359239394 < 0. e 3/2 (1 + 2e1+ 3/2 )(2 + e1+ 3/2 )

(A-3)

Thus one fixed admissible system contains states with both derivative signs. This proves the stated no-go for E+ and −E+ , without making a claim about other candidate Lyapunov functions.

A.3

Consensus and twisted configurations

Proof of theorem 4.1. At consensus, the value average in each row equals x⋆ , so its tangent projection vanishes. On rotation plane ℓ, D

E

Rpi x⋆(ℓ) , Rpj x⋆(ℓ) = x⋆(ℓ)

2



cos ωℓ (pi − pj ) .

Summing over the mutually orthogonal planes proves (14), including its dependence only on the energies aℓ . Proof of theorem 4.2. At (15), every adjusted score phase is zero, hence Aij = 1/n and Fi := θ̇i = Let Gn (ω) =

X  1 n−1 sin ω(i − j) . n j=0

(A-4)

Pn−1 −iωj . Then nFi = Im(eiωi Gn ). If ω ∈ / 2πZ, the geometric-sum formula gives j=0 e

Gn (ω) = e−iω(n−1)/2 and therefore

sin(nω/2) , sin(ω/2)

sin(nω/2) n−1 sin ω i − n sin(ω/2) 2  

Fi =



.

(A-5)

This proves the two bounds in (17). If ω ∈ πZ, every summand in (A-4) vanishes. Suppose instead that sin ω = ̸ 0 and Fi = 0 for all i. The equations for two consecutive indices are two real linear constraints on the real and imaginary 19

parts of Gn ; their determinant is sin ω ̸= 0. Hence Gn = 0, which is equivalent to nω ∈ 2πZ. Conversely, Gn = 0 makes every Fi zero. This proves the equivalence (16). The frequency dependence in the final estimate is essential. Taking ωn = π/n and i = 0 gives F0 = −

X 1 n−1 πj 2 sin −→ − , n j=0 n π

so no O(n−1 ) bound can hold uniformly for frequencies approaching 2πZ with n.

A.4

Consensus linearization

Proof of theorem 4.3. Write yi (x) =

X

Fi (x) = Px⊥i yi (x).

Aij (x)xj ,

j

For tangent perturbations ξi ⊥ x⋆ , the first variation at consensus is δyi =

X

(δAij )x⋆ +

X

A⋆ij ξj .

j

j

P

Row stochasticity implies j δAij = 0, so δyi = Px y = y − ⟨x, y⟩ x at x = y = x⋆ now gives

⋆ ⋆ j Aij ξj , which is tangent at x . Differentiating

P

δ(Px y) = δyi − ξi . This proves (18) and (19). Detailed balance implies (π ⋆ )⊤ A⋆ = (π ⋆ )⊤ . Consequently, along the fixed linearized system, !

X X X d X ⋆ πi⋆ A⋆ij ξj − πi⋆ ξi = 0. πi ξi = dt i j i i

The weights here are frozen at the equilibrium, which is why this calculation does not yield a nonlinear conservation law.

B

Spectral and twisted-state proofs

B.1

Bessel aliasing on a resonant ring

Proof of theorem 5.1. At consensus and under (21), Wij⋆ = eβ cos(2πm(i−j)/n) . Thus W ⋆ is circulant and has a constant row sum Z; consequently A⋆ = W ⋆ /Z is circulant, symmetric, and doubly stochastic. Use the absolutely convergent expansion eβ cos ϑ =

X k∈Z

20

Ik (β)eikϑ .

(B-1)

Substituting (B-1), the unnormalized eigenvalue at Fourier mode q is cq = W

n−1 X

eβ cos(2πmh/n) e−2πiqh/n

h=0

=

X

Ik (β)

k∈Z

n−1 X

e2πi(mk−q)h/n

h=0

X

=n

Ik (β).

k∈Z mk≡q (mod n)

c0 , division proves (22). Since Z = W The congruence mk ≡ q (mod n) is solvable exactly when g = gcd(m, n) divides q. After writing m = gm′ and q = gs, the reachable solutions form one residue class modulo L = n/g, because m′ is invertible modulo L. For β > 0, every Ik (β) is strictly positive, so the L reachable Fourier modes have positive eigenvalues and the remaining n − L modes have eigenvalue zero. At β = 0, only I0 (0) = 1 survives, leaving the Perron eigenvalue one and n − 1 zeros.

B.2

Positive spectrum and the exact equality case

Proof of theorem 5.2. For each sampled position define ϕi =

r √ aℓ cos(ωℓ pi ), aℓ sin(ωℓ pi ) ℓ=1 .

Then Ka (pi − pj ) = ⟨ϕi , ϕj ⟩, so the matrix K = [Kij ] is positive semidefinite. By the Schur product theorem every Hadamard power K ◦k is positive semidefinite, and hence W ⋆ = exp◦ (βK) =

∞ X βk

k! k=0

K ◦k ⪰ 0.

(B-2)

With D = diag(W ⋆ 1), D1/2 A⋆ D−1/2 = D−1/2 W ⋆ D−1/2 =: S ⪰ 0. Thus A⋆ is similar to a symmetric positive semidefinite matrix. Its entries are strictly positive and it is stochastic, so Perron–Frobenius gives a simple eigenvalue one and all remaining eigenvalues strictly below one. This proves (24). At β = 0, W ⋆ = 11⊤ , so γ = 1. Let β > 0. Because the spectrum is nonnegative, γ = 1 exactly when W ⋆ has rank one. Its diagonal is the constant eβ . Vanishing of every 2 × 2 principal minor then gives 0 = e2β − e2βKij , hence Kij = 1 for all i, j. The converse is immediate. Finally, 1 − Kij =

X





aℓ 1 − cos ωℓ (pi − pj ) .

All summands are nonnegative, so Kij = 1 for all pairs precisely when every active plane satisfies (25).

21

B.3

Fixed-period low-temperature asymptotics

Proof of theorem 5.3. The positive spectrum from theorem 5.1 is, up to a permutation, the spectrum of the L-point phase grid. Scaling all entries of its kernel by e−β does not alter the normalized matrix. Put 2πk ∆k = 1 − cos , bk = e−β∆k . L For L ≥ 2, the exact gap is L−1 X

γL (β) =

min

2πsk 1 − cos L



bk

k=0

L−1 X

1≤s≤L−1



.

(B-3)

bk

k=0

If L = 1, the kernel has rank one and γ = 1. If L = 2, the scaled weights are b0 = 1 and b1 = e−2β , so the only nontrivial positive eigenvalue is λ1 =

1 − e−2β = tanh β, 1 + e−2β

which proves (26). Now fix L ≥ 3. The two nearest nonzero phase points, k = 1 and k = L − 1, have the common deficit ∆L = 1 − cos(2π/L); every other nonzero point has strictly larger deficit. Therefore L−1 X

bk = 1 + 2e−β∆L + o(e−β∆L ),

k=0

and, for each s = 1, . . . , L − 1, L−1 X k=0



bk

2πsk 1 − cos L



2πs −β∆L = 2 1 − cos e + o(e−β∆L ). L 



Because the set of s is finite, the remainder is uniform over it, and 

min

1≤s≤L−1

2πs 1 − cos L



= ∆L .

Substitution into (B-3) proves (27). The proof is pointwise in fixed L and therefore supplies no uniform error estimate for the joint limit described after the theorem.

B.4

Linear spectra of twisted configurations

At a twisted configuration, the first variation of every attention score vanishes because the adjusted phase is zero and the derivative of cosine at zero is zero. Differentiating the value torque therefore gives, for a phase perturbation η, (Jη)i =

1X Cij (ηj − ηi ), n j

22



Cij = cos ω(i − j) .

(B-4)

Proof of theorem 7.1. For a nontrivial resonant frequency, the root-of-unity sum gives C1 = 0, so (B-4) reduces to (40). Let zj = e2πimj/n . If 2m̄ ̸≡ 0 (mod n), the Fourier vectors z and z̄ are orthogonal, and   z z∗ 1 1 z̄ z̄ ∗ √ √ √ √ . C= + n 2 n n n n This proves (41). If m̄ = n/2, then Cij = (−1)i−j has rank one and eigenvalue n, which proves (42). In both cases J1 = 0, as also follows from global phase equivariance. Proof of theorem 7.2. Write n = 2r + 1 and sj = (−1)j for j = 0, . . . , n − 1, so 1⊤ s = 1. Then C = ss⊤ , C1 = s, and (B-4) becomes J=

 1 ss⊤ − diag(s) . n

(B-5)

By (B-5), the indices with sj = 1 number r + 1. Vectors supported on them and with coordinate sum zero form an r-dimensional eigenspace of eigenvalue −1/n. The r indices with sj = −1 analogously supply an (r − 1)-dimensional eigenspace of eigenvalue 1/n. Finally, 1 1 J s − 1 = s − 1. n n 



J1 = 0,

These invariant subspaces have total dimension n and prove (43). Removing the unique rotation mode leaves no zero eigenvalue and leaves r stable and r unstable directions.

B.5

Multi-frequency geometry and analytic counterexamples

Proof of theorem 7.3. The Gram–Schur argument in the proof of theorem 5.2 applies verbatim to Ka , proving W ⪰ 0, the stated similarity, and the eigenvalue ordering. The calculation in the proof of theorem 4.3 gives the Kronecker tangent Jacobian. Finally, the rank-one argument following (B-2) proves the equality condition (25). Proof of theorem 7.4. For lag h, the two-plane kernel is Ka (h) = a cos(πh/2) + (1 − a) cos(πh). Hence Ka (1) = a − 1 and Ka (2) = 1 − 2a. After scaling all weights by e−β , set y = e−2βa .

x = eβ(a−2) , The kernel and its degree vector are 

1 x y   W = x 1 x , y x 1

W 1 = (1 + x + y, 1 + 2x, 1 + x + y)⊤ .

The reversal-odd vector (1, 0, −1)⊤ gives the first non-Perron eigenvalue in (46). The second follows from tr(D−1 W ) = 1 + λ− + λ+ , proving that formula. For β = 34 log 2, direct substitution at a = 0, 2/3, 1 gives exactly (45). The three values are 0.564435 . . ., 0.75, and 0.668175 . . ., respectively, which proves the strict interior value above both endpoints claimed in the proposition.

23

Proof of theorem 7.5. After scaling by the common diagonal weight, the two-token kernel is !

1 q , q 1

W =

q = eβ[cos(ω∆)−1] .

Its non-Perron eigenvalue is (1 − q)/(1 + q), proving (47). For β > 0 and ∆ = 1, the low and high frequencies give qlow = e−β and qhigh = e−2β ; since the gap is increasing in q, the high frequency is slower. For ∆ = 2, they give qlow = e−2β and qhigh = 1, so the high-frequency gap is one and the ordering reverses.

C

Proofs of the global geometric estimates

All functions below are evaluated along the unique solution from theorem 3.1. We repeatedly use the uniform lower bound Aij ≥ aβ,n > 0. Lemma C.1 (Finite active-set calculus). Let f1 , . . . , fN be continuously differentiable functions and set fmin = minα fα and fmax = maxα fα . If Imin and Imax are the sets attaining the respective extrema at time t, then D+ fmin (t) =

min

α∈Imin (t)

f˙α (t),

D+ fmax (t) =

max

α∈Imax (t)

f˙α (t).

(C-1)

In particular, if D = maxi θi − mini θi , then D+ D = max θ̇i − min θ̇k . i∈I+

(C-2)

k∈I−

Proof. Because the family is finite, the expansions fα (t + h) = fα (t) + hf˙α (t) + o(h) are uniform in α. Every inactive index has a positive gap from the extremal value and remains inactive for all sufficiently small h > 0. Taking the minimum or maximum of the remaining first-order expansions proves (C-1); subtraction gives (C-2).

C.1

Fixed hemispheres

Proof of theorem 6.1. Let m = mw and take any active minimizing index i. Differentiating hi = ⟨xi , w⟩ gives X ḣi = Aij (hj − ⟨xi , xj ⟩ m) . (C-3) j

Since hj ≥ m, 

hj − ⟨xi , xj ⟩ m ≥ m 1 − ⟨xi , xj ⟩ . If m ≥ 0, the right-hand side is nonnegative. If m < 0, then 0 ≤ 1 − ⟨xi , xj ⟩ ≤ 2, so it is at least 2m. The active-set lemma therefore gives the global inequality D+ m ≥ 2 min{m, 0}.

(C-4)

The function m is absolutely continuous. For its negative part m− := max{−m, 0}, (C-4) implies, almost everywhere, ṁ− ≤ 2m− . If m(0) ≥ 0, Grönwall’s inequality forces m− (t) = 0 for all time. Once m ≥ 0, the active derivatives in (C-3) are nonnegative, so m is nondecreasing and m(t) ≥ m(0).

24

C.2

Pairwise non-obtuse data

For gij = ⟨xi , xj ⟩, direct differentiation of (5) gives ġij =

X

Aik (gkj − gik gij ) +

k

X

Ajk (gik − gjk gij ).

(C-5)

k

Proof of theorem 6.2. Let (i, j) be any active pair with gij = c. First suppose c < 0. Using gkj ≥ c, gik ≤ 1, and the analogous inequalities in the second sum, gkj − cgik ≥ c(1 − gik ) ≥ 2c,

gik − cgjk ≥ c(1 − gjk ) ≥ 2c.

Thus every active pair obeys ġij ≥ 4c, and D+ c ≥ 4c whenever c < 0. At c = 0, every summand in (C-5) for an active pair is nonnegative. Applying Grönwall to c− := max{−c, 0} therefore proves that c(0) ≥ 0 implies c(t) ≥ 0 for all t. Now assume c ≥ 0. For every k, gkj − cgik ≥ c(1 − gik ) ≥ 0,

gik − cgjk ≥ c(1 − gjk ) ≥ 0.

Retaining only k = j in the first sum of (C-5) and k = i in the second gives ġij ≥ (Aij + Aji )(1 − c2 ) ≥ 2aβ,n (1 − c2 ). Taking the minimum over active pairs proves (30). The scalar comparison equation also gives c(t) ≥ tanh(2aβ,n t) when c(0) = 0. At almost every time, differentiation of (31) yields ρ̇ = −

2ċ 1 − c2 ≤ −4a = −4aβ,n ρ. β,n (1 + c)2 (1 + c)2

Grönwall proves (32). Finally, tan2

DS 1 − cos DS 1−c = = = ρ, 2 1 + cos DS 1+c

which proves (33).

C.3

Strict semicircles and the sharp boundary

Proof of theorem 6.3. As long as D < π, take an active maximum i and active minimum k. Every phase difference θj − θi lies in [−D, 0], so all its sines are nonpositive and θ̇i ≤ −Aik sin D. Similarly, every θj − θk lies in [0, D], giving θ̇k ≥ Aki sin D. The active-set identity (C-2) and the attention floor now imply D+ D ≤ −(Aik + Aki ) sin D ≤ −2aβ,n sin D. On the maximal interval where D < π, this makes D nonincreasing, so D(t) ≤ D0 < π and no finite exit can occur. Thus the estimate holds globally. For q = tan(D/2), almost-everywhere differentiation gives q̇ ≤ −aβ,n sec2 (D/2) sin D = −2aβ,n q. One final application of Grönwall proves (35). 25

Proof of theorem 6.4. For xi = σi u, every row average is 

 X

Aij xj = 

X

j

Aij σj  u,

j

which is parallel to xi . Its tangent projection at xi vanishes. When d = 2 and both signs occur, the two phases differ by π, proving the boundary assertion.

C.4

Convergence to one point and observable rates

Proof of theorem 6.5. Both contraction theorems give q(t) := tan(D(t)/2) ≤ Q(t). The maximal chordal distance is therefore L(t) = 2 sin

D(t) 2q(t) 2Q(t) p =p ≤ . 2 1 + q(t)2 1 + Q(t)2

(C-6)

Using row stochasticity and the operator norm of an orthogonal projection, ∥ẋi ∥ ≤

X

Aij Px⊥i (xj − xi ) ≤

j

X

Aij ∥xj − xi ∥ ≤ L(t).

j

For s > t, the geodesic distance is bounded by the Riemannian length of the trajectory segment, so (C-6) gives dS (xi (t), xi (s)) ≤ ≤

Z s

∥ẋi (r)∥ dr

t

1 [arsinh Q(t) − arsinh Q(s)] . aβ,n

Thus each trajectory is Cauchy. The sphere is complete, and D(t) → 0, so all token limits coincide at some x∞ . Sending s → ∞ and using arsinh p z ≤ z for z ≥ 0 proves (36). Finally, D = 2 arctan q ≤ 2q, L = 2q/ 1 + q 2 ≤ 2q, and Y = 1 − c = L2 /2. Substitution of q(t) ≤ q0 e−2aβ,n t proves all three estimates in (37) and explains the factor two between the logarithmic rates of Y and L.

D

Numerical methods and parameter grids

The calculations independently cross-check the analytic statements and do not supply proof steps for any theorem. All computations use double precision, literal finite-chain position differences, and a periodic representation only for the exactly resonant circulant problem.

D.1

Error conventions and time integration

For spectra, the reported error is the largest absolute discrepancy between the compared eigenvalues after sorting, or the absolute discrepancy between two gap formulas. Structural zeros are instead identified from exact congruence reachability. For a measured decay rate γb and a predicted gap γ > 0, the relative error is |γb − γ| . γ One-sided Dini residuals are signed so that the proved inequality corresponds to a nonnegative value. 26

Circle trajectories are integrated in phase coordinates with DOP853. The local-rate fits start from a perturbation of amplitude 10−4 in the slow mode, use 360 output times, and fit log transverse norm over 0.75 ≤ γt ≤ 3.75 with relative and absolute solver tolerances 10−10 and 10−12 . Sphere trajectories use classical RK4 with every stage and accepted state retracted row-wise to the sphere. The multi-frequency rate calculation uses step size 10−2 , the same perturbation amplitude, and the same scaled-time fitting window. A separate order check compares step sizes 0.1, 0.05, and 0.025 with a phase-coordinate DOP853 reference computed at tolerances 10−13 and 10−15 ; both successive error ratios exceed 14, consistent with fourth order.

D.2

Resonant spectra and local rates

The resonant grid uses n ∈ {8, 16, 32, 64, 128}, β ∈ {0.5, 1, 2, 4, 8, 16}, and the distinct modes in {0, 1, 2, n/4, n/2}. For each of the resulting 144 configurations, the complete spectrum from direct diagonalization of the normalized kernel is compared with the congruence-filtered Bessel sum (22); the gap is also evaluated with a cancellation-free Fourier sum. Scaled modified Bessel functions prevent overflow. The exact unreachable-mode mask is determined by gcd(m, n), so an analytically positive mode below working precision is not misclassified as a structural zero. The nonlinear rate schedule consists of the triples (n, m, β) (8, 0, 2), (8, 1, 2), (8, 2, 2), (8, 4, 2), (16, 1, 1), (16, 4, 4), (16, 8, 4), (32, 2, 2), (32, 8, 2), (32, 16, 2), (64, 16, 1), and (128, 64, 1). This schedule covers invisible, coprime, non-coprime, anti-aligned, and multiple-size regimes without treating a dense parameter scan as statistical evidence.

D.3

Geometric-region checks

The pairwise non-obtuse and closed-hemisphere calculations use n ∈ {2, 5, 10}, d ∈ {2, 4, 8}, and β ∈ {0, 0.5, 2, 4, 8}, with 40 fixed seeded configurations per setting and the standard RoPE frequencies 10000−2ℓ/d . Acute states are formed around a random unit anchor with tangent perturbation magnitudes in [0.05, 0.45]. Hemisphere states use margins in [0.02, 0.8], with one token placed exactly on the boundary in half of the configurations. To exercise nonsmooth active sets, equiangular states use n ∈ {3, 5, 7} in two orientations relative to the RoPE planes. The grid combines c0 ∈ {0, 10−12 , 10−6 , 0.1, 0.5} with β ∈ {0, 2, 8, 16}. Every offdiagonal pair is active in these 120 configurations. The circle Dini grid uses n ∈ {2, 5, 10}, the same six inverse temperatures as the resonant grid, ω ∈ {0, 0.7, π}, initial diameters {π/2, 0.9π, π − 10−6 }, and 20 fixed seeded configurations per setting. Integrated comparison checks use n ∈ {2, 6}, β ∈ {0, 1, 2, 4}, the same three frequencies, and initial diameters {π/2, 0.9π, π − 10−3 }. Orthogonal and antipodal two-token states, together with the sharp β = 16 softmax-floor configuration, are evaluated separately.

D.4

Twisted states and nonresonant drift

For each n = 3, . . . , 12 at β = 2, every nonzero resonant mode is tested by comparing the direct right-hand side and the analytic Jacobian with a centered finite-difference Jacobian of step 2 × 10−7 . The odd antipodal family at ω = π is included separately for odd n. Nonresonant checks use √ ω ∈ {0.37, 0.73, 1.11, 2, 2.17} and compare the direct drift with (A-5), its sharp finite-sum bound, and the coarser 1/[n |sin(ω/2)|] envelope.

27

D.5

Multi-frequency spectra and nonlinear rates

The multi-frequency calculations use frequencies (1, 0.01). Three-way gap comparisons are performed for n ∈ {128, 256, 384}, β = 4, and high-frequency plane energies a ∈ {0, 1/2, 1}. Positive semidefiniteness is evaluated for n ∈ {8, 16, 32, 64} and β ∈ {0, 1, 4, 16} using 20 fixed seeded energy allocations per pair. Exact 2π frequency aliasing, block permutation, and invisibility of a nonzero 2π frequency on integer positions are checked separately. Nonlinear slow-mode rates use n = 16, β = 4, and a ∈ {0, 1/4, 1/2, 3/4, 1}. The finite-context observations in figure 4 evaluate endpoint gaps for 18 context lengths between 32 and 629 and five inverse temperatures between 0.5 and 8. Plane-energy curves use 41 equally spaced allocations for n ∈ {128, 256, 384} at β ∈ {2, 4}, together with five allocations for n ∈ {628, 629} at β = 4. These rows describe finite contexts and are not counted as independent theorem evidence. Energy derivative and discrete-time limit. The analytic directional derivative (A-1) is compared with a centered difference of the energy along the vector field using step 10−6 ; the two agree at relative and absolute tolerance 2 × 10−8 . The exact states (A-2)–(A-3) are evaluated directly and retain opposite signs. For the normalized residual update (3), the discrete velocity is compared with (5) at step sizes 2 × 10−4 and 10−4 on a six-token, four-dimensional configuration. Halving the step reduces the maximum discrepancy by at least the expected first-order factor, and the smaller-step discrepancy is below 2 × 10−4 .

28

Record · ID 405665 · SHA-256 6fa456a96217035f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.