ConceptioArchivearXiv CS
arXiv CSopen access

A Wasserstein Geometric Framework for Hebbian Plasticity

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

A Wasserstein Geometric Framework for Hebbian Plasticity

arXiv:2604.16052v1 [math.OC] 17 Apr 2026

Ulrich Tan ORCID: https://orcid.org/0009-0001-2907-501X April 11, 2025 Abstract We introduce the Tan–HWG framework (Hebbian–Wasserstein–Geometry), a geometric theory of Hebbian plasticity in which memory states are modeled as probability measures evolving through Wasserstein minimizing movements. Hebbian learning rules are formalized as Hebbian energies satisfying a sequential stability condition, ensuring well-posed fiberwise JKO updates, optimal-transport realizations, and an energy descent inequality. This variational structure induces a fundamental separation between internal and observable dynamics. Internal memory states evolve along Wasserstein geodesics in a latent curved space, while observable quantities—such as effective synaptic weights—arise through geometric projection maps into external spaces. Simplicial projections recover classical affine schemes (including exponential moving averages and mirror descent), while revealing synaptic competition and pruning as geometric consequences of mass redistribution. Hilbertian projections provide a geometric account of phase alignment and multi-scale coherence. Classical neural networks appear as flat projections of this curved dynamics, while the framework naturally accommodates richer distributional representations, including structural weights and embedding memories, and their spectral extensions in complex internal spaces. Under mild Lipschitz regularity assumptions, including a quasi-stationary “sleep-mode” regime, we establish the existence of continuous-time limit curves. This yields a variational formulation of memory consolidation as a perturbed Wasserstein gradient flow. The framework thus provides a unified geometric foundation for synaptic plasticity, representation dynamics, and context-dependent computation.

Introduction Hebbian plasticity [12] remains one of the central principles for understanding how neural systems adapt to experience. Classical formulations describe how synaptic strengths change as a function of co-activation, while modern approaches increasingly seek geometric or variational interpretations of synaptic dynamics [11, 8, 31]. Yet, despite the diversity of existing models, a unified geometric formulation of Hebbian learning has remained elusive. Optimal transport and Wasserstein geometry provide a natural language for describing the evolution of probability distributions [33, 29]. The JKO scheme [16], originally introduced to discretize gradient flows in Wasserstein space, offers a principled variational framework for modeling distributional dynamics. This makes Wasserstein geometry a compelling candidate for describing synaptic plasticity, although its relevance to Hebbian learning has not been systematically explored. This work introduces the Tan–HWG framework (Hebbian–Wasserstein–Geometry), a geometric theory of Hebbian plasticity in which memory states are represented as probability measures evolving through fiberwise Wasserstein proximal steps. Plasticity unfolds in a latent curved space, while 1

a global contextual signal encodes the functional state of the system and drives the local updates. This formulation reveals a structural separation between internal dynamics in Wasserstein space and observable dynamics obtained through geometric projection maps. Within this framework, classical neural networks appear as flat projections of a richer distributional geometry. The probabilistic representation naturally induces two complementary levels of structure: structural weights, encoding the allocation of synaptic influence, and embedding memories, encoding internal synaptic states. In complex internal spaces, these embeddings give rise to spectral memories that capture phase information. These phenomena emerge from the geometry itself rather than from additional modeling assumptions. The main contributions of this work are as follows: 1. A geometric formulation of Hebbian plasticity. We identify a broad class of Hebbian energies for which the minimizing-movement scheme is well posed, admits fiberwise optimal-transport realizations, and satisfies an energy descent inequality. 2. A dual internal/observable dynamics. Internal memory states evolve along geodesics in the Wasserstein space, while observable quantities arise through projections onto metric external spaces (e.g. simplex or Hilbert spaces). This duality provides geometric interpretations of synaptic competition, pruning, multi-scale coherence, neuronal assemblies, and phase alignment. 3. A variational account of memory consolidation. Under mild Lipschitz regularity assumptions, including a quasi-stationary “sleep-mode” regime, we establish the existence of continuoustime limit curves, yielding a perturbed Wasserstein gradient-flow formulation of memory consolidation. Overall, the framework provides a unified geometric foundation for synaptic plasticity, representation dynamics, and context-dependent computation, revealing how observable learning rules arise as projections of an underlying Wasserstein geometry.

1

Mathematical foundations

1.1

Internal state space

Let (X, µ) be a probability space, or more generally a finite-measure space, representing the set of (postsynaptic) neurons or readout units, µ serving as a reference measure on X. Let (Y, dY ) be a complete, separable, geodesic metric space (i.e. a geodesic Polish space), the internal state space, representing synaptic internal states. This assumption ensures that P2 (Y ) is itself a geodesic Polish space, which will be crucial for the variational dynamics introduced later. Example 1.1 (Compatible internal state spaces). Typical admissible examples of (Y, dY ) include Euclidean spaces Rd , and complete geodesic metric graphs, namely graphs whose edges are finite-length closed segments (ensuring geodesicity) and which have no open ends, missing vertices, or incomplete rays (ensuring completeness). More generally, CAT(0) spaces [3], which are also Polish, are admissible internal state space.

2

1.2

Memory fields

For each x ∈ X, the local memory state is a probability measure ρx ∈ P2 (Y ), and we denote by X := ρ : x 7→ ρx ∈ P2 (Y ) Borel measurable 

the space of memory fields, P2 (Y ) being naturally equipped with Wasserstein–2 metric and its induced Borel σ-algebra. Definition 1.2 (µ-compatible fields). A field ρ ∈ X is said to be µ-compatible if Z Z X

dY (y, y0 )2 dρx (y) dµ(x) < +∞

Y

for some (hence any) y0 ∈ Y . We denote by Xµ (X, Y ), or simply Xµ , the set of all such µ-compatible fields. By definition of P2 (Y ), for ρ ∈ X , for each x, we know that the second moment Z Y

dY (y, y0 )2 dρx (y) < +∞

is finite. The µ-compatibility for a memory field hence states that the µ-average of all fiber moments of the fields is finite. Measure-compatibility ensures tightness and compactness in the product space X with respect to the base measure µ. This property will be crucial later in the paper, in particular for the compactness of the discrete trajectories and the passage to the continuous-time limit in the minimizingmovement scheme. Remark 1.3 (Geometry induced by the memory field). Although not used in the developments below, the memory field ρ = (ρx )x∈X canonically induces a pseudo-metric on X via dρ (x, y) = W2 (ρx , ρy ), endowing X with a data-dependent geometry. This viewpoint highlights that the similarity structure between units is not fixed a priori but emerges from their local memory states. One may further equip (X, dρ ) with its Hausdorff measure µρ , yielding a complete, separable induced metric-measure structure (X, dρ , µρ ). While this evolving geometry will not be exploited in the sequel, it provides a useful conceptual interpretation of the framework as defining a dynamic cognitive geometry shaped by the memory field.

1.3

Example: synaptic relative weight representation

Let X denote a set of neurons. We consider two finite subsets: ′ XM := {x′1 , . . . , x′M } ⊂ X,

XN := {x1 , . . . , xN } ⊂ X,

representing two populations of neurons. We assume that each neuron in XN is connected to each ′ , as in a simple two-layer feed-forward architecture (if X ∩ X ′ ̸= ∅, the network no neuron in XM N M longer has a feed-forward topology, but the construction below still applies). ′ to X , we introduce a complete To represent the space of possible synaptic configurations from XM N geodesic metric graph Y with the following structure:

3

• a distinguished subset of leaves YM := {y1 , . . . , yM } ⊂ Y ′ ; corresponding to the neurons XM

• for each yi , an edge connecting yi to a central hub y o ∈ Y , so that Y is a star-shaped tree with center y o ∈ Y (not associated with any neuron) and leaves YM ; • the metric on Y is the natural path metric induced by the tree structure. y1 y6

y2 yo

y5

y3 y4

Y is a star-shaped geodesic tree with center y o and leaves YM = {y1 , . . . , yM } ′ . Figure 1: Star-shaped metric tree Y representing the synaptic targets XM

XN is naturally equipped with the normalized measure N 1 X δx , N j=1 j

µ=

and for each postsynaptic neuron xj ∈ XN , we represent its synaptic profile by a probability measure ρxj =

M X

pi,j δyi ∈ P2 (Y ),

i=1

supported on the leaves YM , where ρxj ({yi }) = pi,j ∈ [0, 1] denotes the relative synaptic strength from x′i (or yi ) to xj , with the constraint M X

pi,j = 1.

i=1

This representation captures only the relative profile of synaptic weights, not their absolute magnitudes. In particular, two nontrivial positive measures are considered equivalent whenever they differ by a global scaling factor: ∀ ν, ν ′ ∈ M+ (Y ),

ν ∼ ν′

⇐⇒

∃ c > 0, ν = c ν ′ .

Thus, the Wasserstein geometry acts on equivalence classes of synaptic weight profiles, reflecting the fact that only the relative distribution of synaptic strengths is modeled here. This normalization is 4

not intended as a biological assumption, but arises naturally from the Wasserstein formalism: the space P2 (Y ) provides a geometrically well-structured setting in which synaptic configurations can evolve along continuous geodesics. The Wasserstein distance W2 (ρxj , ρ′xj ) quantifies the minimal transport cost required to transform one synaptic profile into another. Because Y is a geodesic metric tree, this transport is uniquely defined and proceeds along the edges of the star-shaped graph. Intermediate points along such a Wasserstein geodesic between two profiles correspond to “virtual” synaptic configurations in which mass may temporarily reside along the edges of the tree or at the central hub y o . These intermediate states do not correspond to actual neurons, but they arise naturally from the requirement that synaptic evolution be continuous in P2 (Y ). 1/3

1/6

1/3

1/3

1/6

1/6 2/3

1/6 1/6

1/3

Intermediate state ρtx

State ρx

1/6

State ρ′x

Figure 2: Illustration of the Wasserstein transport between two discrete synaptic profiles. Mass is moved from ρx (left) to ρ′x (right) along the edges of the graph. The intermediate configuration ρtx (center) shows partial transport along each branch (blue points), representing the geodesic interpolation in the Wasserstein space, with t = 0.25. Thus, the representation

xj 7−→ ρxj ∈ P2 (Y )

provides a geometrically coherent model of synaptic relative weights, in which plasticity can be described as a continuous trajectory in Wasserstein space, while biological observations correspond to projections onto the discrete set of leaves YM . In particular, we can consider that along the geodesic (ρtxj )t∈[0,1] between ρ0xj = ρxj and ρ1xj = ρ′xj , the network is in a mixed state: with probability (1 − t) one observes the original profile ρxj , and with probability t one observes the updated profile ρ′xj . Averaging over multiple observations yields h

i

E ρtxj = (1 − t) ρxj + t ρ′xj . Equivalently, since each distribution is supported on the leaves YM , the expected weights satisfy h

i

E pti,j = (1 − t) pi,j + t p′i,j . Thus, the expected synaptic profile follows a straight line segment in the (M − 1)-dimensional simplex between the points pj = (p1,j , . . . , pM,j )

and

p′j = (p′1,j , . . . , p′M,j ).

Remarkably, this affine evolution is entirely independent of the metric structure of the underlying graph Y : although the internal Wasserstein geodesic depends on the lengths of the edges, the observable mixed state produces a universal barycentric interpolation in the simplex. 5

1.4

Geometry on the space of memory fields

We now introduce a natural geometry on the space of memory fields X . This will allow us to treat the evolution t 7→ ρ(t) as a trajectory in a metric space. Given two memory fields ρ, ρ′ ∈ Xµ , we define their distance by W(ρ, ρ′ ) :=

1/2

Z X

W22 (ρx , ρ′x ) dµ(x)

.

This is the standard product Wasserstein metric on the bundle P2 (Y )X , weighted by the reference measure µ on X (see [32, 1]). The space (Xµ , W) is complete and separable whenever (Y, dY ) is complete and separable. This equips Xµ with a canonical metric structure in which memory fields can be compared, interpolated, and evolved. Remark 1.4 (Domain of the metric.). Although each fibre ρx belongs to P2 (Y ), the product Wasserstein metric Z W22 (ρx , σx ) dµ(x) W(ρ, σ)2 = X

is finite only when the µ-average of the second moments is finite. For this reason, the natural domain of the metric W is the subspace n

Xµ := ρ ∈ X :

Z Z X

Y

o

dY (y, y0 )2 dρx (y) dµ(x) < ∞ ,

on which W is a well-defined finite metric. Working in Xµ ensures tightness, compactness of discrete trajectories, and stability of the Wasserstein geometry under limits. Remark 1.5 (Cognitive interpretation). Geometrically, W(ρ, ρ′ ) measures the total amount of transport required to transform the entire memory field ρ into ρ′ , aggregating the fiberwise Wasserstein costs across X. Two global states are close when, for most units x ∈ X, the corresponding internal distributions encode similar information. Thus W measures the dissimilarity between two “memory states” in a geometrically meaningful way. Moreover, µ-compatibility of memory fields is actually stable under Wasserstein geometry. More precisely, we have the following proposition: Proposition 1.6 (Geodesic stability of µ-compatible fields). Assume ρ0 , ρ1 ∈ Xµ are µcompatible fields, i.e. Z Z X

Y

dY (y, y0 )2 dρix (y) dµ(x) < ∞,

i = 0, 1,

for some (hence any) y0 ∈ Y . Let (ρt )t∈[0,1] be a constant-speed geodesic between ρ0 and ρ1 in (X , W). Then, for every t ∈ [0, 1], the field ρt is µ-compatible. In particular, the set of µ-compatible fields Xµ is geodesically convex in (X , W). Proof. By the product structure of (X , W), a constant-speed geodesic (ρt )t∈[0,1] between ρ0 and ρ1 is given fiberwise by Wasserstein geodesics: for µ-a.e. x ∈ X, (ρtx )t∈[0,1] is a W2 -geodesic in P2 (Y ) between ρ0x and ρ1x . 6

Fix y0 ∈ Y . In P2 (Y ), along a W2 -geodesic (ν t )t∈[0,1] we have Z Y

dY (y, y0 ) dν (y) ≤ (1 − t) t

2

Z

dY (y, y0 ) dν (y) + t 2

Y

0

Z Y

dY (y, y0 )2 dν 1 (y),

t ∈ [0, 1].

Applying this with ν t = ρtx for µ-a.e. x and integrating with respect to µ yields Z Z X

Y

dY (y, y0 )2 dρtx (y) dµ(x) ≤ (1 − t) +t

Z Z

Z

X Z

X

Y

Y

dY (y, y0 )2 dρ0x (y) dµ(x)

dY (y, y0 )2 dρ1x (y) dµ(x)

Since the right-hand side is finite by µ-compatibility of ρ0 and ρ1 , the left-hand side is finite for all t ∈ [0, 1], so each ρt is µ-compatible. This proves the claim. Between two µ-compatible memory fields, it will be useful to consider a µ-measurable family of optimal transport couplings between local (fiberwise) memory states. Hence the following lemma. Lemma 1.7 (Measurable selection of optimal couplings). Let ρ, ρ′ ∈ Xµ . For µ-a.e. x ∈ X, let Π(ρx , ρ′x ) denote the set of optimal couplings between ρx and ρ′x . Then, the set-valued map x 7→ Π(ρx , ρ′x ) admits a measurable selection. Proof. By standard optimal transport theory, this set is nonempty and closed in P2 (Y × Y ). Moreover, since x 7→ ρx and x 7→ ρ′x are measurable (by definition of Xµ ⊂ X ), the set-valued map x 7→ Π(ρx , ρ′x ) admits a measurable selection by the Kuratowski–Ryll-Nardzewski theorem.

1.5

The context space

We now introduce the notion of context, which encodes the functional state of the system. Let (Γ, m) be a measured space representing the internal activation space. A context is encoded by a µmeasurable family of finite complex measures ϕ = {ϕx }x∈X ∈ M(Γ)X , not necessarily normalized. We consider the Banach space of complex measures (M(Γ), ∥ · ∥), where ∥ · ∥ is the total variation norm. We say that a context ϕ is µ-compatible if Z X

∥ϕx ∥2 dµ(x) < +∞,

and we denote by Cµ the set of µ-compatible contexts i.e. 

Cµ := ϕ ∈ M(Γ)X x 7→ ϕx is measurable,



Z X

∥ϕx ∥2 dµ(x) < +∞ .

We endow Cµ with the following distance dC (ϕ, ϕ′ ) :=

Z X

∥ϕx − ϕ′x ∥2 dµ(x)

7

1/2

,

ϕ, ϕ′ ∈ Cµ .

Remark 1.8 (Interpretation of the internal activation space and complex-valued contexts.). The internal activation space represents the state of the world, at least as perceived by the agent. When the world is perceived solely through neuronal activation, it is natural to take the internal activation space to coincide with the neuronal representation, i.e. X and Y . This is analogous to a human perceiving the world only through sensory pathways translated into neural encodings, or to an artificial system encoding only what its sensors capture (visual signals, kinetic sensors etc.). In this view, ∥ϕx ∥ can be seen as the “intensity” of context on x: it is not normalized. The complex-valued nature of contexts will reveal itself as natural in the case of generalized neural networks (Section 6), where complex-valued synaptic weights allow convenient representations of phase synchronization and alignment mechanisms. Moreover, it endows the context space with a natural metric (via total variation norm), which provides a clean setting to define regularity assumptions relative to the contexts (Section 8). In Section 5.7 we will see an example where Γ = Y × Y , and in Section 6, an example where Γ = Y × X.

1.6

External representation space

Let (Z, dZ ) be a metric space, representing the external state space. Z is the space where observables evolve. Z is endowed with its borelian measure. A context ϕ is projected to Z via a global signal (or signal) S(ρ, ϕ), where S : Xµ × M(Γ)X → Z X couples a memory field ρ ∈ Xµ with the context.

2

Hebbian–Wasserstein compatible energies

The intuition behind the framework is that Hebbian memory updates behave like a Wasserstein minimizing-movement dynamics. We therefore introduce a class of energies whose fiberwise structure and stability properties ensure that a JKO-type scheme is well posed.

2.1

Hebbian energies: structural definition

We interpret Hebbian principle [12] as follows: each unit updates its local memory state based solely on its own activity and on a global co-activation signal summarising the collective configuration. In our formalism, the “weight” of x′ on x is represented by the measure ρx around x′ . Thus the energy driving the update of ρx should depend only on the pair (ρx , Sx (ρ, ϕ)), where Sx (ρ, ϕ) encodes the global context. Definition 2.1 (Hebbian energy). An energy functional E : Xµ × M(Γ)X → R ∪ {+∞} is Hebbian if there exist: • a metric space (Z, dZ ), the external state space, equiped with its Borel σ-algebra; • a family of maps S = {Sx }x∈X , the global signal, with each Sx : Xµ × M(Γ)X → Z,

8

such that for every (ρ, ϕ) ∈ Xµ × M(Γ)X , x 7→ Sx (ρ, ϕ) is Borel measurable as a map from X to Z; • a family of maps E = {Ex }x∈X , the local energy density, with each Ex : P2 (Y ) × Z → R ∪ {+∞}, such that the map

(x, η, z) 7→ Ex (η, z)

is Borel measurable on X × P2 (Y ) × Z; such that for all (ρ, ϕ): E(ρ, ϕ) =

Z

Ex ρx , Sx (ρ, ϕ) dµ(x). 

X

This definition captures our structural Hebbian requirement: the energy decouples across fibres, and each fibre (i.e. each unit) interacts with the rest of the system only through the global signal.

2.2

Sequential stability: analytic admissibility

When the global signal is held fixed, each unit should respond in a stable and well-behaved manner to this signal. This reflects the sequential nature of synaptic plasticity, where updates occur locally and asynchronously under a slowly varying global influence. Definition 2.2 (Sequential stability). A Hebbian energy E, with local energy density E, is sequentially stable if for every reference configuration (ϕref , ρref ) ∈ M(Γ)X × Xµ , the frozen energy, defined by Eϕref ,ρref (ρ) :=

Z X

Ex ρx , hx dµ(x), 

where h = S(ρref , ϕref ) is fixed,

satisfies the AGS assumptions, namely: 1. Properness: Eϕref ,ρref is not identically +∞ and never takes the value −∞. 2. Lower semicontinuity: Eϕref ,ρref is lower semicontinuous with respect to W-convergence on Xµ . 3. Geodesic convexity: Eϕref ,ρref is λ-geodesically convex in (Xµ , W), λ > −∞. Notations. At a fixed reference configuration (ϕref , ρref ), we denote by Sϕref ,ρref := S(ρref , ϕref ) the frozen global signal (or simply frozen signal) and by Eϕref ,ρref (ρ) :=

Z X

Ex ρx , Sϕref ,ρref , x dµ(x) 

the frozen energy. The AGS assumptions [1] guarantee the well-posedness of minimizing-movement schemes with the frozen energy as we will see later. Nevertheless, in practice, the definition of a Hebbian energy is

9

naturally given by defining its local energy density. In such a case, establishing sequential stability is more natural at the level of the local energy itself. Hence the following definition. Definition 2.3 (Local sequential stability). A Hebbian energy E is locally sequentially stable if for every reference configuration (ϕref , ρref ) ∈ M(Γ)X × Xµ , for µ-a.e. x ∈ X, when fixing the global signal z = Sx (ρref , ϕref ) ∈ Z, the local energy density Ex (·, z) : P2 (Y ) → R ∪ {+∞} satisfies: 1. Properness: Ex (·, z) is not identically +∞ and never takes the value −∞. 2. Lower semicontinuity: Ex (·, z) is lower semicontinuous with respect to W2 -convergence on P2 (Y ). 3. Geodesic convexity: Ex (·, z) is λ-geodesically convex in (P2 (Y ), W2 ), λ > −∞. 4. Lower bound: there exists a function c : X × Z → R ∪ {+∞} such that for all z ∈ Z: c(·, z) ∈ L1 (X, µ),

and

∀x ∈ X, ∀η ∈ P2 (Y ),

Ex (η, z) ≥ c(x, z).

5. Existence of a finite-energy field: for every reference configuration (ϕref , ρref ), there exists at least one ρ ∈ Xµ such that Z X

Ex (ρx , Sx (ρref , ϕref )) dµ(x) < +∞.

The lower bound property and the existence of a finite-energy field property are key to ensure properness of the total energy, so that we can have the following proposition: Proposition 2.4 (Local sequential stability implies sequential stability). A locally sequentially stable Hebbian energy is sequentially stable. Proof. Fix a reference configuration (ϕref , ρref ) ∈ M(Γ)X × Xµ and set h := S(ρref , ϕref ) ∈ Z X . Define the frozen energy Eϕref ,ρref (ρ) :=

Z X

Ex (ρx , hx ) dµ(x),

ρ ∈ Xµ .

Properness. The lower bound implies Eϕref ,ρref (ρ) ≥

Z X

c(x, hx ) dµ(x) > −∞

for all ρ, so the functional never takes the value −∞. Moreover, the existence of a finite-energy field property ensures that the frozen energy is not identically +∞. Hence the properness of the frozen energy. Lower semicontinuity. Let (ρn )n∈N ⊂ Xµ converge to ρ in (Xµ , W). Up to a subsequence, we may assume ρnx → ρx in W2 for µ-a.e. x. By fiberwise lower semicontinuity, Ex (ρx , hx ) ≤ lim inf Ex (ρnx , hx ) n→∞

10

for µ-a.e. x.

Subtracting the integrable lower bound c(x, hx ) and applying Fatou’s lemma to the nonnegative functions Ex (ρnx , hx ) − c(x, hx ) yields Z X

that is,

Ex (ρx , hx ) dµ(x) ≤ lim inf n→∞

Z X

Ex (ρnx , hx ) dµ(x),

Eϕref ,ρref (ρ) ≤ lim inf Eϕref ,ρref (ρn ), n→∞

so Eϕref ,ρref is W–lower semicontinuous on Xµ . Geodesic convexity. By the product structure of (X , W), a constant-speed geodesic (ρt )t∈[0,1] between ρ0 and ρ1 is given fiberwise by Wasserstein geodesics: for µ-a.e. x ∈ X, (ρtx )t∈[0,1] is a W2 -geodesic in P2 (Y ) between ρ0x and ρ1x . By fiberwise λ-geodesic convexity (λ > −∞), for every t ∈ [0, 1] and µ-a.e. x, Ex (ρtx , hx ) ≤ (1 − t) Ex (ρ0x , hx ) + t Ex (ρ1x , hx ) −

λ t(1 − t)W22 (ρ0x , ρ1x ). 2

Integrating over X gives Eϕref ,ρref (ρt ) ≤ (1 − t) Eϕref ,ρref (ρ0 ) + t Eϕref ,ρref (ρ1 ) −

λ t(1 − t)W 2 (ρ0 , ρ1 ), 2

showing λ-geodesic convexity in (Xµ , W) (λ > −∞). All AGS conditions are therefore satisfied, and the frozen energy is sequentially stable.

2.3

Tan–HWG energy class

Notations. To avoid confusion with existing vocabulary, we will regularly use the tag “Tan–HWG”, standing for Hebbian–Wasserstein–Geometry, to denote specific properties of our framework. This leads to the definition of the class of energies that will be considered in our framework. Definition 2.5 (Tan–HWG energy). A Hebbian energy satisfying sequential stability is called a Tan–HWG energy. Section 5 provides examples of Tan–HWG energies. Beyond biological motivations (i.e. sequential nature of the updates), freezing the signal at each step improves computational practicality, while preserving enough stability for a well-posed dynamics. The next section details and justifies the update mechanism.

3

Discrete Tan–HWG dynamics

We now construct the dynamics associated with Tan–HWG energies. The key idea is that at each time step, the global signal is frozen, and each fibre performs a Wasserstein proximal step with respect to the frozen energy.

11

3.1

The Tan–HWG minimizing-movement scheme

Given a Tan–HWG energy and an initial configuration ρ0 ∈ Xµ , the discrete-time Tan–HWG evolution is defined by the JKO-type scheme [16, 1] 

ρn+1 ∈ argmin Eϕn ,ρnτ (ρ) + τ ρ ∈ Xµ

1 2τ



Z X

W22 (ρx , ρnτ,x ) dµ(x) .

(1)

for a timestep τ > 0. The sequential stability conditions of Definition 2.2 ensure that, at each step, the minimisation problem is well-posed. In particular, every step produces a stable Hebbian update of the memory field. This motivates the definition of the τ -Tan–HWG update operator 

TτE (ρ | ϕ) := argmin Eϕ,ρ (ρ′ ) + ρ′ ∈ X µ

1 2τ

Z X



W22 (ρ′x , ρx ) dµ(x) .

As is standard in minimizing-movement schemes, the operator may be multivalued. We write TτE (ρ | ϕ) to emphasize that the Tan–HWG update acts on the memory field ρ while the context ϕ is frozen and plays the role of an external parameter.

3.2

Well-posedness of Tan–HWG minimizing-movement scheme

The following Theorem formally states that Tan–HWG energies—that is, Hebbian and sequentially stable energies—generate a well-posed Tan–HWG scheme, with an energy descent inequality (EDI). Theorem 3.1 (Discrete Tan–HWG dynamics). Let τ > 0 be a fixed time step. Let E be a Tan–HWG energy with local energy density E and global signal S. Let TτE be the τ -Tan–HWG update operator defined by TτE (ρ | ϕ) := argmin ρ′ ∈ Xµ 

Let (ϕnτ )n≥0 ∈ M(Γ)X

N



1 Eϕ,ρ (ρ ) + W 2 (ρ′ , ρ) . 2τ 

be a set of contexts.

Then for every ρ0τ ∈ Xµ , the sequence ρn+1 ∈ TτE (ρnτ | ϕnτ ), τ

for all n ≥ 0,

is well-defined. Moreover: • the update is fiberwise: for µ-a.e. x, ρn+1 τ,x solves an independent convex minimisation problem; • the scheme satisfies the energy descent inequality (EDI) Eϕnτ ,ρnτ (ρn+1 )+ τ

1 2 n+1 n W (ρτ , ρτ ) ≤ Eϕnτ ,ρnτ (ρnτ ). 2τ

12

Thus the Tan–HWG scheme defines a well-posed discrete-time dynamics on (Xµ , W), in the sense of a minimizing-movement scheme. Proof. Fix ρ0τ ∈ Xµ and n ≥ 0. For ρ ∈ Xµ , define 1 Fn (ρ) := Eϕnτ ,ρnτ (ρ) + W 2 (ρ, ρnτ ) = 2τ where Fn,x (η) := Ex (η, Sϕnτ ,ρnτ , x ) +

Z X

Fn,x (ρx ) dµ(x),

1 2 W (η, ρnτ,x ). 2τ 2

Step 1: existence of the update. By sequential stability, the frozen energy map ρ 7→ Eϕnτ ,ρnτ (ρ) satisfies the AGS assumptions on (Xµ , W). Hence each Fn admits at least one minimizer ρn+1 ∈ Xµ . τ E (ρn | ϕn ) for all n ≥ 0 is well-defined. Hence the sequence ρn+1 ∈ T τ τ τ τ Step 2: fiberwise reduction. Since Fn is an integral of decoupled fiberwise functionals, any minimizer ρ of Fn must satisfy, for µ-a.e. x, ρx ∈ argmin Fn,x (η). η∈P2 (Y )

Otherwise, replacing ρx on a set of positive µ-measure by a measurable selection of fiberwise minimizers would strictly decrease Fn , while preserving µ–compatibility, a contradiction. Step 3: energy descent inequality. Since ρn+1 minimises Fn , we have τ Fn (ρn+1 ) ≤ Fn (ρnτ ), τ which expands as

1 2 n+1 n W (ρτ , ρτ ) ≤ Eϕnτ ,ρnτ (ρnτ ). 2τ This is exactly the stated energy descent inequality. Eϕnτ ,ρnτ (ρn+1 )+ τ

Thus the Tan–HWG update is well defined, with fiberwise minimisation and EDI, and the discrete dynamics is well posed. Remark 3.2 (First variation). For a frozen signal Sϕ,ρ , the frozen energy admits a metric subdifferential which decomposes fiberwise: ξ ∈ ∂Eϕ,ρ (ρ′ )

⇐⇒

ξx ∈ ∂1 Ex (ρ′x , Sϕ,ρ,x )

for µ-a.e. x.

Whenever E is a Tan–HWG energy, we write  δE (ρ, ϕ) ∈ ∂1 Ex ρx , Sx (ρ, ϕ) δρx

for any choice of fiberwise subgradient. The formal identity  δE ∂Ex (ρ, ϕ) = ρx , Sx (ρ, ϕ) δρx ∂ρ

is therefore understood in this subdifferential sense, with the global signal frozen. In other words, the first variation of E is defined relative to the subdifferential of the frozen energy. 13

Remark 3.3 (Global signal and continuous Tan–HWG dynamics). Since the global signal is frozen during each fixed time step τ > 0, no additional regularity assumption on the signal S is required for the well-posedness of the discrete Tan–HWG dynamics. In contrast, the identification of a vanishing time-step limit τ → 0 as a Wasserstein gradient flow requires further assumptions, in particular on the regularity, stability, and coercivity properties of the signal S and of the associated energy. Even for fixed context ϕ, the energy E(ρ, ϕ) =

Z

Ex ρx , Sx (ρ, ϕ) dµ(x) 

X

must satisfy the structural conditions of the AGS minimizing-movement theory [1]. In this work, we focus on discrete Tan–HWG dynamics, which are both mathematically wellposed and computationally practical (e.g. see Section 5.6). Nevertheless, Section 8 will explore a continuous-time limit of the Tan–HWG scheme under the special case of a quadratic energy with Lipschitz global signal and compatible set of contexts.

3.3

Brenier and metric graph realizations of Tan–HWG updates

In the Euclidean case, the fiberwise JKO update admits a classical optimal-transport representation. In the special case Y = Rd and the ρnτ,x admit densities, we can deduce a Brenier realization of the update: Corollary 3.4 (Brenier realization of Tan–HWG updates). Assume the setting of the τ -Tan– HWG scheme, with Y = Rd . Suppose moreover that for each n ≥ 0 and µ-a.e. x ∈ X, the d measures ρnτ,x and ρn+1 τ,x admit densities with respect to the Lebesgue measure on R . Then for µ-a.e. x ∈ X and each n ≥ 0: 1. There exists a convex function Ψnx : Rd → R such that the optimal transport map Txn := ∇Ψnx pushes ρnτ,x forward to ρn+1 τ,x , i.e. n n ρn+1 τ,x = Tx,# ρτ,x .

2. The Wasserstein geodesic between ρnτ,x and ρn+1 τ,x is given by ρnτ,x (t) := (1 − t)Id + tTxn # ρnτ,x ,

t ∈ [0, 1],



and satisfies the continuity equation ∂t ρnτ,x (t) + ∇ · ρnτ,x (t) vxn (t) = 0, 

for a suitable velocity field vxn (t) derived from Txn . In particular, each discrete Tan–HWG update ρnτ 7→ ρn+1 admits a realization as a Brenier τ optimal transport in each fibre, and the Hebbian dynamics can be interpreted as a potentialdriven mass flow in the internal space.

14

Proof. Fix n ≥ 0 and x ∈ X such that ρnτ,x and ρn+1 τ,x admit densities with respect to the Lebesgue d n d measure on R . Since ρτ,x ∈ P2 (R ) is absolutely continuous, Brenier’s theorem applies and yields the claims. The conclusion follows for µ-a.e. x and all n. This is the standard McCann interpolation. Additional mathematical details on fixed points and stability of Tan–HWG dynamics are provided in Appendix A.1. While the Brenier–McCann representation provides a natural description of Tan–HWG updates in Euclidean internal spaces, the applications developed in this work rely on internal representations whose geometry is non-Euclidean. In particular, several of our examples use internal spaces Y that are geodesic metric graphs. In this setting, optimal transport remains well defined, but the Brenier map is replaced by a geodesic interpolation along the branches of the graph. Corollary 3.5 (Geodesic realization of Tan–HWG updates on metric graphs). Assume the setting of the τ -Tan–HWG scheme, and let Y be a compact geodesic metric graph endowed with its shortest-path distance. For each n ≥ 0 and µ-a.e. x ∈ X, let ρnτ,x , ρn+1 τ,x ∈ P2 (Y ) be the fiberwise minimizers produced by the Tan–HWG update. Then for µ-a.e. x ∈ X and each n ≥ 0: 1. There exists an optimal transport plan πxn ∈ Opt ρnτ,x , ρn+1 τ,x



for the quadratic cost on Y . 2. The constant-speed Wasserstein geodesic between ρnτ,x and ρn+1 τ,x is given by ρnτ,x (t) := (et )# πxn ,

t ∈ [0, 1],

where et : Geo(Y ) → Y evaluates a geodesic at time t. 3. This interpolation satisfies the continuity equation on the graph, ∂t ρnτ,x (t) + ∇Y · ρnτ,x (t) vxn (t) = 0, 

where vxn (t) is the metric velocity field induced by the geodesics in the support of πxn . In particular, each discrete Tan–HWG update admits a realization as a geodesic mass flow on the metric graph Y , and the Hebbian dynamics can be interpreted as a redistribution of mass along shortest paths in the internal space. Proof. Since Y is a compact geodesic metric space, (P2 (Y ), W2 ) is geodesic and every pair of measures admits at least one optimal transport plan for the quadratic cost. Let πxn be such a plan between ρnτ,x and ρn+1 τ,x . The geodesic interpolation

ρnτ,x (t) := (et )# πxn

is the standard McCann interpolation on geodesic spaces. It is well known that such interpolations satisfy a continuity equation with a velocity field given by the metric derivative of the underlying geodesics. The conclusion follows for µ-a.e. x. 15

Remark 3.6 (Interpretation). The theorem guarantees the existence of the discrete sequence (ρn )n , but it does not describe the mechanism by which one passes from ρn to ρn+1 . The corollary shows that this update is not a mere jump between two abstract measures: it can always be realized as a continuous displacement of mass along a Wasserstein geodesic. This continuous interpolation is essential for interpreting the Tan–HWG dynamics as an internal evolution of mass on the graph Y , rather than as a purely discrete update rule. Remark 3.7 (Euclidean vs. graph geometry). In the Euclidean case Y = Rd , optimal transport between absolutely continuous measures is induced by a unique map ∇Ψ for a convex potential Ψ (Brenier’s theorem). The Wasserstein geodesic is therefore obtained by pushing mass along straight lines in Rd . In contrast, when Y is a geodesic metric graph, optimal transport is generally not induced by a map but by an optimal plan supported on geodesics of the graph. Mass travels along shortest paths, which may branch or merge, and the Wasserstein geodesic is obtained by interpolating these paths. Thus, while the Euclidean case yields a “potential flow”, the graph case yields a “geodesic redistribution of mass” along the combinatorial structure of Y . Mass on the edges and motivation for a leaf-level representation. In the Tan–HWG scheme on a metric graph Y , each fiberwise update ρnx 7→ ρn+1 is obtained by minimizing over the x whole space P2 (Y ). In particular, there is no reason for the iterates (ρnx )n to remain supported on the leaves of Y : whenever transporting mass through the interior edges reduces the total cost, the optimal states ρnx may (and typically will) place positive mass on the segments of the graph. By optimality of transport plans between discrete distributions, each state remains a discrete measure, but the movement of mass occurs continuously along edges. This phenomenon is intrinsic to the Wasserstein geometry on Y and reflects the fact that the Tan–HWG update corresponds to a genuine displacement of mass in the internal space, rather than a purely discrete jump between leaf configurations. For this reason, it is natural to introduce a leaf-level representation of the dynamics, obtained by projecting the local memory states onto distributions on the set of leaves. The resulting projected evolution lives in the simplex P(YM ) and provides a convenient observable description of the model. The next section formalises this idea in a general setting, leading to the geodesic projection principle in Principle 4.4. Remark 3.8. We write P2 (Y ) for the W2 –Wasserstein space over Y , and P(YM ) for the Euclidean simplex of probability measures over the finite observable space YM .

4

Extension of the dynamics to finite internal support systems

Unstability of Tan–HWG dynamics on finite support. Finite spaces YM = {yi , . . . , yM } cannot be admissible as internal state spaces since they are not geodesic. A natural way to circumvent this is to lift YM into an admissible space Y i.e. YM ⊂ Y . In this case, consider a memory set (ρn )n≥0 following a Tan–HWG minimizing-movement scheme. Assume a state ρn has all its supports on YM i.e. for µ-a.e. x, ρnx is a probability distribution with support exclusively on YM . ′ In the general setting, nothing prevents the following states (ρn )n′ ≥n to have supports beyond YM : the induced distribution dynamics is not stable on the finite support.

16

4.1

Finite internal support systems

We start by generalizing a simplex representation for distributions that have a finite support on an arbitrary internal state spaces Y . Definition 4.1. We say that a neural system is a finite neural system with M internal degrees of freedom on Y , if, for each x ∈ X, its local memory state can be represented as a finitely supported probability measure ρx =

M X

wi,x δyi ,

i=1

for some fixed family {yi }i∈J1,M K ⊂ Y and weights wx = (w1,x , . . . wM,x ) ∈ ∆M −1 . In such settings, memory fields ρ admits a natural simplicial representation w = {wx }x∈X , with wx = (w1,x , . . . wM,x ) ∈ ∆M −1 . Remark 4.2 (Intuition). Assume a finite neural system evolves according to a discrete τ -Tan–HWG minimizing-movement scheme, where each ρn admits a representation on the simplex. For a fiber x, the simplicial representation of the trajectory (ρnx )n≥0 would be a discontinuous set of points on the simplex (due to the finite timestep τ of the scheme). A continuous simplicial curve representation occurs when a continuous limit curve of the trajectory exists at the limit τ → 0. Two challenges arise: 1. first, representing ρn on the simplex is not trivial when ρn has not all its support on YM , which is generally the case; 2. second, the existence of a limit curve when τ → 0 for a Tan–HWG minimizing-movement scheme is not guaranteed. The geodesic projection principle introduced below provides a way to overcome the first difficulty by defining a canonical projection of local memory states onto measures on a finite set. We present a projection principle based on geodesics in the next section.

4.2

Geodesic projection principle

We first define the notion of observable space and observable states on this space to represent accessible information about the internal state space Y . Definition 4.3 (Observable space and observable states). An observable space is a finite (or countable) set of internal states S ⊂ Y . An observable state relative to S is a probability measure ρ̂ ∈ P(S), that is, a probability distribution over the elements of S. Such a state represents the coarse, externally accessible information about an internal distribution on Y . We denote by P2 (S) the space of observable states relative to S when such observables are lifted to P2 (Y ). This represents the idea that an agent has only access to a limited set of internal states. A neural network for instance, cannot “see” synaptic weights that are not tied to an actual synapse. They can only be influenced by accessible weights. When a probability measure in P2 (Y ) (i.e. a synaptic 17

strength profile) attributes weights outside S, those weights remain inaccessible to the agent. A probability measure in P2 (Y ) has to be projected into a probability measure in P(S) (a simplex). Principle 4.4 (Stochastic geodesic projection principle). Let ρ̂, ν̂ ∈ P(S) be two observable states relative to the observable space S, and let (ρt )t∈[0,1] denote the constant-speed Wasserstein geodesic between their internal lifts ρ, ν ∈ P2 (Y ). For each t ∈ [0, 1], we define a random observable state ρ̂t taking values in {ρ̂, ν̂} by P(ρ̂t = ρ̂) = 1 − t,

P(ρ̂t = ν̂) = t.

The stochastic geodesic projection of the local memory state ρt onto the observable space S is then defined as the expectation Πρ̂→ν̂ (ρt ) := E[ρ̂t ]. This is a stochastic projection. It expresses the fact that a mass particle travelling along the internal geodesic from ρ to ν is observed at geodesic parameter t as belonging to the origin state with probability 1 − t and to the target state with probability t. Remark 4.5 (Affine form). Since ρ̂t takes only the two values ρ̂ and ν̂, its expectation is simply the convex combination Πρ̂→ν̂ (ρt ) = (1 − t) ρ̂ + t ν̂. Thus the linear interpolation between observable states arises as the expectation of a binary random observation of the internal geodesic state. This projection acts as a bridge between the internal Wasserstein geodesic structure on P2 (Y ) and the linear geodesic structure of the simplex P(S) endowed with the Euclidean metric. In particular, it is the only admissible map that sends the internal constant-speed geodesic (ρt )t∈[0,1] to a constant-speed geodesic in P(S). Indeed, since geodesics in the Euclidean simplex are exactly affine segments, any geodesicity-preserving projection must coincide with the linear interpolation t 7→ (1 − t)ρ̂ + tν̂. Remark 4.6 (Relativity of the stochastic projection). The stochastic geodesic projection is not a global map P2 (Y ) → P(S). It is defined relatively to a pair of observable states (ρ̂, ν̂), via the Wasserstein geodesic between their internal lifts (ρ, ν). In other words, the projection is contextualized by (ρ̂, ν̂), which can be seen as a global signal (e.g. S(·, ϕ) := ϕ with context ϕ := (ρ̂, ν̂)). Remark 4.7 (Extension to combinatorial graphs). Any finite combinatorial graph G = (V, E) can be canonically thickened into a metric graph Y by replacing each edge (i, j) ∈ E with an interval of length ℓij > 0. This construction endows Y with a geodesic metric, and therefore with a well-defined Wasserstein space (P2 (Y ), W2 ) on which the Tan–HWG update is formulated. The Tan–HWG minimization is then performed over P2 (Y ), and the resulting iterates (ρn )n≥0 naturally live on the whole metric graph, not only on the vertices. In Section 5.1, we will see an illustration where the signal plays the role of a target distribution at each step, stabilized by its frozen nature. This will lead to an update ρn+1 located on the geodesic between the previous state ρn and the target frozen signal distribution ν n . The stochastic geodesic 18

projection then gives a natural representation of observable weights on the combinatorial graph, given by a recursive relation detailed later. Remark 4.8 (Perspectives). Beyond its conceptual interest, the framework developed here may have applications in areas where probability distributions supported on a finite set evolve dynamically. A rapidly growing line of work in machine learning studies continuous-time flows on the categorical simplex for generative modeling, distillation, or accelerated sampling [19, 10, 21, 13]. These approaches typically operate directly in the simplex equipped with a Riemannian or informationgeometric structure, without modeling any latent geometry underlying the discrete states. In contrast, our construction provides a principled way to introduce such a latent geometry: a distribution evolves internally on a metric graph according to a Wasserstein gradient-flow scheme, and the observable categorical distribution arises through stochastic geodesic projection. This viewpoint could be useful in several ways. First, it offers a geometrically structured latent space for categorical flows, potentially improving interpretability or stability. Second, it provides a natural regularisation mechanism when the discrete states possess semantic or hierarchical relations representable as a graph. Finally, it suggests a new class of generative models in which the internal dynamics is Wasserstein-geodesic while the observable dynamics remains linear on the simplex. Investigating these directions lies beyond the scope of the present work but constitutes a promising avenue for future research. The next section will show how, in Tan–HWG minimizing-movement schemes, the signal plays the role of a “target” measure at each step, placing the update on the geodesic between the previous state and the target frozen signal.

5

Quadratic energies and constant contraction update rule

In this section, we consider the case where the external state space is Z = P2 (Y ), so that both local states and the global signal fiber components live in the same Wasserstein space. This setting preserves the full optimal-transport geometry and provides a natural benchmark for the framework. In this regime, we show how the global signal acts as an “attractor” in the sense of Principle 5.3, and that a stable, state-independent contraction update rule emerges (Theorem 5.8). We apply the stochastic geodesic projection principle (Section 5.5) and recover classical update rules as special cases of the framework (Sections 5.6 and 5.7).

5.1

W2 –isotropic energy and global signal as an attractor

We show the attractor role of the global signal in the case of isotropic (or radial) energies. Definition 5.1 (W2 –isotropic Tan–HWG energy). Consider a local energy density depending only on the Wasserstein distance to the global signal fiber component z ∈ P2 (Y ) i.e. a local energy density of the general form Ex (·, z) := Fx W2 (·, z) . 

Assume Ex (·, Sx (ρ, ϕ)) is such that it defines a Tan–HWG energy E with global signal S. We say that such energy is a W2 –isotropic Tan–HWG energy.

19

Remark 5.2 (Signal-driven isotropy). Isotropy is defined relative to the global signal, which serves as the reference point of the geometry. The energy treats all directions pointing toward the signal equivalently, so that the signal behaves as an attractor. This reflects the fact that isotropy is not an intrinsic property of the space, but a property induced by the signal that organises the geometry of the dynamics. When Fx is strictly increasing for µ-a.e. x, decreasing the energy E(ρ, ϕ) forces the Wasserstein distances W2 (ρx , Sx (ρ, ϕ)) to decrease fiberwise, so that local state ρx is geometrically attracted towards the fiber component Sx (ρ, ϕ) of the global signal in the Wasserstein space for µ-a.e. x. This yields the following principle. Principle 5.3 (Geometric attraction in the isotropic case). Assume E is a W2 –isotropic Tan–HWG energy, with global signal S and local energy density E, with Ex (ρx , Sx (ρ, ϕ)) = Fx W2 (ρx , Sx (ρ, ϕ)) , with Fx strictly increasing for µ-a.e. x. Then, along the Tan–HWG minimizing-movement scheme dynamics, for µ-a.e. x, the local memory ρx is attracted toward the global signal fiber component Sx (ρ, ϕ), in the sense that whenever the energy decrease is strict, one has Eϕn ,ρn (ρn+1 ) < Eϕn ,ρn (ρn )

=⇒

n W2 (ρn+1 x , Sϕn ,ρn , x ) < W2 (ρx , Sϕn ,ρn , x )

for µ-a.e. x ∈ X. In Section 5.3, we give an explicit example of such isotropic Tan–HWG energy i.e. satisfying conditions of Definition 2.5.

5.2

W2 –anisotropic internal geometries

If the internal space Y is equipped with another distance dA that is bi-Lipschitz equivalent to the reference metric dY , then the associated Wasserstein distance WA is equivalent to W2 . In this case, an energy that is isotropic with respect to WA can be interpreted as a W2 –anisotropic energy. The anisotropy is therefore not intrinsic: it arises from expressing the Tan–HWG dynamics in a different but equivalent internal geometry. In this sense, the “anisotropic” case is simply the Tan–HWG dynamics written with respect to a bi-Lipschitz deformation of the internal metric. The underlying geometric structure is unchanged; only the coordinate representation of the dynamics differs. Remark 5.4. Anisotropy therefore reflects the choice of internal coordinates rather than a change in the underlying variational structure. This viewpoint is consistent with the interpretation of observable deviations from affine dynamics as signatures of curvature in the internal geometry.

5.3

W2 –quadratic energy

Whenever Y is CAT(0), the Wasserstein space (P2 (Y ), W2 ) is CAT(0) (see [30]), and the map η 7→ W22 (η, z) is geodesically convex for z ∈ P2 (Y ). In particular, for any convex nondecreasing function fx : R+ → R, the functional Ex (·, z) := fx W22 (·, z)



20

is geodesically convex along W2 -geodesics. This is the most natural case among isotropic energies, where we can guarantee such convexity. Proposition 5.5. Assume that (Y, dY ) is a Polish CAT(0) space. Let f : x 7→ fx be a measurable map with, for all x ∈ X, fx : R+ → R ∪ {+∞} proper, lower semicontinuous, and convex nondecreasing, and such that fx (0) < +∞,

for µ-a.e. x and

(x 7→ fx (0)) ∈ L1 (X, µ).

Define the local energy density E by Ex (ρx , z) := fx W22 (ρx , z) ,

z ∈ P2 (Y ).



Then, for any measurable global signal S : Xµ × M(Γ)X → Xµ , the energy E(ρ, ϕ) :=

Z X

Ex ρx , Sx (ρ, ϕ) dµ(x) 

is a locally sequentially stable Hebbian energy, hence a Tan–HWG energy. We call such energy a W2 –quadratic Tan–HWG energy. Proof. E is Hebbian in the sense of Definition 2.1. We prove that E is locally sequentially stable (Definition 2.3): • the local energy density is proper and lower semicontinuous since fx is proper, l.s.c.; • geodesic convexity of the local energy density comes from the composition of the squared Wasserstein distance in the CAT(0) space with fx convex nondecreasing; • the lower bound is ensured by taking c(x, z) = fx (0) and by monotony of x 7→ fx W22 (ρx , z) ; 

• Existence of a finite-energy field: for every reference configuration (ϕref , ρref ), by taking ρ = S(ρref , ϕref ) ∈ Xµ we have Z X

Ex (ρx , Sx (ρref , ϕref )) dµ(x) =

Z X

fx (0) dµ(x) < +∞.

Then, by Proposition 2.4, E is sequentially stable, hence is a Tan–HWG energy. Remark 5.6 (CD(0, ∞)). More generally, when Y is equiped with a reference mesure mY , and (Y, dY , mY ) is a Polish space satisfying the synthetic curvature–dimension condition CD(0, ∞) in the sense of Lott–Sturm–Villani [22], then (P2 (Y ), W2 ) is a geodesic space and the squared Wasserstein distance is geodesically convex along W2 -geodesics. In this case, Proposition 5.5 still holds. To our knowledge, this is the most general case where we can guarantee convexity of the isotropic energy. Other forms of W2 -isotropic Tan–HWG energies may exist: we leave the study of those for future work. Remark 5.7. If each fx is λ-convex and nondecreasing on R+ , with λ ∈ (−∞, 0) independent of x, then on any bounded subset {ρ : W2 (ρ, z) ≤ R} the map ρ 7→ fx (W22 (ρ, z)) is λR –geodesically convex along W2 –geodesics, with λR = 4R2 λ < 0. A global quantitative λ–geodesic convexity

21

estimate, independent of R, would require stronger structural assumptions and is beyond the scope of this work. Nevertheless, in practice, it is enough to have such local λR –geodesic convexity on a bounded subset that is invariant under the JKO scheme. In that case, each minimisation step is well-posed, and the discrete-time dynamics is well-defined.

5.4

State-independent contraction factor

Within the class of W2 –isotropic Tan–HWG energies, the following result shows that the quadratic case is the unique situation in which the update reduces to a universal Wasserstein geodesic interpolation. Theorem 5.8 (State-independent contraction factor). Let Y be a CAT(0) Polish space, and E a W2 –isotropic Tan–HWG energy in the sense of Definition 5.1, with signal S : Xµ × M(Γ)X → Xµ , Z E(ρ, ϕ) =

Fx W2 (ρx , Sx (ρ, ϕ)) dµ(x). 

X

Assume Fx is a continuously differentiable and not constant for µ-a.e. x ∈ X. Let (ϕn )n≥0 be a countable set of contexts. Let τ > 0. Assume the Tan–HWG scheme: n+1

ρ

∈ TτE (ρn | ϕn ) = argmin ρ∈Xµ

Z X

 Fx W2 (ρx , hnx ) dµ(x) +

1 2τ

Z X



W22 (ρx , ρnx ) dµ(x)

,

(*)

where at step n, hn := Sϕn ,ρn = S(ρn , ϕn ) is the frozen global signal. Then the Tan–HWG scheme (*) admits a geodesic update of the form ρn+1 = ρnx,tτ,x , x

W2 (ρnx,tτ,x , ρnx ) = tτ,x W2 (ρnx , hnx ),

for µ-a.e. x ∈ X, where • (ρnx,t )t∈[0,1] is the W2 –geodesic from ρnx to hnx , • and tτ,x ∈ (0, 1) is a state-independent contraction factor i.e. independent of ρnx and hnx , if and only if

αx 2 r + const, with fiberwise constant αx > 0. 2 We say that E is an purely quadratic energy with parameter α = {αx }x∈X . In particular, we have αx τ tτ,x = , 1 + αx τ Fx (r) =

and the purely quadratic Tan–HWG scheme is unique given an initial condition ρ0 ∈ Xµ . Among W2 –isotropic Tan–HWG energies, the purely quadratic case is the unique one that yields a state-independent Tan–HWG contraction factor on each fibre. In other words, only W2 quadratic energies with local energy density affine in W22 yield such fiberwise state-independent contraction factor tτ,x . Proof. Fix n, and set hn := Sϕn ,ρn = S(ρn , ϕn ) ∈ Xµ the frozen signal. By the Hebbian structure 22

and the definition of Eϕn ,ρn , the JKO functional at step n reads J (ρ) =

Z  X

Fx W2 (ρx , hnx ) + 

1 2 W (ρx , ρnx ) dµ(x). 2τ 2 

The integral is a sum of independent fiberwise contributions, so the minimization decouples: for µ-a.e. x, ρn+1 minimizes x Jx (ρx ) :=

 1 2 W2 (ρx , ρnx ) + Fx W2 (ρx , hnx ) 2τ

over P2 (Y ). Assume

ρn+1 = ρnx,tτ,x , x

W2 (ρnx,tτ,x , ρnx ) = tτ,x W2 (ρnx , hnx ),

for µ-a.e. x ∈ X, where (ρnx,t )t∈[0,1] is the W2 –geodesic from ρnx to hnx , and with a coefficient tτ,x ∈ (0, 1) independent of ρnx . Fix such x and denote D := W2 (ρnx , hnx ), n n n On the geodesic, we have W2 (ρn+1 x , ρx ) = W2 (ρx,tτ,x , ρx ) = tτ,x D, and for all t ∈ [0, 1]

W2 (ρnx,t , ρnx ) = t D

and W2 (ρnx,t , hnx ) = (1 − t) D.

minimizes Jx , it minimizes it along the geodesic (ρnx,t )t∈[0,1] i.e. it minimizes Since ρn+1 x f (t) := Jx (ρnx,t ) =

 1 2 n n 1 2 2 W2 (ρx,t , ρx ) + Fx W2 (ρnx,t , hnx ) = t D + Fx (1 − t) D). 2τ 2τ

The minimum is reached at t = tτ,x where ρnx,t = ρn+1 x , so, with Fx continuously differentiable, we have tτ,x 2 0 = f ′ (tτ,x ) = D − DFx′ (1 − tτ,x ) D). τ Since the coefficient tτ,x ∈ (0, 1) is independent of ρnx , this holds true for any D = W2 (ρnx , hnx ). If W2 (ρnx , hnx ) = 0 for all x and all n, then ρnx = hnx for all n and the trajectory is constant. In this trivial case the statement is empty, so we may assume that there exists at least one x and one step n such that D := W2 (ρnx , hnx ) > 0. And, for all D > 0, F ′ (1 − tτ,x ) D) tτ,x = x . τ D This has to hold true for any ρn and hn , therefore for any distance D = W2 (ν, ν ′ ) with ν ̸= ν ′ e := t D, arbitrary, as well as any distribution on the geodesic between ν and ν ′ , hence for any D t

F ′ (1−tτ,x ) D)

x t ∈ [0, 1]. Thus τ,x τ − D linear in D i.e. of the form

= 0 on any segment [0, D] for any D. Thus Fx′ (1 − tτ,x ) D) is Fx′ (r) = αx r,

with αx > 0 since Fx is convex non constant. We can compute αx from tτ,x = αx (1 − tτ,x ) τ yielding αx =

tτ,x . τ (1 − tτ,x ) 23

Then integrating gives Fx (r) =

αx 2 r + const. 2

Conversely, if Fx (r) = α2x r2 + const, with αx > 0, the JKO functional at step n reads J (ρ) =

Z  X

αx 2 1 W2 (ρx , hnx ) + W22 (ρx , ρnx ) dµ(x). 2 2τ 

The integral is a sum of independent fiberwise contributions, so the minimization decouples: for µ-a.e. x, ρn+1 minimizes x 1 2 αx 2 W2 (ρx , ρnx ) + W2 (ρx , hnx ) 2τ 2

Jx (ρx ) :=

over P2 (Y ). Fix x and denote D := W2 (ρnx , hnx ). Let (ρnx,t )t∈[0,1] be the W2 –geodesic from ρnx to hnx . Along this geodesic we have W2 (ρnx,t , ρnx ) = tD,

W2 (ρnx,t , hnx ) = (1 − t)D,

so that

1 2 2 αx t D + (1 − t)2 D2 2τ 2 is a strictly convex quadratic polynomial in t. Its unique minimizer is f (t) := Jx (ρnx,t ) =

tτ,x = and the corresponding value is

αx τ , 1 + αx τ

min f (t) = Jx (ρnx,tτ,x ).

t∈[0,1]

We now show that this minimizer on the geodesic is actually a minimizer on all P2 (Y ). For any ρx with r := W2 (ρx , hnx ), the triangle inequality gives W2 (ρx , ρnx ) ≥ |D − r|, hence

1 αx 2 (D − r)2 + r =: ΨD (r). 2τ 2 A direct computation shows that ΨD is strictly convex in r and attains its unique minimum at r∗ = (1 − tτ,x )D, with ΨD (r∗ ) = Jx (ρnx,tτ,x ). Jx (ρx ) ≥

Therefore, for any ρx , Hence,

Jx (ρx ) ≥ ΨD (r) ≥ ΨD (r∗ ) = Jx (ρnx,tτ,x ). min

ρx ∈P2 (Y )

so

Jx (ρx ) ≥ Jx (ρnx,tτ,x ),

ρnx,tτ,x ∈ argmin Jx (ρx ). ρx ∈P2 (Y )

Thus, the Tan–HWG scheme admits a solution of the form ρn+1 = ρnx,tτ,x , x

W2 (ρnx,tτ,x , ρnx ) = tτ,x W2 (ρnx , hnx ), 24

for µ-a.e. x ∈ X, where (ρnx,t )t∈[0,1] is the W2 –geodesic from ρnx to hnx , and with a coefficient tτ,x =

αx τ ∈ (0, 1) independent of ρnx . 1 + αx τ

This proves the converse. Given an initial condition ρ0 ∈ Xµ , the scheme is uniquely defined, since for each x the functional Jx is strictly convex on P2 (Y ) when Y is CAT(0), hence admits a unique minimizer. This completes the proof. Remark 5.9 (Geometric intuition). Quadratic energies are the only ones whose Wasserstein gradient flow naturally follows geodesics. Indeed, for a functional of the form ρ 7→ 12 W22 (ρ, z), the Wasserstein gradient is exactly the initial velocity of the geodesic from ρ to z. As a consequence, the JKO update balances two geodesic velocities—one pointing toward the signal and one toward the previous state—and the minimizer lies at a fixed proportion along the geodesic. For any nonquadratic energy, the gradient is no longer aligned with the geodesic direction, and the update cannot reduce to a universal geodesic interpolation. Notation. For each fiber x, the parameter αx is a constant depending only on x. We may write α : x 7→ αx . Similarly, the contraction factor tτ,x depends only on x and on the timestep τ , and we may write tτ (x). Remark 5.10 (Clustering and reduction to a constant parameter). The Hebbian structure ensures that the scheme decouples fiberwise. We may therefore group together fibers that share the same local energy density. Such a group, or clusters, is a subset X ′ ⊂ X on which all Fx coincide. On each cluster, the parameter αx is constant, say α > 0, and so is the contraction factor tτ,x , which we denote simply by tτ : ατ . tτ = 1 + ατ In this case, the update can be written compactly at the field level as ρn+1 = ρntτ . Unless otherwise stated, we will work within a single cluster, which allows us to simplify notation without loss of generality. Table 1 summarizes what we have seen so far in the present section, with the following inclusions: {Purely quadratic energies} ⊂ {W2 –quadratic energies} ⊂ {W2 –isotropic energies}. Purely quadratic energies form the rigid core of the isotropic class: they are the only ones whose gradient flow follows Wasserstein geodesics with a state-independent speed. Quadratic energies in W22 retain convexity but lose this rigidity, while general isotropic energies preserve only the attractor structure.

25

Energy class

Local density form

Main property

W2 –isotropic Tan–HWG W2 –quadratic Purely quadratic

Fx ◦ W2 fx ◦ W22 α 2 2 W2 + const

S is an attractor Tan–HWG energy ρn+1 = ρntτ & uniqueness

Table 1: Summary of main properties by isotropic energy class.

5.5

Stochastic Tan–HWG geodesic projector

In this section, we apply the stochastic geodesic projection principle seen in Section 4.2 along the trajectory of a purely quadratic Tan–HWG scheme. Let Y be a CAT(0) Polish space. Consider a finite neural system with M internal degrees of freedom on Y . We denote by YM ⊂ Y the family of internal states supporting the local memory states ρx of the system. Consider a family (ρn )n∈N following a purely quadratic Tan–HWG scheme, with parameter α, time step τ > 0, signal S and associated frozen signal family (hn )n∈N ∈ XµN (for a countable set of contexts (ϕn )n≥0 ). Assume that h0x ∈ P2 (YM ),

and

ρ0x = ρx ∈ P2 (YM ),

i.e. the initial memory states and initial global signal are observable states relative to the observable space YM in the sense of Definition 4.3. Assume that for all (η, ψ) ∈ Xµ × M(Γ)X supp(Sx (η, ψ)) ⊂ YM , i.e. the signal has values in the space of observable states P2 (YM ), hence hnx ∈ P2 (YM ) for all n. Intuition. Theorem 5.8 yields that the update ρn+1 is on the constant-speed geodesic between ρnx x n and hx , at parameter tτ (we reduce to the study of clusters in the sense of Remark 5.10). We apply stochastic geodesic projection principle to ρn+1 lying on the geodesic between ρnx and hnx . Then, x by recursively applying such stochastic geodesic projections, we obtain a projected observable state relative to YM . n ) More precisely: for all x ∈ X, we denote by (Sx,τ n∈N the family of finite subsets of Y defined recursively by n+1 n 0 Sx,τ := supp(ρn+1 and Sx,τ = YM . x ) ∪ Sx,τ ,

In particular,

n supp(hnx ) ⊂ Sx,τ ,

by stability of support under the signal map S. n and let π n = (π n ) Fix x ∈ X and n ∈ N. Let {y1 , . . . , yNn } := Sx,τ x kl 1≤k,l≤Nn be an optimal transport n n n > 0, we denote by plan between ρx and hx . For each pair (yk , yl ) with πkl

γkl (t)

(t ∈ [0, 1])

the W2 –geodesic interpolation between yk and yl .

26

Step 1. Couples generating a point.

For any point n+1 n y ∈ Sx,τ \ Sx,τ ,

we define the set of generating couples, i.e. the pairs (k, l) whose geodesic interpolation at time tτ produces the new point y,  n (y) := (k, l) γkl (tτ ) = y , ζx,τ as well as the sets of such couples originating from, and arriving to, a given i: Ti (y) := l γil (tτ ) = y ,

Oi (y) := k γki (tτ ) = y .





n+1 \ S n , and for all couple (k, l) ∈ J1, N K2 : Lemma 5.11. For any point y ∈ Sx,τ n x,τ n (k, l) ∈ ζx,τ (y) =⇒ k ̸= l,

And we have

[ n+1 n y∈Sx,τ \Sx,τ

n ζx,τ (y) = {(k, l) ∈ J1, Nn K2 | k ̸= l}.

n . n , contradicting y ∈ / Sx,τ Proof. Assume the first claim is not true, then y = γkk (tτ ) = yk ∈ Sx,τ n n+1 n Now fix (k, l) ∈ y∈Sx,τ n+1 n ζx,τ (y), then there exists y ∈ Sx,τ \ Sx,τ such that y = γkl (tτ ) and by \Sx,τ the first claim, k ̸= l.

S

Conversely, for any k ̸= l in J1, Nn K, taking y = γkl (tτ ) suffices to prove the converse inclusion. n+1 ). We define a random measure ν̂ Step 2. Stochastic backward projector. Let ν ∈ P2 (Sx,τ n n+1 with values in P2 (Sx,τ ), via a random partition of Sx,τ :

n+1 Sx,τ =

N Gn

Yi ∩ Yj = ∅ if i ̸= j,

Yi ,

i=1 n+1 : such that, ν̂({yi }) = ν(Yi ), and for all y ∈ Sx,τ n , the mass ν({y }) stays at y i.e. • if y = yi ∈ Sx,τ i i

P (yi ∈ Yi ) = 1; n+1 \ S n , then: • if y ∈ Sx,τ x,τ





n+1 n P y ∈ Yi | y ∈ Sx,τ \ Sx,τ = (1 − tτ )

n X πiln πki + t . τ ρn+1 ({y}) ρn+1 ({y}) l∈T (y) x k∈O (y) x

X i

i

n+1 n n+1 The second item is well defined since ρn+1 x ({y}) ̸= 0 with y ∈ Sx,τ \ Sx,τ ⊂ supp(ρx ), and it corresponds to the following draws:

27

n (y) with probability 1. we choose a pair (k, l) ∈ ζx,τ

P (k, l) is chosen | y = 

n πkl , ρn+1 x ({y})

i.e. the probability of selecting (k, l) is the portion of mass transported along the geodesic γkl that contributes to y; 2. and we send the mass ν({y}) to yk with probability (1 − tτ ),

yl with probability tτ ,

which corresponds to the stochastic geodesic principle. We call the map Πnx,τ : ν 7→ ν̂, the stochastic backward projector. The deterministic expectation operator

Step 3. Expectation operator.

n+1 n Enx,τ : P2 (Sx,τ ) −→ P2 (Sx,τ )

is defined by

Enx,τ (ν) := E[ν̂].

n , one has Lemma 5.12. For every yi ∈ Sx,τ

Enx,τ (ν)({yi }) = ν({yi }) +

X

ν({y})

h

(1 − tτ )1{yk =yi } + tτ 1{yl =yi }

X n (y) (k,l)∈ζx,τ

n+1 n \Sx,τ y∈Sx,τ

i

n πkl . ρn+1 x ({y})

n+1 ) to P (S n ). For all t ∈ [0, 1], and all In particular Enx,τ is a linear operator from P2 (Sx,τ 2 x,τ ′ n+1 n+1 (ν, ν ) ∈ P2 (Sx,τ ) × P2 (Sx,τ ),

Enx,τ ((1 − t)ν + tν ′ ) = (1 − t)Enx,τ (ν) + tEnx,τ (ν ′ ). Moreover, for all n, Enx,τ acts as the identity on P2 (YM ), and Enx,τ (hkx ) = hkx , and

for all k,

n n Enx,τ (ρn+1 x ) = (1 − tτ ) ρx + tτ hx .

Proof. The first claim is a direct calculation of the expected value of ν̂, and the second is a direct consequence. n , we have η({y}) = 0 for y ∈ n , hence En (η) = η For all η ∈ P2 (YM ), η has support in YM ⊂ Sx,τ / Sx,τ x,τ k by direct application of the first claim. In particular, since hx ∈ P2 (YM ) for all k, Enx,τ (hkx ) = hkx . n , For every yi ∈ Sx,τ n+1 Enx,τ (ρn+1 x )({yi }) = ρx ({yi }) +

X n+1 n Sx,τ \Sx,τ

= ρn+1 x ({yi }) +

i X h ρn+1 x ({y}) n (1 − t )1 + t 1 τ {yk =yi } τ {yl =yi } πkl ρn+1 ({y}) x ζ n (y) x,τ

X

X

n (y) n+1 n (k,l)∈ζx,τ y∈Sx,τ \Sx,τ

28

h

i

n (1 − tτ )1{yk =yi } + tτ 1{yl =yi } πkl .

Since Lemma 5.11 yields [ n+1 n y∈Sx,τ \Sx,τ

n ζx,τ (y) = {(k, l) ∈ J1, Nn K2 | k ̸= l},

the previous equality reduces to n+1 Enx,τ (ρn+1 x )({yi }) = ρx ({yi }) +

Xh

i

n (1 − tτ )1{yk =yi } + tτ 1{yl =yi } πkl

k̸=l

= ρn+1 x ({yi }) +

X

(1 − tτ ) πiln +

l̸=i

X

n tτ πki .

k̸=i

By construction of πxn and γ(·,·) , and by Theorem 5.8: ρnx =

Nn Nn X X

hnx =

n πkl δ yk ,

Nn Nn X X

ρn+1 = x

n πkl δyl ,

n πkl δγkl (tτ ) .

k=1 l=1

l=1 k=1

k=1 l=1

Nn Nn X X

In particular, evaluating at yi yields ρnx ({yi }) =

Nn X

πiln ,

hnx ({yi }) =

l=1

Nn X

n πki ,

n ρn+1 x ({yi }) = πii ,

k=1

since δγkl (tτ ) (yi ) = 0 whenever k ̸= i and l ̸= i, because γkl (tτ ) cannot coincide with yi unless k = l = i. Thus (1 − tτ ) ρnx ({yi }) + tτ hnx ({yi }) = (1 − tτ )

Nn X

πiln + tτ

Nn X

n πki

l=1

= πiin +

k=1 X n n (1 − tτ ) πil + tτ πki . l̸=i k̸=i

X

Comparing with the previous evaluation of Enx,τ (ρn+1 x )({yi }) yields the final claim. Note that the restriction of each expectation operator Enx,τ on P2 (YM ) is the identity. This technical construction ensures the well-posedness of the following definition. Definition 5.13 (Stochastic Tan–HWG geodesic projector). Consider the setting described n ) −→ P (Y ) defined for all (x, n) ∈ X × N by above. The projector Projnx,τ : P2 (Sx,τ 2 M n n Projn+1 x,τ := Projx,τ ◦ Ex,τ ,

Proj0x,τ = IdP2 (YM ) ,

is called a stochastic Tan–HWG geodesic projector for the purely quadratic Tan–HWG scheme. We also denote Projτ := ({Projnx,τ }x∈X )n∈N . Remark 5.14 (Relativity of projector and non uniqueness). A stochastic Tan–HWG geodesic pron ) → P (Y ) defined subsequently to a purely quadratic Tan–HWG jector is a family of maps P2 (Sx,τ 2 M scheme.

29

The construction above show that such projector always exists for a purely quadratic Tan–HWG scheme. The stochastic Tan–HWG projector is unique exactly when both the optimal transport plan πxn and all induced geodesics γkl are unique. This is automatic in Euclidean spaces with quadratic cost when ρnx is absolutely continuous (Brenier), but not in general CAT(0) spaces. For the purpose of this paper, we do not require uniqueness of the projector. Since uniqueness is not guaranteed, we may write p ∈ Projτ to denote the choice of such a projector. We have and

0 1 n Projn+1 x,τ = Ex,τ ◦ Ex,τ · · · ◦ Ex,τ , n+1 n n n Projn+1 x,τ (ρx ) = (1 − tτ )Projx,τ (ρx ) + tτ hx ,

n . This is the key identity showing since Projnx,τ (hnx ) = hnx for hnx ∈ P2 (YM ) with support in YM ⊂ Sx,τ that the projected dynamics follows an exponential moving average of the signal (see Section 5.6).

5.6

Closed-form EMA observable dynamics

The projector mechanism of the previous section induces a closed-form observable dynamics. Theorem 5.15 (Observable dynamics induced by stochastic geodesic projection). Let Y be a CAT(0) Polish space, and YM ⊂ Y the support of a finite neural system with M internal degrees of freedom on Y . Consider a purely quadratic Tan–HWG energy E with parameter α, and signal S : Xµ × M(Γ)X → Xµ . Let τ > 0. Let (ϕnτ )n≥0 be a countable set of contexts. Assume Sx (Xµ , ϕ(t)) ⊂ P2 (YM ) for all t ≥ 0, and consider an initial state ρ0 such that for all x ∈ X, ρ0x ∈ P2 (YM ) (observability of signal and initial state). Let (ρnτ )n∈N follow the τ -Tan–HWG scheme relative to E with initial condition ρ0 , under context ϕ: ρn+1 ∈ TτE (ρnτ | ϕnτ ), n ∈ N. τ We denote (hnτ )n∈N the associated frozen signal family, and Projτ = ({Projnx,τ }x∈X )n∈N an associated stochastic Tan–HWG geodesic projector. The observable sequence (ρ̂nτ )n∈N defined by ρ̂nτ,x := Projnx,τ (ρnτ ) satisfies the affine recurrence ρ̂n+1 = (1 − tτ ) ρ̂nτ + tτ hnτ . τ

(2)

Consequently, the observable sequence admits the explicit closed-form expression ρ̂nτ = (1 − tτ )n ρ̂0 + tτ

n−1 X

(1 − tτ )n−1−k hkτ .

(3)

k=0

Assume ρnτ and ϕnτ admit limit curves ρ : t → ρ(t) and ϕ : t → ϕ(t) when τ → 0 with nτ → t. In the continuous-time limit τ → 0 with nτ → t, the observable converges to the solution of

30

the ODE

d ρ̂(t) = −α ρ̂(t) + α h(t), dt

namely ρ̂(t) = e−αt ρ̂(0) + α

Z t 0

e−α(t−s) h(s) ds,

(4) (5)

where h(t) = S(ρ(t), ϕ(t)). Proof. Applying recursively Lemma 5.12 yields (2). Iterating the recurrence gives (3). The continuous-time limit follows from the standard convergence of the discrete exponential filter to the ODE (4), whose solution is (5). Thus, in the purely quadratic case, the observable dynamics is an exponential filtering of the signal. In the general setting, this result holds by homogeneous cluster with corresponding parameters α(x) and contraction factors tτ (x) (see Remark 5.10). Remark 5.16 (Bridge OT-emerging stochastic process-exponential filtering). The observable dynamics (2)–(4) can thus be interpreted as the macroscopic effect of a particle-wise stochastic geodesic projection along the optimal transport plans generated by the Tan–HWG update. At the observable level, this induces an exponential moving average (EMA) of the past signal values, with decay factor (1 − tτ ). Conversely, any EMA dynamics on a simplex arises as the observable dynamics of the stochastic Tan–HWG geodesic projection applied to a purely deterministic Tan–HWG dynamics (convexity of the purely quadratic energy yields uniqueness of the minimizer at each update1 ). Since exponential moving averages are ubiquitous in signal processing, machine learning, and computational neuroscience, this geometric derivation reveals a surprisingly universal mechanism underlying a widely used update rule. To our knowledge, this is the first framework in which an OT-induced stochastic process yields an exact exponential moving average at the observable level. Remark 5.17 (On the structural role of the contraction factor). A notable feature of the Tan–HWG update is that the contraction factor ατ tτ = 1 + ατ is not an approximation of a continuous-time dynamics, but the exact minimizer of the discrete variational problem associated with the quadratic Wasserstein energy E(ρx , hx ) =

α 2 W (ρx , hx ). 2 2

As a consequence, the observable update ρ̂n+1 = (1 − tτ )ρ̂nx + tτ hnx x 1

This determinism is specific to the purely quadratic case: the strict convexity of the quadratic energy ensures uniqueness of the minimizer at each JKO step. For general W2 –isotropic energies, only weak or local convexity is available, and the JKO update may admit multiple minimizers, preventing a deterministic underlying dynamics.

31

is exactly linear in ρ̂x , with a contraction factor that is constant and independent of the state. This linearity is the direct reflection of the state-independent geodesic contraction established in Theorem 5.8. Such a fixed-ratio interpolation along Wasserstein geodesics can only arise from a purely quadratic isotropic energy, as shown in Theorem 5.8. In the continuous-time limit, this yields the exponential filtering equation d ρ̃x = −α(ρ̃x − hx ), dt which therefore appears not as an external modeling assumption, but as the unique macroscopic law compatible with the geometric structure of the purely quadratic Tan–HWG dynamics. Moreover, the exact expression of tτ contrasts with the usual Euler-type approximations, where one replaces tτ by ατ . Such approximations may lose stability (e.g. when ατ > 1), whereas the Tan–HWG factor always satisfies 0 < tτ < 1 for all τ > 0, ensuring unconditional stability of the discrete dynamics. The Tan–HWG scheme is therefore not a numerical surrogate of a continuous equation, but a geometrically exact discrete mechanism, from which the continuous-time dynamics emerges only as a limit. Taken together, these observations provide an indirect but compelling confirmation of the relevance of the Wasserstein geometric framework: the most natural energy in this geometry yields the most natural observable plasticity law, and the discrete update is exact, stable, and structurally constrained by the geometry rather than by numerical choices. Remark 5.18 (Markovian observable dynamics). Interestingly, the affine recurrence ρ̂n+1 = (1 − τ tτ ) ρ̂nτ + tτ hnτ shows that, in the purely quadratic case, the observable dynamics is Markovian. This property is tied to the quadratic Wasserstein geometry and its step-independent contraction factor. For general isotropic energies, the observable update may depend on the full internal structure. Remark 5.19 (Memory field interaction). This theorem allows for an interesting interaction mechanism: if we consider another memory field ρ′ following its own dynamics, we may take h = ρ′ ∈ Xµ as the signal for ρx . In this case, the observable expectation satisfies d E[ρ̂x (t)] = −α E[ρ̂x (t)] + α E[ρ̂′x (t)]. dt Thus, ρx tends to align with ρ′x , regardless of the specific dynamics followed by ρ′x . If, moreover, ρ′x follows the same purely quadratic Tan–HWG dynamics, the system becomes symmetric and we obtain   d E[ρ̂x (t)] − E[ρ̂′x (t)] = −2α E[ρ̂x (t)] − E[ρ̂′x (t)] , dt

that is, E[ρ̂x (t)] − E[ρ̂′x (t)] = e−2α t E[ρ̂x (0)] − E[ρ̂′x (0)] . 



The two observable fields therefore follow a consensus-type dynamics, exponentially aligning their distributions. Such alignment mechanisms are reminiscent of classical models of multi-agent consensus and synchronization [9, 25, 26, 15, 28].

32

In Section 8, we will see that, under additional regularity assumptions on the global signal S, and on the contexts, in the continuous-time limit τ → 0 with nτ → t, the internal Tan–HWG scheme admits a limiting curve (Theorem 8.7), which can be interpreted as a perturbed Wasserstein gradient flow due to the frozen signal at each step. This will complete the previous theorem.

5.7

Mirror descent on the simplex as a special case of Tan–HWG dynamics

As an application example, we now show that mirror descent on the simplex is a special case of Tan–HWG observable dynamics. Proposition 5.20 (Mirror descent as a special case of Tan–HWG observable dynamics). Consider a neural network with input layer YM = {y1 , . . . yM }, output layer XN = {x1 , . . . xN }, and weights w = {wi,j }(i,j)∈J1,M K×J1,N K such that, for all i, j: M X

wk,j = 1,

and

wi,j ≥ 0.

k=1

Consider a differentiable loss function Loss : RM ×N → R. Assume that there exists Cmax > 0 such that ∂Loss < Cmax , ∂wi,j

for all (i, j) ∈ J1, M K × J1, N K.

Then, the mirror descent update scheme on w relative to Loss n+1 w̃i,j

n+1 wi,j := P

∂Loss(wn ) n+1 n with w̃i,j = wi,j exp −η , ∂wi,j !

n+1 , k w̃k,j

is equivalent to a purely quadratic Tan–HWG observable scheme on the observable state w relative to YM . Proof. Step 1: identification of weight matrices with memory fields. For a vector wj = (wk,j )k=1...M on the M − 1 dimensional simplex ∆M −1 , we identify such vector with the probability distribution ρ(wj ) :=

M X

wk,j δyk ∈ P2 (YM ).

k=1

This is the explicit equivalence ∆M −1 ≃ P2 (YM ). Then we identify w = {wi,j }(i,j)∈J1,M K×J1,N K with the memory field {ρ(wj )}j=1...N (i.e. (∆M −1 )N ≃ Xµ with X = XN finite, hence X = Xµ for finite reference measure µ). Step 2: signal on the simplex. Fix τ > 0 and set ρ0 = w0 . For a given n ∈ N, define hn :=

 1  n+1 w − (1 − tτ )wn , tτ

33

tτ ∈ (0, 1).

Since wjn and wjn+1 lie in the simplex, we immediately have M X i=1

hni,j =

M X

1 tτ

n+1 wi,j − (1 − tτ )

i=1

M X

! n wi,j

= 1.

i=1

To ensure hni,j ≥ 0, we control the variation of wn under mirror descent. Using the update n exp(−η ∂ n wi,j wi,j Loss(w )) n+1 wi,j = P n , n k wk,j exp(−η ∂wk,j Loss(w ))

and the bound |∂wi,j Loss| ≤ Cmax , we obtain n+1 n wi,j ≥ wi,j exp(−2ηCmax ).

Hence

n+1 n n   wi,j − (1 − tτ )wi,j wi,j ≥ exp(−2ηCmax ) − 1 + tτ . tτ tτ n Therefore hi,j ≥ 0 for all i, j provided

hni,j =

tτ > 1 − exp(−2ηCmax ). Under this condition, each hnj = (hni,j )i=1...M lies in the simplex ∆M −1 ≃ P2 (YM ). 2 Step 3: observable dynamics. Hence, Theorem 5.15 applies for the purely quadratic Tan–HWG observable scheme with affine recurrence ρ̂n+1 = (1 − tτ )ρ̂nj + tτ hnj , j tτ > 0. Thus with parameter α := τ (1−t τ)

− (1 − tτ )ρ̂nj = wjn+1 − (1 − tτ )wjn , ρ̂n+1 j and with initial condition ρ̂0 = ρ0 = w0 , this yields ρ̂n = wn

for all n ∈ N.

This proves the proposition. Note that, in this setting, the context at time t can be defined as follows. For each discrete step n, define the empirical distribution on YM φnj :=

M X

n wi,j δyi ,

(j = 1 . . . N ).

i=1

We extend this piecewise-constantly in time by setting φj (t) := φnj

for t ∈ [nτ, (n + 1)τ ).

The context ϕ(t) = {ϕj (t)}j=1...N ∈ M(YM × YM )XN is then defined by ϕj (t) := φj (t + τ ) ⊗ φj (t) ∈ M(YM × YM ), 2

For η < 1/(2Cmax ), we can even take tτ ∈ (2η Cmax , 1).

34

which encodes the pair of successive observable states on the input layer. With this choice, the Tan–HWG signal takes the explicit form S(ρ, ϕ(t)) := so that the observable update

1 (φ(t + τ ) − (1 − tτ )φ(t)) , tτ

tτ ∈ (0, 1),

ρ̂n+1 = (1 − tτ )ρ̂n + tτ hn

coincides exactly with the mirror-descent update on the simplex. Remark 5.21 (The role of the signal as a weight target). In the Tan–HWG observable dynamics, the signal hn acts as an instantaneous target towards which the observable state is attracted through an exponential moving average. The mirror descent update shows that the gradient of the loss function plays exactly the same role: it determines the next iterate wn+1 , and therefore determines the target hn through the relation wn+1 = (1 − tτ )wn + tτ hn . Thus, in this interpretation, gradient descent does not appear as a learning rule per se, but rather as a mechanism that generates a sequence of targets (hn )n for the Tan–HWG dynamics. Remark 5.22 (Generality of the Tan–HWG framework). Although in this proposition the signal hn is obtained from a mirror descent step, the purpose of the construction is not to suggest that one must compute a gradient in order to define a valid signal. Instead, the result shows that any mirror descent scheme on the simplex can be realized as a Tan–HWG observable dynamics. In particular, the Tan–HWG framework strictly contains gradient-based updates as a special case, while allowing for much more general classes of signals, including local, activation-dependent, or non-variational mechanisms. The identification above highlights that Tan–HWG dynamics provides a richer modeling space than classical gradient descent. The signal hn may be generated by biological or local plasticity rules, by external feedback, or by dynamics that do not derive from a global loss. Moreover, the Tan–HWG formalism naturally accommodates observable states represented by general probability measures rather than Dirac masses, allowing for distributed or uncertain synaptic amplitudes. In this sense, gradient descent corresponds to a highly constrained choice of signal within a broader and more flexible geometric framework. In the next section, we give examples of how our framework can extend classic neural network models.

6

Generalized neural networks

In this section, we introduce a neural network architecture that extends classic feedforward neural networks. It illustrates how classic artificial neural networks can be seen as special cases of our framework, while introducing new features, in particular: • a distinction between structural weights (attention weights) and embedding weights (semantic or informational weights) (Section 6.2.1), which emerges from the framework applied at the synaptic connection level, recovering a form of attention mechanism; 35

• a synaptic embedding memory as a distribution, which naturally encodes multi-semanticity and stochasticity, combined with a time-dependent contextualisation (Remark 6.7); • a spectral memory, derived from the synaptic embedding memory, which allows phase synchronization and phase alignment dynamics, reminiscent of Kuramoto interactions [17] and Hebbian population alignments (Section 6.3); • a multi-timescale mechanism ensuring controlled geodesic convexity of the energy, yielding a geometric interpretation of neuronal assemblies and their functional role in consolidation dynamics (Section 6.2.5); • a selection dynamics, with sparsity and structural pruning dynamics (Section 6.3.3); • a geometrically stable and robust dynamics, living in the Wasserstein space. In the following, the finite set YM is implicitly lifted to a CAT(0) Polish space as previously seen in Section 4.

6.1

General setting and external mapping

So far, the framework models local memory states as positive, normalized weights. This is natural for probabilistic representations, but classical machine-learning architectures allow negative parameters, and biological neural systems critically rely on inhibitory synapses and phase synchronisation to regulate activity and shape their computational repertoire. Capturing such signed or oscillatory effects requires enriching the geometric structure of the model. This motivates the introduction of geometric representation maps, of the form G : Xµ → Z X , which lift probability-valued memory fields into the product external state space Z X , where inhibition, interference, and linear or oscillatory interactions can be represented. The cases Z = L2 (Y, K), with K = R or C are particularly of interest since they endow Z with an inner product. Product of internal state spaces. Y =

Y

Y k,

The countable product of internal state spaces Y k admissible internal state space,

k

endowed with distance !1/p

dp (y, y ′ ) =

X

dY k (y k , y ′k )p

p ∈ N∗

,

k

is an admissible internal state space (i.e. a geodesic Polish space). However, only the case p = 2 yields a CAT(0) space (provided each Y k is CAT(0)). Hence in the following, we restrict to this case ! 1/2

dY (y, y ′ ) =

X

dY k (y k , y ′k )2

.

k

In this paper, we will consider example cases where Y = YM × R and Y = YM × C, where YM is implicitly lifted to a CAT(0) Polish space as previously seen in Section 4. 36

External observables. We now give examples of generic forms of geometric representation maps, interacting naturally with the underlying Wasserstein geometry. Definition 6.1 (External observable map). We say that the geometric representation map G : Xµ → Z X is an external observable map if it decomposes fiberwise as G(ρ) = {Gx (ρx )}x∈X with Z Gx (ρx ) = Vx (y) dρx (y), ρx ∈ P2 (Y ), Vx : Y → Z; Y

with x 7→ Vx and x 7→ Gx assumed µ-measurable. The image of a memory state by such a representation map is called an external observable. Remark 6.2 (Affine projections vs geodesically affine projections). External observables of the form Gx (ρx ) =

Z Y

Vx (y) dρx (y)

are affine in ρx , but they are not, in general, geodesically affine in the Wasserstein space. This distinction is crucial: geodesic affinity (or at least controlled geodesic deviation) is key to build geodesically λ–convex energies by composing Gx with λ–convex Euclidean nonlinearities such as standard activation functions. When Vx is λ-convex, the map Gx is geodesically λ-convex (it has a McCann potential energy form [24]). However, λ–convexity alone does not guarantee that the composition with λ-convex Euclidean functions remains geodesically λ–convex. For this, one typically needs the stronger property that Gx be geodesically affine (or effectively affine along the relevant geodesics), in addition to being Lipschitz - a condition that generally fails. This issue becomes particularly acute for contextualized geometric maps introduced later (Section 7), such as Z u(ρ) = z ψ(y ′ ) dρ(y ′ , z), which are never geodesically affine when transport occurs in the y ′ component. As shown in Section 6.2.5, this obstruction can be overcome through a natural multi–timescale mechanism that restores effective geodesic affinity by separating fast and slow variables.

6.2 6.2.1

Generalized Neural Networks with real value weights Internal synaptic space and attention weights

Let YM = {y1 , . . . , yM } denote the finite input set, and XN = {x1 , . . . xN } the finite output set. We define the internal synaptic space as the product Y := YM × R, endowed with the reference measure mY :=

M 1 X δy ⊗ ν, M i=1 i

where ν is a finite positive measure on R. 37

Each output neuron xj is represented by a probability measure ρxj ∈ P2 (Y ),

ρxj =

M X

pi,j δyi ⊗ µi,j ,

i=1

where: •

i=1 pi,j δyi is a probability distribution on YM , i.e. pi,j ≥ 0 and

PM

i=1 pi,j = 1,

PM

• µi,j ∈ P2 (R) is a synaptic amplitude distribution associated with the synapse (yi , xj ). Definition 6.3 (Structural weights and synaptic embedding memory). The coefficients pi,j , which describe the relative influence of the synapse from yi to xj within the network architecture, are called the structural weights or attention weights. The probability distribution µi,j , encoding the internal synaptic state in the Wasserstein space, is called the synaptic embedding memory. Remark 6.4 (Interpretation). Each synapse (i, j) stores two complementary forms of information: • structural information, given by pi,j , which specifies the relative strength of the connection in the network topology; • embedding information, encoded in the synaptic embedding memory µi,j , which represents the internal synaptic state, including its preferred amplitude, variability, and multi-semantic structure. The synaptic embedding memory µi,j is a latent geometric representation of the synapse, from which observable weights arise as projections (see next section). This representation generalizes the classical setting, recovered when each µi,j is a Dirac mass. 6.2.2

Synaptic amplitude samples and observable synaptic weights

For each synapse (i, j), we introduce the synaptic amplitude sample in L0 (Ω, R) as the random variable following µi,j : Zi,j ∼ µi,j . The observable synaptic weight is the random variable defined by the product of the synaptic amplitude sample with the structural weight of the synapse Wi,j := Zi,j pi,j ∈ L0 (Ω, R), defining the observable weight Wj :=

M X

M X

Wi,j 1yi =

i=1

pi,j Zi,j 1yi .

i=1

The external observable weight is given by the external observable wj (ρxj ) :=

Z YM ×R

z 1y′ dρxj (y ′ , z) = 

M X

Z

pi,j R

i=1

38

z dµi,j 1yi =

M X i=1

pi,j E[Zi,j ] 1yi ∈ L2 (YM ),

which corresponds to the expectation of the observable weight (relative to the synaptic embedding memory). We denote the weighted first moment of Zi,j by wi,j (ρxj ) := pi,j E[Zi,j ]. Remark 6.5 (Excitatory/inhibitory representation). When the amplitude Zi,j is restricted to Zi,j ∈ P PM {−1, +1}, we have M i=1 |Wi,j | = i=1 pi,j = 1: Zi,j gives the direction (the sign), and pi,j gives the relative weight. It generalizes the normalized weight representation {pi,j }i=1...M by allowing positive and negative weights, representing excitatory or inhibitory synapses. More generally, with the equivalence R ≃ R+ × {−1, +1}, the information stored in the synapse memory can be richer, encoding amplification/reduction on top of the nature of the synapse (excitatory/inhibitory). Using R also conveniently gives a Euclidean internal space, which is a geodesic Polish space. Remark 6.6 (Classical neural networks as a special case). If each synaptic distribution is a Dirac mass, µi,j = δzi,j , then the weight wi,j := Wi,j = zi,j pi,j ∈ R, is deterministic, and the representation reduces exactly to a standard neural network with deterministic weights. The stochastic formulation is biologically coherent: synaptic transmission is inherently stochastic, with probabilistic vesicle release and variable postsynaptic response amplitudes. 6.2.3

Forward computation with random synaptic transmission

Given an input activation state ψ :=

M X

ψi 1yi ∈ L2 (YM , R) ≃ RM ,

i=1

the synaptic transmission from yi to xj uses the random weight Wi,j . The postsynaptic input to neuron xj is therefore the random variable Psj (ψ | ρxj ) :=

M X

ψi Wi,j ∈ L0 (Ω, R).

i=1

This quantity corresponds to the pre-activation state of neuron xj . We can define the global postsynaptic input state by Ps(ψ | ρ) :=

N X

Psj (ψ | ρxj ) 1xj = ψ W,

j=1

where the product is understood as matrix-vector multiplication in RM ×N , with W :=

N X

Wj 1xj =

j=1

X i,j

39

Wi,j 1(yi ,xj ) .

We consider the expectation h

i

psj (ρxj , ψ) := E Psj (ψ | ρxj ) =

M X

ψi E [Wi,j ] ,

i=1

which corresponds to the deterministic external observable pre-activation state of neuron xj . Note that this pre-activation can be seen as a “contextualized” geometric map in the following form: psj (ρxj , ψ) =

Z

ψ(y ′ )z dρxj (y ′ , z).

YM ×R

This form will prove convenient in later discussions. At the global level, we may write compactly ps(ρ, ψ) := ψ w(ρ)

where w(ρ) =

N X

wj (ρxj ) 1xj .

j=1

Activation functions. Consider the forward output activation function, or prediction function, F : RN → L2 (XN , R) ≃ RN defined by N X

F : u 7→ F(u) :=

Fj (uj ) 1xj ,

j=1

where Fj is the activation function of neuron xj . This gives the random forward prediction Pred(ψ | ρ) := F(Ps(ψ | ρ)) = F(ψ W ), and the effective prediction, the prediction computed from the external observable weights: pred(ρ, ψ) := F(ps(ρ, ψ)) = F(ψ w(ρ)). Note that, unless Fj affine, in the general case, we have pred ̸≡ E [Pred]. Remark 6.7 (Induced stochastic embeddings and time-dependent contextualisation). Considering Wi,· :=

N X

Wi,j 1xj

j=1

as the embedding vector of yi , the stochastic nature of this embedding can be interpreted as encoding the multi-semanticity of a given activation pattern. We also note that the activation vector ψ actually embeds the residual activation of each neuron. For instance, consider an activation filtered by an exponential kernel memory ψ(t) =

Z t 0

ψ(s) exp(−

t−s ) ds, τm

with forgetting parameter τm (whose modeling is compatible with a Tan–HWG observable dynamics as seen in Section 5.4)3 . The memory dynamics encodes time-dependency and multi-semanticity. 3

The forgetting parameter can also be fiberwise dependent as τm,xj .

40

This stochastic embedding interpretation highlights that synaptic distributions encode multi semantic structure, and that the geometry of P2 (Y ) naturally induces context-dependent representations, in contrast with fixed Euclidean embeddings used in classical neural networks. Note that the observable synaptic weight being the product of the synaptic amplitude sample and the attention weight, we may consider the raw embedding vector of yi : Zi,· :=

N X

Zi,j 1xj ,

j=1

which is not weighted by the attention weights pi,j . Remark 6.8 (Synaptic activation vector). Note that we may consider ψ to be dependent on the neuron xj X ψ := ψi,j 1(yi ,xj ) , i,j

so that the activation vector embeds the residual activation of each synapse, and not just the presynaptic neurons. With ψj :=

M X

ψi,j 1yi ,

i=1

the postsynaptic input would be given by Psj (ψj | ρxj ) :=

M X

ψi,j Wi,j ∈ L0 (Ω, R),

i=1

and the global postsynaptic input state by Ps(ψ | ρ) := Diag(ψ W ). This synaptic activation modeling yields a finer representation, biologically more plausible. For each neuron xj , we consider its set of synapses, M being the maximum number of such synapses per neuron. In such representation, each postsynaptic neuron xj can be linked to its own specific set of presynaptic neurons {x′i }i=1...M which could differ from one xj to another. This opens the representation beyond simple feedforward network topology. This illustrates the versatility of the framework. 6.2.4

Feedback activation and alignment energy

Given a feedback output activation ϕ :=

N X

ϕj 1xj ∈ L2 (XN , R) ≃ RN ,

j=1

we can define the feedback error energy, or stochastic alignment energy, by Ê(ψ, ϕ | ρ) : =

1 1 2 2 Pred(ψ | ρ) − ϕ = F(ψ W ) − ϕ 2 2

41

and the alignment energy by E(ρ, ψ, ϕ) : =

1 1 2 2 pred(ρ, ψ) − ϕ = F(ψ · w(ρ)) − ϕ 2 2

N M X 1X = Fj ψi pi,j E [Zi,j ] − ϕj 2 j=1 i=1

!

!2

.

We can rewrite the energy in the following form N    2 1X E(ρ, ψ, ϕ) = Fj ψ wj (ρxj ) − ϕj 2 j=1 N 1X Fj = 2 j=1



Z



ψ(y )z dρxj (y , z) − ϕj

2

,

explicitly showing the fiberwise dependence on ρxj . Remark 6.9 (Interpretation: classic ML correlation and Hebbian alignment). This expression of the alignment energy is reminiscent of classic machine-learning loss computations. It is as well reminiscent of Hebbian alignment in the sense that stimulating the output layer via ϕ will adjust the synaptic weights (i.e. ρ) so that the forward activation aligns accordingly. Maximazing a scalar product (a correlation between vectors) is actually equivalent to minimizing the ℓ2 distance, provided the vectors are normalized. Hence, here, we could also define a correlation energy by Ẽ(ρ, ψ, ϕ) := −

ϕ Pred(ψ | ρ) , . ∥ϕ∥ ∥Pred(ψ | ρ)∥

This expression is similar to the alignment energy, modulo normalization, where the products ϕj Fj (Psj (ψ)) appear explicitly, recovering the fact that neurons that “fire together, wire together”. It provides an intuitive interpretation of the alignment energy as a normalized correlation objective, making explicit the Hebbian structure of the update. Remark 6.10 (Generalized energy combination). The alignment energy is a special case of a combination of more general types of energies (potential, interaction, and internal). Indeed, expanding the ℓ2 product in the alignment energy gives a local energy of the general form Ex (ρx |(ψ, ϕ)) = Fx where EV (ρx |ψ) =

Z YM ×R



1 EV (ρx |ψ), EV2 (ρx |ψ) ϕ , 2 

V (y, z|ψ) dρx ,

and

V (y, z|ψ) = ψ(y)z.

This is a special case of a local energy of the form Ex (ρx |(ψ, ϕ)) = Fx EV (ρx |ψ), EK (ρx |ψ) ϕ , 

where the term EK (ρx |ψ) =

ZZ (YM ×R)2

K((y, z), (y ′ , z ′ )|ψ) dρx (y, z) dρx (y ′ , z ′ ), 42

generalizes the quadratic term. This energy can be seen as a combination of energies, with potential energy EV and interaction energy EK . Adding a McCann internal energy [24] EU (ρx |ψ) =

Z YM ×Rd

A(ρx (y, z)|ψ) dρx ,

where λd A(λ−d ) convex non-increasing on λ ∈ (0, +∞) with A(0) = 0, d = 1, provides a general form of energy Ex (ρx |(ψ, ϕ)) = Fx EV (ρx |ψ), EK (ρx |ψ), EU (ρx |ψ) ϕ . 

This naturally generalizes to the case Y = YM × Rd , with d ≥ 1. 6.2.5

Plasticity dynamics and λ-geodesic-convexity of the energy

We assume that the memory plasticity dynamics follows a variational dynamics in the Wasserstein space. For the alignment energy E(ρ, ψ, ϕ) =

N 1X Fj 2 j=1



Z



ψ(y ′ )z dρxj (y ′ , z) − ϕj

2

,

this means minimizing the energy for ρxj evolving in the Wasserstein space. Note that the subsequent dynamics gives both: • the structural weight dynamics of the pi,j , • and the synaptic embedding memory dynamics of the µi,j . E has a Hebbian form, but in order to have a well-posed dynamics, the challenge is to have a Tan–HWG energy, in particular, it means having E λ–geodesically-convex (with λ > −∞, possibly negative). Standard activation functions used in machine-learning (e.g. ReLU, Leaky ReLU, Softplus, ELU, sigmoid, tanh), are at least weakly convex. We can show that standard activation functions lead to having the fiberwise map 1 x 7→ (Fj (x) − ϕj )2 2 at least weakly convex. However, this does not guarantee λ-geodesic-convexity of the alignment energy. Such λ-geodesic-convexity can be obtained when the map ρxj 7→ psj (ρxj , ψ) =

Z

ψ(y ′ )z dρxj (y ′ , z),

is geodesically affine and Lipschitz. We check that psj (·, ψ) is Lipschitz when ψ is bounded, but geodesic affinity is the main obstruction. This is due to the rigidity of geodesic affinity in Wasserstein spaces.

43

Intuition. Intuitively, ψ “breaks” geodesic affinity unless it is constant or when the structural weights pi,j are constant (i.e. no mass transport between synapses). As a consequence, the associated energy fails to be geodesically λ–convex, and a Tan–HWG variational dynamics cannot be defined directly. However, geodesic affinity can be controlled through a natural multi-timescale mechanism, consisting of: • a fast “active” regime, where activations ψ can evolve, but the structural weights are stable, • and a slow “consolidation” regime, where, conversely, activations are stable (i.e. “quiet” contexts), but structural weights can evolve. Remark 6.11. The map psj (·, ψ) is actually a contextualized geometric representation map (see Section 7). We see that studying geodesic affinity and Lipschitz continuity of such a map is key to ensure sequential stability of induced Hebbian energies (when externally composed with a λ-convex function). A systematic analysis is beyond the scope of this article. However we can elaborate on the case of psj (·, ψ).

Proposition 6.12 (Rigidity of the pre-activation observable geodesic affinity). Let (Y1 , dY ) be a geodesic CAT(0) space, let Y2 = Rd with its Euclidean metric, and consider the product space Y = Y1 × Y2 endowed with the product metric and the 2-Wasserstein distance W2 on P2 (Y ). Let ψ : Y1 → R be a bounded Borel function and define u : P2 (Y ) → R , d

u(ρ) :=

Z Y1 ×Y2

z ψ(y ′ ) dρ(y ′ , z).

Assume that u is geodesically affine on a subset A ⊂ P2 (Y ) in the following sense: for every W2 –geodesic (ρt )t∈[0,1] contained in A, the map t 7→ u(ρt ) is affine. If A contains two measures whose Y1 –marginals have different supports connected by a nontrivial geodesic in Y1 , then ψ must be constant on the union of these supports. In particular, if A is stable under W2 –geodesics and contains measures with arbitrary Y1 –marginals, then ψ is constant on Y1 . Proof. Fix two points (y0 , z0 ), (y1 , z1 ) ∈ Y1 ×Y2 with y0 ̸= y1 and z0 ̸= z1 , and consider the measures ρ0 := δ(y0 ,z0 ) ,

ρ1 := δ(y1 ,z1 ) .

The unique W2 –geodesic (ρt )t∈[0,1] between ρ0 and ρ1 is given by ρt = δ(yt ,zt ) ,

yt := γy0 →y1 (t),

zt := (1 − t)z0 + tz1 ,

where γy0 →y1 is the constant-speed geodesic in Y1 from y0 to y1 . Along this geodesic,

u(ρt ) = zt ψ(yt ) = (1 − t)z0 + tz1 ψ γy0 →y1 (t) . 



By assumption, t 7→ u(ρt ) is affine. In particular, for each fixed pair (y0 , z0 ), (y1 , z1 ), the map t 7→ (1 − t)z0 + tz1 ψ γy0 →y1 (t) 

44



must be affine on [0, 1]. Now fix y0 ̸= y1 and vary z0 , z1 ∈ Rd . By choosing z0 , z1 linearly independent (or simply noncollinear in d ≥ 2, or distinct scalars in d = 1), one checks that the only way for the product of an affine function in t (namely zt ) with the scalar function t 7→ ψ(γy0 →y1 (t)) to remain affine for all choices of z0 , z1 is that t 7→ ψ(γy0 →y1 (t)) is constant on [0, 1]. Hence ψ γy0 →y1 (t) = ψ(y0 ) = ψ(y1 ) 

for all t ∈ [0, 1].

Therefore, whenever two points y0 , y1 ∈ Y1 are connected by a geodesic appearing as the Y1 – component of a W2 –geodesic in A, the values of ψ at y0 and y1 must coincide. If A contains measures whose Y1 –marginals have supports connected by nontrivial geodesics in Y1 , this implies that ψ is constant on the union of these supports. If, moreover, A is geodesically rich enough to contain measures with arbitrary Y1 –marginals, we conclude that ψ is constant on Y1 . Corollary 6.13 (Multi-scale coherence). In the setting above, suppose one considers an energy functional of the form  E(ρ) := F u(ρ) , with F : Rd → R λ–convex. If E is required to be geodesically λ–convex on a class A ⊂ P2 (Y ) that is stable under W2 –geodesics and contains measures with nontrivial variation in their Y1 –marginals, then necessarily ψ must be constant on the relevant region of Y1 . Consequently, any nontrivial spatial modulation of synaptic activation (i.e., any non-constant ψ) is compatible with a geodesically convex gradient-flow structure only on subsets where the Y1 –marginal is effectively fixed (i.e. no transport occurs in Y1 ). This yields a natural separation of scales: a fast dynamics acting on the Rd –component at fixed Y1 –marginals, and a slow structural dynamics (consolidation regime) that can only be made variationally coherent after suitable averaging or coarse-graining of ψ. Remark 6.14 (Functional role of neuronal assemblies in consolidation regime). This gives a geometric interpretation of the functional role of neuronal assemblies in consolidation regime: groups of neurons with similar activation patterns are clustered in order to obtain an “effectively constant” activation state ψ, allowing proper Tan–HWG dynamics of structural weights. 6.2.6

Discussion

Activation functions and non-linearity. In classical artificial neural networks, activation functions are introduced to provide non-linear computation capacities to otherwise “flat” Euclidean representational networks. In our framework, we can rethink their functional role: the underlying geometry already encodes rich non-linear features, and activation functions can be seen as representational maps that have to respect such geometry (i.e. preserving geodesicity via λ-convexity). In other words, they act as compatible projections of the underlying latent space. They do not shape an absolute latent space: they provide observable projections, from which we approximate a fundamentally curved geometry, which, in turn, is the agent’s representation of the world. In contrast with classical neural networks, non-linearity does not come from the activation function alone, but from the underlying Wasserstein geometry.

45

Spiking neural networks and sequentiality of context capture. In biological neural networks, activation functions are spiking functions. Again, this is compatible with our framework. For instance, using sigmoid or tanh activation functions of the form x 7→ σ(βx),

or

x 7→ tanh(βx),

β ≫ 1,

would give “quasi-spiking” neural networks with non-convex, but still, Tan–HWG energies. Moreover, it also provides a biological justification for the sequential capture of external context. Since spikes have refractory periods, this can be seen as an active freezing of the signal. A multi-scale coherence mechanism. Our model naturally separates synaptic plasticity into two distinct dynamical regimes: a fast functional dynamics acting on the synaptic weights (µi,j ), and a slow structural dynamics acting on the distribution of synaptic connections (pi,j ). This separation is not imposed a priori; rather, it emerges from the geometric constraints of the Wasserstein space in which the dynamics evolves. At the fast time scale, the synaptic activation function ψ may vary across synapses, leading to a heterogeneous and locally irregular evolution of the synaptic weights. In this regime, the energy functional remains geodesically convex when the structural weights (pi,j ) are held fixed, allowing a well-posed Tan–HWG dynamics. At the slow time scale, however, the same geometric constraints imply that a coherent variational formulation is only possible when the influence of ψ is effectively averaged or homogenized (to obtain ψ effectively constant). In other words, the slow structural dynamics can only be geometrically consistent if it “sees” a coarse-grained version of the fast synaptic activity. Biologically, this mechanism resonates with well-documented observations. Fast synaptic plasticity (e.g., LTP/LTD, STDP) depends on local, heterogeneous activity, whereas slow structural plasticity (e.g., spine growth and elimination, long-term consolidation) is driven by global or averaged signals, such as those emerging during sleep or during the coordinated reactivation of neuronal assemblies. Thus, the mathematical structure of the model mirrors the known separation between rapid functional plasticity and slow structural remodeling in neural circuits. Non-uniqueness of dynamics. Note that λ-convexity with λ < 0 does not guarantee uniqueness of the JKO step, but the Tan–HWG framework naturally accomodates multivalued updates. In other words, the dynamics can be non-deterministic: agents facing exactly the same contexts and initial conditions may show different memory dynamics. This is not to be confused with the stochastic nature of random variables like the amplitude Zi,j , which stochasticity comes from the distribution µi,j . Instead, the non-deterministicity here is relative to the evolution of ρ, thus µi,j itself. This evolution can be deterministic (e.g. when λ ≥ 0) while Zi,j remains a random value. Global signal. as

Note that in the Tan–HWG alignment energy, we implicitly identified the context

Φ :=

M X

ψi δyi ⊗

i=1

N X

ϕj δxj ∈ M(YM × XN ) ⊂ M((YM × R) × XN )XN ,

j=1

or, in the more general synaptic activation case (see Remark 6.8), Φ := {Φj }j=1...N ,

Φj :=

M X i=1

ψi,j δyi ⊗

N X

ϕk δxk ,

k=1

46

Φ ∈ M(YM ×XN )XN ⊂ M((YM ×R)×XN )XN ,

and the global signal as S(ρ, Φ) := (Pin (Φ), Pout (Φ)),

Pin (Φ) =

(M X

)

, and Pout (Φ) =

ψi,j 1yi

i=1

j=1...N

(N X k=1

)

.

ϕ k 1 xk j=1...N

Hence the external space is actually Z = L2 (Y ) × L2 (XN ), and we implicitly used the canonical embedding L2 (YM ) ,→ L2 (YM × R) ,→ Z = L2 (Y ) × L2 (XN ). for the geometric representation maps. Extension. The previous construction, based on the internal space YM × R, admits a natural geometric reinterpretation. Indeed, using the polar decomposition of real numbers, R ≃ R+ × {−1, 1}, we may rewrite YM × R ≃ YM × R+ × {−1, 1}, where the discrete set {−1, 1} can be viewed as a degenerate phase space restricted to angles 0 and π. This observation suggests a natural extension: replace the discrete phase set by the continuous unit circle. Introducing the internal space Y := YM × R+ × S1 ≃ YM × C, where S1 denotes the unit circle, allows us to encode both amplitudes and continuous phases. Although S1 is not a CAT(0) space, the identification (r, θ) 7−→ reiθ ∈ C embeds amplitude–phase pairs into the Euclidean plane, which is CAT(0). Hence the extended internal synaptic space Y remains compatible with the Tan–HWG geometric framework. This extension is the object of the next section.

6.3

Extension to complex valued weights: spectral memory and phase locking

Biological and artificial neural systems often rely on oscillatory interactions to coordinate activity across populations of units. Phase synchronisation plays a central role in attention, memory consolidation, and large-scale communication between brain areas. Oscillatory couplings also appear in machine-learning architectures such as Hopfield networks [14], reservoir systems, and oscillatorbased computing. Internal oscillatory state space. To incorporate such effects, the internal synaptic space must be enriched so as to encode phases or complex amplitudes. The representation Y := YM × C provides a natural geometric setting for this purpose, enabling the introduction of spectral memory and phase locking mechanism within the Tan–HWG framework. For each (y, z) ∈ YM × C, the complex coordinate z ∈ C encodes an oscillatory state with amplitude and phase. 47

Each output neuron xj is now represented by a probability measure ρxj ∈ P2 (YM × C),

ρxj =

M X

pi,j δyi ⊗ µi,j ,

i=1

where each synapse (yi , xj ) has a complex synaptic embedding memory µi,j ∈ P2 (C). This distribution is an internal structural state: it encodes a preferred phase profile, possibly multimodal, with uncertainty. 6.3.1

Spectral memory

The complex synaptic amplitude sample is now a complex random variable Zi,j ∼ µi,j , with polar decomposition Zi,j = ri,j eiθi,j ,

where

ri,j ∈ L0 (Ω, R+ ),

and

θi,j ∈ L0 (Ω, R/(2πZ)).

We also consider the polar decomposition of the first moment E [Zi,j ] = r̂i,j eiθ̂i,j . We define: Definition 6.15 (Spectral memory). For a given memory field ρ, the spectral memory of synapse (i, j) is the probability distribution ξi,j of the random phase θi,j . This spectral memory is a structural phase memory: it is an internal property of the synapse, distinct from the instantaneous activation phases of the units. Activation states now become complex-valued, ϕ ψ ψi = |ψi | eiθi , ϕj = |ϕj | eiθj , and interact with the internal spectral memory through the observable synaptic weight Wi,j := pi,j ri,j eiθi,j . Postsynaptic input and spectral filtration. random variable Psj (ψ) :=

M X

The postsynaptic input becomes the complex  

|ψi | pi,j ri,j exp i θi,j + θiψ



.

i=1

The real part (or any projection) of this signal is modulated by the spectral memory profile. Depending on the shape of µi,j , the synapse implements: • a deterministic phase shift (Dirac spectral memory); • multimodal phase selectivity (e.g. finite or countable support); • robust spectral filtering (e.g. Gaussian or mixture distributions). Thus, the synapse acts as a functional spectral filter whose selectivity is encoded in its internal memory distribution. 48

6.3.2

Alignment energy and explicit Hebbian structure

The alignment energy associated with a memory field ρ and activation states (ψ, ϕ) is E(ρ, ψ, ϕ) :=

1 2 F(ψ · w(ρ)) − ϕ . 2

Expanding the Hermitian product yields 1 1 2 F(ψ · w(ρ)) + ∥ϕ∥2 − ℜ⟨F(ψ · w(ρ)), ϕ⟩ 2 2 !# " M N X X 1 1 2 2 F(ψ · w(ρ)) + ∥ϕ∥ − ψi pi,j E[Zi,j ] = ℜ ϕj Fj 2 2 i=1 j=1

E(ρ, ψ, ϕ) =

M N X X 1 1 2 F(ψ · w(ρ)) + ∥ϕ∥2 − ψi pi,j E[ri,j eiθi,j ] = ℜ ϕj Fj 2 2 i=1 j=1

!#

"

.

Locally affine activation functions. Assume that each activation function Fj is locally affine on a non-negligible open neighbourhood U ⊂ C containing the values M X

ψi pi,j E[ri,j eiθi,j ],

i=1

and takes the form

Fj (g) = aj g + bj ,

aj , bj ∈ C,

for instance when each Fj is affine by parts. In this region, the alignment energy becomes E(ρ, ψ, ϕ) = where A :=

N X

1 2 ψ · w(ρ) · A − (ϕ − B) , 2 B :=

aj 1(xj ,xj ) ,

N X

bj 1xj .

j=1

j=1

Remark 6.16 (Interpretation). The vector B plays the role of a bias, while A acts as an inertia √ operator, in analogy with the coefficients αj in the purely quadratic Tan–HWG case and the role of αj in the associated EMA observable dynamics. Expanding the Hermitian product on U gives E(ρ, ψ, ϕ) =

1 1 2 ψ · w(ρ) · A + ∥ϕ − B∥2 − ℜ⟨ψ · w(ρ) · A, ϕ − B⟩ . 2 2

The quadratic term expands as N M X 1 1X 2 ψ · w(ρ) · A = |aj |2 ψi wi,j (ρxj ) 2 2 j=1 i=1

=

2

N   X 1X |aj |2 |ψi ||ψi′ | pi,j pi′ ,j r̂i,j r̂i′ ,j cos (θ̂i,j + θiψ ) − (θ̂i′ ,j + θiψ′ ) . 2 j=1 i,i′

49

The linear term becomes −ℜ⟨ψ · w(ρ) · A, ϕ − B⟩ = −

X





|ϕj − bj | |aj | |ψi | pi,j r̂i,j cos θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj ) .

i,j

Thus, for each fixed j, the fiberwise energy reads   X 1 |ψi ||ψi′ | pi,j pi′ ,j r̂i,j r̂i′ ,j cos (θ̂i,j + θiψ ) − (θ̂i′ ,j + θiψ′ ) Ej (ρxj , ψ, ϕ) := |aj |2 2 i,i′

− |aj | |ϕj − bj |

X





|ψi | pi,j r̂i,j cos θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj )

(6)

i

+ |ϕj − bj | . 2

Remark 6.17 (Explicit Hebbian structure of the energy). By the energy equation (6), we see that the synaptic memory weights pi,j , r̂i,j and θ̂i,j are directly influenced by co-activations ψi ψi′ and ψi (ϕj − bj ): this is an explicit manifestation of the Hebbian nature of the energy. In particular, the output residual activation is only present in the linear term −|aj | |ϕj − bj |

X



|ψi | pi,j r̂i,j cos θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj )



i

which has a direct Hebbian interpretation. More precisely, when both the residual presynaptic activity |ψi | and the residual postsynaptic activity |ϕj − bj | (corrected by bias) are simultaneously large, it decreases the energy proportionally to 



pi,j r̂i,j cos θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj ) . Hence it will favor simultaneously (i) larger weights pi,j r̂i,j as well as (ii) external field alignment i.e. when θ̂i,j leads to an alignment of input phase θiψ with the output phase arg(ϕj − bj ) − arg(aj ) (including correction by bias and inertia). In other words, the synapse adjusts its internal phase so as to compensate the phase mismatch between presynaptic and postsynaptic oscillatory activity. This is a phase–Hebbian mechanism: structural phase alignment occurs when the two units fire together. Remark 6.18 (Phase-locking by large output stimulation and spectral memory as a repository of stable phase offsets.). When the output residual activation (ϕj ) is large, the alignment mechanism above contributes to shaping the spectral memory. The spectral memory acts as a repository of phase offsets that repeatedly and consistently occur with large input/output coactivations. As learning progresses and predictions become more accurate, the effective phase differences stabilise, and the structural phases θi,j converge accordingly. The plasticity dynamics therefore implements a structural phase-locking mechanism, learning and storing stable phase relationships between input and output activity patterns. The explicit form of the energy provides a convenient basis for a local parametric analysis, which is the purpose of the following section.

50

6.3.3

Local parametric analysis: pruning, selectivity, pressure equalisation, synchronization, alignment

We now analyse the fiberwise structure of the energy to reveal mechanisms governing pruning, selectivity, synchronization and phase alignment. The following analysis is qualitative and aims at revealing the mechanisms encoded in the energy landscape, rather than providing a full classification of equilibria. A reasonable assumption is to assume each Fj sufficiently regular to locally admit an affine approximation in an open neighbourhood U ∈ C of the form Fj (g) = Fj (g0 ) + F′j (g0 ) (g − g0 ) + o(g − g0 ),

for fixed g0 ∈ U

and all g ∈ U .

In the following analysis, we neglect the o(g) term (which actually vanishes in the case of Fj affine by parts). In this setting, the structure of the energy reveals three fundamental mechanisms: i. Amplitude equilibrium and pruning.

The partial derivative with respect to r̂i,j is

X ∂E = |aj |2 |ψi |2 p2i,j r̂i,j + |aj |2 |ψi |pi,j |ψi′ |pi′ ,j r̂i′ ,j cos((θ̂i,j + θiψ ) − (θ̂i′ ,j + θiψ′ )) ∂ r̂i,j i′ ̸=i

− |aj | |ϕj − bj | |ψi | pi,j cos(θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj )). Either aj = 0 (the synapse is ignored in U), or pi,j = 0 (structural pruning), or yi is not activated (ψi = 0, hence no Hebbian phenomenon involved), or the equilibrium magnitude is theoretically eq = r̂i,j

Since

|ϕj − bj | cos(θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj )) |aj | |ψi | pi,j X |ψi′ |pi′ ,j − r̂i′ ,j cos((θ̂i,j + θiψ ) − (θ̂i′ ,j + θiψ′ )). |ψ |p i i,j ′ i ̸=i ∂2E 2 2 2 2 = |aj | |ψi | pi,j ≥ 0, ∂ r̂i,j

eq eq r̂i,j is a local minimum and if r̂i,j < 0, and unless the joint evolution of the other variables leads eq eq to lift r̂i,j above 0, the constraint r̂i,j ≥ 0 enforces pruning (r̂i,j → r̂i,j until reaching 0). Note that a stricter constraint could be given by considering the frontier of U: we do not detail such consideration in this paper.

ii. Pressure equalization, pruning and selectivity. For fixed amplitudes and phases, the P optimization of the structural weights (pi,j )i is constrained by the simplex condition i pi,j = 1, pi,j ≥ 0. Introducing a Lagrange multiplier λj for the equality constraint, the stationarity condition reads ∂E = −λj for all i with pi,j > 0, ∂pi,j together with the complementary slackness condition pi,j = 0 ⇒

∂E + λj ≥ 0. ∂pi,j 51

Since

∂2E 2 = |aj |2 |ψi |2 r̂i,j ≥ 0, ∂p2i,j

the energy is convex in each pi,j , and interior equilibria satisfy ∂E ∂E = ∂pi,j ∂pi′ ,j

for all active indices i, i′ .

Thus, active synapses equalize their “energy pressure” ∂E/∂pi,j , while inactive synapses lie on the boundary of the simplex. In particular, if the unconstrained equilibrium value peq i,j obtained from ∂E/∂pi,j = 0, i.e. peq i,j =

|ϕj − bj | cos(θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj )) |aj | |ψi | r̂i,j X |ψi′ |r̂i′ ,j − pi′ ,j cos((θ̂i,j + θiψ ) − (θ̂i′ ,j + θiψ′ )), |ψ |r̂ i i,j i′ ̸=i

satisfies peq i,j < 0, the Karush-Kuhn-Tucker condition forces pi,j towards 0, corresponding to structural pruning. Conversely, if peq i,j > 1, the simplex constraint forces all other weights towards vanishing, yielding selectivity. Interior equilibria occur only when all active weights share the same value of ∂E/∂pi,j . eq iii. Synchronization and alignment. peq i,j and r̂i,j actually follow the same equation, which gives the following Hebbian equilibrium equation: eq |ψi | peq i,j r̂i,j =

|ϕj − bj | eq + θiψ + arg(aj ) − arg(ϕj − bj )) cos(θ̂i,j |aj | −

X

ψ eq ψ eq eq |ψi′ |peq i′ ,j r̂i′ ,j cos((θ̂i,j + θi ) − (θ̂i′ ,j + θi′ )),

i′ ̸=i eq eq satisfied by the theoretical equilibrium point (peq i,j , r̂i,j , θ̂i,j ) (“theoretical” in the sense that it may eq < 0). This equation can be rewritten as / [0, 1] or r̂i,j never be attained if peq i,j ∈ M X

eq ψ ψ eq eq |ψi′ |peq i′ ,j r̂i′ ,j cos((θ̂i,j +θi ) − (θ̂i′ ,j +θi′ )) =

i′ =1

|ϕj −bj | eq cos(θ̂i,j +θiψ +arg(aj )−arg(ϕj −bj )). |aj |

(7)

(iii.a.) If the internal phases satisfy cos((θ̂i,j + θiψ ) − (θ̂i′ ,j + θiψ′ )) > 0

for all i, i′ ,

the synapses are said to be globally synchronized relative to a fixed activation input ψ. By the eq Hebbian equilibrium equation (7), the equilibrium phase θi,j satisfies eq cos(θ̂i,j +θiψ +arg(aj )−arg(ϕj −bj )) ≥ 0.

Hence the dynamics pushes the spectral memory to compensate the input phase towards alignment with the output field (corrected by bias and inertia). In other words, synchronized groups pushes towards phase alignment, which recovers a phase locking phenomenon. 52

Note that, if

cos(θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj )) < 0,

then both

∂E > 0, ∂ r̂i,j

and

∂E > 0, ∂pi,j

and since we have the constraints r̂i,j ≥ 0 and pi,j ≥ 0, this forces both structural weights and amplitude towards vanishing i.e. pruning: the regime is unstable as expected. (iii.b) Conversely, having cos((θ̂i,j + θiψ ) − (θ̂i′ ,j + θiψ′ )) < 0

for all i, i′ ,

is only possible in a group of at most 3 synapses, since on the circle, a set of points with all pairwise negative cosine differences cannot exceed three elements. In such a limited group of “anti”-synchronized neurons, we have eq cos(θ̂i,j +θiψ +arg(aj )−arg(ϕj −bj )) ≤ 0,

and the group can be said to be “anti”-aligned. We can check that group alignment, i.e. cos(θ̂i,j + θiψ + arg(aj ) − arg(ϕj − bj )) > 0, may push some peq i,j > 1, pushing towards selectivity, pruning the other connections. In large groups of active synapses (≫ 3), stable non-degenerate equilibria would therefore generally correspond to synchronization + alignment. Remark 6.19 (Spectral memory interpretation). This generalizes the phase-locking and stable phase offsets memory mechanisms discussed in Remark 6.18, adding synchronization as a joint phenomenon. Remark 6.20 (From frequency mismatch to frequency locking in dynamic spectral memory: a Kuramoto dynamics). Spectral memory only stores phase distribution and not frequencies. Consider now the case in which the input and output phases evolve as θiψ (t) = ωin t + θi0 ,

θjϕ (t) = ωout t + θj0 ,

with possibly distinct frequencies ωin ̸= ωout . The effective phase difference that the Tan–HWG plasticity attempts to compensate is then target ∆θi,j (t) = θjϕ (t) − θiψ (t) = (ωout − ωin ) t + Ci,j .

When ωout = ωin , this target offset is constant and the structural phase θi,j can converge to a stable value, yielding a genuine phase memory. In contrast, when ωout ̸= ωin , the target offset drifts linearly in time. A static spectral memory cannot encode such a drift: the phase that the plasticity attempts to learn is no longer a fixed point but a continuously rotating quantity. In this regime, the synapse would need to “rotate” its structural phase at a rate ωi,j = ωout − ωin ,

53

which effectively defines a learned intrinsic frequency. However, such a frequency cannot be stored as a static synaptic parameter: it must be maintained dynamically by an internal oscillatory mechanism. A biologically plausible solution is the presence of a recurrent loop or microcircuit generating a self0 . This internal oscillator provides sustained oscillation whose phase evolves as θi,j (t) = ωi,j t + θi,j the synapse with a continuously rotating reference phase, allowing the spectral memory to encode stable offsets relative to this internally generated drift. Such self-sustained oscillations are well documented in cortical, thalamic, hippocampal and cerebellar circuits [5, 34, 20, 18], where recurrent excitation–inhibition loops maintain intrinsic frequencies that interact with external inputs through phase-locking mechanisms. In summary, a frequency mismatch ωout ̸= ωin cannot be stored as a static spectral memory. Instead, it requires an internally generated oscillation that continuously maintains the appropriate drift. The Tan–HWG plasticity then performs structural phase-locking relative to this internal oscillator, rather than memorising a frequency as a fixed synaptic parameter. This “internal” reference frequency, together with synchronization mechanisms, is reminiscent of a Kuramoto-like dynamics. However, we did not need to introduce an additional oscillatory interaction equation. In inference mode (e.g. when no feedback output activation signal is given), the spectral memory acts as a spectral filter favoring compatible input phases θiψ . Hence, in this perspective, the synchronization of the θiψ comes from phase filtration rather than from an explicit synchronization force (although such an additional force would not be incompatible with the framework). Remark 6.21 (Functional analysis.). At the level of the memory field ρ, the plasticity dynamics could be rigorously described as a Tan–HWG dynamics. In the simple affine case, the local alignment energy takes the form ZZ 1 2 Ex (ρx |(ψ, ϕ)) := |A(x)| K((y, z), (y ′ , z ′ )|ψ) dρx (y, z)dρx (y ′ , z ′ ) 2 2 (YM ×C) 

− ℜ (ϕ(x)−B(x))A(x)

Z YM ×C

V (y, z|ψ) dρx



+ |ϕ(x) − B(x)|2 , where K((y, z), (y ′ , z ′ )|ψ) := ψ(y)ψ(y ′ )z z ′ ,

and

V (y, z|ψ) := ψ(y)z.

The qualitative analysis in terms of amplitudes, phases, and structural weights can be seen as a parametric reading of this underlying variational structure. An explicit Wasserstein gradient flow formulation of ρx may be given, leading to a Hebbian plasticity equation   δEx ∂t ρx = ∇ · ρx ∇ , (8) δρx where, in the affine case, we have Z h i δEx (y, z) = |A(x)|2 K((y, z), (y ′ , z ′ )|ψ) dρx (y ′ , z ′ ) − ℜ (ϕ(x)−B(x))A(x)V (y, z|ψ) . δρx YM ×C A full analytical characterization of the dynamics would require additional structural assumptions. In this work, we therefore focus on the qualitative mechanisms revealed by the energy landscape (pruning, selectivity, synchronization and alignment), rather than on a complete classification of trajectories. 54

6.3.4

Discussion

On the functional role of the structural weights. The structural weights pi,j have distinguished features, which are otherwise not covered when solely considering amplitudes Zi,j : • Synaptic competition. The simplex constraint i pi,j = 1 gives a representation of synaptic competition over resources: the postsynaptic neuron can only provide limited attention to its synapses. Amplitudes and structural weights both show pruning capacities, but solely structural weights are subject to selectivity, i.e. competition pushed to its degenerate limit when only one synapse dominates, pushing the others to vanish. Note that selectivity here refers to functional dominance, not anatomical elimination. P

• Semantic bumper for stability and robustness of informational embedding. Since the energy is a function of the product pi,j · r̂i,j eiθ̂i,j , small perturbations of the input or of the synaptic amplitudes can be absorbed either by adjusting the fibre laws µi,j (via r̂i,j , θ̂i,j ) or by redistributing mass across inputs (via pi,j ), which enhances robustness to noise. If we interpret amplitudes (r̂i,j eiθ̂i,j ) as semantic or informational embeddings, the layer of structural weights pi,j can act as a stabilizer of such embeddings. When facing repeated activations (ψ, ϕ) that would otherwise tend to modify the embedding (e.g. presenting incoherences or small perturbations with previous activation profiles), the discrepancy can be partially absorbed by the structural weights. It allows robust embeddings, maintaining coherence across inputs/outputs activations. This absorption capacity has limits, in particular since pi,j ∈ [0, 1], and since active synapses are pushed towards respecting the stabilized energy pressure attractor ∂E/∂pi,j = −λj at postsynaptic neuronal level. These limits can be seen as gating transitions from a semantic value to another, which can be linked to the phase transition feature (see next item). • Phase transitions. In this setting, a phase transition refers to a qualitative change in the structure of the minimizer of the local energy Ex . More precisely, it occurs when the optimal structural weights (pi,j )i=1...M move from an interior point of the simplex (all synapses active) to a boundary point (some pi,j = 0), or conversely. Such transitions correspond to changes in the active support of ρx and are induced by the geometry of the energy landscape. They are invisible when optimising only over amplitudes, since the simplex constraint is essential for the emergence of these structural bifurcations. Comparison with existing architectures. The separation between structural weights pi,j and semantic complex amplitudes zi,j = r̂i,j eiθ̂i,j may evoke mechanisms found in several existing models, but the proposed construction differs from all of them in essential ways. • Mixture models and EM-type architectures. Gaussian mixture models and related EM algorithms also separate mixture weights from component parameters. However, these models lack any notion of synaptic synchronization or alignment. • Attention mechanisms and mixture-of-experts. Transformer attention and MoE architectures use routing weights (softmax scores or gating coefficients), which is reminiscent of the structural weights pi,j forming a probability distribution over input locations and modulating the contribution of semantic synaptic states zi,j = r̂i,j eiθ̂i,j . Yet these weights do not induce pruning or selectivity through simplex boundary equilibria, nor do they account for synchronization and alignment phenomena.

55

• Sparse coding and dictionary learning. Sparse coding separates activation coefficients from dictionary atoms, but the activations are continuous and do not represent structural mass allocation. There is no phase dynamics, no synchronization mechanism, and no structural pruning analogous to pi,j → 0. • Bayesian nonparametrics. Dirichlet-process mixtures also feature structural weights, but these arise from stochastic priors rather than from a deterministic energy landscape. In summary, while several families of models exhibit partial analogies (routing vs. content, mixture weights vs. component parameters), none of them combine: (i) structural mass allocation on a simplex, (ii) semantic complex amplitudes with phase interactions, (iii) a Wasserstein gradient flow on a mixed discrete–continuous space, and (iv) emergent pruning and selectivity through KKT boundary equilibria. Practical advantages The geometric and dynamical structure of the proposed model leads to several practical advantages: • Intrinsic stability and convergence. The present architecture is governed by a well posed minimizing scheme relative to an explicit energy, yielding intrinsic stability, convergence to equilibrium states, and robustness to noise through mass redistribution and phase synchronization. • Emergent sparsity and potential computational benefits. Attention weights in Transformers are dense and rarely reach exact zeros, unless sparsity is imposed externally. Here, the simplex constraint on (pi,j )i and the associated KKT conditions naturally produce structural pruning and selectivity: entire synaptic branches may vanish at equilibrium. This leads to adaptive sparsity, reduced computational cost, and improved interpretability. This opens the door to adaptive, data-driven sparsification without additional regularization or architectural heuristics. • Geometric coherence and phase interactions. Classic neural networks operate in a flat Euclidean space and lacks an intrinsic geometric structure. The proposed model evolves in a rich, intrinsically curved, CAT(0) Polish space, P2 (YM × C) (YM being implicitly lifted into a CAT(0) Polish space, see Section 4), combining mass transport with complex-valued phase interactions. This ensures uniqueness of geodesics and convexity properties that are absent from classical Euclidean embeddings. This induces geometric coherence, synchronization phenomena, and alignment with external fields, which can improve robustness and temporal consistency in applications involving oscillatory or structured signals.

7

Mapped energies

In Section 6 we introduced geometric representation maps, and we used a contextualized geometric representation map of the form ρ 7→ ψ w(ρ) to define the alignment energy. In this section, we generalize this construction. Definition 7.1 (Mapped energy). A Hebbian energy is said to be a mapped energy if its local density has the form Ex (ρx , z) = Eex (Hx (ρx , z), z),

56

ρx ∈ P2 (Y ), z ∈ Z,

where H is said to be a contextualized geometric representation map or Hebbian map H = {Hx }x∈X where Hx : P2 (Y ) × Z → Z, x 7→ Hx measurable. When Z is a Hilbert space, we say that the mapped energy is a Hilbertian mapped energy. The term “Hebbian” in the name Hebbian map, reflects the fact that ρx 7→ Hx (ρx , z) is a geometric representation map contextualized by a Hebbian global signal z. Mapped energies provide a unifying framework for expressing Hebbian-type energies as compositions of geometric maps with classical Euclidean costs. This makes it possible to analyze the variational structure of complex neural architectures by studying the regularity of the underlying geometric maps.

7.1

Examples: energies induced by activation functions

Hilbertian mapped energies conveniently map probability distributions ρx to vectors in Z. Then, this vectorial representation allows us to bridge the underlying Wasserstein geometry with classical alignment or correlation “cost” functions. Alignment energy. For instance, the alignment energy, defined in Section 6, is a special case of Hilbertian mapped energy with E(ρ, S(ρ, ϕ)) =

N  2 1X Hxj (ρxj , (ψ·,j , ϕj )) − Sxj (ρ, (ψ, ϕ)) , 2 j=1

Sxj (ρ, (ψ, ϕ)) = (ψ·,j , ϕj ),

and 



Hxj (ρxj , (ψ·,j , ϕj )) = ψ·,j , Fj (ψ·,j · wj (ρxj )) =

M X

ψ·,j , Fj

!!

ψi,j · wi,j (ρxj )

.

i=1

where we used the general synaptic activation vector ψ·,j = Remark 6.8).

i=1 ψi,j 1yi (which depends on xj ; see

PM

Remark 7.2 (Elementary contextualized map). The Hebbian map H is not unique: the alignment energy can also be written N 1X E(ρ, ψ, ϕ) = Fj 2 j=1



Z



ψ(y )z dρxj (y , z) − ϕj

2

= Eex (psj (ρxj , ψ), z),

using psj (ρxj , ψ) =

Z

ψ(y ′ )z dρxj (y ′ , z),

as the contextualized geometric representation map. More generally, we may factor “elementary” contextualized maps within mapped energies. Elementary features of such building blocks, like geodesic affinity and Lipschitz continuity, inform us on the behavior of the memory dynamics (e.g. see Section 6.2.5). Mapped energies allow us to decompose complex neural architectures into elementary geometric maps, whose regularity properties directly control the behavior of the memory dynamics.

57

7.2

dZ –isotropic, dZ -anisotropic and dZ -quadratic energies

The present section can be seen as a generalization of Section 5, where we had Z = P2 (Y ) and Hx (·, z) ≡ IdP2 (Y ) . In particular: • dZ -isotropic energy: We can define a class of dZ -isotropic energies, with densities of the form Ex (·, z) := Fx (dZ (Hx (·, z), z)),

z ∈ Z.

• dZ -anisotropic energy: Use a bi-Lipschitz equivalent distance on Z, to define a dZ -anisotropic energy. • dZ -quadratic energy: When Z is CAT(0) and when Hx (·, z) is geodesically convex for µ-a.e. x, then we can define dZ -quadratic Tan–HWG energies with densities of the form Ex (·, z) := fx (d2Z (Hx (·, z), z)), where fx is proper, l.s.c. and convex nondecreasing, with x 7→ fx measurable, and fx (0) < +∞ µ-a.e. and (x 7→ fx (0) ∈ L1 ). • State-independent contraction factor: Assuming Hx (·, z) is geodesically affine and 1-Lipschitz, we can obtain a generalized state-independent contraction factor theorem for generalized quadratic energy (see Definition 7.3), with a coefficient tτ,x ∈ (0, 1) independent of ρnx such that n n Hx (ρn+1 x , hx ) = hx,tτ,x ,

dZ (Hx (ρnx , hnx ), hnx,tτ,x ) = tτ,x dZ (Hx (ρnx , hnx ), hnx ),

for µ-a.e. x ∈ X, where (hnx,t )t∈[0,1] is the geodesic from Hx (ρnx , hnx ) to hnx . This also assumes Hx (·, hnx )−1 (hnx ) ̸= ∅. Note that, having Hx (·, z) geodesically affine and 1-Lipschitz means Hx (·, z) is geodesically an isometry, which is a strong assumption: Hx (·, z) strictly respects the underlying Wasserstein geometry. Such assumptions are rarely satisfied in practice, but they illustrate the structural behavior of mapped energies. In particular, if we require convexity of the energy, dZ -quadratic Tan–HWG energies provide a wide class of compatible energies. However this construction excludes the use of weakly convex mapped energies. 7.2.1

Generalized quadratic energy (GQE)

We now present the class of energies that we will consider to establish a continuous-time limit curve in Section 8. It is a generalization of the purely quadratic Tan–HWG energy class, using geometric representation maps. Definition 7.3 (Generalized quadratic energy). A generalized quadratic energy (GQE), with parameter α > 0, and mapping H, is a mapped energy of the form α E(ρ, ϕ) := 2

Z X

d2Z Hx (ρx , Sx (ρ, ϕ)), Sx (ρ, ϕ) dµ(x). 

When the energy is Tan–HWG, we say that it is a Tan–HWG GQE. Note that, in the general case, we cannot guarantee the λ-convexity of a GQE. 58

8

Vanishing time step limit curve

Up to this point, the global signal—and in particular the context—has been treated as an arbitrary external input, since the Tan–HWG minimizing-movement scheme freezes it at each step. In particular, with a fixed timestep τ , the context could be considered as a given fully external process, taken as an input for the system. In this section we will show that, under suitable structural assumptions on the signal and a coupling constraint between memory fields and contexts, the system may enter an autonomous, or quasi-autonomous regime, with a “free evolution” of the signal (i.e. under stationary or quasi-stationary external context). Remarkably, in this regime, Tan–HWG schemes admit continuous-time limit curves (Theorem 8.7). This suggests that such limit curves are key features of memory consolidation, which is more likely to occur in quasi-stationary external context (e.g. “sleep”).

8.1

Regularity assumptions

We now give three regularity assumptions —two structural and one contextual, under which, Tan– HWG minimizing movement schemes can admit continuous-time limits. Lipschitz continuity is the key structural assumption, while “sleep-mode” is the key contextual assumption: both will allow us to control the freezing error in the Tan–HWG scheme. 8.1.1

First structural assumption: Lipschitz signal

Consider a Hebbian energy E(ρ, ϕ) :=

Z X

Ex ρx , Sx (ρ, ϕ) dµ(x). 

Assumption 8.1 (Lipschitz signal). The global signal is Lipschitz, in the sense that there exists LS > 0 such that for µ-a.e. x ∈ X, 1/2



dZ Sx (ρ, ϕ), Sx (ρ′ , ϕ′ )) ≤ LS W 2 (ρ, ρ′ ) + d2C (ϕ, ϕ′ )

,

for all ρ, ρ′ ∈ Xµ , and all ϕ, ϕ′ ∈ Cµ i.e. for µ-compatible contexts. 8.1.2

Second structural assumption: Lipschitz map

Consider a mapped energy E(ρ, ϕ) :=

Z X

Eex Hx (ρx , Sx (ρ, ϕ)), Sx (ρ, ϕ) dµ(x). 

Assumption 8.2 (Lipschitz map). The contextualized geometric representation map is Lipschitz, in the sense that there exists LH > 0 such that for µ-a.e. x ∈ X, 

1/2

dZ Hx (ρx , z), Hx (ρ′x , z ′ )) ≤ LH W22 (ρx , ρ′x ) + d2Z (z, z ′ )

59

,

for all ρx , ρ′x ∈ P2 (Y ), and all z, z ′ ∈ Z. 8.1.3

Contextual assumption: the sleep-mode assumption

Consider a Tan–HWG energy E(ρ, ϕ) :=

Z X

Ex ρx , Sx (ρ, ϕ) dµ(x). 

We consider that when the system is left in free evolution, with no external stimulation, it presents an autonomous regime, where activations (contexts) depend solely on internal parameters, mainly the memory field. The discrete evolution of the context can be written n n n+1 ϕn+1 := ΦE , τ τ (ρτ |ϕτ ) + ϵτ n+1 represents externally driven where the operator ΦE τ describes the purely internal evolution, and ϵτ activations (e.g. sensory driven activations).

The sleep-mode assumption formalizes a regime in which internally driven, stable neural dynamics dominate external perturbations (ϵn+1 ). τ Assumption 8.3 (The sleep-mode assumption). There exists LC > 0 such that, for every timestep τ > 0 and every discrete Tan–HWG sequence (ρnτ )n≥0 associated with E, and with a µ-compatible context sequence (ϕnτ )n≥0 ∈ CµN , we have dC (ϕnτ , ϕn+1 ) ≤ LC W(ρnτ , ρn+1 ), τ τ

ρn+1 ∈ TτE (ρnτ | ϕn ). τ

We also say that ϕ presents quasi-stationary external context. This is a coupling constraint between (ρnτ )n and (ϕnτ )n . This assumption formalizes the idea that consolidation occurs when external perturbations are weak relative to intrinsic neural dynamics. In this regime, the external context perceived by the network is quasi-stationary, so that sensory-driven fluctuations are minimal. As a consequence, neuronal activations evolve in a calm, weakly perturbed dynamical state. The network operates in a nearly unconstrained “free evolution” mode, where the intrinsic flow of neural activity dominates over externally imposed variations. Remark 8.4. If the context depends solely on the memory field, i.e. ϕ = ϕ(ρ), then the sleep-mode assumption reduces to a Lipschitz regularity condition on ϕ. This corresponds to a fully memory driven free evolution regime: the system is isolated from external stimuli (e.g. sleep), and internal contextual states evolve smoothly with the memory. In the general case, ϕ may include external influences, and other internal influences, such that the expression using the operator ΦE τ is more general. In any case, the assumption enforces a quasi-free evolution, independently of the dependencies of ϕ: external variations must be sufficiently regular so as not to disrupt the internal Tan–HWG dynamics.

8.2 8.2.1

Continuous-time limit of generalized quadratic Tan–HWG dynamics Control of the freezing error propagation

Intuition. When the energy is Tan–HWG, the frozen energy is AGS compatible, which ensures that it is well behaved for a continuous-time limit, but only locally, relative to a fixed frozen 60

signal. Structural Lipschitz assumptions and the sleep-mode assumption will ensure that the global signal’s evolution is sufficiently smooth. From there, the key is to control the freezing error, and the propagation of this error. Indeed, when τ → 0, the number of steps increases relative to a given time horizon T (nτ ≤ T ). Therefore, we start by establishing a result allowing us to control the freezing error. Lemma 8.5 (Control of the freezing error). Assume the generalized quadratic energy α E(ρ, ϕ) := 2

Z X

d2Z Hx (ρx , Sx (ρ, ϕ)), Sx (ρ, ϕ) dµ(x) 

is a Tan–HWG energy. Under the regularity assumptions: • Lipschitz signal (Assumption 8.1): 

dZ Sx (ρ, ϕ), Sx (ρ′ , ϕ′ )) ≤ LS W 2 (ρ, ρ′ ) + d2C (ϕ, ϕ′ )

1/2

,

µ-a.e.,

• Lipschitz map (Assumption 8.2): 1/2



dZ Hx (ρx , z), Hx (ρ′x , z ′ )) ≤ LH W22 (ρx , ρ′x ) + d2Z (z, z ′ )

,

µ-a.e.,

there exists L > 0 such that, for every timestep τ > 0 and every discrete Tan–HWG sequence (ρnτ )n≥0 associated with E, and with a µ-compatible context sequence (ϕnτ )n≥0 ∈ CµN , under sleep-mode assumption • sleep-mode (Assumption 8.3): dC (ϕnτ , ϕn+1 ) ≤ LC W(ρnτ , ρn+1 ), τ τ we have the freezing error control inequality E(ρn+1 , ϕn+1 ) − Eϕnτ ,ρnτ (ρn+1 ) ≤ τ τ τ

1 2 n+1 n W (ρτ , ρτ ) + α L2 τ (E(ρn+1 , ϕn+1 ) + E(ρnτ , ϕnτ )). τ τ 4τ

Moreover, we have the pre-telescopic inequality E(ρn+1 , ϕn+1 )+ τ τ

1 2 n+1 n W (ρτ , ρτ ) ≤ E(ρnτ , ϕnτ ) + α L2 τ (E(ρn+1 , ϕn+1 ) + E(ρnτ , ϕnτ )). τ τ 4τ

and, for τ small enough, we have the Grönwall inequality E(ρn+1 , ϕn+1 ) ≤ (1 + 4αL2 τ )E(ρnτ , ϕnτ ). τ τ Proof. Fix τ > 0. To simplify notations, we write ρn := ρnτ , ϕn := ϕnτ , E n := E(ρn , ϕn ) and S n := S(ρn , ϕn ). The frozen energy at step n is Eϕn ,ρn (ρ) :=

α 2

Z X

d2Z Hx (ρx , Sxn ), Sxn dµ(x). 

61

Step 1: Error control. We estimate the difference between the true and frozen energies at ρn+1 : ∆n,n+1 := E n+1 − Eϕn ,ρn (ρn+1 ). We have  α  2 n+1 n+1  2 n+1 n n dµ(x) dZ Hx (ρn+1 , S ), S − d H (ρ , S ), S x x x x x Z x x 2 X Z  α n+1 n n  ≤ dZ Hx (ρn+1 ), Sxn+1 +dZ Hx (ρn+1 x , Sx x , Sx ), Sx δn,n+1 dµ(x), 2 X Z

|∆n,n+1 | =

n+1 ), S n+1 −d H (ρn+1 , S n ), S n . By triangle inequalities we have with δn,n+1 := dZ Hx (ρn+1 x x Z x , Sx x x x





n+1 n n+1 ), Sxn+1 −dZ Hx (ρn+1 δn,n+1 ≤ dZ Hx (ρn+1 x , Sx x , Sx ), Sx





n n+1 n n −dZ Hx (ρn+1 + dZ Hx (ρn+1 x , Sx ), Sx x , Sx ), Sx



n+1 n n+1 ≤ dZ Hx (ρn+1 ), Hx (ρn+1 , Sxn x , Sx x , Sx ) +dZ Sx







Then, by Assumptions 8.1 and 8.2 δn,n+1 ≤ LH dZ (Sxn+1 , Sxn )+dZ Sxn+1 , Sxn

(by Assumption 8.2)





≤ (1 + LH )LS W 2 (ρn+1 , ρn ) + d2C (ϕn+1 , ϕn )

1/2

(by Assumption 8.1)

1/2



≤ (1 + LH )LS W 2 (ρn+1 , ρn ) + L2C W 2 (ρn+1 , ρn )

(by Assumption 8.3)

q

≤ (1 + LH )LS 1 + L2C W(ρn+1 , ρn ) ≤ L̃ W(ρn+1 , ρn ), q

with L̃ := LS (1 + LH ) 1 + L2C . Thus, we obtain |∆n,n+1 | ≤

 α n+1 n n  dZ Hx (ρn+1 ), Sxn+1 +dZ Hx (ρn+1 L̃ W(ρn+1 , ρn ) x , Sx x , Sx ), Sx dµ(x), 2 X Z

By Cauchy–Schwarz, α |∆n,n+1 | ≤ L̃W(ρn+1, ρn ) 2

Z

α ≤ LW(ρn+1, ρn ) 2

X

Z X

 n+1 dZ Hx (ρn+1 ), Sxn+1 +dZ x , Sx

1

2 1 n n 2 Hx (ρn+1 dµ(x) µ(X)2 x , Sx ), Sx

n+1 n n dZ Hx (ρn+1 ), Sxn+1 +dZ Hx (ρn+1 x , Sx x , Sx ), Sx



2

1

2

dµ(x) ,

1 2

with L := L̃ µ(X) . Using (a + b)2 ≤ 2(a2 + b2 ), we obtain |∆n,n+1 | ≤

αL

q

E n+1 + Eϕn ,ρn (ρn+1 ) W(ρn+1, ρn ).

Moreover, by the discrete energy inequality, Eϕn ,ρn (ρn+1 ) ≤ E n . Hence √

p

E n+1 + E n W(ρn+1 , ρn ). √ √ 1 2 Using Young’s inequality ab ≤ 4τ a + τ b2 with a = W(ρn+1 , ρn ), b = α L E n+1 + E n , we get |∆n,n+1 | ≤

αL

p

αL

E n+1 + E n W(ρn+1 , ρn ) ≤

1 2 n+1 n W (ρ , ρ ) + α L2 τ (E n+1 + E n ). 4τ 62

Thus E n+1 − Eϕn ,ρn (ρn+1 ) ≤

1 2 n+1 n W (ρ , ρ ) + α L2 τ (E n+1 + E n ), 4τ

which proves the first claim. Step 2: Pre-telescopic inequality. Adding with the discrete energy inequality Eϕn ,ρn (ρn+1 ) + gives

1 2 n+1 n W (ρ , ρ ) ≤ Eϕn ,ρn (ρn ) = E(ρn , ϕn ) = E n , 2τ

1 2 n+1 n W (ρ , ρ ) ≤ E n + α L2 τ (E n+1 + E n ). 4τ

E n+1 + which proves the second claim.

Step 3: Grönwall inequality. We rewrite the previous inequality (1 − α L2 τ )E n+1 +

1 2 n+1 n W (ρ , ρ ) ≤ (1 + α L2 τ )E n . 4τ

For τ small enough, α L2 τ ≤ 21 , and we can write E n+1 +

1 1 + αL2 τ n 2 n+1 n W (ρ , ρ ) ≤ E . 4τ (1 − αL2 τ ) 1 − αL2 τ

2

1+αL τ 1 2 2 Since 1−αL 2 τ ≤ 1 + 4αL τ whenever αL τ ≤ 2 , we obtain

E n+1 +

1 W 2 (ρn+1 , ρn ) ≤ (1 + 4αL2 τ )E n . 4τ (1 − αL2 τ )

Then,

E n+1 ≤ (1 + 4αL2 τ )E n ,

which proves the final claim. This result allows us to control the propagation of the second moments: Lemma 8.6 (Stability of the dynamics). Let E be a Tan–HWG GQE, and assume Assumptions 8.1 and 8.2 are true. Let τ > 0, and (ρnτ )n≥0 be the discrete Tan–HWG sequence associated with E, and with a µ-compatible context sequence (ϕnτ )n≥0 ∈ CµN , under sleep-mode assumption 8.3, with initial conditions ρ0τ = ρ0 ∈ Xµ and ϕ0τ = ϕ0 ∈ Cµ independent of τ . Then, for every T > 0, there exists a constant CT > 0, independent of τ , such that for all n ∈ N with nτ ≤ T , E(ρnτ , ϕnτ ) ≤ CT

(Energy stability),

(9)

1 2 k+1 k W (ρτ , ρτ ) ≤ CT τ k=0

(Action stability),

(10)

(Moment stability),

(11)

n−1 X

Z Z X

Y

dY (y, y0 )2 dρnτ,x (y) dµ(x) ≤ CT

for some (hence any) y0 ∈ Y , and for τ small enough.

63

Proof. We use the same notations as in the previous proof. Fix T > 0. Step 1: Energy stability. By the Grönwall inequality of Lemma 8.5, we have 2

E n+1 ≤ (1 + 4αL2 τ )E n ≤ e4αL τ E n , which gives, for all nτ ≤ T , 2

(1)

2

(1)

with CT := e4α L T E 0 independent of τ.

E n ≤ e4L nτ E 0 ≤ CT ,

Step 2: Action stability. Summing the pre-telescopic inequality of Lemma 8.5 gives En +

n−1 X 1 X 1 n−1 W 2 (ρk+1 , ρk ) ≤ E 0 + α L2 τ (E k+1 + E k ) 4 k=0 τ k=0 2

≤ E 0 + 2α L2 nτ e4α L nτ E 0 . Hence, for all nτ ≤ T , n−1 X

1 2 k+1 k (2) W (ρ , ρ ) ≤ CT τ k=0

2

(2)

with CT := 4E 0 (1 + 2α L2 T e4α L T ) independent of τ.

Step 3: Propagation of second moments. Fix y0 ∈ Y and define M (ρ) :=

Z Z X

Y

dY (y, y0 )2 dρx (y) dµ(x).

n n+1 and ρn . Since Fix n ∈ N and x ∈ X. Let πx ∈ Π(ρn+1 x , ρx ) be an optimal coupling between ρx x

dY (y, y0 )2 ≤ dY (y ′ , y0 )2 + dY (y, y ′ )2 + 2dY (y ′ , y0 )dY (y, y ′ )   1 ′ 2 ≤ (1 + τ ) dY (y , y0 ) + 1 + dY (y, y ′ )2 , (using Young’s inequality) τ we obtain 1 dY (y , y0 ) dπx (y, y ) + 1+ dY (y, y0 ) dπx (y, y ) ≤ (1+τ ) τ Y×Y Y×Y

Z

Z

2



2

Z Y×Y

dY (y, y ′ )2 dπx (y, y ′ ).

Which, by definition of the marginals of πx and the Wasserstein distance, yields Z Y

dY (y, y0 )

2

dρn+1 x (y) ≤ (1+τ )

Z

Y

dY (y , y0 )

2

dρnx (y ′ ) +



1 n W22 (ρn+1 1+ x , ρx ) τ 

Integrating over X gives M (ρ

n+1

1 ) ≤ (1+τ ) M (ρ ) + 1+ W 2 (ρn+1 , ρn ). τ 

n



Summing this inequality and using the action stability result of Step 2 yields, for nτ ≤ T , n−1 X

1 M (ρ ) ≤ M (ρ ) + τ M (ρ ) + 1+ τ k=0 n

0

≤ M (ρ0 ) + τ

n−1 X

k



 n−1 X

W 2 (ρk+1 , ρk )

k=0 (2)

M (ρk ) + (τ + 1)CT

k=0



(2)

≤ M (ρ0 ) + 2CT



+

n−1 X k=0

64

M (ρk )τ,

for τ ≤ 1.

By Grönwall lemma we obtain 

(2)

M (ρn ) ≤ M (ρ0 ) + 2CT (1)

(2)



(3)

enτ ≤ CT

(3)



(2)

with CT := M (ρ0 ) + 2CT



eT independent of τ.

(3)

Setting CT = max{CT , CT , CT } proves the claim. 8.2.2

Existence of subsequential limits

We now show the existence of subsequential limits to the Tan–HWG scheme when τ → 0, in the case of Tan–HWG GQE, under the three regularity assumptions of Section 8.1. The moment stability result of the previous lemma is key to ensure the existence of such limits. Moreover, we show that the limit curve is locally absolutely continuous with respect to the Wasserstein metric and satisfies a perturbed energy–dissipation inequality. This provides a rigorous continuous-time formulation of the dynamics. Theorem 8.7 (Continuous dynamics for Tan–HWG GQE). Assume Γ is a finite set. Let E be a Tan–HWG GQE, and assume Assumptions 8.1 and 8.2 are true. Let (ρnτ )n≥0 be the Tan–HWG minimizing-movement scheme with timestep τ > 0, associated with µ-compatible contexts (ϕnτ )n≥0 , under sleep-mode assumption 8.3, with fixed initial conditions ϕ0τ = ϕ0 ∈ Cµ , and ρ0τ = ρ0 ∈ Xµ independent of τ . ⌊t/τ ⌋

Let ρτ be the piecewise-constant interpolation ρτ (t) := ρτ ⌊t/τ ⌋ interpolation ϕτ (t) := ϕτ

, and ϕτ be the piecewise-constant

Then, as τ → 0, the family (ρτ )τ admits limit curves along subsequences, with associated limit curves of (ϕτ )τ . Moreover, for any limit curves ρ : [0, ∞) → Xµ and ϕ : [0, ∞) → Cµ : • ρ ∈ ACloc ([0, ∞); (Xµ , W)) (i.e. ρ locally absolutely continuous); • for all t ≥ 0, we have the perturbed energy–dissipation inequality E(ρ(t), ϕ(t)) +

1 4

Z t 0

|ρ̇|2 (s) ds ≤ E(ρ0 , ϕ0 ) + 2αL2

Z t 0

E(ρ(s), ϕ(s)) ds;

• in particular, t 7→ E(ρ(t), ϕ(t)) is locally absolutely continuous and obeys the Grönwall bound 2 E(ρ(t), ϕ(t)) ≤ E(ρ0 , ϕ0 ) e2αL t . Proof. Step 1: existence of limit curve by equicontinuity and compactness. Fix T > 0. For s, t ∈ [0, T ] with s < t, let n = ⌊s/τ ⌋, and m = ⌊t/τ ⌋. By the triangle inequality, Cauchy–Schwarz and

65

stability equation (10) of Lemma 8.6, n W(ρτ (t), ρτ (s)) = W(ρm τ , ρτ ) ≤

m−1 X

W(ρk+1 , ρkτ ) τ

k=n

m−1 X

!1/2

m−1 X

1 2 k+1 k W (ρτ , ρτ ) τ k=n

τ

k=n

≤ ((m − n)τ ) ≤

p

1/2

!1/2

!1/2

m−1 X

1 2 k+1 k W (ρτ , ρτ ) τ k=n

2CT |t − s|1/2

(for τ small enough).

Thus (ρτ )τ >0 is equicontinuous in (Xµ , W) on [0, T ]. The uniform second-moment bound (11) of Lemma 8.6 and Prokhorov’s theorem yield tightness in each fibre, hence relative compactness in the narrow topology on Xµ . By a standard Arzelà–Ascoli/diagonal argument, there exists a subsequence τk ↓ 0 and a limit curve ρ : [0, T ] → Xµ such that ρτk (t) ⇀ ρ(t)

narrowly in Xµ

for all t ∈ [0, T ], and W(ρτk (t), ρ(t)) → 0. Extending this construction to [0, ∞) by a diagonal argument yields the claimed convergence. Moreover, the uniform second-moment bound (11) implies uniform integrability of dY (·, y0 )2 ; hence, along the extracted subsequence, narrow convergence of (ρτk (t))x together with convergence of second moments yields W2 (ρτk (t))x , ρx (t) → 0 for µ-a.e. x, and therefore W(ρτk (t), ρ(t)) → 0 by dominated convergence. Moreover, by sleep-mode assumption 8.3, dC (ϕτ (t), ϕτ (s)) ≤ LC

m−1 X

W(ρk+1 , ρkτ ) ≤ LC τ

2CT |t − s|1/2

(for τ small enough),

p

k=n

and (ϕτ )τ >0 is also equicontinuous in (Cµ , dC ) on [0, T ]. We also have ϕτ (t) − ϕ0

2

= d2C (ϕτ (t), ϕ0 ) ≤ L2C 2CT t,

t ∈ [0, T ],

so that ∥ϕτ ∥ is uniformly bounded on [0, T ]. Since Γ is a finite set, we have uniform tightness of total variations of (ϕτk )τk in each fiber. Hence, we can apply Prokhorov’s theorem extension to the family of complex finite measures (ϕτk )τk . Following same arguments as above for (ρτ )τ , applied to (ϕτk )τk , we can extract a subsequence τk′ from τk to obtain both limit curves ρ and ϕ as claimed. Step 2: absolute continuity and metric derivative. The discrete action bound (10) implies Z T τ

W(ρτ (t), ρτ (t − τ )) τ

2

⌊T /τ ⌋−1

dt ≤

X

τ

k=0

W(ρk+1 , ρkτ ) τ τ

!2

≤ CT .

Passing to the limit along τk ↓ 0 and using the lower semicontinuity of the metric derivative (see e.g. [1]), we obtain Z T 0

|ρ̇|2 (t) dt ≤ CT , 66

where |ρ̇|(t) is the metric derivative of ρ relative to W. So ρ ∈ AC([0, T ]; (Xµ , W)) for all T > 0, hence ρ ∈ ACloc ([0, ∞); Xµ ). Step 3: energy-dissipation inequality. By Lemma 8.5: E(ρn+1 , ϕn+1 )+ τ τ

1 2 n+1 n W (ρτ , ρτ ) ≤ E(ρnτ , ϕnτ ) + αL2 τ (E(ρnτ , ϕnτ ) + (1 + 4αL2 τ )E(ρnτ , ϕnτ )) 4τ ≤ (1 + 2αL2 τ + 4α2 L4 τ 2 )E(ρnτ , ϕnτ ).

Summing over n from 0 to ⌊t/τ ⌋ − 1 E(ρτ (t), ϕτ (t)) +

1 4

⌊t/τ ⌋−1

X k=0

⌊t/τ ⌋−1 X 1 2 k+1 k W (ρτ , ρτ ) ≤ E(ρ0 , ϕ0 ) + 2αL2 (1 + 2αL2 τ ) E(ρkτ , ϕkτ )τ. τ k=0

Passing to the limit along τk ↓ 0, using the lower semicontinuity of the metric derivative and the convergence ρτk (t) → ρ(t), we obtain the following continuous energy-dissipation inequality (perturbed EDI) 1 E(ρ(t), ϕ(t)) + 4

Z t 0

|ρ̇| (s) ds ≤ E(ρ , ϕ ) + 2αL 2

0

0

2

Z t 0

E(ρ(s), ϕ(s)) ds,

for all t ≥ 0. Step 4: Grönwall bound. By Grönwall’s lemma we infer 2

E(ρ(t), ϕ(t)) ≤ E(ρ0 , ϕ0 )e2αL t for all t ≥ 0. Remark 8.8 (Structural non-deterministic dynamics). Since geodesic convexity of E is not established, the limiting Tan–HWG dynamics need not be unique. In the AGS framework, uniqueness of the gradient flow is guaranteed only for geodesically λ-convex energies with λ ≥ 0. When convexity fails, multiple curves of maximal slope may coexist, even under identical initial conditions. This implies that an agent may follow different limiting trajectories starting from the same memory state. Such structural non-determinism is not due to any stochastic component of the model: a genuine Wasserstein gradient flow would be fully deterministic in the absence of external randomness. Here, the multiplicity of limit curves stems from the sequential freezing of the global signal at each Tan–HWG step. This phenomenon is consistent with neuroscientific observations: identical stimuli can elicit distinct neural trajectories depending on the internal state of the network [2, 6, 4, 23, 27]. In this sense, the framework naturally captures the intrinsic variability of cortical dynamics.

Discussion: Geometry of Memory The framework developed in this work proposes a geometric reinterpretation of synaptic plasticity, neural computation, and memory consolidation. By embedding memory states in Wasserstein spaces and by introducing contextualized geometric representation maps, we obtain a unified variational formulation of Hebbian learning that naturally incorporates stochasticity, multi-semanticity, oscillatory structure, and multi-timescale dynamics. This section discusses the conceptual implications of the model, its relationship to biological and artificial neural systems, and its limitations. 67

Geometry as a flexible and unifying cognitive language. A central contribution of this work is the shift from fixed Euclidean embeddings to geometric memories represented as probability measures. This representation generalises existing models: classical neural networks appear as degenerate cases in which internal distributions collapse to Dirac masses. The formulation allows synapses to encode not only scalar weights but full distributions, capturing uncertainty, multimodality, and latent semantic structure. Observable synaptic weights arise as projections of these internal states, and their geometry is governed by Wasserstein transport. This perspective provides a principled explanation for several phenomena typically introduced heuristically in machine learning, such as attention mechanisms, pruning, sparsity, synchronization and representational drift. A duality between rich internal dynamics and flat observable projections. The introduction of mapped energies generalizes the idea that complex architectures can be decomposed into elementary geometric maps whose regularity properties determine the global behavior of the system. This decomposition clarifies the role of activation functions, which no longer serve as the primary source of non-linearity but are instead practical building blocks for geometrically compatible projections of a fundamentally curved latent space. Hence, the framework offers a principled geometric account of cognitive computation, suggesting that cognition may be more naturally described in geometric terms (distances, curvatures, geodesics) rather than algebraic ones (vectors, scalars, matrices). In particular, algebraic descriptions of neural computation may represent flattened projections of richer geometric dynamics. The quadratic–affine regime as an anchor point. The quadratic–affine regime plays the role of an anchor point where the geometric and classical views coincide. When the internal energy is quadratic, the dynamics projects to geodesics, and when the projection map is affine, internal geodesics project to affine observable dynamics. This regime recovers classical update rules such as exponential moving averages, mirror descent, and state-independent contraction factors, and admits a clean continuous-time limit. As such, it provides a consistency check for the framework: in the geometrically ideal case, the model reduces to well-known computational schemes. Deviations from affine behavior can then be interpreted as signatures of underlying geometric curvature. Classical indicators of nonlinearity (e.g. discrete second differences, orthogonal update components, or angles between successive gradients) acquire a geometric interpretation as measures of deviation from geodesic flatness. Mathematical constraints as organisational principles. A key insight of the framework is the rigidity of geodesic affinity in Wasserstein spaces. Contextualized maps are rarely geodesically affine, which prevents the associated energies from being Tan–HWG. This geometric constraint naturally supports a multi-timescale coherence mechanism: fast synaptic dynamics operate under fixed structural weights, while slow structural plasticity requires coarse-grained or quasi-stationary activations. This provides a mathematical explanation for the long-standing observation that memory consolidation occurs preferentially during sleep or quiet wakefulness, when external perturbations are minimal and neural activity is dominated by internally generated patterns. Moreover, non-uniqueness of non-convex dynamics (discrete or continuous) supports structural nondeterminism, again in coherence with neuroscientific observations. Hence, geometry acts not merely as a descriptive tool, but as a generative principle shaping the structure of learning dynamics.

68

Limitations and computational challenges. Several aspects of the framework remain interpretative and require empirical validation. The biological relevance of multi-semantic embedding, robustness, and stability in generalized neural networks remains to be demonstrated in practice. Computationally, Wasserstein updates are expensive, and fast closed-form solutions exist only in restricted cases. While entropic regularisation and Sinkhorn-type methods [7] accelerate the computation of optimal transport, the scalability of Tan–HWG dynamics for large-scale architectures remains an open question. Moreover, nothing guarantees that generalized networks constructed within this framework will match the performance of modern artificial neural networks, which currently excel across a wide range of tasks. Future work may explore: • Dynamics: explicit gradient-flow equations for the full synaptic distribution and a systematic classification of equilibria or convergence rates; the study of attractors, bifurcations, cycles, stability regimes, and long-term behavior of Tan–HWG flows may reveal new forms of memory organisation; • Representations: the impact of internal-space topology (e.g. different CAT(0) structures or metric trees, including forbidden connections, directed edges, or structured sparsity) on learning dynamics; • Structures: pruning and self-organised sparsity-natural consequences of the simplex constraint—suggest promising directions for performance optimisation; • Cognition: connection between geometric dynamics and cognitive phenomena, such as consolidation, abstraction, or compositionality; • Applications: continual learning, meta-learning, or neuromorphic computing.

Conclusion The Tan–HWG framework reframes Hebbian plasticity as a geometric process rather than an algebraic update of synaptic coefficients. By lifting memory states to Wasserstein spaces and expressing plasticity as a minimizing movement, learning appears as the evolution of probability measures constrained by curvature, geodesic rigidity, and the regularity of contextualized representation maps. This perspective reveals that many qualitative features of learning—stability, competition, consolidation, synchronization, sparsity—arise not from additional mechanisms but from the geometric structure in which memory evolves. This shift challenges the traditional separation between “weights” and “computation”. In the Tan– HWG view, observable weights are projections of a richer internal geometry, and classical update rules emerge as flat limits of curved dynamics. The rigidity of geodesic affinity, the emergence of multi-timescale coherence, and the structural non-determinism induced by frozen signals highlight a deeper principle: learning systems are shaped less by the rules they implement than by the geometry that makes these rules possible. The framework also brings to light a structural tension inherent to synaptic plasticity. Internal dynamics must remain geometrically coherent, while observable behavior must remain simple, stable, and low-dimensional. The coexistence of curved internal trajectories with affine observable recursions is not an artefact of the model but a consequence of compressing rich internal representations into actionable outputs. This duality provides a geometric explanation for the stability of 69

observable behavior despite the complexity of internal dynamics. Under quasi-stationary contexts, the existence of continuous-time limit curves shows that consolidation corresponds to the regime in which the geometry can unfold without external disruption. In this sense, memory is the trace left by a trajectory in a curved space: its organisation reflects both the information acquired and the geometric constraints under which learning takes place. Rather than closing the theory, the framework delineates a landscape in which learning, representation, and consolidation are unified by a single geometric principle. It provides a foundation for analysing how internal geometries shape the dynamics of plasticity and how observable behavior emerges from the curvature of latent spaces.

70

A

Self consistency and stability of fixed points in Tan-HWG dynamics

This appendix gathers the arguments underlying the fixed-point structure and stability properties of Tan–HWG dynamics. It is self-contained and independent of the examples presented in the main text.

A.1

Fixed Points and Self-Consistency

A distinctive feature of Tan–HWG dynamics is that the global signal driving the evolution of the local states is itself induced by the current configuration of the memory field. As a consequence, equilibrium configurations are characterized by a self-consistency condition linking the local states and the global signal they collectively generate. We formalize this notion below and relate fixed points of the dynamics to Wasserstein critical points of the energy. Proposition A.1 (Fixed points as critical points). Let E be a Tan–HWG energy, and fix a E denote the τ -Tan–HWG update operator. Let stationary context ϕ ∈ M(Γ)X . Let τ > 0, Tϕ,τ n ∗ (ρ ) be the associated dynamics. Then ρ is a fixed point of the dynamics if and only if it is a Wasserstein critical point of E(·, ϕ): ρ∗ ∈ Tϕ,τ (ρ∗ ) (fixed point)

0 ∈ ∂W E(ρ∗ , ϕ),

⇐⇒

where ∂W denotes the Wasserstein subdifferential. Remarkably, the characterization of fixed points is independent of the time step τ > 0: the parameter τ affects the transient dynamics but not the location of equilibrium configurations, which are solely determined by the critical points of the energy. In this sense, τ plays a role analogous to a learning rate in machine learning. This result shows that equilibrium configurations of the Tan–HWG dynamics are precisely the self-consistent states where the Wasserstein subdifferential vanishes.

A.2

Stability of fixed points

We now consider the Brenier setting: quadratic cost on Rd with absolutely continuous measures. In this case, the Wasserstein subdifferential reduces to a singleton, and the condition 0 ∈ ∂W E(ρ, ϕ) is equivalent to ∇W E(ρ, ϕ) = 0. Linearising the dynamics around a fixed point ρ∗ in the Otto calculus yields ∂t δρ(t) = −HE (ρ∗ , ϕ) δρ(t), where HE (ρ⋆ , ϕ) denotes the Wasserstein Hessian of E(·, ϕ) at ρ⋆ . The solution is δρ(t) = exp −t HE (ρ∗ , ϕ) δρ(0). 

This yields the following lemma:

71

E at a fixed point ρ∗ is given by Lemma A.2. The linearisation of Tϕ,τ E DTϕ,τ (ρ∗ ) = exp −τ HE (ρ∗ , ϕ) .



Proposition A.3 (Spectral stability). Let E be a Tan–HWG energy, ϕ a fixed context, and ρ∗ a fixed point. Let (µi )i be the eigenvalues of the Wasserstein Hessian HE (ρ∗ , ϕ) and (λi )i those of the Jacobian of the time–1 map of the gradient flow (or of the 1-Tan–HWG dynamics) at ρ∗ . Then λi = e−µi for all i. In particular: µi > 0 ⇒ |λi | < 1,

µi = 0 ⇒ |λi | = 1,

µi < 0 ⇒ |λi | > 1.

The choice τ = 1 is made for simplicity, since the stability properties are independent of the time step. This gives a simple classification of the dynamics: • µi > 0 ⇐⇒ |λi | < 1: stable directions (local attractor); • µi = 0 ⇐⇒ |λi | = 1: neutral directions (flat manifolds or symmetries); • µi < 0 ⇐⇒ |λi | > 1: unstable directions (saddles, bifurcations). Corollary A.4 (Geometric consolidation in stationary context). Let E be a Tan–HWG energy and ϕ ∈ M(Γ)X . If ρ∗ is a fixed point such that the Wasserstein Hessian HE (ρ∗ , ϕ) is strictly positive definite, then ρ∗ is a locally asymptotically stable fixed point of the Tan–HWG dynamics. In particular, for all initial conditions ρ0 sufficiently close to ρ∗ , the iterates E ρn+1 ∈ Tϕ,1 (ρn )

converge to ρ∗ as n → ∞.

A.3

Interpretation and dynamical consequences

Remark A.5 (Cognitive interpretation and model-based prediction). In stationary contexts, Hebbian plasticity can be interpreted as a geometric consolidation process, whereby the dynamics converge towards local minima of the energy E(·, ϕ) and stabilise the associated internal geometry. E (ρ∗ ) = Although the set of equilibrium configurations is independent of τ > 0, the linearisation DTϕ,τ ∗ exp(−τ HE (ρ , ϕ)) shows that τ controls the contraction rate along stable directions. This suggests a possible link between the timescale of Hebbian updates and the temporal organisation of cognitive processes.

Computational constraints and effective equilibrium selection. The dynamics may effectively select different equilibria depending on the available time horizon and numerical stability. Very small τ may lead to slow convergence, potentially too slow to be achieved within a finite time window, while excessively large τ may induce oscillatory or unstable behavior around otherwise stable equilibria. Thus, τ controls which equilibria are effectively reachable and consolidated. If discrete updates are gated by oscillatory cycles, the model suggests that different gating frequencies modulate consolidation rates and the selection of computationally compatible equilibria, rather than 72

the location of the equilibria themselves. Nested updates. Multi-scale updates involving several time steps τ , such as τ1 ≪ τ2 , may increase the speed and robustness of convergence compared to the use of a single time step. The family E } {Tϕ,τ τ >0 can be viewed as a temporal decomposition of the Hebbian update operator, analogous to a spectral decomposition. Remark A.6 (Bifurcations and phase transitions). As the context ϕ varies, the structure of the energy landscape E(·, ϕ) may undergo qualitative changes. Bifurcations occur when eigenvalues of the Wasserstein Hessian cross zero, leading to the loss of stability of equilibria and the emergence of new stationary states. These events can be interpreted as geometric phase transitions in the space of representations. While the location of these transitions is determined by the energy landscape, the time step τ controls their dynamical manifestation, with critical slowing down near bifurcation points and enhanced sensitivity to perturbations. Remark A.7 (Effective dimensionality of representations). The spectral structure of the Wasserstein Hessian HE (ρ∗ , ϕ) provides a natural notion of effective dimensionality of the representation around an equilibrium. Directions associated with large positive eigenvalues are rapidly contracted and effectively frozen, while directions corresponding to small or vanishing eigenvalues remain dynamically accessible. As a result, even when the ambient representation space is high-dimensional, as commonly encountered in deep learning models, the dynamics effectively evolves on a lower-dimensional manifold whose dimension depends on the local geometry of the energy landscape. Changes in this effective dimensionality naturally occur at bifurcation points, when eigenvalues cross zero.

73

References [1] Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré. Gradient Flows in Metric Spaces and in the Space of Probability Measures. Birkhäuser, 2008. [2] Amos Arieli, Alexander Sterkin, Amiram Grinvald, and Ad Aertsen. Dynamics of ongoing activity: explanation of the large variability in evoked cortical responses. Science, 273(5283): 1868–1871, 1996. [3] Martin R. Bridson and André Haefliger. Metric Spaces of Non-Positive Curvature, volume 319 of Grundlehren der mathematischen Wissenschaften. Springer, 1999. [4] Dean V Buonomano and Wolfgang Maass. State-dependent computations: spatiotemporal processing in cortical networks. Nature Reviews Neuroscience, 10(2):113–125, 2009. [5] György Buzsáki. Rhythms of the Brain. Oxford University Press, 2006. [6] Mark M Churchland, Byron M Yu, John P Cunningham, Leo P Sugrue, Marlene R Cohen, Greg S Corrado, William T Newsome, Andrew M Clark, Paymon Hosseini, Benjamin B Scott, et al. Stimulus onset quenches neural variability: a widespread cortical phenomenon. Nature Neuroscience, 13(3):369–378, 2010. [7] Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transportation distances. NeurIPS, 2013. [8] Peter Dayan and L. F. Abbott. Theoretical Neuroscience: Computational and Mathematical Modeling of Neural Systems. MIT Press, 2001. [9] Morris H DeGroot. Reaching a consensus. Journal of the American Statistical Association, 69 (345):118–121, 1974. [10] Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, and Yaron Lipman. Discrete flow matching. NeurIPS, 2024. [11] Wulfram Gerstner, Werner M. Kistler, Richard Naud, and Liam Paninski. Neuronal Dynamics: From Single Neurons to Networks and Models of Cognition. Cambridge University Press, 2014. [12] Donald O. Hebb. The Organization of Behavior. Wiley, 1949. [13] Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré, and Max Welling. Argmax flows and multinomial diffusion: Learning categorical distributions. In NeurIPS, 2021. [14] John J. Hopfield. Neural networks and physical systems with emergent collective computational abilities. Proceedings of the National Academy of Sciences, 79(8):2554–2558, 1982. [15] A. Jadbabaie, J. Lin, and A.S. Morse. Coordination of groups of mobile autonomous agents using nearest neighbor rules. In Proceedings of the 41st IEEE Conference on Decision and Control, 2002., volume 3, pages 2953–2958 vol.3, 2002. doi: 10.1109/CDC.2002.1184304. [16] Richard Jordan, David Kinderlehrer, and Felix Otto. The variational formulation of the fokker– planck equation. SIAM Journal on Mathematical Analysis, 29(1):1–17, 1998. [17] Yoshiki Kuramoto. Self-entrainment of a population of coupled non-linear oscillators. In H. Araki, editor, International Symposium on Mathematical Problems in Theoretical Physics, volume 39 of Lecture Notes in Physics, pages 420–422. Springer, 1975.

74

[18] Ferster D. Lampl I, Reichova I. Synchronous membrane potential fluctuations in neurons of the cat visual cortex. Neuron, 22(2):361–374, 1999. [19] Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. ICLR, 2023. [20] John E. Lisman and Ole Jensen. The theta-gamma neural code. Neuron, 77(6):1002–1016, 2013. [21] Xuehai Liu et al. Sfm: Stochastic flow matching for discrete and categorical data. NeurIPS, 2024. [22] John Lott and Cédric Villani. Ricci curvature for metric-measure spaces via optimal transport. Annals of Mathematics, 169(3):903–991, 2009. [23] Valerio Mante, David Sussillo, Krishna V Shenoy, and William T Newsome. Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature, 503(7474):78–84, 2013. [24] Robert J. McCann. A convexity principle for interacting gases. Advances in Mathematics, 128 (1):153–179, 1997. doi: 10.1006/aima.1997.1634. [25] R. Olfati-Saber and R.M. Murray. Consensus problems in networks of agents with switching topology and time-delays. IEEE Transactions on Automatic Control, 49(9):1520–1533, 2004. doi: 10.1109/TAC.2004.834113. [26] Reza Olfati-Saber, J. Alex Fax, and Richard M. Murray. Consensus and cooperation in networked multi-agent systems. Proceedings of the IEEE, 95(1):215–233, 2007. doi: 10.1109/JPROC.2006.887293. [27] Misha Rabinovich, Ramón Huerta, and Gilles Laurent. Transient dynamics for neural processing. Science, 321(5885):48–50, 2008. [28] Wei Ren and Randal W Beard. Consensus seeking in multiagent systems under dynamically changing interaction topologies. IEEE Transactions on Automatic Control, 50(5):655–661, 2005. [29] Filippo Santambrogio. Optimal Transport for Applied Mathematicians, volume 87 of Progress in Nonlinear Differential Equations and Their Applications. Birkhäuser, 2015. [30] Karl-Theodor Sturm. Probability measures on metric spaces of nonpositive curvature. Heat Kernels and Analysis on Manifolds, Graphs, and Metric Spaces, pages 357–390, 2003. [31] Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, 2 edition, 2018. [32] Cédric Villani. Optimal Transport: Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften. Springer, 2009. [33] Cédric Villani. Topics in Optimal Transportation. AMS, 2003. [34] Xiao-Jing Wang. Neurophysiological and computational principles of cortical rhythms in cognition. Physiological Reviews, 90(3):1195–1268, 2010.

75

Record · ID 31281 · SHA-256 3859c8b7071a6351
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.