ConceptioArchivearXiv CS
arXiv CSopen access

Free energy landscape of Dense Associative Memory

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Free energy landscape of Dense Associative Memory Sumedha,1,2∗ and Abhishek Singh1,2 1

School of Physical Sciences, National Institute of Science Education and Research, Bhubaneswar, P.O. Jatni, 752050, India and 2 Homi Bhabha National Institute, Training School Complex, Anushakti Nagar 400094, India

arXiv:2607.19195v1 [cond-mat.dis-nn] 21 Jul 2026

Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative memories, including dense associative memories. We illustrate the method by reproducing classical results for the Hopfield model. For a finite number of patterns, we derive the temperature-dependent free energy functional for dense associative memories featuring polynomial interactions and Log-Sum-Exponential (LSE) activation. We also evaluate the disorderaveraged ground-state energy of these systems in the extensive limit. Our analytical framework reveals how memory retrieval depends on the initial state in higher-order dense networks, and gives the exact full-retrieval threshold for the LSE model. This method provides a systematic procedure for analyzing diverse, complex architectures in associative memory.

Dense Associative Memories (DenseAMs) are energybased neural architectures that vastly surpass traditional Hopfield networks by using higher-than-quadratic neuron interactions to expand memory storage capacity [1–3]. In these networks, stored patterns represent low-energy configurations, allowing the system to perform powerful error correction by mapping partial or corrupted inputs back to the correct pattern. This process closely mirrors the biological error-correction mechanisms of the human brain, which makes it possible for it to retrieve memories from incomplete information [4]. Because of this high-capacity retrieval and error-correction, DenseAMs have become useful tools in areas like modern deep learning, generative AI, pattern recognition, and transformer architectures [5–8]. In his seminal work, Hopfield mapped neuronal states to Ising spins and synaptic weights to the Hebb rule [9]. This established a spin-glass-like model where specific configurations serve as stored memories. As a result, the model was successfully solved using the replica method [10, 11]. In associative memory (AM) models, gradient descent dynamics drive the retrieval of stored memories by guiding the system toward the nearest local minimum on a free-energy landscape. The recent work in DenseAMs though have primarily concentrated on energy dynamics rather than the full free-energy framework. In this paper, we derive the free energy landscape for a general class of associative memories and use it to study higher DenseAMs[1, 12–14]. We develop a method of solution that makes use of the tilted property of the measure of exponential functions to obtain the rate function. We had earlier used similar methods to study random field quenched disorder for ferromagnets [15–17]. We consider models with neurons as Ising spins. Each neuron exists in two states : ±1. A pattern is a configuraµ tion of N binary random variables {ξ1µ , ...ξN } , where µ takes P values for P stored patterns. For a configuration

C : {s1 , s2 , . . . si , . . . , sN }, a random variable N

mµ =

1 X µ ξ si N i=1 i

captures the overlap between the µth memory and C. If the configuration C matches with the pattern µ, then mµ = 1. We define a P dimensional vector m = {m1 , m2 .....mP } to represent the overlap between a configuration and the patterns. A general DenseAM [1] is defined by an energy function/Hamiltonian of the form: ! P N X X µ H(m) = − F ξi si = −N f (m) (2) µ=1

i=1

where F (x) is a smooth function. The gradient descent dynamics then works in the direction of lowering the energy and stops when energy can no longer be decreased via single spin flip[5]. Hence H can be considered the cost function of the memory retrieval which is to be minimised. In this work we develop a formalism using large deviations to get the free energy for any function f . For P patterns represented by the vectors ξ µ , the joint probability distribution of the order parameter (m), i.e., Q(m) in the absence of cost function is given by ! P N X Y X µ Q(m) = p({si }) δ ξi si = N mµ (3) {si }

µ=1

i=1

here p({si }) is the probability of having a configuration C = {si }. Using the integral representation of the delta function by introducing a set of variables λ = {λ1 , ...λµ ..., λP } we get Z Q(m) =

dP λ exp(−N λ · m) (2π)P

 X  ∗ [email protected]

(1)

{si }

p({si }) exp

X µ

! X µ λµ ξi si  i

(4)

2

!

gives the value of the free energy at a given β. Let m∗ ∂I be a fixed point of Iβ , then at m∗ , since ∂mβµ = 0,

(5)

∂f ∂I0 =β µ ∂m m∗ ∂mµ m∗

The sum over spins si can be done easily to give Z Q(m) =

Z =

N Y X dP λ exp(−N λ · m) cosh λµ ξiµ P (2π) µ i=1

dP λ exp [−N (λ · m − Φ(λ))] (2π)P 1 N

µ i=1 log(cosh(λµ ξi )). In the limit

where Φ(λ, ξ) = N → ∞, via the law of large numbers we get

(6)

where, Eξ [· · · ] is the average over the distribution from which the patterns ξ have been drawn. The contours of integration in Eq. 5 are understood to be analytically deformed, so that they pass the saddle point. The probability Q(m) satisfies large deviation principle with rate function I0 (m) , i.e Q(m) ≍ exp[−N I0 (m)]. The saddle point approximation then gives I0 (m) = sup [λ · m − Φ(λ)]

(7)

λ∈RP

Let λ∗ be the supremum of the function on the right. It is a solution of a set of P equations of the kind mµ∗ =

∂ϕ(λ) ∂λµ m⃗∗

where A is the subset of all possible configurations compatible with the value m for P patterns. For QH,β (m) ∼ exp(−N I), the rate function I can be calculated using the tilted large deviations principal [18, 19] that connects I and I0 through the relation

Substituting back in Eq. 11, we get: λ∗µ = β

(10)

where I0 (m) = λ ·m−Φ(λ ). The function Iβ (m) is the free energy landscape for the system whose global minima

∂f ∂mµ m∗

(13)

For our case of binary spins we use Eq. 6 to arrive at the P self-consistency equations for fixed point {mµ∗ }. They are " !# X ∂f µ µ ν m∗ = Eξ ξ tanh β ξ (14) ∂mν m∗ ν The rate function, which is the generalised free energy functional of the DenseAMs comes out to be IH,β (m) = βF (m) = β

A

∂I0 = ∂ [λ∗ (m) · m − Φ(λ∗ (m))]/∂mµ (12) ∂mµ X ∂λν X ∂Φ(λ) ∂λν − = λ∗µ + mν µ µ ∂m ∂m ∂λν λ∗ λ∗ λ∗ ν ν

(8)

In general, in order to find λ∗ , we need to find the solution of the above P dimensional array of equations. For P = 1 it is easy and I0 (m) is the rate function of a set of N non-interacting Ising spins. For higher P we circumvent the step where the hard inversion need to be performed by developing a procedure similar to the one used by us for random field ferro-magnets [16, 17]. The probability of a configuration C of neurons under Gibbs measure is proportional to exp(−βH(m)) with β as the inverse temperature that controls the noise. Since m is a random variable drawn from a distribution Q(m) that is defined on {−1, 1}P , the probability QH,β (m), in the presence of the cost function f (m) is given by Z QH,β (m) = Q(m) exp(N βf (m)) (9)

Iβ (m) = I0 (m) − βf (m)

(11)

Hence, we have

PN

Φ(λ) = Eξ [log cosh(λµ ξ µ )] ,

m∗

X

mν∗

ν

" − Eξ log cosh

X ν

∂f ∂mν m∗

∂f β ξν ∂mν m∗

(15) !# − βf (m∗ ).

We have defined F(m) = βIH,β , as the free energy functional of the system and the free energy is β1 infm IH,β . Dense Associative memory Hopfield Model with polynomial interaction: The Hamiltonian is P

HN = −

N X µ k (m ) k! µ=1

(16)

PP β 1 µ k ∗ µ k−1 The f (m) = k! . µ=1 (m ) and λµ = (k−1)! (m ) Substituting in Eq. 14 and 15 we get " !# X β mµ = Eξ ξ µ tanh (mµ )k−1 ξ ν (17) (k − 1)! ν P

F(m) =

k−1 X µ k (m ) k! µ=1

(18)

" !# X 1 β ν k−1 ν − Eξ log cosh (m ) ξ β (k − 1)! ν

3 Note that while for even k, the above equations hold for −1 < mµ < 1, for odd k they are valid only for positive values 0 < mµ < 1. This is because for odd k, negative values of mµ results in a higher energy than the mirror state with positive mµ . The Eq. 17 matches with the stochastic dynamics fixed point equation for DenseAMs obtained recently[20]. For k = 2, the expression of F(m) derived above matches with the classic results in [10]. The F(m) for k > 2 however was not known from the earlier studies. Our method gives the exact expression of the free energy functional for any k. The large deviations approach allowed us to handle higher order interactions with ease, which is not possible with the standard Hubbard-Stratonovich transformation. We can study any finite P using Eqs. 17 and 18. For P = 1, the F and mµ equation are similar to that of a k-spin Ising model. Hence single pattern DenseAMs has continuous transition for k = 2 and first order transition for k ≥ 3. The transition temperatures for example for k = 2, 3, 4 are βc = 1, 0.23, 0.04 respectively. The stability analysis can also be performed straightforwardly by finding the eigenvalues of the Hessian for P = 2. We would though focus on the extensive limit in the rest of this section. The capacity of a network is defined as the maximum number of patterns that can be stored such that a randomly chosen pattern is retrieved with small error as N → ∞, P → ∞ at zero temperature. We study the F(m) by taking N → ∞, P → ∞ along with β → ∞ for the retrieval of a randomly chosen pattern. Each of (mµ )k corresponding to other patterns is taken as a random variable distributed with mean 0 and standard deviation ek [1, 20], where ek is a constant dependent on k. N (k−1)/2 PP Then by central limit theorem, the z = µ=2 (mµ )k−1 ξµ can be taken as a Gaussian random variable with mean 0 and variance P e2k /N k−1 . For P = αN k−1 , γ = e2k α is the 2 1 variance and p(z) = √2πγ e−z /2γ . Separating m1 = m from other patterns and performing the average over ξ 1 and z in the limit of β → ∞, we get the disorder averaged ground state energy of the system as r   −m2(k−1) k−1 k 1 2γ m − exp R(m) = k! (k − 1)! π 2γ (19)   mk−1 mk−1 + erf √ (k − 1)! 2γ PP We dropped the term µ=2 (mµ )k while writing the above expression as its mean is 0 and variance decays as 1/N . The Eq. 19 gives the landscape of the gradient dynamics [21]. The fixed points of Eq. 19 are given by  k−1  m m = erf √ (20) 2α For k = 2 this equation is the same as the fixed point equation for m obtained in [11], with γ = rα, where

k γg γl m∗ 2 0.66 0.66 0 3 0.18 0.26 0.85 4 0.1 0.2 0.92 5 0.063 0.17 0.95 10 0.015 0.13 0.98

Table I: For γ < γl a local minima at m∗ in R(m) appears which becomes a global minima at γg .

r is mean square random overlap obtained via replica calculation. Let us study the consequence of R(m) for DenseAMs in a bit more detail: for k = 2, the function R(m) has a minima at m = 0 for high γ, which continuously changes into a double well as γ is lowered as shown in Fig. 1. The system falls into one of the two minima spontaneously with decreasing γ at γc = 2/π. The local and global minima are the same and gradient descent always reaches the global minima. For k > 2 the system undergoes a first order transition from m = 0 to m = ̸ 0 state. The m = 0 stops being a global minima at γg but continues to be a local minima as γ is decreased. This can be seen by evaluating the second derivative of R(m) at m = 0. χ(m) = ∂ 2 R(m)/∂m2 = p −2(k−1) /2γ 1 − (k − 1) 2/πγmk−2 em . For k > 2, χ(0) = 1 and hence m = 0 is a minima for all values of γ, though it stops being a global minima at a certain γg (see Fig. 1). As a result, the steady state of gradient descent depends on the initial starting state of the system. If the initial state is in the basin of attraction of m = 0, the pattern is not retrieved for any γ. At a threshold γl < γg , the m= ̸ 0 minima shows up for the first time. As k increases the basin of attraction of this non zero fixed point shrinks, though the minima gets closer to m = 1. As a result there is less error in retrieval, provided the starting state is in the basin of the rerieval state. Since basin reduces, choices of initial state for successful memory retrieval gets more restricted. The threshold on αk for reliable retrieval of memory depends on the percentage of allowed error (given by (1 − m)/2). This for 1.5% error for Hopfield model is α ≈ 0.14 [3, 11]. A bound for 0.5% was obtained in [1]. This threshold is different than the transition threshold discussed above as the m at the transition though non zero approaches 1 only in the limit of k → ∞ (see Table I). We show next for LSE the two thresholds match completely due to full pattern retrieval at and below the transition. Log-Sum-Exponential (LSE) model: The cost function for binary uncorrelated patterns can be written as ! P X NX µ 2 1 µ H(m, λ) = (m ) − log exp(N λm ) 2 µ λ µ=1 (21) where λ is the interaction strength.

4

Figure 1: Plot of disorder averaged ground state energy of retrieval(R(m)) of a random memory when N, P → ∞. We have taken ek = 1 and hence γ = α. a) for Hopfield model,b) k = 4 DenseAM. In both cases the function R(m) is plotted near, above and below the transition, In (c) k = 4 basin of m = 0 state is shown to illustrtate that it is always present. In (d) R(m) for LSE model above and below the transition is shown.

1 We define ϕ = λN ln [12]. Then we get

PP

µ=2 exp(λN m

µ

) as was done in

P

f=

  1 1X µ 2 1 (m ) − log eλN m + eλN ϕ 2 µ=1 Nλ

(22)

For small N , substituting the f above in Eqs. 14 and 15 gives the exact F(m). We instead focus on the large N limit here. Using the law of large numbers we approximate ϕ ∼ P ⟨exp(λN mµ )⟩mµ . The mµ for µ = ̸ 1 are random variables drawn from a Gaussian √ distribution µwith mean 0 and standard deviation σ/ N . Hence p(m ) ∼ exp(−N I0 (mµ )) with I0 = (mµ )2 /(2σ 2 ). Using the tilted LDP, we then get, ⟨exp(λN mµ )⟩p(mµ ) ∼ exp(−N λ2 σ 2 /2). Assuming P ∼ exp(αN ), we get   1 λ2 σ 2 ϕ= α+ (23) λ 2 √ The extrema of ϕ occurs at λ∗ = 2α/σ. For λ < λ∗ , ϕ increases rapidly with increasing λ and is a slow increasing function for λ > λ∗ . The f hence takes two different forms for ϕ > m1 and for ϕ < m1 as follows: P X

fϕ<m1 =

1 2 (mµ ) − m1 2 µ=1

fϕ>m1 =

1X µ 2 (m ) − ϕ 2 µ=2

(24)

P

(25)

This gives P

Fϕ<m1 (m) =

1 X µ2 m − m1 2 µ=1

(26)

" !# P X  1 1 1 µ µ − Eξ log cosh β m − 1 ξ + β m ξ β µ=2 " !# P P X 1 X µ2 1 µ µ Fϕ>m1 (m) = m − Eξ log cosh β m ξ 2 µ=1 β µ=1 (27)

For ϕ > m1 the model behaves just like the standard Hopfield model. Since P ∝ exp(αN ) and the capacity of Hopfield model is O(N ), there is no retrieval in this case. The memory retrieval is possible only for ϕ < m1 . For β → ∞, we get r   1 1 2 2s (m1 − 1)2 1 Rϕ<m1 (m) = (m ) − m − exp − 2 π 2s (28)  1  m −1 √ − (m1 − 1)erf 2s with 1

m = 1 + erf where s ∝ we get



m1 − 1 √ 2s

 (29)

P . Taking the s → ∞ limit for P ∝ exp(αN ),

Rϕ<m1 (m1 ) =

1 1 2 (m ) − m1 2

(30)

with a fixed point at m1∗ = 1. This is the free energy landscape of the retrieval phase of LSE. As shown in Fig. 1, the landscape has tilted towards the m = 1 state and the retrieval of the memory is error free. Equating ϕ = 1 then gives us the threshold αc (λ) below which the pattern is completely retrieved with no error. We fix σ = 1. Then αc (λ) ≤ 1/2 for real λ. Since λ = 1 at αc = 1/2   λ αc = λ 1 − (31) 2 for λ ≤ λ∗ = 1. We have hence reproduced the threshold for retrieval derived recently using the random energy model[12]. The treatment of LSE here is done assuming exponential number of stored patterns, but the method can be applied to any number of patterns. Unlike polynomial interaction DenseAMs, for LSE the αc gives the threshold of retrieval as m∗ = 1 for α < αc . Discussion: Our method can evaluate any cost function defined by Eq. (1). Although currently demonstrated using binary variables, the framework generalizes effortlessly to other discrete or continuous variables. We validated

5 our approach by analyzing two prominent AM models. We showed that the disorder-averaged ground state energy, R(m), is useful for decoding gradient dynamics. In DenseAMs, we showed a strong initial-state dependence during memory retrieval, as the basin at m = 0 persists

for all α. One would require a mechanism like a stochastic noise [20] to come out of the basin of m = 0 state. Finally, for LSE, our results align with the findings by Lucibello et al. [12].

[1] D. Krotov and J. J. Hopfield, Dense Associative Memory for Pattern Recognition Advances in Neural Information Processing Systems, vol. 29, (2016). [2] E. Gardner, Multiconnected neural network models, J. Phys. A: Math. Gen. 20 3453 (1987). [3] J. J. Hopfield, Neural networks and physical systems with emergent collective computational abilities, Proceedings of the National Academy of Sciences 79, 2554–2558 (1982). [4] J. R. Anderson and G. H. Bower, Human Associative Memory (Psychology Press, 1974). [5] D. Krotov, B. Hoover, P. Ram and B. Pham, Modern methods in Associative memory, arXiv:2507.06211 (2025). [6] J. Simon, D. Kunin, A. Atanasov, E. Boix-Adserà, B. Bordelon, J. Cohen, N. Ghosh, F. Guth, A. Jacot, M. Kamb, D. Karkada, E. J. Michaud, B. Ottlik, J. Turnbull, There will be a scientific theory of Deep Learning arXiv:2604.21691 [7] B. Hoover, Y. Liang, B. Pham, R. Panda, H. Strobelt, D. H. Chau, M. Zaki, and D. Krotov. Energy transformer, Advances in Neural Information Processing Systems, 36, (2024). [8] M. Yampolskaya and P. Mehta, Hopfield Networks as Models of Emergent Function in Biology, Annual Review of Biophysics 55 323-342(2026). [9] Donald O. Hebb, The organization of behavior, new york: Wiley,The first stage of perception: growth of the assembly, In Neurocomputing, Volume 1: Foundations of Research. The MIT Press, (1949). [10] D. J. Amit, H. Gutfreund, and H. Sompolinsky, Physical Review A 32, 1007–1018 (1985). [11] D. J. Amit, H. Gutfreund, and H. Sompolinsky, Physical Review Letters 55, 1530–1533 (1985). [12] C. Lucibello and M. M´ezard, Exponential Capacity of Dense Associative Memories, Physical Review Letters 132 077301 (2024).

[13] M. Demircigil, J. Heusel, M. Löwe, S. Upgang and F. Vermet, On a Model of Associative Memory with Huge Storage Capacity, Journal of Statistical Physics 168 288(2017). [14] H. Ramsauer, B. Sch¨afl, J. Lehner, P. Seidl, M. Widrich, T. Adler, L. Gruber, M. Holzleitner, M. Pavlovi´c, G. K. Sandve, V. Greiff, D. Kreil, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter, Hopfield networks is all you need, International Conference on Learning Representations (2021); arXiv:2008.02217. [15] Sumedha, and S. K. Singh, Effect of random field disorder on the first order transition in p-spin interaction model Physica A: Statistical Mechanics and its Applications 442, 276–283 (2016). [16] S. Mukherjee and Sumedha, Phase Transitions in the Blume–Capel Model with Trimodal and Gaussian Random Fields, Journal of Statistical Physics, 188 22(2022). [17] Sumedha, and M. Barma Phase transitions in XY models with randomly oriented crystal fields, Phys. Rev. E, 105 105, 024111(2022). [18] Frank den Hollander, Large Deviations, Fields Institute Monographs, AMS (2000) Theorem III.17. [19] Sumedha, and Nabin K Jana, Absence of first order transition in the random crystal field Blume–Capel model on a fully connected graph (see Appendix), J. Phys. A: Math. Theor. 50 015003 (2017). [20] S. Rooke, D. Krotov, V. Balasubramanian, and D. Wolpert, Stochastic thermodynamics of associative memory, arXiv:2601.01253 (2026). [21] Gradient dynamics is same as the zero temperature Glauber dynamics. For zero temperature Glauber dynamics we have recently shown this for a different model, namely the random field Blume Capel model in Sumedha and Aldrin B E, Glauber dynamics phase transitions in athermal random field Blume-Capel and Blume-EmeryGrifitths models, arXiv:2607:16561.

Record · ID 386916 · SHA-256 94445b6d85142180
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.