Two-dimensional Hyperbolic RNN Neural Quantum State H. L. Dao∗ June 25, 2026
arXiv:2606.25600v1 [quant-ph] 24 Jun 2026
Abstract In the first part of this work, we construct the first type of two-dimensional (2D) hyperbolic neural quantum state (NQS) in the form of the Lorentz 2DRNN (Recurrent Neural Network) and benchmark its performance against the Euclidean 2DRNN in the paradigmatic N × N 2D Transverse Field Ising Model (2DTFIM) setting with different lattice sizes up to N = 12 and at different transverse magnetic field strengths. We find that hyperbolic Lorentz 2DRNN NQS definitively outperform Euclidean 2DRNN NQS when the system is at the phase transition point when the physics can be described by a conformal field theory (CFT), which is known to be dual to an Anti-de-Sitter (AdS) space whose spatial geometry is hyperbolic. In the second part of this work, we benchmark the performances of the recently introduced one-dimensional (1D) hyperbolic NQS including Poincaré RNN/GRU and Lorentz RNN/GRU against their Euclidean NQS versions in N ×N 2DTFIM, which has to be converted to a one-dimensional setting to allow for the use of 1D NQS. The findings in this case extend our previous results that 1D hyperbolic NQS definitively outperform 1D Euclidean NQS, thanks to the combined effects of the hierarchical structure comprising the first and N th neighbor interactions present in the 1D system arising from the 2D lattice and the CFT physics at the critical point. While more studies with larger system sizes are required, our work serves as a proof-of-concept for the utility, effectiveness as well as the superior performances of one- and two-dimensional hyperbolic NQS ansatzes compared to the existing Euclidean NQS in many-body quantum physics systems, especially when these systems exhibit structural hierarchy or when they are at criticality, or a combination of both.
Contents 1 Introduction
2
2 Poincaré and Lorentz hyperbolic RNN/GRU 2.1 One-dimensional networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.2 Two-dimensional networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3 3 4
3 Hyperbolic RNN-based NQS wavefunctions
5
4 2D TFIM VMC experiments with 2D hyperbolic NQS 7 4.1 Different lattice sizes, fixed magnetic strength . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 4.2 Fixed lattice sizes, different magnetic strengths . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 5 2D TFIM VMC experiments with 1D hyperbolic NQS
16
6 Concluding remarks
26
A Appendix A.1 Hierarchical structures of neighbor interactions in spin systems . . . . . . . . . . . . . . . . . . . A.2 Poincaré disk model of hyperbolic space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A.3 Lorentz hyperboloid model of hyperbolic space . . . . . . . . . . . . . . . . . . . . . . . . . . . .
28 28 29 30
References
32
1
1
Introduction
Since the seminal works of [1]-[4] that established the viability and utility of the simplest types of neural networks such as Restricted Boltzmann Machine and feedforward networks as wavefunction approximations in quantum many-body systems, steady progresses have followed with ever-increasing complex architectures involving CNNs (Convolutional Neural Networks), RNNs (Recurrent Neural Networks) and transformers [5] - [10]. In our recent works [12], [13], we introduced the first non-Euclidean neural quantum states in the form of the hyperbolic recurrent networks, including Poincaré RNN/GRU (Gated Recurrent Unit) and Lorentz RNN/GRU, whose underlying geometries are the Poincaré and Lorentz models of hyperbolic space instead of Euclidean space like all other NQS constructions that came before. In these papers, we showed that these hyperbolic NQS could achieve superior performances compared to their Euclidean counterparts in various many-body quantum system settings that exhibit some hierarchical structure in the form of the different degrees of nearest neighbor interactions such as the Transverse Field Ising models (2DTFIM) as well as the Heisenberg J1 J2 and Heisenberg J1 J2 J3 models. In constructing and proposing these hyperbolic NQS, we were inspired by results from NLP (Natural Language Processing) [14], [19], [22], [24] as well as graph-related tasks (graph embedding, link prediction, node classification, etc.) [49] that showed definitive outperformances delivered by different types of hyperbolic neural networks compared to their Euclidean versions in settings where hierarchical, tree-like structures exist in the training datasets. These impressive performances by hyperbolic neural networks can be attributed to the fact that the exponential growth of hyperbolic space (area and volume) allows for a very low distortion embedding of tree-like, hierarchical structures that are not possible with Euclidean space (whose area and volume only grow polynomially with distance) [26], [27], [28], [30], [31]. While the work [13] enlarged the class of non-Euclidean hyperbolic NQS to include new constructions like Lorentz RNN/GRU and Poincaré RNN besides the first hyperbolic NQS - Poincaré GRU - that was introduced in [12], the scope of the work [13] was restricted to the Heisenberg quantum systems. In [12], we also studied the performances of Poincaré GRU in the N ×N 2DTFIM context (up to N = 9) and reported its notable outperformance over the Euclidean GRU NQS when the two dimensional problem is mapped to a one dimensional setup, the process of which introduces a hierarchical structure in the form of the first and N th neighbor interactions. In this work, we revisit the problem of applying hyperbolic NQS to the 2D TFIM setting at different magnetic field strengths using larger lattice sizes, ranging from 8 × 8, 10 × 10 to 12 × 12, using two different approaches. In the first approach, we introduce and construct the first type of two-dimensional hyperbolic NQS ansatz, the Lorentz 2DRNN, to use in the 2DTFIM setting without mapping to one dimension. In the second approach, we repeat the 2D-1D mapping experiment (effectively unrolling the 2D setting into a 1D one) of [12] with the new hyperbolic NQS constructions from [13] including Poincaré RNN, Lorentz RNN/GRU, in addition to the original Poincaré GRU already studied in [12]. In both cases, we benchmark the performances of all hyperbolic NQS against their Euclidean versions. • In the native 2D settings with only horizontal and vertical nearest neighbors where no hierarchical structure exists, we find that Lorentz 2DRNN consistently and definitively outperformed Euclidean 2DRNN when the transverse magnetic field is at the critical value when the system is at the phase transition point and can be described as a conformal field theory (CFT). A possible explanation for this observed outperformance of hyperbolic NQS in this case might have to do with the fact that from the AdS/CFT correspondance [45], D-dimensional CFTs are known to be dual to (D + 1)-dimensional AdS space whose spatial geometry is hyperbolic, which means that the hyperbolic geometry underlying the construction of Lorentz 2DRNN might be exactly what is required to provide a better ansatz approximation for TFIM at criticality. • In the unrolled 2D settings with one-dimensional NQS, similar to our results in [12], we find that the one-dimensional hyperbolic NQS emerged as the better performing ansatzes compared to the Euclidean ones thanks to a combination of the induced hierarchical structures determined by the lattice size N and the CFT physics of the system at the critical point. This paper is organized as follows. In Section 2, we describe the constructions of the Poincaré and Lorentz hyperbolic RNN/GRU, both in one dimension (Section 2.1) and in two dimensions (Section 2.2). In Sec 3, we summarize the variational Monte Carlo method, the construction of the neural quantum states based on either Euclidean or hyperbolic RNN/GRU, as well as some remarks regarding the training of various NQS. The main results of this work are reported in Section 4 for two-dimensional NQS and in Section 5 for one-dimensional NQS. In Section 6, some concluding remarks and future work directions are discussed. In the Appendix, we illustrate the different hierarchical structures representing the different degrees of neighbor interactions in the forms of graphs in Section A.1. We include the definitions of all mathematical operations in the Poincaré disk in Section A.2 and in the Lorentz hyperboloid in Section A.3 that are used to construct the hyperbolic NQS in this work.
2
The Python codes (Pytorch implementation) used in this work to construct the hyperbolic 1D hyperbolic Poincaré/Lorentz RNN/GRU real NQS and 2D hyperbolic Lorentz 2DRNN can be found at https://github.com/lorrespz/hypnqs tfim. The construction of the various Lorentz NQS makes use of the hypercore library [15] with some modifications (files lmath.py and lorentzian.py from the subdirectory manifolds of the main directory hypercore main), similar to our previous work [13].
2
Poincaré and Lorentz hyperbolic RNN/GRU
In this section, we first recall our constructions of the one-dimensional hyperbolic Poincaré and Lorentz RNN/GRU in [12], [13], and then proceed to present our custom construction of the two-dimensional hyperbolic Poincaré and Lorentz RNN.
2.1
One-dimensional networks
The constructions of the Poincaré and Lorentz variants of hyperbolic RNNs were described in detail in our recent work [13], and are based on the following definitions of Euclidean RNN and GRU: Starting with an input vector ⃗xi ∈ Rdx at step i of size dx , the conventional version of RNN is a function that relates the hidden state vector ⃗hi ∈ Rdh of size dh to the input ⃗xi and the same hidden state vector ⃗hi−1 at the previous time step (i − 1) ⃗hi = f (Wh⃗hi−1 + Uh ⃗xi + ⃗bh ) ,
RNN :
(1)
where f is a nonlinear activation function (often ‘tanh’), Wh is a dh × dh weight matrix, Uh is a dh × dx weight matrix, and bh ∈ Rdh is a vector of size dh known as the bias. GRU [17] is a more sophisticated version of RNN with additional structures comprising the reset gate ⃗ri ∈ Rdh and the update gate ⃗zi ∈ Rdh . The defining equations of the conventional or Euclidean GRU are GRU : ⃗ri = σ Wr⃗hi−1 + Ur ⃗xi + ⃗br ⃗zi = σ Wz⃗hi−1 + Uz ⃗xi + ⃗bz h i ⃗ h̃i = f Wh (⃗ri ⊙ ⃗hi−1 ) + Uh ⃗xi + ⃗bh ⃗hi = (1 − ⃗zi ) ⊙ ⃗hi−1 + ⃗zi ⊙ ⃗h̃i
(2)
⃗ In Eq.(2), ⃗hi is the final hidden state at step i while h̃i is the new state computed in the same time step, f is a nonlinear activation function (often taken to be ‘tanh’), σ is the sigmoid activation, ⃗xi is the input vector of length dx , Wh , Wz , Wr are the dh × dh weight matrices, Uh , Uz , Ur are the dh × dx weight matrices, ⃗bh , ⃗bz , ⃗br ∈ Rdh are the bias vectors, and ⊙ is the pointwise multiplication operation. Note that for GRU, the last equation appearing in Eq.(2) that defines the final hidden state ⃗hi can also be written as follows ⃗hi = ⃗hi−1 + ⃗zi ⊙ (−⃗hi−1 + ⃗h̃i ) ,
(3)
⃗ where the update step involves the previous state ⃗hi−1 and the new state h̃i , which is a function of the reset ⃗ gate ⃗ri , the previous state ⃗hi−1 and the input ⃗xi . When ⃗zi ≈ 0, the final state ⃗hi is almost entirely h̃i - the new state. When ⃗zi ≈ 1, the final state ⃗hi is almost entirely its previous state ⃗hi−1 . When ⃗ri ≈ 1 and ⃗zi ≈ 1, the GRU essentially reduces to the RNN. The Poincaré hyperbolic RNN and hyperbolic GRU [12], [14] are defined by the following formulae h i ⃗hi = f ⊗c (Wh ⊗c ⃗hi−1 ) ⊕c (Uh ⊗c ⃗xi ) ⊕c ⃗bh Poincaré RNN : Poincaré GRU :
n h io logc0 (Wr ⊗c ⃗hi−1 ) ⊕c (Ur ⊗c ⃗xi ) ⊕c ⃗br , n h io ⃗zi = σ logc0 (Wz ⊗c ⃗hi−1 ) ⊕c (Uz ⊗c ⃗xi ) ⊕c ⃗bz , h i ⃗ h̃i = f ⊗c Wh ⊗c (⃗ri ⊙c ⃗hi−1 ) ⊕c (Uh ⊗c ⃗xi ) ⊕c ⃗bh , ⃗hi = ⃗hi−1 ⊕c ⃗zi ⊙c −⃗hi−1 ⊕c ⃗h̃i .
(4)
⃗ri = σ
3
(5)
In Eqs.(4), (5), the various weight matrices Wh,r,z and Uh,r,z have the same meanings as in the Euclidean RNN/GRU case - Eqs.(1), (2), but the hidden state vector ⃗hi ∈ Ddc h , input vector ⃗xi ∈ Ddc i , and bias vectors ⃗bh,r,z ∈ Ddh are all vectors on the Poincaré disk. Unlike the Euclidean case where a nonlinear activation function c f (often ‘tanh’) is always applied, we have the choice of either using f ⊗c or leaving it out completely because hyperbolic space itself provides some degree of nonlinearity. The meanings of the Poincaré mathematical operations ⊗c , ⊕c , ⊙c , f ⊗c , expc0 , logc0 are described in detail in Section A.2. Other variants of hyperbolic RNN networks, the Lorentz RNN and Lorentz GRU, are based on the Lorentz hyperboloid model of hyperbolic space [13]. Their defining equations are also based on the Euclidean definitions of RNN and GRU (Eqs.(1), (2)) h i ⃗hi = f ⊗L Wh ⊗L ⃗hi−1 ⊕L (Uh ⊗L ⃗xi ) ⊕L ⃗bh (6) Lorentz RNN : h io n Lorentz GRU : ⃗ri = σ log0L Wr ⊗L ⃗hi−1 ⊕L (Ur ⊗L ⃗xi ) ⊕L ⃗br h io n (7) ⃗zi = σ log0L Wz ⊗L ⃗hi−1 ⊕L (Uz ⊗L ⃗xi ) ⊕L ⃗bz h i ⃗ h̃i = f ⊗L Wh ⊗L (⃗ri ⊙L ⃗hi−1 ) ⊕L (Uh ⊗L ⃗xi ) ⊕L ⃗bh , ⃗hi = ⃗hi−1 ⊕L ⃗zi ⊙L −⃗hi−1 ⊕L ⃗h̃i . where the weight matrices Wh,r,z , Uh,r,z and bias vectors have the same meaning as previously defined. In Eqs.(6), (7), the RNN/GRU hidden state ⃗hi at the end of each computation is hyperbolic and lives on the Lorentz hyperboloid. The weight matrices Wh,r,z , Uh,r,z are Euclidean, while the biases ⃗bh,r,z can be hyperbolic or Euclidean (in which case, they are projected onto the hyperboloid using the exponential map exp0L ). For the Lorentz GRU, the last equation in Eq.(7) is the update step for the hidden state ⃗hi , which was done on the hyperboloid in an identical manner to the Poincaré case. However, it must be noted that in Lorentz hyperbolic space, the state −⃗hi−1 is not simply the state ⃗hi−1 with all components multiplied by −1 (as is the case in Euclidean space as well as in the Poincaré disk). Instead, the −⃗hi−1 state has the components (x0 , −⃗xi ) if ⃗hi−1 = (x0 , ⃗xi ), since we need ⃗hi−1 ⊕L (−⃗hi−1 ) = 0L . The meanings of the Lorentz mathematical operations ⊗L , ⊕L , ⊙L , f ⊗L , log0L are described in detail in Section A.3.
2.2
Two-dimensional networks
Having recalled the constructions of the 1D hyperbolic NQS in the preceding section, we will now describe the new construction of the two-dimensional hyperbolic Poincaré and Lorentz 2DRNN networks, which are based on the custom construction of the Euclidean 2D RNN described in [7] ⃗ht+1 = f Uh ⃗x0 + Wh⃗h0 + Uv ⃗x1 + Wv⃗h1 + ⃗b Euclidean 2DRNN : (8) t t t t where f is a nonlinear activation function, which can be ‘tanh’ or ‘elu’. The input ⃗xt = (⃗x0t , ⃗x1t ) and RNN hidden state ⃗ht = (⃗h0t , ⃗h1t ) at time step t in the computational process are two-dimensional vectors. The number of weight matrices in this case have doubled from two to four Wh , Wv , Uh , Uv , to take into account the horizontal as well as vertical previous time steps. The hyperbolic versions of Eq.(8) are Poincaré 2DRNN : Lorentz 2DRNN :
i Uh ⊗c ⃗x0t ⊕c Wh ⊗c ⃗h0t ⊕c Uv ⊗c ⃗x1t ⊕c Wv ⊗c ⃗h1t ⊕c ⃗b (9) i h ⃗ht+1 = f ⊗L Uh ⊗L ⃗x0 ⊕L Wh ⊗L ⃗h0 ⊕L Uv ⊗L ⃗x1 ⊕L Wv ⊗L ⃗h1 ⊕L ⃗b (10) t t t t
⃗ht+1 = f ⊗c
h
The Poincaré mathematical operations ⊗c , ⊕c , f ⊗c and the Lorentz mathematical operations ⊗L , ⊕L , f ⊗L are defined in Section A.2 and A.3, respectively. In this work, we will only focus on the Lorentz 2DRNN due to our limited computing resources, and the fact that in general, Lorentz hyperbolic networks tend to outperform Poincaré networks [13]. Furthermore, we note that while it is possible to also construct the hyperbolic Poincaré/Lorentz versions of the two-dimensional Euclidean GRU as discussed in [12], the training of such networks will be very computationally intensive and exceed our computational resources at the moment.
4
3
Hyperbolic RNN-based NQS wavefunctions
In this section, we summarize the main points regarding the variational Monte-Carlo method, as well as the construction of the real NQS wavefunction using either Euclidean or hyperbolic RNN/GRU in discrete Hamiltonian spin systems. The detailed description can be found in [7], [12], [13]. Variational Monte Carlo (VMC): The variational Monte Carlo (VMC) method, often used to train NQS, involves the process of sampling from a probability distribution represented by the square of the trial wavefunction/the variational ansatz and subsequently using these generated samples to calculate some obervables such as the ground state energy. In what follows, we recall the derivation of the local energy formula in the Variational Monte Carlo (VMC) method [16]. Given a quantum Hamiltonian H, in VMC, the local energy Eloc (x) of samples |x⟩ generated by an NQS Ψ is given by Eloc (x) =
X
⟨x|H|x′ ⟩
x′
⟨x′ |Ψ⟩ . ⟨x|Ψ⟩
(11)
Eloc (x) is non-zero only for those non-zero Hamiltonian elements ⟨x|H|x′ ⟩ ̸= 0. Corresponding to each generated sample |x⟩ is a probability |Ψ(x)|2 2 |x⟩ |Ψ(x)|
Ploc (x) = P
(12)
In the variational problem of interest, given a Hamiltonian H and a trial wavefunction Ψ, the VMC task is to estimate the ground state energy E = ⟨Ψ|H|Ψ⟩ which can be written in terms of the local energy Eloc and the probability Ploc (x) X E= Ploc (x)Eloc (x) . (13) |x⟩
⃗ with trainable parameters θ, ⃗ at each When the trial wavefunction is a neural network quantum state |Ψ(θ)⟩ ⃗ ⃗ training iteration i, the variational energy E(θ) is calculated from the local energy Eloc (θ) of the Monte Carlo ⃗ using an optimizer (such as samples generated from the NQS using Eq.(13). The process of minimizing E(θ) SGD - Stochastic Gradient Descent or Adam - Adaptive Moment Estimation) updates the trainable parameters θ⃗ until convergence is reached. For the ensuing discussion below regarding the construction of the real NQS ansatzes, the basis state of the Hamiltonian H is denoted by |⃗σ ⟩ = (σ1 , . . . , σN ) where N is the dimensionality of H, and each component σi (1 ≤ i ≤ N ) assumes discrete values of either 0 or 1. The real NQS wavefunction is defined as Xp |Ψ⟩ = P (⃗σ )|⃗σ ⟩ (14) ⃗ σ
where P (⃗σ ), the probability of a particular configuration |⃗σ ⟩, is the output of the real NQS. P (⃗σ ) = P (σ1 )P (σ2 |σ1 ) . . . P (σN |σ1 , σ2 , . . . , σN −1 )
5
(15)
Ψ(⃗σ )
P (⃗σ1 )
P (⃗σ2 )
P (⃗σ3 )
Dense (Softmax)
Dense (Softmax)
Dense (Softmax)
h1 h0
h2
Euclidean/ Hyperbolic RNN/GRU
h1
Euclidean/ Hyperbolic RNN/GRU
σ0
σ1
h3 h2
Euclidean/ Hyperbolic RNN/GRU σ2
p Figure 1: Schematic of the process of calculating the RNN wavefunction Ψ(⃗σ ) = P (⃗σ )|⃗σ ⟩ from the probability P (⃗σ ) of the sample ⃗σ . Here P (⃗σ ) = P (σ1 )P (σ2 |σ1 ) . . . P (σN |σN −1 ). For a compact representation, N = 3 in the schematic. The recurrent network in the diagram can be either Euclidean RNN/GRU or Hyperbolic RNN/GRU. Figure adapted from [12]. The autoregressive RNN-based real NQS wavefunction consists of a layer of RNN/GRU followed by a dense layer of 2 units with the Softmax activation function. The RNN/GRU layer can be Euclidean or hyperbolic, with the latter choice includes both the options of Poincaré and Lorentz models of hyperbolic space. The structure of the RNN-based real NQS function is illustrated in Fig.1. At a step i where 0 ≤ i ≤ N , the RNN/GRU cell takes in two arguments, one being the previous RNN/GRU hidden state ⃗hi−1 and the input ⃗σi−1 (the one-hot-encoded (i − 1)th component of the generated spin sample ⃗σ = (σ1 , . . . , σN )), to calculate the current RNN/GRU hidden state ⃗hi . This is then passed into the Dense layer with the Softmax activation function1 to obtain an output ⃗yi , which is multiplied with the one-hot-encoded sample σi to calculate the probability P (⃗σi ). The total probability P (⃗σ ) for the single spin sample ⃗σ is the product of all individual P (⃗σi ). In this paper, we only work with the real NQS wavefunctions as we are interested in applying the newly constructed 2D hyperbolic NQS Eq.(10) as well as the 1D hyperbolic NQS Eqs.(6), (7) to the 2D TFIM system. Further remarks: Before moving on to describing the results of the VMC experiments, we note the following details that are common to all experiments done in this work (many of these details are similar or identical to our recent works [12], [13]). • Because of our limited computational resources, in order to run a large number of VMC experiments, the sizes of all neural networks used in this work are relatively small, and on the order of a few thousand parameters only. As the fundamental natures of this and our earlier works ([12], [13]) are proof-of-concept works, our focus is on show-casing the capability of our newly constructed hyperbolic NQS in a qualitative manner. With more computational resources, the problem of scaling up the hyperbolic NQS introduced in our works can be readily tackled. • For training, we used 80 samples for all NQS ansatzes, while for inference with trained NQS, we use 104 samples. The small number of samples used during training is justified by the fact that these are autoregressive NQS, which allows for exact and independent sample generations, unlike non-autoregressive NQS ansatzes utilizing Markov-chain Monte Carlo sampling, which typically requires thousands of samples. During training, the saving of a model’s weights is done contingent on certain strict criteria being met regarding the improvement of the network. These criteria include the lowering of the mean energy reachable by the network as well as the energy variance being under a specified tolerance threshold. • The common points concerning the training of all NQS include the use of gradient clipping, Adam optimizer and the scheduled learning rate decay ReduceLROnPlateau that automatically reduces the learning rate 1 Recall that the Softmax function is defined as
exp(vk ) Softmax(vk ) = P i exp(vi )
6
(16)
by half when an energy plateau is detected within 40 epochs. In particular, Adam optimizer is used for all Euclidean parameters (of Euclidean networks and hyperbolic networks). • As noted in detail in [13], both Lorentz and Poincaré hyperbolic NQS took much longer and were more complicated to train than Euclidean NQS, because their defining mathematical operations are much more complex and there are more hyperparameters that have to be taken into account when dealing with hyperbolic networks. These additional hyperparameters include the hyperbolic learning rates and hyperbolic spatial constraint parameter (Rmax for Poincaré and Lmax for Lorentz). When not chosen carefully, the resulting performances of the hyperbolic NQS can be very poor compared to their Euclidean counterparts. For Poincaré networks, Riemannian SGD optimizer is used on the hyperbolic parameters, while for Lorentz networks, Adam optimizer is used on the hyperbolic parameters as well, since these are actually exponentiated Euclidean parameters. Since hyperbolic networks are more prone to numerical instabilities, the hyperbolic learning rates applicable to the hyperbolic parameters have to be chosen carefully, and on a case-by-case basis. Furthermore, it must be noted that the optimization landscapes of hyperbolic neural networks are much steeper and more difficult for an optimizer to navigate compared to Euclidean NQS, because hyperbolic space has an exponential growth in volume and constant negative curvature, compared to the flat and zero curvature of Euclidean space.
4
2D TFIM VMC experiments with 2D hyperbolic NQS
The Hamiltonian system of interest to us is the two-dimensional transverse field Ising model (2DTFIM) with open boundary conditions of the form X X H = −J σiz σjz − Bx σix (17) i
⟨i,j⟩
where J = 1.0, and the first sum in Eq.17 runs over pairs of vertical and horizontal neareast neighbors. This system exhibits a phase transition at Bc = 3.044, separating an ordered, ferromagnetic phase with ⟨σ z ⟩ ̸= 0 from a disordered, paramagnetic phase where ⟨σ z ⟩ = 0 [29]. At Bx ≈ Bc near the critical magnetic field strength Bc , the 2DTFIM is described by a 3D conformal field theory (3DCFT)2 [46]. At the critical point, the system, described the 3DCFT physics, exhibits a long range, power-law spin spin correlation ⟨σ0z σrz ⟩ of the form3 ⟨σ0z σrz ⟩ =
A , rD−2+η
(18)
where D is the number of spatial dimensions, η is a constant that varies depending on the number of spatial dimensions, A is a proportionality constant chosen depending on the system under study. For D = 3 (corresponding to 3D classical Ising model dual to 2D TFIM), η = 0.0363 [38]. The critical exponent γ = D − 2 + η is a universal quantity that defines the phase transition, but the proportionality constant A is a non-universal prefactor that depends on the microscopic cutoff and operator normalization of specific discrete lattice simulations, so A can be treated as a free scaling parameter to vertically align the profiles of the spin-spin correlation curves in systems under study. In this work, we will carry out different sets of VMC experiments using different types of NQS ansatzes. These experiments are designed to benchmark the performances of our newly constructed hyperbolic NQS (both one-dimensional and two-dimensional) ansatzes against the performances of the established Euclidean RNN ansatzes. In this section, we report the results of the VMC experiments in which we apply the newly constructed hyperbolic Lorentz 2DRNN NQS in Eq.(10) in the 2DTFIM setting on a square lattice4 . There are two subsets of experiments, one involving different lattice sizes (N, N ) = (8, 8), (10, 10), (12, 12) at fixed magnetic field strength Bx = 3.0 (almost at criticality), and one involving the same lattice size (N, N ) = (12, 12) at different magnetic field strengths Bx = 2.0, 3.0, 4.0. For all experiments, our NQS ansatzes and their corresponding parameters are listed in Table 1. For (N, N ) = (8, 8), the RNN hidden vector size of all ansatzes is 50, while for (N, N ) = (10, 10), (12,12), the hidden vector size is 60. For Lorentz 2DRNN, in some experiments, we considered several different ansatzes corresponding to different choices of the spatial constraint hyperparameter Lmax . 2 Recall that 2DTFIM at criticality where B = B is equivalent to a 3D classical Ising model at criticality where T = T [47], x c c which can be described by a 3D conformal field theory [46]. For the correspondance between D-dimensional classical Ising model at criticality and (D + 1)-dimensional CFT, see the book [46] Chapter 12. 3 Note that away from the critical point, this spin-spin correlation has the exponential decay form ⟨σ z σ z ⟩ ∼ e−r/ξ , where ξ is a 0 r characteristic length scale depending on the system. 4 In this work, we use N × N and (N, N ) interchangeably to denote the square lattice size of the 2DTFIM.
7
Ansatz Euclidean 2DRNN
Hidden dimension size 50
Parameters 5352
60
7622
50
5352
60
7622
Lorentz 2DRNN
Experiment (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12) (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12)
Table 1: The two-dimensional NQS ansatzes used to run the 2DTFIM VMC experiments involving different lattice sizes (N, N ). For (N, N ) = (8, 8), the NQS ansatzes have a fixed RNN hidden dimension size of 50, while for (10,10) and (12,12), the RNN hidden dimension size is 60.
4.1
Different lattice sizes, fixed magnetic strength
The results of the VMC experiments with different (N, N ) at fixed Bx = 3.0 are shown in Tables 2, 3, 4 and Fig.6. In these tables and figures, we list both the best and averaged achievable mean energy, where ‘best’ refers to the best result chosen from all different individual VMC runs, while ‘averaged’ refers to the averaged result taken from all VMC runs. In Fig.3 and Fig.4, we also plotted the ⟨σ0z σrz ⟩ spin-spin correlation length decay ⟨σ0z σrz ⟩ for all 2D NQS ansatzes for the case of (N, N ) = (12, 12) using the averaged results from different VMC runs and the individual results from each run, respectively5 (N, N ) = (8, 8), Bx = 3.0 Best Averaged −202.4230 −202.3914 0.0099 0.0108
NQS Ansatz Euclidean 2DRNN
−202.4597 0.0074
Lorentz 2DRNN (Lmax = 2.0) DMRG (not exact)
−202.4538 0.0076
-202.5077
Table 2: VMC results (both best and averaged from different runs) of 2D TFIM with the lattice size (N, N ) = (8,8) at the fixed magnetic field strength Bx = 3.0. For each entry, we first recorded the mean energy in the upper line, followed by the corresponding standard error in the lower line. The best-performing results are noted in bold.
(N, N ) = (10, 10), Bx = 3.0 Best Averaged −316.8991 −316.8717 0.0089 0.0103
NQS Ansatz Euclidean 2DRNN
−316.9353 0.0077
Lorentz 2DRNN (Lmax = 2.0) DMRG (not exact)
−316.8919 0.0099
-316.9770
Table 3: VMC results (both best and averaged from different runs) of 2D TFIM with varied size lattice (N, N ) = (10,10) at the fixed magnetic field strength Bx = 3.0. For each entry, we first recorded the mean energy in the upper line, followed by the corresponding standard error in the lower line. The best-performing results are noted in bold.
5 Unlike the unrolled case considered in the next section where we have a 2D to 1D mapping in which each lattice size N represents a different hierarchical neighbor interaction structure ⟨i, i + 1⟩-⟨i, i + N ⟩, in the natively two-dimensional setting, only nearest neighbor interactions (both horizontal and vertical) are present, and the ‘hierarchical’ structure (or the lack thereof) is the same for all lattice sizes. It suffices, therefore, to only consider the spin-spin correlation curves of all 2D NQS ansatzes for the largest lattice size of N = 12.
8
(N, N ) = (12, 12), Bx = 3.0 Best Averaged −456.8887 −456.8643 0.0122 0.0123
NQS Ansatz Euclidean 2DRNN
Lorentz 2DRNN (Lmax = 2.0)
−456.9595 0.0084
-456.9121 0.0104
Lorentz 2DRNN (Lmax = 5.0)
-456.9670 0.0088
−456.8606 0.0124
DMRG (not exact)
-457.0416
Table 4: VMC results (best and averaged from different runs) of 2D TFIM with the lattice size (N, N ) = (12,12) at the fixed magnetic field strength Bx = 3.0. For each entry, we first recorded the mean energy in the upper line, followed by the corresponding standard error in the lower line. The best-performing results are noted in bold. From the results recorded in Tables 2 - 4, the following observations are made. • (N, N ) = (8,8), (10,10): – Lorentz 2DRNN (with Lmax = 2.0) definitively outperformed Euclidean 2DRNN for both lattice sizes, both in terms of the best and averaged mean energy reachable. • (N, N ) = (12,12): – In this case, we considered two different Lorentz 2DRNN NQS, one with Lmax = 2.0 and one with Lmax = 5.0. In terms of the best energy reachable from multiple individual runs, Lorentz 2DRNN with Lmax = 5.0 is the best ansatz, but in terms of the averaged energy reachable from all individual runs taken together, Lorentz 2DRNN with Lmax = 2.0 is the best ansatz. – In terms of the averaged spin-spin correlation decay length ⟨σ0z σrz ⟩ (see the log-log plots in Fig.3), Lorentz 2DRNN with Lmax = 5.0 shows the best performance, in the form of an averaged curve that conforms the best to the 3D CFT line6 compared to both Euclidean 2DRNN and Lorentz 2DRNN with Lmax = 2.0. The latter two show very close spin-spin correlation decay curves, with Euclidean 2DRNN performing slightly better than Lorentz 2DRNN with Lmax = 2.0. In Fig.4, where individual correlation curves resulting from different VMC runs are shown, the Lorentz 2DRNN (Lmax = 5.0) curve shows a slower correlation decay, with a slope that better matches the power-law exponent in Eq.(18) than the corresponding Euclidean 2DRNN’s curve for every single random seed considered. In Fig.5, when comparing the Lmax = 2.0 against the Lmax = 5.0 for each VMC run (using the same seed), the curves belonging to the Lmax = 5.0 curve lie well above those belonging to Lmax = 2.0 for two out of three runs (in the remaining run, the two curves are identical). – While both Lorentz 2DRNN considered in this case could outperform the Euclidean 2DRNN in terms of lower reachable energy, their behaviors are very different when it comes to capturing the long range correlation near the phase transition point, with one Lorentz 2DRNN clearly performing much better than the other. This highlights the importance of choosing the right Lmax to ensure the optimal performance of the hyperbolic Lorentz 2DRNN. • Overall, in this series of experiments with fixed magnetic field strength Bx = 3.0 at criticality, Lorentz 2DRNN has shown a consistent and definitive outperformance compared to Euclidean 2DRNN. This is interesting because this probably has very little to do with the structure of the neighbor interactions, which is not that hierarchical (since it only comprises nearest neighbor interactions) compared to the unrolled 2D case studied in [12] and also in the next section. The natural question that arises is, if not because of the hierarchical advantage associated with hyperbolic space that has led to the outperformance of hyperbolic NQS compared to their Euclidean counterpart as seen previously in [12], [13], what can account for this observed performance trend ? A possible answer to this question might be found in the fact that at criticality, the physics of a (D − 1)dimensional quantum TFIM system can be described by D-dimensional CFT [46], [47]. CFTs in D 6 On a log-log scale, the power-law relation Eq.(18) becomes a straight line and the exponent γ = D − 2 + η in the denominator becomes the slope of this line.
9
spatial dimensions, on the other hand, are known from the celebrated AdS/CFT correspondance [45], to be dual to an Anti-de-Sitter (AdS) space in (D + 1) dimensions, whose spatial geometry is that of a D-dimensional hyperbolic space. In this particular case under study, we have 2DTFIM at criticality described by a 3DCFT dual to an AdS4 space with a hyperbolic H3 spatial geometry. In this sense, the two-dimensional hyperbolic space underlying the construction of the Lorentz 2DRNN 7 might be the exact reason why hyperbolic Lorentz 2DRNN NQS outperforms its Euclidean counterpart at a point where the system displays conformal field theory behaviors8 . Other supporting evidences for the connection between the underlying hyperbolic geometry of Lorentz 2DRNN NQS playing a decisive role in its outperformance over Euclidean 2DRNN NQS in TFIM at criticality come from the seminal works [39], [40] that proposed a type of variationial ansatz known as MERA (Multiple Entanglement Renormalization Ansatz) that was shown to be particularly suited to represent quantum ground states at criticality [40]. In particular, MERA is a type of tensor network ansatz consisting of two types of tensors - isometries (w) and disentanglers (u) - arranged in a hierarchical, tree-like manner. Its winning point over other types of tensor network ansatzes such as DMRG in 1D and PEPS (Projected Entangled Pair States) or TTN (Tree Tensor Network) in 2D lies in the fact that it was specifically designed to represent quantum states at criticality with long range entanglement and power-law correlation [40], [41]. In [42], it was realized that the discrete minimal-cut paths through a MERA network is mathematically identical to the spatial geodesic in a continuous 2D hyperbolic/AdS geometry. This means that MERA is a discrete, skeletal version of an emergent holographic spacetime, where the vertical layer depth (z) is the extra radial dimension of AdS space [43]. In other words, given that MERA, the well-established tensor ansatz designed to simulate critical quantum states, is fundamentally discrete hyperbolic geometry, it stands to reason that our hyperbolic NQS constructions with their natively continuous hyperbolic geometry are able to automatically exploit the natural exponential volume expansion of the underlying space to capture the power-law correlations of TFIM critical quantums states much more efficiently than Euclidean NQS ever could. Furthermore, while MERA runs into the problem of increasing computational complexity in 2D because of the exponential number of tensors required in their construction, 1D and 2D hyperbolic NQS ansatzes constructed in this work and in our previous works [12], [13] do not have this problem since they leverage continuous neural network optimization while capturing the same hyperbolic geometry advantage. While more studies involving 2D hyperbolic NQS in larger 2DTFIM systems (such as (14,14), (16,16) or even (20,20) and beyond) are required to verify this hypothesis, at this point, given the results of this work as well as related evidences involving MERA in the literature, this possibility seems to offer a compelling explanation.
7 Despite the apparent 2D-3D mismatch in the dimensionality of the hyperbolic space H2 in which the Lorentz 2DRNN is constructed and the hyperbolic space H3 that is the spatial section of the AdS4 space dual to a 3DCFT, hyperbolic Lorentz 2DRNN still possesses the ‘correct’ geometry to represent the ground state wavefunction of the 2DTFIM at criticality much better than Euclidean 2DRNN. 8 Recall that in the 1DTFIM VMC experiments performed at the critical field strength B = 1.0 in [12], for N = 80 spins, 1D x Poincaré GRU NQS did not outperform Euclidean GRU NQS but for N = 100 spins Poincaré GRU NQS did. However, these results are not complete in the sense that other types of 1D hyperbolic NQS ansatzes such as Lorentz RNN/GRU and Poincaré RNN were not included at that time since the experiments were performed prior to their proposed constructions in [13]. Also, the construction of the Poincaré GRU NQS in [12] was restricted to the case where the nonlinear activation f ⊗c in Eq.5 is the identity function while other possibilities like ‘tanh’ and ‘elu’ exist.
10
Figure 2: A comparison of the performances of Euclidean 2DRNN and Lorentz 2DRNN (both best and averaged from different runs) in the VMC experiments involving 2DTFIM with the lattice sizes of (N, N )=(8,8), (10,10), (12,12) at a fixed magnetic field strength Bx = 3.0. For (8, 8) and (10,10) cases, the Lorentz 2DRNN has Lmax = 2.0 while for (12, 12), two different Lorentz 2DRNN NQS, with Lmax = 2.0, 5.0 were considered.
11
Figure 3: The spin-spin correlation decay length curves versus distance r for Lorentz 2DRNN (with Lmax = 2.0 and Lmax = 5.0) and Euclidean 2D RNN NQS averaged across three different seeds in 2DTFIM VMC experiments with the lattice size of (N, N )= (12,12) at a fixed magnetic field strength Bx = 3.0. The red dashed line corresponds to the 3D CFT power-law spin-spin correlation ⟨σ0z σrz ⟩ ∝ 1/|r|D−2+η where D = 3 and η = 0.0363. For the p 2D lattice setting, the distance between two spins (i1 , j1 ) and (i2 , j2 ) on the 2D lattice is calculated as r = (i1 − i2 )2 + (j1 − j2 )2 . We restricted the r range to 101 to filter out the fluctuations due to finite size effects and open boundary conditions.
12
Figure 4: The individual spin-spin correlation decay length curves for Lorentz 2DRNN and Euclidean 2D RNN NQS in 2DTFIM VMC experiments with the lattice size of (N, N )= (12,12) at a fixed magnetic field strength Bx = 3.0. The red dashed line corresponds to the 3D CFT power-law spin-spin correlation ⟨σ0z σrz ⟩ ∝ 1/|r|D−2+η where D = 3 and η = 0.0363. When seed=111, the two curves are identical and overlap. For the 2D p lattice setting, the distance between two spins (i1 , j1 ) and (i2 , j2 ) on the 2D lattice is calculated as r = (i1 − i2 )2 + (j1 − j2 )2 . We restricted the r range to 101 to filter out the fluctuations due to finite size effects and open boundary conditions. Note that while seed=333 displays a significantly flatter correlation profile than the other initializations, its final converged energy remains highly accurate and close to the true ground state.
Figure 5: The individual spin-spin correlation decay length curves for Lorentz 2DRNN (with Lmax = 2.0 and Lmax = 5.0) in 2DTFIM VMC experiments with the lattice size of (N, N )= (12,12) at a fixed magnetic field strength Bx = 3.0. The red dashed line corresponds to the 3D CFT power-law spin-spin correlation ⟨σ0z σrz ⟩ ∝ 1/|r|D−2+η where D = 3 and η = 0.0363. For the p 2D lattice setting, the distance between two spins (i1 , j1 ) and (i2 , j2 ) on the 2D lattice is calculated as r = (i1 − i2 )2 + (j1 − j2 )2 . We restricted the r range to 101 to filter out the fluctuations due to finite size effects and open boundary conditions.
13
4.2
Fixed lattice sizes, different magnetic strengths
While it is shown above that Lorentz 2DRNN outperform Euclidean 2DRNN quite definitively at criticality when Bx = 3.0, we are also interested in the performances of Lorentz 2DRNN when it comes to different magnetic field strengths away from the critical point. As such, we fix the lattice size of (N, N )=(12,12) and look at the cases where Bx = 2.0, 3.0, 4.0. The results of these experiments (both best and averaged) are recorded in Table 5 for Bx = 2.0 and Table 6 for Bx = 4.0. For Bx = 3.0, the results were already listed in Table 4 in the previous section. (N, N ) = (12, 12), Bx = 2.0 Best Averaged −346.8500 −346.7618 0.0087 0.0111
NQS Ansatz Euclidean 2DRNN
Lorentz 2DRNN (Lmax = 2.0)
−346.3433 0.0183
−346.3065 0.0188
Lorentz 2DRNN (Lmax = 6.0)
−346.4035 0.0172
−346.3426 0.0185
DMRG (not exact)
-346.9828
Table 5: VMC results (best and averaged from different runs) of 2D TFIM with the fixed size lattice (N, N ) = (12,12) at the transverse magnetic field strength Bx = 2.0. For each entry, we first recorded the mean energy in the upper line, followed by the corresponding standard error in the lower line. The best-performing results are noted in bold.
(N, N ) = (12, 12), Bx = 4.0 Best Averaged −593.5006 −593.3540 0.0072 0.0144
NQS Ansatz Euclidean 2DRNN
Lorentz 2DRNN (Lmax = 2.0)
−593.4953 0.0077
−593.3216 0.0139
Lorentz 2DRNN (Lmax = 1.5)
−593.4719 0.0068
−593.4648 0.0079
DMRG (not exact)
-593.5389
Table 6: VMC results (best and averaged from different runs) of 2D TFIM with the fixed size lattice (N, N ) = = (12,12) at the transverse magnetic field strength Bx = 4.0. For each entry, we first recorded the mean energy in the upper line, followed by the corresponding standard error in the lower line. The best-performing results are noted in bold. From Tables 4, 5, 6, the following observations are made. • At Bx = 3.0 (Table 4): When Bx = 3.0, the system is at the phase transition point with long-range spinspin correlation Eq.(18). As discussed in detail in the previous section, Lorentz 2DRNN outperformed Euclidean 2DRNN definitively, probably thanks to the fact that the physics of the system at this point displays conformal symmetry and describable by a CFT, which is dual to an AdS space whose spatial geometry is hyperbolic. Thus, the hyperbolic geometry underlying the construction of the Lorentz 2DRNN provides an inherently better match for NQS than Euclidean space. • At Bx = 2.0 (Table 5): When Bx = 2.0, the ground state is an ordered ferromagnetic system, and Euclidean 2DRNN outperformed Lorentz 2DRNN definitively, both in terms of best and averaged reachable ground state energy. In this case, where the entire ground state is ordered and uniform, the flat Euclidean geometry wins over the hyperbolic one, since there is no hierarchical structure nor CFT physics that offers hyperbolic geometry an advantage. • At Bx = 4.0 (Table 6): When Bx = 4.0, the system is a disordered paramagnet, and interestingly, the performances of Euclidean 2DRNN and Lorentz 2DRNN are comparable. In terms of the absolute best 14
energy obtained from individual VMC runs, Euclidean 2DRNN emerged as the better NQS, while in terms of the averaged best energy obtained from all VMC runs taken together, Lorentz 2DRNN emerged as the more stable and better NQS.
Figure 6: A comparison of the performances of Euclidean 2DRNN and Lorentz 2DRNN in the VMC experiments involving 2DTFIM with the lattice size of (N, N ) = (12, 12) at three different magnetic field strengths Bx = 2.0, 3.0, 4.0.
15
5
2D TFIM VMC experiments with 1D hyperbolic NQS
In this section, we report the results of the VMC experiments involving the recently constructed one-dimensional hyperbolic Poincaré and Lorentz RNN/GRU NQS ansatzes in [13] when these are applied to the problem of two-dimensional TFIM in the same manner to what was done in [12]. Our main aim in carrying out this type of experiments is to understand the representational capacity of hyperbolic NQS versus Euclidean ones in capturing the hierarchical neighbor interaction structure in the 2DTFIM setting unrolled in one dimension. As described in detail in [12], in translating or unrolling from two dimensions to one, the nearest horizontal and vertical neighbor interactions of the 2D N × N square lattice become the (i, i + 1) and (i, i + N ) interactions in one dimension. Specifically, if we start with a two-dimensional N × N square lattice of spins whose sites are labeled (i2D , j2D ) with 1 ≤ i2D , j2D ≤ N , in two dimensions, the horizontal and vertical nearest neighbor pairs are ⟨(i2D , j2D ), (i2D + 1, j2D )⟩ ,
⟨(i2D , j2D ), (i2D , j2D + 1)⟩ .
(19)
When treated as a one-dimensional system, this 2D square lattice become a spin chain of length N 2 = (N − 1)N + N , with the following mapping of site location: (i2D , j2D ) → [(i2D − 1)N + j2D ]
(20)
which leads to the 2D nearest neighbor interactions becoming the following 1D neighbor interactions, upon substituting i1D = (i2D − N ) + j2D 9 : ⟨(i2D , j2D ), (i2D + 1, j2D )⟩
→
⟨i1D , i1D + 1⟩
⟨(i2D , j2D ), (i2D , j2D + 1)⟩
→
⟨i1D , i1D + N ⟩
(22)
This means that the mapping of 2DTFIM into the setting of 1D TFIM results in the original 2D nearest neighbor interactions becoming the first neighbor interaction ⟨i, i + 1⟩ and N th neighbor interactions ⟨i, i + N ⟩ in one dimension. Thus, different 2D lattice size N results in inherently different 1D hierarchical structures10 , in constrast to keeping the 2D setting intact where different 2D lattice size N does not change the nature of the interaction being fundamentally first neighbor interactions only (along horizontal and vertical dimensions). In [12], we hypothesized that this hierarchy of interactions between the first and N th neighbor interactions was the main reason that 1D hyperbolic NQS - in the form of Poincaré GRU - outperformed its Euclidean version, the Euclidean GRU NQS in 2DTFIM systems of up to (N, N ) = (9, 9). In this work, we expand this result to larger systems including (N, N ) = (10, 10), (12, 12), with additional types of 1D hyperbolic NQS ansatzes including Poincaré RNN, Lorentz RNN and Lorentz GRU, in addition to the originally introduced Poincaré GRU in [12]. The six types of 1D NQS used to run the 2DTFIM VMC experiments are listed in Table 7. Similar to [13], these six NQS types encompasses two architecture variants: RNN and GRU, with two underlying geometries: Euclidean and hyperbolic (which further divided into the subgeometry of Poincaré disk and Lorentz hyperboloid). For these 1D NQS ansatzes, we are only interested in the experiments with varying lattice sizes at the fixed magnetic field strength of Bx = 3.0, which is practically at the phase transition point (since Bc = 3.044). In Table 8, we list the best VMC results corresponding to each these NQS (where the best is chosen from different VMC runs of one particular NQS ansatz), while in Table 9, we list the averaged results obtained from all different VMC runs. In Fig.7, we show their best performances in ascending order, while in Fig.8, we show their performances averaged over different VMC runs in ascending order. In the same manner as the 2D NQS case studied above, in addition to computing the mean ground state energy that each 1D NQS ansatz is capable of reaching, we also computed the spin-spin correlation length decay ⟨σ0z σrz ⟩ of each of the six 1D NQS ansatzes for each lattice size N = 8, 10, 12 (shown in Figs.9, 10, 11). It is important to note that for this particular situation where we have a one-dimensional setting arising from artificially unrolling the two dimensional system using Eq.(20), a ‘good’ 1D NQS is not one that is capable of producing a straight-line decaying curve conforming to Eq.(18), but one that is capable of learning the inherent two-dimensional geometry of the lattice. In other words, a good 1D NQS ansatz in this case would be one that is capable of capturing the N th -neighbor interactions ⟨i, i + N ⟩, and this shows up in the spin-spin correlation curve as a periodic pattern with the period corresponding exactly to the 2D lattice size N . In fact, the power-law 9 For later convenience, we also included the reverse mapping from 1D spin chain to 2D lattice:
i2D = (i1D //N ) + 1
j2D = (i1D mod N ) + 1 .
(21)
10 In Figs.12 and 13 in the Appendix, we illustrate different neighbor interaction hierarchical structures of different spin systems in the form of graphs.
16
correlation Eq.(18), from the 3D CFT that is equivalent to the 2DTFIM at criticality, when translated to the unrolled 1D setting, is ⟨σoz σrz ⟩1D =
A 2 ]1.0363/2 [i22D + j2D
,
(23)
where A is a proportionality constant chosen depending on the system under study, i2D and j2D are the 2D lattice coordinates written in terms of 1D spin chain site i1D as defined in Eq.(21). Eq.(23) describes an oscillating, periodic pattern whose period exactly coincides with the 2D lattice size N . More specifically, Eq.(23) is not a smooth sine-wave-like curve but rather a sawtooth waveform because of the discreteness of the spin site locations. In Figs. 9, 10, 11, the 3DCFT power relation in 1D form Eq.(23) is also included as a reference to gauge the performances of different NQS ansatzes in terms of their abilities to capture the correct spin-spin correlation decay characteristics. Ansatz Euclidean RNN
Euclidean GRU
Poincaré RNN
Poincaré GRU
Lorentz RNN
Lorentz GRU
Hidden dimension size 60
Parameters 3902
70
5252
60
11462
70
15472
60
3902
70
5252
60
11462
70
15472
60
3902
70
5252
60
11462
70
15472
Experiment (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12) (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12) (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12) (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12) (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12) (N, N ) = (8, 8) (N, N ) = (10, 10) (N, N ) = (12, 12)
Table 7: Six types of 1D Euclidean and hyperbolic Poincaré/Lorentz RNN/GRU NQS ansatzes used in the 2D TFIM VMC experiments. For (N, N ) = (8, 8), all ansatzes have the hidden RNN/GRU dimension of 50, while for (N, N ) = (10, 10), (12,12), the hidden dimension is 60. With the same hidden dimension, RNN variants have almost three times fewer parameters as GRU variants.
17
Ansatz Euclidean RNN Poincare RNN Lorentz RNN Euclidean GRU Poincare GRU Lorentz GRU DMRG (not exact)
(N, N ) = (8, 8) -201.4058 0.0314 -201.7627 0.0269 -201.5256 0.0315 -200.9939 0.0441 -201.7624 0.0296 -200.8296 0.0445 -202.5077
(N, N ) = (10, 10) -309.1322 0.0780 -314.8574 0.0517 -315.1620 0.0412 -314.1182 0.0611 -315.3490 0.0449 -314.6966 0.0537 -316.9770
(N, N ) = (12, 12) -447.3467 0.0852 -445.6974 0.0944 -445.4792 0.0938 -452.3138 0.0763 -454.2120 0.0636 -453.5092 0.0691 -457.0416
Table 8: This table lists the best results for each NQS ansatz (chosen from different runs) of 2D TFIM VMC experiments using 1D Euclidean and hyperbolic Poincaré/Lorentz RNN/GRU NQS ansatzes for different square lattice sizes (N, N ) at Bx = 3.0. The number of samples used for inference is 104 . For each NQS ansatz, we first list the mean energy in the first line, followed by the standard error directly below in the second line. For each (N, N ) experiment setting, the two best performing NQS ansatzes are noted in bold. Note that in all three cases, the top performing ansatzes are always hyperbolic NQS.
Ansatz Euclidean RNN Poincare RNN Lorentz RNN Euclidean GRU Poincare GRU Lorentz GRU DMRG (not exact)
(N, N ) = (8, 8) -197.8471 0.0528 -201.3019 0.0328 -200.5496 0.0409 -200.9467 0.0441 -201.6585 0.0312 -200.7991 0.0457 -202.5077
(N, N ) = (10, 10) -308.6596 0.0801 -311.4692 0.0705 -312.2401 0.0637 -313.9553 0.0617 -314.9809 0.0506 -314.1184 0.0591 -316.9770
(N, N ) = (12, 12) -444.7609 0.0889 -444.7816 0.0996 -445.2729 0.0929 -452.0062 0.0772 -453.8204 0.0668 -452.9004 0.0735 -457.0416
Table 9: This table lists the averaged results for each NQS ansatz across multiple different runs of 2D TFIM VMC experiments using 1D Euclidean and hyperbolic Poincaré/Lorentz RNN/GRU NQS ansatzes for different square lattice sizes (N, N ) at Bx = 3.0. The number of samples used for inference is 104 . For each NQS ansatz, we first list the mean energy in the first line, followed by the standard error directly below in the second line. For each (N, N ) experiment setting, the two best performing NQS ansatzes are noted in bold. Note that in all three cases, the top performing ansatzes are always hyperbolic NQS. From Table 8 Fig.7, Table 9 and Fig.8, as well as Figs.9, 10, 11 the following observations are noted. • (N, N ) = (8, 8): – In terms of the best results chosen from different VMC runs (see the first subfigure of Fig.7), the best three performing NQS ansatzes are Poincaré RNN (at −201.7627), Poincaré GRU (at −201.7624) and Lorentz RNN (−201.5256). Interestingly, Euclidean RNN could reach a best value of -201.4058, outperforming the best reachable by Euclidean GRU (−200.9939) and Lorentz GRU (−200.8296). The underperformance of Lorentz GRU in this case can be attributed to a sub-optimal selection of hyperparameters rather than the inherent ability of Lorentz GRU as an NQS ansatz. In terms of the performances averaged over different VMC runs (see top subfigure of Fig.8), Poincaré GRU is the best NQS, followed by Poincaré RNN, each with the energy average in the range (−201.70, −201.30). Euclidean RNN, however, is the worst ansatz, with an energy average of −197.8471, underperfroming both Lorentz GRU and Euclidean GRU. – In terms of the spin-spin correlation decay ⟨σ0z σrz ⟩, Fig.9, all curves show a peak of correlation at r = 8, signaling their abilities to capture the pattern of 2D-1D mapping Eq.(20) when the 2D lattice size is 18
N = 8. Among the RNN-based variants, Euclidean RNN has the flattest curve, followed by Lorentz RNN and Poincare RNN. The GRU-based variants have almost identical curves which all show a clear periodic pattern that is much more pronounced than the RNN-based variants, demonstrating their superior ability to capture the hierarchical interaction patterns ⟨i, i + 1⟩ - ⟨i, i + 8⟩. Note that while the 3D CFT reference curve shows a relatively flat decay at large distance, the more accurate picture is one in which the correlation drops sharply near the edges due to the open boundary conditions. This is what is observed for the curves produced by the GRU variants, which accurately capture this sharp decay at large distances. • (N, N ) = (10, 10): – In terms of the best value reachable by each NQS (see the second subfigure of Fig.7), Poincaré GRU (−315.3490) is the best NQS, followed by Lorentz RNN (−315.1620) and Poincaré RNN (−314.8574) and Lorentz GRU at the fourth place. In terms of the averaged results over different VMC runs (see the second subfigure of Fig.8), the best performing ansatzes are Poincaré GRU, Lorentz GRU and Euclidean GRU. – In terms of the spin-spin correlation ⟨σ0z σrz ⟩, Fig.10, all curves, except Euclidean RNN’s, show a peak at r = 10. However, the curves corresponding to the GRU-based variants (which are again almost identical among the three variants of Euclidean, Poincaré and Lorentz GRU) are much more periodic than the RNN-based ones which are almost flat by comparison. This again signals the better ability of the GRU variants at capturing the 2D-1D mapping Eq.(20) and interaction hierachy ⟨i, i + 1⟩ ⟨i, i + 10⟩. • (N, N ) = (12, 12): – In terms of best energy value chosen from different runs, Poincaré GRU (−454.2120) is again the best performing ansatz, followed by Lorentz GRU (−453.5092), and Euclidean GRU (−452.3138), far exceeding the RNN variants where Euclidean RNN (at −447.3467) is the best performing variant. In terms of the averaged performances across different runs (Table 9), the top three spot remain exactly the same, but among the bottom three spots, Lorentz RNN ranked first, followed by Poincaré RNN and Euclidean RNN. It is interesting to note that among the RNN variants, Euclidean RNN actually outperformed Lorentz and Poincaré RNN in terms of best E reachable, but underperformed both in terms of averaged E reachable. The underperformance of Lorentz and Poincaré RNN does not signal their lack of inherent expressivity as NQS ansatz but rather is due to a lack of comprehensive hyperparameter tuning to select the optimal configuration. – In terms of the spin-spin correlation decay in Fig.11, the three NQS in the RNN variants failed completely to capture the 12th neighbor interaction, as their curves are completely flat. while the GRU variants still successfully captured this 2D-1D mapping in this case. Among the GRU variant, hyperbolic Poincaré and Lorentz GRU outperformed Euclidean GRU, as evidenced by their corresponding curves - the hyperbolic ones clearly show the periodic structure with the r = 12 period, even at large distances while the Euclidean one flattened out quickly with distance. In particular, at intermediate distances (20 < r < 70), the Poincaré and Lorentz GRUs traced both the peaks and the valley depths of the translated 3D CFT baseline with high fidelity, whereas the Euclidean GRU failed to decay deeply enough in the valleys due to an unphysical long-range correlation tail. At extreme distances (r > 80), the Euclidean curve appeared to match the continuum CFT line better, but this is a known artifact of under-optimization. The hyperbolic GRUs correctly captured the finite-size boundary effect by plunging at the tail, while the Euclidean network over-estimated long-range alignment. • Overall, across three different metrics (absolute best value reachable, averaged performance, spin-spin correlation decay length), at three different lattice sizes (N, N )=(8,8),(10,10), (12,12), hyperbolic NQS ansatzes clearly outperformed Euclidean ones. In particular, Poincaré GRU emerged as the best performer for all three 2DTFIM settings, outperforming all other NQS ansatzes, including Euclidean GRU and Lorentz GRU. This observation agrees with and extends our previous result from [12] where we showed that Poincaré GRU consistently outperformed Euclidean GRU NQS with different lattice sizes from (N, N ) = (5,5) to (9,9). In this work, as the lattice size increases from (8, 8) to (10, 10) to (12, 12), the GRU variants increasingly demonstrated their superior abilities over the RNN variants to capture the ever-larger hierarchical N th neighbor interaction where N goes from 8, 10, to 12 as seen from the series of spin-spin correlation curves in Figs.9, 10, 11. Among the GRU variants, as N increases to 12, hyperbolic Lorentz and Poincaré GRU showed a clear dominance over Euclidean GRU.
19
While the probable main reason for the clear outperformance achieved by 1D hyperbolic NQS over Euclidean ones in this unrolled 2DTFIM setting is the presence of the structural hierarchy ⟨i, i + 1⟩-⟨i, i + N ⟩ of neighbor interactions, one must also not discount the fact that these VMC experiments are performed at the critical magnetic field strength Bx = 3.0 where the system displays CFT physics. As pointed out in the previous section, the connection to the AdS/CFT correspondance can also be a determining factor in the observed outperformance by hyperbolic 2D NQS over Euclidean 2D NQS in the lack of structural hierarchy. In this case, we have both structural hierarchy and CFT physics which can both contribute to the advantage of hyperbolic 1D NQS. Furthermore, we hypothesize that if the VMC experiments in the unrolled 2DTFIM setting were to be done at Bx = 2.0, 1D hyperbolic NQS would still outperform 1D Euclidean NQS despite the lack of a CFT physics description because the long range, ordered nature of the 2D ground state at Bx = 2.0 demands that this ordered correlation be preserved in the hierarchical structure ⟨i, i + 1⟩-⟨i, i + N ⟩ in the 1D setting. For Bx = 4.0, it is less certain that 1D hyperbolic NQS would outperform 1D Euclidean NQS in the unrolled 2DTFIM because the disordered nature of the 2D ground state at Bx = 4.0 weakens the ⟨i, i + 1⟩-⟨i, i + N ⟩ structure.
20
Figure 7: The best performances of 1D Euclidean and hyperbolic Poincaré/Lorentz NQS ansatzes, chosen from different runs, ranked in ascending order for each of the 2D TFIM VMC experiment setting (N, N ) = (8, 8), (10,10), (12,12).
21
Figure 8: The performances of 1D Euclidean and hyperbolic Poincaré/Lorentz NQS ansatzes, averaged over different runs, ranked in ascending order for each of the 2D TFIM VMC experiment setting (N, N ) = (8, 8), (10,10), (12,12).
22
Figure 9: The ⟨σ0z σrz ⟩ spin-spin correlation decay length curves for all six types of 1D NQS ansatzes from the 2D TFIM with (N, N ) = (8, 8). From top to bottom: In the top subfigure, the curves for all six ansatzes are shown together, in the middle subfigure, the zoomed-in curves of the three RNN-based NQS variants are shown, in the bottom subfigure, the zoomed-in curves of the three GRU-based NQS variants are shown. Note that due to the open boundary conditions, the correlation length curves should decay much faster at larger distances (due to the decrease of neighbors for spin sites nearing the lattice edge). The fluctuations seen at large r are due to finite size effects of the lattice. The black, dashed line is the 3D CFT power relation Eq.18 translated to one-dimensional setting as in Eq.23. The exact 3D CFT baseline exhibits a sharp sawtooth profile because it maps a continuous, isotropic physical power-law (r−d ) onto a discrete 1D unrolled sequence of 2D lattice coordinates.
23
Figure 10: The ⟨σ0z σrz ⟩ spin-spin correlation decay length curves for all six types of 1D NQS ansatzes from the 2D TFIM with (N, N ) = (10, 10). From top to bottom: In the top subfigure, the curves for all six ansatzes are shown together, in the middle subfigure, the zoomed-in curves of the three RNN-based NQS variants are shown, in the bottom subfigure, the zoomed-in curves of the three GRU-based NQS variants are shown. Note that due to the open boundary conditions, the correlation length curves should decay much faster at larger distances (due to the decrease of neighbors nearing the lattice edge). The fluctuations seen at large r are due to finite size effects of the lattice. The black, dashed line is the 3D CFT power relation Eq.18 translated to one-dimensional setting as in Eq.23. The exact 3D CFT baseline exhibits a sharp sawtooth profile because it maps a continuous, isotropic physical power-law (r−d ) onto a discrete 1D unrolled sequence of 2D lattice coordinates.
24
Figure 11: The ⟨σ0z σrz ⟩ spin-spin correlation decay length curves for all six types of 1D NQS ansatzes from the 2D TFIM with (N, N ) = (12, 12). From top to bottom: In the top subfigure, the curves for all six ansatzes are shown together, in the middle subfigure, the zoomed-in curves of the three RNN-based NQS variants are shown, in the bottom subfigure, the zoomed-in curves of the three GRU-based NQS variants are shown. Note that due to the open boundary conditions, the correlation length curves should decay much faster at larger distances (due to the decrease of neighbors nearing the lattice edge). The fluctuations seen at large r are due to finite size effects of the lattice. The black, dashed line is the 3D CFT power relation Eq.18 translated to one-dimensional setting as in Eq.23. The exact 3D CFT baseline exhibits a sharp sawtooth profile because it maps a continuous, isotropic physical power-law (r−d ) onto a discrete 1D unrolled sequence of 2D lattice coordinates.
25
6
Concluding remarks
In this work, we introduce the Lorentz 2DRNN, which is the first two dimensional non-Euclidean NQS construction. We benchmark the performances of Lorentz 2DRNN against its Euclidean counterpart, the twodimensional Euclidean 2DRNN - a custom construction introduced in [7] in a variety of different 2DTFIM VMC experiment settings involving different lattice sizes (N, N ) at different magnetic field strengths Bx . In addition to the newly constructed two-dimensional hyperbolic Lorentz 2DRNN, we also expanded the results of [12] concerning the use of one-dimensional hyperbolic NQS in the 2DTFIM setting where the two dimensional system is translated/unrolled into one dimension. In this work, the 1D NQS includes not only the original Poincaré GRU first introduced in [12], but also new 1D hyperbolic NQS such as Lorentz RNN, Lorentz GRU and Poincaré GRU introduced in [13]. Our main findings are summarized below. • Two-dimensional NQS : Lorentz 2DRNN NQS defnitively outperform Euclidean 2DRNN in all three system sizes 8 × 8, 10 × 10, 12 × 12 at Bx = 3.0 (at the phase transition point) both in terms of best and averaged reachable mean ground state energy (see Table 2, Table 3, Table 4). In terms of the spin-spin correlation ⟨σ0z σrz ⟩ in the case of 12×12 2DTFIM (see Fig.3), the power-law correlation decay is best reproduced by a Lorentz 2DRNN NQS with Lmax = 5.0 among the three NQS considered: Euclidean 2DRNN, Lorentz 2DRNN with Lmax = 2.0 and Lorentz 2DRNN with Lmax = 5.0. Away from the phase transition point, when Bx = 2.0, Euclidean 2DRNN NQS definitively outperform hyperbolic Lorentz 2DRNN (see Table 5), while for Bx = 4.0, the performances of the two are comparable where neither one emerge as a definitive better NQS (see Table 6). As described in detail in Section 4.1, a possible reason for the observed outperformance of Lorentz 2DRNN compared to Euclidean 2DRNN at the phase transition point has to do with the fact that the physics of the 2DTFIM system at criticalily is described by a conformal field theory that is known to be dual to an AdS space whose spatial geometry is none other than hyperbolic space. Thus, the hyperbolic space underlying the construction of the Lorentz 2DRNN naturally endows it with a representational capacity that is well-matched to the physics of the critical system. This is reminiscient of the way another nonNQS variational ansatz, MERA (Multiple Entanglement Renormalization Ansatz), is known to efficiently represent quantum states at criticality [40], [42], [43]: MERA, with its exponential tree-like structure of isometries and disentanglers is fundamentally a discrete version of hyperbolic space whose geometry matches the CFT physics of the critical quantum states. While both our hyperbolic NQS and MERA share a hyperbolic geometry advantage when it comes to representing critical quantum states, their main difference, apart from the obvious fact that one is neural-network-based and one is tensor-based, lies in the continuous versus discrete nature of their hyperbolic constructions. As MERA builds a discrete, layer-by-layer hyperbolic tree to represent a state, our hyperbolic NQS utilizes a continuous hyperbolic manifold to achieve a structurally parallel representation of critical spatial scaling. While more studies are required for even larger system sizes beyond N = 12, we hypothesize that this might be the reason that hyperbolic 2D NQS ansatzes (such as the Lorentz 2DRNN constructed in this work, and other yet-to-beconstructed hyperbolic 2D NQS variants such as Lorentz 2DGRU Poincare 2DRNN/2DGRU) outperform their Euclidean 2D NQS counterparts at the phase transition point when there is a CFT description of the physical system. • One-dimensional NQS : 1D hyperbolic NQS definitively outperform Euclidean NQS in all three lattice sizes considered (N, N ) = (8,8), (10,10), (12,12) at Bx = 3.0 (see Table 8 and Table 9, as well as Fig.7 and Fig.8). In terms of the mean energy values, Poincaré GRU emerge as the best overall NQS ansatz in all three cases under study, with the second best ansatz varies in each case: Lorentz GRU for N = 12, Lorentz RNN for N = 10 and Poincaré RNN for N = 8. The use of 1D NQS necessitates the conversion of the N × N 2DTFIM into a one-dimensional setting where the vertical nearest neighbors in the 2D lattice become the N th neighbor in the 1D spin chain, which means that different lattice size N leads to a different type of hierarchical structure such that the larger N is, the harder it becomes for the 1D NQS ansatz to capture this structure. This is clearly seen in the spin-spin correlation length ⟨σ0z σrz ⟩ plots, Fig.9, Fig.10 and Fig.11, where a successful 1D NQS is one that displays the periodic structure in ⟨σ0z σrz ⟩ with the period corresponding to the lattice size N . When (N, N ) = (8,8), all ansatzes - including 1D Euclidean RNN - display the periodic pattern in their curves with the period of r = 8. When (N, N )=(10,10), all ansatzes - excluding Euclidean RNN - display the periodic pattern with the period of r = 10. The GRU variants display a much more pronounced periodic structure in their correlation curves than the RNN variants, whose curves are far flatter by comparison, for both these system sizes. However, as the system size increases to (N, N )=(12,12), only the GRU variants 26
display the periodic pattern with the period of r = 12, while the RNN variant curves are completely flattened out. Even among the three GRU variants, the Euclidean GRU curve is noticeably flatter compared the Poincaré’s and Lorentz GRU’s. Apart from the obvious structural hierarchy ⟨i, i + 1⟩-⟨i, i + N ⟩ in the neighbor interactions that can account for the outperformance of 1D hyperbolic NQS compared to their Euclidean counterparts as pointed out in [12], in this work, we also note that the presence of the CFT physics at the critical point where the VMC experiments are carried out also plays an important part. In particular, it is a combination of both these factors (structural hierarchy and CFT physics) that contribute to the observed outperformance of 1D hyperbolic NQS at the critical point in the unrolled 2DTFIM setting. While we did not perform the VMC experiments for the unrolled 2D setting at other magnetic field strengths besides Bx = 3.0, we hypothesize that at Bx = 2.0 when the ground state is an ordered ferromagnet, in the absence of CFT physics, 1D hyperbolic NQS would continue to outperform their Euclidean versions thanks to the fact that the ‘orderedness’ nature of the state can still be captured effectively through structural hierarchy. For Bx = 4, when the system is in a disordered state, it is uncertain whether 1D hyperbolic NQS would have an advantage over 1D Euclidean NQS since structural hierarchy does not play a role in conveying the ‘disorderedness’ information of the system. Given the results of this work, several interesting future directions emerge. As already noted in [12], [13], of immdiate interests might be the problem of extending the Lorentz 2DRNN construction in this work to the more complex 2DGRU version or even more broadly, exploring other constructions of both 1D and 2D hyperbolic NQS based on different types of architectures such as CNN (Convolutional Neural Network)11 or transformers and applying these new constructions to the paradigmatic settings of either TFIM or Heisenberg models. Another immediate direction might involve a new construction of the already-proposed 1D/2D hyperbolic NQS12 in this work and [12], [13] based on a learnable curvature parameter for both the Poincaré disk and the Lorentz hyperboloid, which might remove the need to artificially introduce the spatial constraint hyperparameters Lmax and Rmax . Furthermore, it would be crucial if more efficient and systematic hyperparameter tuning methods for hyperbolic NQS could be implemented so that optimal hyperparameters can be selected rather than relying on a trial-and-error basis as was done in our works, due to a lack of computational resources. Going farther afield beyond the context of spin models, it might also be interesting and instructive to explore the performances of hyperbolic NQS in different quantum systems, one example of which is the SU (N ) matrix models that have recently been explored in the context of variational quantum circuits [37], [44]. We hope to return to these issues in future works.
11 It might also be interesting to explore a hybrid version of CNN-RNN architecture as done in [18] 12 In our series of works on hyperbolic NQS so far, we have always used a fixed curvature constant of c = 1 for the Poincaré networks and k = 1 for the Lorentz networks.
27
A
Appendix
A.1
Hierarchical structures of neighbor interactions in spin systems
In this section, we illustrate the hierarchical neighbor interactions of different spin systems including TFIM and Heisenberg J1 J2 and J1 J2 J3 systems in the form of graphs in Fig.12 and Fig.13. Each type of neighbor interaction, be it first (⟨i, i + 1⟩), second (⟨i, i + 2⟩), third (⟨i, i + 3⟩) or N th (⟨i, i + N ⟩), is represented by an edge connecting vertex i and vertex j (where j = 1, 2, 3, N ) in a graph whose vertices are the spin sites. In this manner, the neighbor interaction structures corresponding to the Heisenberg J1 J2 and J1 J2 J3 spin systems are shown in Fig.12, while the neighbor interaction structures corresponding to the 2DTFIM mapped to the 1DTFIM using Eq.20 are shown in Fig.13. In this case, the 2D lattice sizes (N, N ) = (8, 8), (10,10), (12,12) correspond to the 8th - ⟨i, i + 8⟩, 10th - ⟨i, i + 10⟩ and 12th - ⟨i, i + 12⟩ neighbor interactions in the 1D spin chain. Collectively, these graphs can be very loosely categorized as Watts-Strogatz [48] type, or ‘small-world’ networks consisting of vertices that are connected to a fixed number of neighboring vertices to form a network. This model is capable of interpolating between regular and random networks by using a certain nonzero ‘rewiring’ probability p that randomly rewires the connections between vertices. In our context, the neighbor interaction structures of spin models correspond to Watts-Strogatz graphs with zero rewiring probability. Incidentally, in the graph-embedding and graph classification context, hyperbolic neural networks have been shown to outperform their Euclidean counterparts in these tasks thanks to their more efficient representational capacity [49].
Figure 12: An illustration of the interaction hierarchies of different degrees of nearest neighbor interactions in the form of graphs in spin models for 1D Heisenberg J1 J2 and J1 J2 J3 spin chain models. For illustration purpose, the number of sites is chosen to be N = 18.
28
Figure 13: An illustration of the interaction hierarchies of different degrees of nearest neighbor interactions in the form of graphs for 8 × 8, 10 × 10 and 12 × 12 2DTFIM mapped to 1DTFIM spin chain model where the vertical nearest neighbor in the N × N lattice becomes the N th neighbor in the 1D spin chain. For illustration purpose, the number of sites is chosen to be 24, 30 and 36.
A.2
Poincaré disk model of hyperbolic space
The Poincaré ball model (DN , g D ) of hyperbolic space is defined by the manifold DN N DN c = x ∈ R : c||x|| < 1
(24)
N N where the parameter √ c is the Poincaré ball’s radius. When c = 0, DN c = R , while when c > 0, Dc is the N open ball of radius 1/ c. In all the computations used in this work, we set c = 1. The space Dc in Eq.(24) is equipped with the metric
gxD = λ2x g E ,
λx =
2 1 − ||x||2
(25)
c (v) of a where g E = 1N is the identity matrix representing the Euclidean metric. The parallel transport P0→x N N vector v ∈ T0 Dc to another tangent space Tx Dc are defined as c P0→x (v) = logcx [x ⊕c expc0 (v)]
(26)
where the operations log and exp are the exponential and logarithmic maps. For any point x ∈ DN c , any vector c M N N v ̸= 0 and any point y ̸= x, the exponential and logarithmic maps expcx : Tx DM → D and log : c c x Dc → Tx Dc between the hyperbolic space and its Euclidean tangent space are defined as √ λcx ||v|| v √ expcx (v) = x ⊕c tanh c (27) 2 c||v|| −x ⊕c y √ 2 logcx (y) = √ c tanh−1 c|| − x ⊕c y|| (28) || − x ⊕c y|| cλx 29
where λcx = 2/(1 − c||x||2 ). When x = 0, the above maps take more compact forms expc0 (v)
=
logc0 (y)
=
√
v c||v|| √ c||v|| y −1 √ c||y|| √ tanh c||y||
tanh
M T0M DM c → Dc
N DN . c → T0N Dc
(29) (30)
The mathematical operations ⊕c , ⊗c , ⊙c appearing in the equations above are the Poincaré hyperbolic analogs of the Euclidean addition, matrix multiplication and pointwise multiplication. • For x, y ∈ DN c , the Mobius addition ⊕c is x ⊕c y ≡
(1 + 2c⟨x, y⟩ + c||y||2 )x + (1 − c||x||2 )y . 1 + 2c⟨x, y⟩ + c2 ||x||2 ||y||2
(31)
N In terms of parallel transport, the Mobius addition for x ∈ DN c with b ∈ Dc can be written as h i c x ⊕c b = expcx P0→x logc0 (b) . M • For x ∈ DN → RN , the Mobius matrix multiplication is defined by c \{0} and W : R √ Wx 1 ||W x|| tanh−1 ( c||x||) W ⊗c x = √ tanh ||x|| ||W x|| c
• For x ∈ DN c \{0} and r ∈ R, the Mobius pointwise multiplication ⊙c , is defined as √ 1 ||rx|| rx r ⊙c x = √ tanh tanh−1 ( c||x||) ||x|| ||rx|| c
(32)
(33)
(34)
• The nonlinear activation f ⊗c (x) where x ∈ DN c is f ⊗c = expc0 (f (logc0 (x)))
A.3
(35)
Lorentz hyperboloid model of hyperbolic space
The n-dimensional Lorentz hyperboloid Hn with constant negative curvature −k where k = 1 is given as Hn ≡ x ∈ Rn+1 : ⟨x, x⟩L = −1, x0 > 0 (36) where ⟨x, y⟩L , the Lorentzian scalar product for (n + 1)-dimensional vectors x, y ∈ Rn+1 , is defined as ⟨x, y⟩L ≡ −x0 y0 +
n X
xi yi .
(37)
i=1
Unlike the Poincaré disk where the origin 0P = (0, 0, . . . , 0), in the Lorentz model of hyperbolic space, the origin is the point 0L = (1, 0, 0, . . . , 0). The distance function between two points x, y ∈ Hn is dH (x, y) = arcosh (−⟨x, y⟩L ) ,
(38)
while the Lorentzian norm of a vector v is ||v||L =
p
⟨v, v⟩L .
(39)
The tangent space at x is the n-dimensional Euclidean vector space approximating Hn around x: Tx Hn ≡ x ∈ Rn+1 : ⟨v, x⟩L = 0
(40)
Similar to the Poincaré model described in the previous section, mappings between the Lorentz hyperboloid and its tangent space are done using the exponential expx (v) and logarithmic logx (y) maps. Tx Hn → Hn :
expx (v) = cosh (||v||L ) x + sinh (||v||L )
Hn → Tx Hn :
logx (y) = dH (x, y) 30
y + ⟨x, y⟩L x ||y + ⟨x, y⟩L x| |L
v ||v||L
(41) (42)
The parallel transport operation that maps a point z ∈ Tx Hn to a point in Ty Hn is defined as Px→y (z) = z +
⟨y, z⟩L (x + y) 1 − ⟨(x, y)⟩L
(43)
P0L →y (z) = z +
⟨y, z⟩L (0L + y) 1 − ⟨(0L , y)⟩L
(44)
When x = 0L ,
The various Lorentz mathematical operations appearing Eqs.(6), (7) are defined in terms of the exponential/logarithm mappings at x = 0L in (41), (42) and parallel transport operations given in (44) as follows. • For x, y ∈ Hn , the Lorentz addition ⊕L is defined as: Lorentz addition :
x ⊕L y = expx P0L →x (log0L (y))
(45)
• The scalar multiplication ⊙L between a point x ∈ Hn and r ∈ R is Lorentz scalar multiplication :
r ⊙L x = exp0L r log0L (x)
• The matrix multiplication ⊗L between a point x ∈ Hn and a matrix M ∈ Rn × Rn is M ⊗L x = exp0L M log0L (x)
(46)
(47)
• The nonlinear activation f ⊗L (x) where x ∈ Hn is f ⊗L = exp0L f (log0L (x))
31
(48)
References [1] G. Carleo and M. Troyer, Solving the Quantum Many-Body Problem with Artificial Neural Networks, Science 355, 602 (2017), arXiv:1606.02318 [cond-mat.dis-nn] [2] L. Huang and L. Wang, Accelerate Monte Carlo Simulations with Restricted Boltzmann Machines, arXiv:1610.02746v2. [3] Z. Cai and J. Liu, Approximating quantum many-body wave-functions using artificial neural networks,Phys. Rev. B 97, 035116 (2018), arXiv:1704.05148 [cond-mat.str-el] [4] H. Saito and M. Kato, Machine learning technique to find quantum many-body ground states of bosons on a lattice, J. Phys. Soc. Jpn. 87, 014001 (2018), arXiv:1709.05468 [cond-mat.dis-nn] [5] X. Liang, Wen-Yuan Liu, Pei-Ze Lin, Guang-Can Guo, Yong-Sheng Zhang, and Lixin He, Solving frustrated quantum many-particle models with convolutional neural networks, arXiv: 1807.09422v2 [6] C. Roth, A. Szabó, and A. H. MacDonald, High-accuracy variational Monte Carlo for frustrated magnets with deep neural networks, Phys. Rev. B 108, 054410 (2023), arXiv:2211.07749v2 [cond-mat.str-el] [7] M. Hibat-Allah, M. Ganahl, L. E. Hayward, R. G. Melko, and J. Carrasquill, Recurrent neural network wave functions, Physical Review Research 2, 023358 (2020). [8] M. Hibat-Allah, R. G. Melko, J. Carrasquilla, Supplementing Recurrent Neural Network Wave Functions with Symmetry and Annealing to Improve Accuracy, Machine Learning and the Physical Sciences, NeurIPS 2021, arXiv:2207.14314v2 [cond-mat.dis-nn] [9] M. Hibat-Allah, E. Merali, G. Torlai, R. G. Melko and J. Carrasquilla, Recurrent neural network wave functions for Rydberg atom arrays on kagome lattice, arXiv:2405.20384v1 [cond-mat.quant-gas] [10] K. Sprague and S. Czischek, Variational Monte Carlo with Large Patched Transformers, Commun Phys 7, 90 (2024), arXiv:2306.03921 [quant-ph] [11] H. Lange, G. Bornet, G. Emperauger, C. Chen, T. Lahaye, S. Kienle, A. Browaeys, A. Bohrdt, Transformer neural networks and quantum simulators: a hybrid approach for simulating strongly correlated systems, Quantum 9, 1675 (2025), arXiv:2406.00091 [cond-mat.dis-nn] [12] H. L. Dao, Hyperbolic recurrent neural network as the first type of non-Euclidean neural quantum state ansatz, Eur. Phys. J. Plus 141:199, arXiv:2505.22083 [quant-ph, cond-mat.dis-nn, cs.LG, physics.comp-ph] (2026) [13] H. L. Dao, New non-Euclidean neural quantum state ansatzes from additional types of hyperbolic neural networks, arXiv:2604.2337 [quant-ph, cs.LG, cond-mat.dis-nn] [14] O.-E. Ganea, G. Becigneul, and T. Hofmann, Hyperbolic Neural Networks, Advances in Neural Information Processing Systems 31, pages 5345–5355. Curran Associates, Inc. arXiv: 1805.09112 [cs.LG] [15] N. He, M. Yang and Rex Ying, HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules, arXiv:2504.08912 [cs.LG] [16] F. Becca and S. Sorella, Quantum Monte Carlo approaches for correlated systems, Cambridge University Press 2017, DOI: 10.1017/9781316417041 [17] J. Chung, C. Gulcehre, K. Cho, Y. Bengio, Gated feedback recurrent neural networks, in: ICML, 2015. [18] H. L. Dao, Deep Learning Calabi-Yau four folds with hybrid and recurrent neural network architectures, Nucl. Phys. B 1013 (2025) 116832, arXiv:2405.17406 [hep-th, cs.LG, math.AG] [19] W. Chen, X. Han, Y. Lin, H. Zhao, Z. Liu, P. Li, M. Sun, J. Zhou, Fully Hyperbolic Neural Networks, in ACL 2022 Main Conference, arXiv:2105.14686 [cs.CL]. [20] W. Peng, T. Varanka, A. Mostafa, H. Shi, G. Zhao, Hyperbolic Deep Neural Networks: A Survey, arXiv:2101.04562 [cs.LG, cs.CV] [21] C. Gulcehre, M. Denil, M. Malinowski, A. Razavi, R. Pascanu, et. al., Hyperbolic Attention Networks, arXiv:1805.09786v1 [cs.NE] [22] F.Lopez and M. Strube, A Fully Hyperbolic Neural Model for Hierarchical Multi-Class Classification, Findings of EMNLP2020, arXiv:2010.02053 [cs.CL]. 32
[23] R. Shimizu, Y. Mukuta and T. Harada, Hyperbolic Neural Networks++, The Ninth International Conference on Learning Representations (ICLR 2021), arXiv:2006.08210 [cs.LG] [24] E. Mathieu, C. Le Lan, C. J. Maddison, R. Tomioka, and Y. W. Teh, Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders, arXiv:1901.06033. [25] G. Bachmann, G. Becigneul, O.-E. Ganea, Constant Curvature Graph Convolutional Networks, arXiv:1911.05076v3. [26] N. Linial, E. London, and Y. Rabinovich. The geometry of graphs and some of its algorithmic applications, Combinatorica, 15(2):215–245, 1995. [27] D. Krioukov, F. Papadopoulos, A. Vahdat, and M. Boguná, Curvature and temperature of complex networks, Physical Review E, 80(3):035101, 2009. [28] D. Krioukov, F. Papadopoulos, M. Kitsak, A. Vahdat, and M. Boguná, Hyperbolic geometry of complex networks, Physical Review E, 82(3):036106, 2010. [29] H. W. J. Blöte and Y. Deng, Cluster Monte Carlo simulation of the transverse Ising model, Phys. Rev. E 66, 066110 (2002). [30] R. Sarkar. Low distortion Delaunay embedding of trees in hyperbolic plane, In Proc. of the International Symposium on Graph Drawing (GD 2011), pages 355–366, Eindhoven, Netherlands, 2011. [31] F. Sala, C. De Sa, A. Gu, and C. Ré. 2018. Representation tradeoffs for hyperbolic embeddings. In International Conference on Machine Learning, pages 4457–4466. [32] S. Bonnabel, Stochastic gradient descent on riemannian manifolds, IEEE Transactions on Automatic Control, 58(9):2217–2229, Sept 2013. [33] O.-E. Ganea, G. Bécigneul, and T. Hofmann. Hyperbolic entailment cones for learning hierarchical embeddings. In Proceedings of the thirty-fifth international conference on machine learning (ICML), 2018 [34] G. Bécigneul and O.-E Ganea, Riemannian Adaptive Optimization Methods, International Conference on Learning Representations (ICLR) (2019), arXiv:1810.00760 [cs.LG] [35] M. Nickel and D. Kiela, Poincaré embeddings for learning hierarchical representations, arXiv:1705.08039 [cs.AI] [36] M. Nickel and D. Kiela, Learning Continuous Hierarchies, in the Lorentz Model of Hyperbolic Geometry, ICML 2018, arXiv:1806.03417 [cs.AI] [37] H. L. Dao, Exploring new variational quantum circuit ansatzes for solving SU (2) matrix models, Eur. Phys. J. C (2025) 85:705, arXiv:2503.13368 [quant-ph, hep-th, physics.comp-ph] [38] S. El-Showk, M. F. Paulos, D. Poland, S. Rychkov, D. Simmons-Duffin, A. Vichi, Solving the 3d Ising Model with the Conformal Bootstrap II. c-Minimization and Precise Critical Exponents, J. Stat. Phys. 157, 869-914 (2014), arXiv:1403.4545 [hep-th] [39] G. Vidal, Entanglement renormalization, Phys. Rev. Lett. 99, 220405 (2007), arXiv:cond-mat/0512165 [cond-mat.str-el] [40] G. Vidal, A class of quantum many-body states that can be efficiently simulated, Phys. Rev. Lett. 101, 110501 (2008) (2007), arXiv:quant-ph/0610099 [41] G. Evenbly and G. Vidal, Entanglement renormalization in two spatial dimensions, Phys. Rev. Lett. 102, 180406 (2009), arXiv:0811.0879 [cond-mat.str-el] [42] B. Swingle, Entanglement Renormalization and Holography, Phys. Rev. D 86, 065007 (2012), arXiv:0905.1317 [cond-mat.str-el] [43] B. Swingle, Constructing holographic spacetimes using entanglement renormalization, arXiv:1209.3304 [hep-th] [44] E. Rinaldi, X. Han, M. Hassan, Y. Feng, F. Nori, M. McGuigan, M. Hanada, Matrix-model simulations using quantum computing, deep learning, and lattice Monte Carlo. PRX Quantum 3, 010324 (2022) [45] J. M. Maldacena, The large N limit of superconformal field theories and supergravity, Adv. Theor. Math. Phys. 2 (1998) 231 [Int. J. Theor. Phys. 38 (1999) 1113] [arXiv:hep- th/9711200]. 33
[46] P. Di Francesco, P. Mathieu, D. Sénéchal, Conformal Field Theory, Springer-Verlag (1997) [47] S. Sachdev, Quantum Phase Transitions, Cambridge University Press (2011). [48] D. J. Watts and S. H. Strogatz. Collective dynamics of ‘small-world’ networks. Nature, 393(6684):440, 1998. [49] Q.Liu, M. Nickel and D. Kiela, Hyperbolic Graph Neural Networks, arxiv:1910.12892 [cs.LG]
34