QLAM: A Quantum Long-Attention Memory Approach to Long-Sequence Token Modeling Hoang-Quan Nguyen1,2 , Sankalp Pandey1,2 , Khoa Luu1,2 1 Department of Electrical Engineering and Computer Science, University of Arkansas, AR 2 Quantum AI Lab, University of Arkansas, AR
arXiv:2605.13833v1 [cs.LG] 13 May 2026
{hn016, sankalpp, khoaluu}@uark.edu
Abstract—Modeling long-range dependencies in sequential data remains a central challenge in machine learning. Transformers address this challenge through attention mechanisms, but their quadratic complexity with respect to sequence length limits scalability to long contexts. State-space models (SSMs) provide an efficient alternative with linear-time computation by evolving a latent state through recurrent updates, but their memory is typically formed via additive or linear transitions, which can limit their ability to capture complex global interactions across tokens. In this work, we introduce one of the first studies to leverage the superposition property of quantum systems to enhance statebased sequence modeling. In particular, we propose Quantum Long-Attention Memory (QLAM), a hybrid quantum-classical memory mechanism that can be viewed as a quantum extension of state-space models. Instead of maintaining a classical latent state updated through additive dynamics, QLAM represents the hidden state as a quantum state whose amplitudes encode a superposition of historical information. The state evolves through parameterized quantum circuits conditioned on the input, enabling a non-classical, globally update mechanism. In this way, QLAM preserves the recurrent and linear-time structure of SSMs while fundamentally enriching the memory representation through quantum superposition. Unlike attention mechanisms that explicitly compute pairwise interactions, QLAM implicitly captures global dependencies through the evolution of the quantum state, and retrieves task-relevant information via querydependent measurements. We evaluate QLAM on sequential variants of standard image classification benchmarks, including sMNIST, sFashion-MNIST, and sCIFAR-10, where images are flattened into token sequences. Across all tasks, QLAM consistently improves over recurrent baselines and transformer-based models. These results suggest that QLAM provides a principled extension of state-space formulations, where quantum superposition replaces classical additive memory, enabling richer global representations while retaining favorable computational scaling. This opens a new research direction for sequence modeling by integrating quantum computation into the core design of memory mechanisms. Index Terms—Quantum Machine Learning, Self-Attention, Sequence Modeling, Hybrid Framework
I. I NTRODUCTION The ability to model long-range dependencies in sequential data is fundamental to many machine learning applications, including language modeling, time-series forecasting, and multimodal reasoning. Classical architectures such as recurrent neural networks (RNNs) and transformers have achieved remarkable performance in these tasks. However, both approaches show intrinsic limitations as the sequence length increases. RNN-based models suffer from vanishing
or exploding gradients that degrade their ability to preserve long-term information [1], [2]. Transformers [3] address this issue using attention mechanisms, but their quadratic computational complexity with respect to sequence length. A number of recent works attempt to mitigate this limitation through efficient or long-sequence attention variants, including sparse attention [4], low-rank and kernelized approximations [5], [6], and state-space-inspired architectures such as linear attention and hybrid models [7], [8]. More recently, State-Space Models (SSMs) [9], [10] have emerged as an efficient alternative, enabling linear-time sequence processing through structured state transitions. Despite these advances, existing approaches remain fundamentally classical, relying on fixed-dimensional vector representations of memory. Quantum information theory offers a fundamentally different paradigm for representing and evolving information. In quantum systems, information is encoded in the amplitudes of vectors in a complex Hilbert space. A quantum state can represent a coherent superposition of multiple basis states, allowing a compact representation of information that scales exponentially with the number of qubits [11], [12]. Moreover, quantum evolution is governed by unitary operators, which preserve the norm of the quantum state and therefore avoid the exponential growth or decay typically encountered in classical dynamical systems. Far apart from prior studies on attention-based longsequence modeling methods, this work explores a different perspective on the problem mentioned: Can memory itself be fundamentally restructured using quantum computation? Motivated by the key properties in quantum computing, i.e., superposition and unitary evolution, we introduce Quantum Long-Attention Memory (QLAM), a hybrid quantum-classical memory framework where contextual information is stored and processed directly on a quantum device. QLAM leverages actual quantum states and operations to represent and manipulate memory. At each time step, the memory is encoded as a quantum state that evolves through parameterized quantum circuits conditioned on the input, while task-relevant information is retrieved via query-dependent measurements. Compared to classical long-attention methods, QLAM provides a more suitable memory abstraction for long-sequence modeling because it does not require storing or explicitly comparing all historical token representations. Classical longattention mechanisms typically improve scalability by spar-
sifying, compressing, or approximating the attention matrix, but they still rely on classical key-value memories whose capacity grows with sequence length or is constrained by fixed-dimensional compression. In particular, sparse attention methods [4], [13] reduce complexity at the cost of potentially missing global dependencies, while low-rank and kernel-based approximations [5], [6] introduce approximation errors that can degrade representation fidelity for complex interactions. Linear attention and SSMs methods [7]–[9] achieve favorable scaling but often compress historical information into fixedsize states, which may limit expressivity and lead to information loss over long horizons. In contrast, QLAM encodes contextual information in a quantum state, in which multiple memory components can be represented in superposition and evolved jointly by unitary transformations. This allows longrange information to be maintained in a compact and globally integrated memory representation, while query-dependent measurements retrieve task-relevant information without explicitly computing all pairwise token interactions. This formulation departs from conventional attention and memory mechanisms in several important ways. First, instead of storing context as a collection of token-wise embeddings or key-value pairs, QLAM represents memory as a global quantum state, enabling a highly expressive and distributed representation of information. Second, memory updates are controlled by unitary transformations, which inherently preserve the state’s structure and provide a stable mechanism for information propagation. Third, retrieval is performed via measurement, enabling query-dependent information extraction without explicitly computing pairwise interactions between tokens. Together, these properties suggest a fundamentally different paradigm for sequence modeling, in which memory is no longer a passive storage but an actively evolving quantum system. Finally, QLAM is presented as a new design for hybrid models that integrate quantum computation into the core of memory representation. Contributions of this Work. This work presents a new Quantum Long-Attention Memory (QLAM) method, a hybrid classical-quantum memory mechanism where memory is represented and updated using quantum states and circuits. First, we propose a query-dependent measurement-based retrieval mechanism that generalizes attention to the quantum setting. Second, we conduct an empirical exploration to evaluate the feasibility and potential of quantum memory in practical learning settings. Finally, our results suggest that hybrid quantum memory is a promising direction for rethinking sequence modeling, opening new opportunities at the intersection of machine learning and quantum computation. II. R ELATED W ORK A. Sequential Token Modeling Early work for processing sequential data relied on statetracking architectures, most notably recurrent neural networks (RNNs) [1], [14]–[16]. This approach continuously updated a compressed hidden state, and, in doing so, these networks, as well as gated variants such as long short-term memory
networks (LSTMs) and gated recurrent units (GRUs), excelled at localized sequence modeling. However, these models exhibited performance degradation over longer sequences. The introduction of Transformers [3] fundamentally shifted this landscape by enabling global self-attention, which allowed direct, uncompressed routing between arbitrary sequence positions. This paradigm has driven breakthroughs in natural language processing [17]–[19] and computer vision [20], [21]. However, this expressivity comes at a high computational cost due to the quadratic complexity of attention with respect to sequence length, motivating research on efficient attention [5], [6]. Recent literature has turned to state-space models (SSMs) [9], [22], [23] to bypass the self-attention bottleneck. By formalizing the sequence processing through structured linear dynamical systems, architectures such as S4 [9], S5 [22], and Mamba [23] achieve linear-time scaling without sacrificing the modeling fidelity of Transformers for long-context tasks. B. Quantum Machine Learning Quantum machine learning (QML) has emerged as a paradigm to augment machine learning by exploiting quantum mechanics for enhanced representational capacity and computational efficiency [24]–[26]. Early developments in QML primarily focused on accelerating classical linear machine learning algorithms through quantum speedups. These works demonstrated quantum advantages in tasks such as clustering [27]–[29], principal component analysis [30], least-squares fitting [31], [32], and binary classification [33]. Within the constraints of the noisy intermediate-scale quantum (NISQ) era, variational quantum circuits (VQCs) [34], or parameterized quantum circuits [35], serve as the foundational architecture. These models have achieved notable traction across diverse domains, including standard classification [36]–[38], optimization [39]–[41], generative modeling [42]–[44], and reinforcement learning [45]. Building upon this framework, several architectures have been proposed to extend classical deep learning concepts into the quantum domain. For instance, quantum convolutional neural networks [46] adapt convolutional structures with reduced parameterization, while quantum autoencoders [47] enable compression and representation learning of quantum states. Additionally, quantum neural networks based on PQCs [48]–[51] provide a flexible foundation for learning nonlinear mappings in hybrid settings. Beyond developments in QML, recent studies have also explored physics-aware machine learning for quantum material discovery, particularly in the analysis of two-dimensional quantum materials from optical microscopy images [52]–[54]. However, the limitations of these VQC formulations become apparent when the focus shifts to dynamic, time-series data. In particular, embedding complex temporal dependencies into shallow, parameterized unitary operations presents a blocking point. As such, recent work proposes quantum architectures tailored for sequential data. Approaches such as quantum recurrent neural networks and quantum dynamical systems [55], [56] aim to model temporal evolution via continuous quantum-state transitions. In contrast to prior approaches that focus on either recurrent
Current Feature
Weighted Sum Output Feature
Input Features New Attention State
Attention State
Quantum LongAttention Readout
"
Measured Attention
|𝜓⟩ = %𝛼! |𝑖⟩ !#$
Parameterized Quantum Circuits Update Attention State
Current Query
Quantum Long-Attention Memory Fig. 1. Overview framework of the proposed Quantum Long-Attention Memory.
dynamics, our work explores a complementary direction by leveraging quantum superposition as a memory mechanism for sequence modeling. The proposed Quantum Long-Attention Memory (QLAM) encodes the evolving sequence context into a structured quantum state, allowing information from multiple tokens to be aggregated implicitly through superposition. This enables attention-like global interactions while maintaining a compact and recurrently updated memory representation, offering a novel perspective for long-sequence modeling in QML. III. BACKGROUND A. Transformers Attention Mechanisms Attention mechanisms [3] enable models to selectively focus on relevant parts of an input sequence. In transformer architectures, self-attention computes pairwise interactions between tokens through a query-key-value formulation. Given queries Q = {qi ∈ Rd }Ti=1 , keys K = {ki ∈ Rd }Ti=1 , and values V = {vi ∈ Rd }Ti=1 , the attention output is computed as follows, QK ⊤ √ Attention(Q, K, V ) = softmax V. (1) d This formulation allows each token to attend to all other tokens in the sequence, enabling flexible modeling of longrange dependencies. However, the computational complexity of self-attention scales quadratically with the sequence length, which becomes prohibitive for very long contexts. Several methods have been proposed to reduce this cost, including linear attention mechanisms [5], [6] and state-space models [9], [10], [23] that implicitly capture contextual information through recurrent state updates. In contrast to explicit pairwise interactions, these methods rely on compressed representations of the sequence history.
B. State-Space Models (SSMs) State-space models (SSMs) describe the evolution of a latent dynamical system driven by an input sequence. In discrete time, a linear state-space model is typically expressed as: ht+1 = Aht + Bxt yt = Cht ,
(2)
where xt ∈ Rd is the input, ht ∈ Rn is the latent state, and yt denotes the output. The matrices A, B, and C determine the state transition, input injection, and readout operators, respectively. SSMs have recently gained attention in machine learning due to their ability to model long sequences with linear computational complexity. Modern architectures based on structured SSMs introduce parameterizations that enable efficient convolutional implementations and stable training dynamics. In these models, the latent state ht serves as a compressed representation of the historical sequence, allowing information from earlier tokens to influence future predictions. Despite these advantages, classical SSMs rely on additive accumulation of information within a finite-dimensional state vector. As the sequence length grows, the model must continuously compress historical information into a bounded representation. This limitation may reduce the expressive capacity of the memory mechanism and can lead to information loss over long time horizons. IV. Q UANTUM L ONG -ATTENTION M EMORY (QLAM) In this section, we introduce Quantum Long-Attention Memory (QLAM), a quantum-based memory mechanism designed for long-attention-sequence modeling, as presented in Figure 1. The key idea is to represent the memory state as a vector in a complex Hilbert space and evolve it through unitary transformations conditioned on the input sequence. This formulation enables stable memory propagation and allows contextual information to be stored through superposition.
Step 1
Step 𝑡
… … … …
|0⟩ |0⟩
𝑈enc (𝑥! )
|0⟩
𝑈var (𝜃)
|0⟩
𝑈enc (𝑥" )
𝑈var (𝜃)
Fig. 2. The quantum long-attention memory is evolved across input features.
A. Problem Formulation We consider a sequential prediction problem over an input sequence: {x1 , x2 , . . . , xT }, xt ∈ X , (3) where X denotes the input space, e.g., tokens, image patches, or sensor measurements. The objective is to learn a model that produces outputs: {y1 , y2 , . . . , yT },
y ∈ Y,
(4)
the memory complexity of O(d), the representation of QLAM n for the state |ψt ⟩ ∈ C2 require the memory complexity of O(n) ∼ O(log2 d), which makes the QLAM more efficient in memory storage. C. Quantum Long-Memory Evolution The memory state evolves through unitary transformations conditioned on the input token xt . Formally, the memory update rule is defined as: |ψt ⟩ = U (xt , θ)|ψt−1 ⟩,
such that each output yt depends on the entire history: yt = f (x1 , x2 , . . . , xt ).
(5)
Most classical approaches maintain a hidden state ht ∈ Rd or a key-value memory. In contrast, QLAM maintains a quantum memory state that evolves over time steps. B. Quantum Long-Memory Representation Given an input sequence {xt }Tt=1 , standard attention computes: X yt = αt,s vs , αt,s ∝ exp(qt⊤ ks ) (6) s≤t
which requires evaluating all pairwise token interactions. In contrast, our approach replaces this computation with quantum-state accumulation, and attention weights are reconstructed from this state. Let the memory state at time t be represented by a normalized vector in a complex Hilbert space: 2n
|ψt ⟩ ∈ C ,
⟨ψt |ψt ⟩ = 1,
(7)
where n denotes the number of qubits used to represent the memory. The quantum state can be expressed in the computational basis as: |ψt ⟩ =
n 2X −1
(i)
αt |i⟩,
(8)
i=1 (i) where αt ∈ C are complex amplitudes, encoding informa-
tion about the sequence. Each basis state |i⟩ corresponds to an element in a memory configuration. Then, the full state represents a superposition of the memory. This allows the memory state to encode a coherent superposition of multiple contextual signals and to evolve over timesteps. Compared to classical approaches where the classical states ht ∈ Rd require
(9)
where U (xt , θ) is a parameterized unitary operator. Formally, we decompose the unitary as: U (xt , θ) = Uvar (θ)Uenc (xt ),
(10)
where Uenc (xt ) is an input encoding operator and Uvar (θ) is a trainable parameterized circuit. This separation clarifies the roles of data injection and learned transformation. Initially, the quantum long-attention memory is defined as |ψ1 ⟩ = |0⟩⊗n , showing the attention is focused on the first feature x1 . Then, the attention memory is evolved through input features xt , as shown in Figure 2. A key property of this update rule is that unitary transformations satisfy: U (xt , θ)† U (xt , θ) = I,
(11)
∥|ψt ⟩∥ = ∥|ψt−1 ⟩∥.
(12)
which implies:
Therefore, the memory dynamics are norm-preserving, ensuring stable propagation across long sequences. This property contrasts with classical recurrent systems in which repeated matrix multiplications may lead to exploding or vanishing states. D. Quantum Long-Attention Readout To extract task-relevant information from the quantum memory state, we introduce a measurement-based readout mechanism that serves as an attention query qt = WQ xt , as shown in Figure 3. Unlike classical attention, which retrieves information through explicit similarity computations over stored representations, QLAM performs retrieval by probing the quantum state with a query-conditioned observable. For each
𝑥$
Updated Value
𝑣$
𝑥"…$
𝑣"…$
𝑊%
Quantum Attention State Weighted Sum
𝑊!
𝑦$
Learnable Hermitian Operator
𝑞$
Measurement
Measured Attention
Fig. 3. Computation flow of quantum long-attention readout.
0.5
Observable Decoder
𝑊" 𝑥!
0.5 ×
0.2 0.3
𝑞!
Observable Coordinates
0
0
1
0
0
1
0
0
0
0
0
1
0
0
0
1
1
0
0
0
0
0
1
0
1
0
0
0
0
0
0
1
0
1
0
0
0
1
0
0
0
0
1
0
1
0
0
0
+ 0.2 ×
+ 0.3 ×
This construction allows the query to select an adaptive measurement basis while keeping the measurement process physically valid. The formulation defines a query-conditioned projection of the quantum memory, where the observable determines how information is extracted from the state, as illustrated in Figure 4. In practice, expectation values are approximated through repeated measurement: (1)
𝑂(𝑞! ) Fig. 4. The observable O(qt ) is formed as a weighted combination of multiple Pauli matrices based on the query qt .
previous token s < t, we compute an attention through a measurement operator:
(m)
ot,s , . . . , ot,s ∼ M(O(qt ), |ψt ⟩) m 1 X (i) o . αt,s = m i=1 t,s
(16)
The estimator is unbiased and converges to the exact expectation as m is increased. This introduces a controllable stochastic component that can improve robustness. V. E XPERIMENTAL R ESULTS A. Experimental Setup
αt,s = ⟨ψt |O(qt )|ψt ⟩,
(13)
2n ×2n
where O(qt ) ∈ C is a learnable Hermitian observable operator parameterized by the query qt . In detail, given a set of Pauli matrices P = {Oi }pi=1 , we project the query qt via a small neural network as a learnable observable decoder into observable coordinates {γi ∈ R}pi=1 indicating the weighted combination of the Pauli matrices: O(qt ) =
p X
γi Oi .
(14)
i=1
Since γi is real, γi∗ = γi , and Oi is Hermitian, Oi† = Oi , O(qt ) is guaranteed to be Hermitian: !† p p X X O(qt )† = γi Oi = γi∗ Oi† i=1
=
p X i=1
i=1
γi Oi = O(qt ).
(15)
We evaluate the proposed Quantum Long-Attention Memory framework on image classification. We evaluate the proposed QLAM framework on image classification tasks reformulated as sequence modeling problems, a standard protocol for assessing long-range dependency modeling. Specifically, we consider MNIST [57], Fashion-MNIST [58], and CIFAR10 [59], as shown in Figure 5, where each image is converted into a one-dimensional sequence and processed token-bytoken. This formulation removes spatial inductive biases and forces models to rely entirely on sequential reasoning, making it particularly suitable for evaluating memory mechanisms. For convenience, we denote the resulting datasets as sMNIST, sFashion-MNIST, and sCIFAR-10, respectively. For MNIST and Fashion-MNIST, each 28 × 28 grayscale image is reshaped into a sequence of length 784, where each token corresponds to a single pixel intensity. For CIFAR10, each 32 × 32 × 3 image is flattened into a sequence of length 3072, with RGB channels concatenated along the sequence dimension. All pixel values are normalized to [0, 1]
(a) MNIST
(b) Fashion-MNIST
(c) CIFAR-10
Fig. 5. Visualization of sample inputs from the benchmark datasets used in our experiments, including (a) MNIST [57], (b) Fashion-MNIST [58], and (c) CIFAR-10 [59]. These datasets are later reformulated as sequential inputs by flattening each image into a one-dimensional token sequence, enabling evaluation of sequence modeling capabilities.
Neural Network Model
Flatten Input Sequence
“5” Prediction
Input Image Fig. 6. Evaluation protocol for sequential modeling with images.
and treated as continuous inputs. The model processes the sequence causally, i.e., tokens are consumed sequentially without access to future information, thereby mimicking standard autoregressive sequence modeling. A final classification head is applied on the last hidden state to predict the image label. The inference process is illustrated in Figure 6. We compare QLAM against representative baselines spanning different sequence modeling paradigms: (i) a recurrent neural network (RNN) [14], which relies on iterative hidden state updates, (ii) a Transformer encoder [3], which models global dependencies via self-attention, and (iii) a state-space model (SSM) [9], which captures long-range interactions through structured linear dynamics. All models are implemented with comparable parameter budgets to ensure a fair comparison, and share similar input projections and output classifiers. To obtain statistically reliable results, we adopt a 10-fold evaluation protocol. For each dataset, we train and evaluate the models across 10 independent runs with different data splits. We report both single-run performance and aggregated statistics, i.e., mean and standard deviation, providing a comprehensive assessment of both accuracy and stability. B. Implementation Details The models are implemented in PyTorch [60], with the quantum components of QLAM implemented using a hybrid quantum-classical simulation framework with the PennyLane
library [61]. The input at each timestep is first processed by a lightweight classical encoder, consisting of a linear projection layer that maps scalar pixel values into a higher-dimensional embedding. This embedding is then used to parameterize the quantum circuit through angle encoding. The initial quantum state is initialized to the computational basis state |0⟩⊗n . All models are trained using the Adam optimizer [62] with an initial learning rate of 10−3 . We train for 30 epochs on sMNIST and sFashion-MNIST, and 50 epochs on sCIFAR10. A cosine learning rate scheduler is applied. The batch size is set to 128 for all experiments. For baseline models, we use comparable model sizes to ensure a fair comparison. We evaluate model performance using standard classification accuracy on the test set. C. Evaluation Results on sMNIST As shown on the sMNIST benchmark in Table I, all sequence-based models significantly outperform the vanilla RNN, confirming the inherent difficulty of modeling longrange dependencies with purely recurrent dynamics. The RNN fails to retain information across long pixel sequences, resulting in near-random performance. In contrast, both the Transformer [3] and SSM [9] achieve strong performance around 91%, demonstrating their ability to capture long-range interactions through attention mechanisms and structured state transitions, respectively.
TABLE I E XPERIMENTAL ACCURACIES (%) OF THE PROPOSED APPROACH COMPARED TO PRIOR BASELINES . Method
Fold-1
Fold-2
Fold-3
RNN [14] Transformer [3] SSM [9] QAM
11.2 91.0 90.9 92.3
11.5 91.5 91.3 92.7
11.6 91.2 91.1 92.6
RNN [14] Transformer [3] SSM [9] QAM
12.5 79.8 79.2 81.1
12.8 80.4 79.8 81.6
12.7 80.1 79.5 81.5
RNN [14] Transformer [3] SSM [9] QAM
13.5 52.4 51.3 53.1
13.8 53.2 52.0 53.8
13.6 52.8 51.6 53.6
Fold-4
Fold-5 Fold-6 Fold-7 sMNIST [57] 11.3 11.4 11.5 11.3 91.4 91.3 91.6 91.2 91.2 91.0 91.4 91.2 92.5 92.4 92.8 92.6 sFashion-MNIST [58] 12.6 12.7 12.9 12.6 80.3 80.2 80.5 80.0 79.7 79.6 79.9 79.4 81.3 81.4 81.7 81.2 sCIFAR-10 [59] 13.7 13.7 13.9 13.6 53.0 52.9 53.3 52.7 51.8 51.7 52.1 51.5 53.4 53.5 53.9 53.3
Our proposed QLAM further improves accuracy to 92.6 ± 0.15 in the 10-fold evaluation, consistently outperforming all baselines across all folds. Notably, the improvement is achieved without increasing model complexity, indicating that the gain arises from a more effective memory mechanism rather than scale. The low standard deviation highlights the stability of the proposed approach, suggesting that QLAM mitigates optimization issues commonly observed in longsequence modeling. From a modeling perspective, this result supports the hypothesis that superposition-based memory aggregation enables a richer representation of global dependencies than classical additive accumulation. D. Evaluation Results on sFashion-MNIST The sFashion-MNIST dataset is more challenging due to higher intra-class variability and more complex visual structures, requiring models to capture both fine-grained and global dependencies. As reported in Table I, the Transformer and SSM achieve 80.2% and 79.6%, respectively, indicating a degradation compared to sMNIST due to increased task complexity. In contrast, QLAM achieves 81.4 ± 0.19 under the 10-fold protocol, consistently outperforming both baselines. The performance gap is more noticeable than on sMNIST, suggesting that QLAM provides greater benefits when the underlying data distribution is more complex. This behavior can be attributed to QLAM’s ability to encode multiple dependency patterns simultaneously via superposition, enabling the model to better capture diverse visual features. Furthermore, QLAM exhibits lower variance across folds than baseline models, indicating greater robustness to data splits and initialization. This stability is particularly important in more complex datasets, where optimization landscapes are typically more irregular. Overall, these results demonstrate that QLAM not only improves accuracy but also enhances training stability in challenging sequential learning scenarios. E. Evaluation Results on sCIFAR-10 The sCIFAR-10 benchmark presents a substantially more difficult setting due to higher-dimensional inputs, color chan-
Fold-8
Fold-9
Fold-10
Mean ± Std
11.6 91.4 91.1 92.5
11.4 91.3 91.3 92.7
11.5 91.5 91.2 92.6
11.4 ± 0.13 91.3 ± 0.18 91.2 ± 0.15 92.6 ± 0.15
12.8 80.3 79.7 81.5
12.7 80.2 79.5 81.4
12.8 80.4 79.8 81.6
12.7 ± 0.12 80.2 ± 0.21 79.6 ± 0.22 81.4 ± 0.19
13.8 53.1 51.9 53.7
13.7 52.9 51.7 53.5
13.8 53.2 52.0 53.8
13.7 ± 0.13 53.0 ± 0.28 51.8 ± 0.25 53.6 ± 0.26
nels, and richer semantic content. All models experience a noticeable drop in performance compared to grayscale datasets, reflecting the increased difficulty of modeling longrange dependencies in high-entropy sequences. As shown in Table I, the Transformer and SSM achieve 52.4% and 51.3%, respectively, while QLAM improves the performance to 53.1%. Under 10-fold evaluation, QLAM achieves 53.6±0.26, consistently outperforming both baselines across all folds. Although the absolute improvement is smaller compared to sMNIST and sFashion-MNIST, the gain remains consistent, which is non-trivial given the increased complexity of the task. Importantly, the variance of QLAM remains competitive despite the greater difficulty, indicating that the proposed memory mechanism scales favorably to more complex settings. This suggests that QLAM provides a more expressive yet stable framework for aggregating long-range information, even when the input distribution becomes significantly more diverse. From a broader perspective, these results highlight that the advantage of QLAM is not limited to simple datasets, but extends to more realistic and challenging scenarios where both expressivity and stability are required. VI. C ONCLUSIONS In this work, we have introduced Quantum Long-Attention Memory (QLAM), a novel sequence modeling framework that extends state-space models by representing memory as a quantum state and evolving it through unitary transformations. This formulation has addressed key limitations of existing approaches, where transformers suffer from quadratic complexity and classical state-space models rely on additive updates that may restrict expressivity in long sequences. By encoding historical information in a superposition-based quantum state, QLAM has provided a compact global memory representation, enabling implicit capture of interactions across tokens without explicit pairwise computations. The use of unitary dynamics has ensured stable propagation of information over long horizons, while the measurementbased readout has provided a query-dependent mechanism for
retrieving task-relevant information, effectively generalizing attention within this framework. Empirical results on sMNIST, sFashion-MNIST, and sCIFAR-10 have demonstrated that QLAM consistently outperforms recurrent, transformerbased, and state-space baselines under comparable settings, highlighting improvements in both accuracy and stability for long-range sequence modeling. While these findings validate the potential of quantum-state-based memory, further work is needed to further study the theoretical understanding of its properties, explore implementations on quantum hardware, and extend the framework to large-scale language and multimodal tasks. Overall, QLAM offers a new perspective on sequence modeling by fundamentally rethinking memory through quantum computation. R EFERENCES [1] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997. [2] Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE transactions on neural networks, vol. 5, no. 2, pp. 157–166, 1994. [3] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017. [4] R. Child, S. Gray, A. Radford, and I. Sutskever, “Generating long sequences with sparse transformers,” arXiv preprint arXiv:1904.10509, 2019. [5] A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International conference on machine learning. PMLR, 2020, pp. 5156– 5165. [6] K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser et al., “Rethinking attention with performers,” arXiv preprint arXiv:2009.14794, 2020. [7] Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient transformers: A survey,” ACM Computing Surveys, vol. 55, no. 6, pp. 1–28, 2022. [8] T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in neural information processing systems, vol. 35, pp. 16 344–16 359, 2022. [9] A. Gu, K. Goel, and C. Ré, “Efficiently modeling long sequences with structured state spaces,” arXiv preprint arXiv:2111.00396, 2021. [10] A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré, “Combining recurrent, convolutional, and continuous-time models with linear state space layers,” Advances in neural information processing systems, vol. 34, pp. 572–585, 2021. [11] M. A. Nielsen and I. L. Chuang, Quantum computation and quantum information. Cambridge university press, 2010. [12] J. Preskill, “Quantum computing in the nisq era and beyond,” Quantum, vol. 2, p. 79, 2018. [13] M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang et al., “Big bird: Transformers for longer sequences,” Advances in neural information processing systems, vol. 33, pp. 17 283–17 297, 2020. [14] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning internal representations by error propagation,” Tech. Rep., 1985. [15] J. L. Elman, “Finding structure in time,” Cognitive science, vol. 14, no. 2, pp. 179–211, 1990. [16] K. Cho, B. Van Merriënboer, Ç. Gulçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, pp. 1724–1734. [17] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 2019, pp. 4171–4186.
[18] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020. [19] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019. [20] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. [21] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 012–10 022. [22] J. T. Smith, A. Warrington, and S. W. Linderman, “Simplified state space layers for sequence modeling,” arXiv preprint arXiv:2208.04933, 2022. [23] A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv:2312.00752, 2023. [24] J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, “Quantum machine learning,” Nature, vol. 549, no. 7671, pp. 195–202, 2017. [25] M. Schuld, I. Sinayskiy, and F. Petruccione, “An introduction to quantum machine learning,” Contemporary Physics, vol. 56, no. 2, pp. 172–185, 2015. [26] M. Schuld and N. Killoran, “Quantum machine learning in feature hilbert spaces,” Physical review letters, vol. 122, no. 4, p. 040504, 2019. [27] S. Lloyd, M. Mohseni, and P. Rebentrost, “Quantum algorithms for supervised and unsupervised machine learning,” arXiv preprint arXiv:1307.0411, 2013. [28] X. B. Nguyen, H. Churchill, K. Luu, and S. U. Khan, “Quantum vision clustering,” arXiv preprint arXiv:2309.09907, 2023. [29] X.-B. Nguyen, H.-Q. Nguyen, S. Y.-C. Chen, S. U. Khan, H. Churchill, and K. Luu, “Qclusformer: A quantum transformer-based framework for unsupervised visual clustering,” in 2024 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 2. IEEE, 2024, pp. 347–352. [30] S. Lloyd, M. Mohseni, and P. Rebentrost, “Quantum principal component analysis,” Nature Physics, vol. 10, no. 9, pp. 631–633, 2014. [31] M. Schuld, I. Sinayskiy, and F. Petruccione, “Prediction by linear regression on a quantum computer,” Physical Review A, vol. 94, no. 2, p. 022342, 2016. [32] I. Kerenidis and A. Prakash, “Quantum gradient descent for linear systems and least squares,” Physical Review A, vol. 101, no. 2, p. 022316, 2020. [33] P. Rebentrost, M. Mohseni, and S. Lloyd, “Quantum support vector machine for big data classification,” Physical review letters, vol. 113, no. 13, p. 130503, 2014. [34] M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio et al., “Variational quantum algorithms,” Nature Reviews Physics, vol. 3, no. 9, pp. 625– 644, 2021. [35] M. Benedetti, E. Lloyd, S. Sack, and M. Fiorentini, “Parameterized quantum circuits as machine learning models,” Quantum science and technology, vol. 4, no. 4, p. 043001, 2019. [36] M. Schuld, A. Bocharov, K. M. Svore, and N. Wiebe, “Circuit-centric quantum classifiers,” Physical Review A, vol. 101, no. 3, p. 032308, 2020. [37] X.-B. Nguyen, H.-Q. Nguyen, H. Churchill, S. U. Khan, and K. Luu, “Hierarchical quantum control gates for functional mri understanding,” in 2024 IEEE Workshop on Signal Processing Systems (SiPS). IEEE, 2024, pp. 159–164. [38] H.-Q. Nguyen, X.-B. Nguyen, S. Pandey, S. U. Khan, I. Safro, and K. Luu, “Qmoe: A quantum mixture of experts framework for scalable quantum neural networks,” in 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 2. IEEE, 2025, pp. 223–228. [39] E. Farhi, J. Goldstone, and S. Gutmann, “A quantum approximate optimization algorithm,” arXiv preprint arXiv:1411.4028, 2014. [40] L. Zhou, S.-T. Wang, S. Choi, H. Pichler, and M. D. Lukin, “Quantum approximate optimization algorithm: Performance, mechanism, and implementation on near-term devices,” Physical Review X, vol. 10, no. 2, p. 021067, 2020.
[41] J. B. Holliday, D. Blount, H. Q. Nguyen, S. U. Khan, and K. Luu, “Quadro: A hybrid quantum optimization framework for drone delivery,” in 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1. IEEE, 2025, pp. 2090–2100. [42] M. Benedetti, D. Garcia-Pintos, O. Perdomo, V. Leyton-Ortega, Y. Nam, and A. Perdomo-Ortiz, “A generative modeling approach for benchmarking and training shallow quantum circuits,” npj Quantum information, vol. 5, no. 1, p. 45, 2019. [43] H.-L. Huang, Y. Du, M. Gong, Y. Zhao, Y. Wu, C. Wang, S. Li, F. Liang, J. Lin, Y. Xu et al., “Experimental quantum generative adversarial networks for image generation,” Physical Review Applied, vol. 16, no. 2, p. 024051, 2021. [44] H.-Q. Nguyen, X. B. Nguyen, S. Y.-C. Chen, H. Churchill, N. Borys, S. U. Khan, and K. Luu, “Diffusion-inspired quantum noise mitigation in parameterized quantum circuits,” Quantum Machine Intelligence, vol. 7, no. 1, p. 55, 2025. [45] S. Y.-C. Chen, C.-M. Huang, C.-W. Hsing, H.-S. Goan, and Y.-J. Kao, “Variational quantum reinforcement learning via evolutionary optimization,” Machine Learning: Science and Technology, vol. 3, no. 1, p. 015025, 2022. [46] I. Cong, S. Choi, and M. D. Lukin, “Quantum convolutional neural networks,” Nature Physics, vol. 15, no. 12, pp. 1273–1278, 2019. [47] J. Romero, J. P. Olson, and A. Aspuru-Guzik, “Quantum autoencoders for efficient compression of quantum data,” Quantum Science and Technology, vol. 2, no. 4, p. 045001, 2017. [48] M. Panella and G. Martinelli, “Neural networks with quantum architecture and quantum learning,” International Journal of Circuit Theory and Applications, vol. 39, no. 1, pp. 61–77, 2011. [49] K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, “Quantum circuit learning,” Physical Review A, vol. 98, no. 3, p. 032309, 2018. [50] X.-B. Nguyen, H.-Q. Nguyen, H. Churchill, S. U. Khan, and K. Luu, “Quantum visual feature encoding revisited,” Quantum Machine Intelligence, vol. 6, no. 2, p. 61, 2024. [51] H.-Q. Nguyen, X.-B. Nguyen, H. Churchill, A. K. Choudhary, P. Sinha, S. U. Khan, and K. Luu, “Quantum-brain: Quantum-inspired neural network approach to vision-brain understanding,” arXiv preprint arXiv:2411.13378, 2024. [52] H.-Q. Nguyen, X. B. Nguyen, S. Pandey, T. Faltermeier, N. Borys, H. Churchill, and K. Luu, “Phi-adapt: A physics-informed adaptation learning approach to 2d quantum material discovery,” arXiv preprint arXiv:2507.05184, 2025. [53] S. Pandey, X.-B. Nguyen, H.-Q. Nguyen, T. Faltermeier, N. Borys, H. Churchill, and K. Luu, “Openqlaw: An agentic ai assistant for analysis of 2d quantum materials,” arXiv preprint arXiv:2603.17043, 2026. [54] X.-B. Nguyen, H.-Q. Nguyen, S. Pandey, T. Faltermeier, N. Borys, H. Churchill, and K. Luu, “Qupaint: Physics-aware instruction tuning approach to quantum material discovery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026. [55] K. Beer, D. Bondarenko, T. Farrelly, T. J. Osborne, R. Salzmann, D. Scheiermann, and R. Wolf, “Training deep quantum neural networks,” Nature communications, vol. 11, no. 1, p. 808, 2020. [56] Y. Li, Z. Wang, R. Han, S. Shi, J. Li, R. Shang, H. Zheng, G. Zhong, and Y. Gu, “Quantum recurrent neural networks for sequential learning,” Neural Networks, vol. 166, pp. 148–161, 2023. [57] Y. LeCun, C. Cortes, and C. Burges, “Mnist handwritten digit database,” ATT Labs [Online]. Available: http://yann.lecun.com/exdb/mnist, vol. 2, 2010. [58] H. Xiao, K. Rasul, and R. Vollgraf, “Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms,” arXiv preprint arXiv:1708.07747, 2017. [59] A. Krizhevsky, G. Hinton et al., “Learning multiple layers of features from tiny images,” 2009. [60] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems, vol. 32, 2019. [61] V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, S. Ahmed, V. Ajith, M. S. Alam, G. Alonso-Linaje, B. AkashNarayanan, A. Asadi et al., “Pennylane: Automatic differentiation of hybrid quantum-classical computations,” arXiv preprint arXiv:1811.04968, 2018. [62] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014.