Variational Quantum Transformer Architecture for Synthetic Language Generation
arXiv:2609.18565v1 [quant-ph] 16 Sep 2026
Julian Hager1[0000−0001−8220−4522] , Michael Kölle1[0000−0002−8472−9944] , Gerhard Stenzel1[0009−0009−0280−4911] , Tobias Rohe1[0009−0003−3283−0586] , Jonas Stein1[0000−0001−5727−9151] , and Claudia Linnhoff-Popien1[0000−0001−6284−9286] Institute of Informatics, LMU Munich, Oettingenstraße 67, 80538 Munich, Germany [email protected]
Abstract. We propose a compact NISQ-compatible quantum transformer architecture for synthetic QNLP sequence modelling. The model preserves the autoregressive next-token interface of a classical transformer, but replaces attention and feed-forward sublayers with variational quantum encoder blocks, connector circuits, decoder blocks and a direct two-qubit measurement readout. Token contexts are angle-encoded into small quantum registers, processed by parallel variational heads and encoder integration circuits and conditioned through decoder ancillae to produce a distribution over a four-token vocabulary. We evaluate several architecture variants on deterministic and lexicographic grammargeneration tasks against a compact classical transformer baseline. The quantum models are trainable end-to-end and learn nontrivial grammar structure, including perfect deterministic generation in individual runs and high lexicographic validity in the strongest variant. The classical baseline remains more accurate and stable and the quantum models are sensitive to initialization. The contribution is therefore not a claim of quantum advantage, but a concrete architecture and evaluation of transformer-inspired QNLP sequence modelling under near-term quantum constraints. Keywords: Quantum Natural Language Processing · Quantum Transformer · Variational Quantum Circuits · NISQ · Sequence Modelling
1
Introduction
Transformers are the dominant architecture for sequence modelling because they expose a practical autoregressive interface while integrating contextual information through trainable internal blocks [1]. Quantum natural language processing (QNLP) offers a complementary perspective in which linguistic structure is represented through tensorial or circuit-like models [6,7]. This raises a natural architectural question: can a transformer-style next-token model be expressed using compact variational quantum circuits while remaining compatible with near-term quantum constraints? This question is nontrivial in the NISQ setting, where useful models must use few
2
J. Hager et al.
qubits, moderate circuit depth and measurement interfaces that can be trained in a hybrid loop [2,3,5]. We address it by proposing a NISQ-compatible quantum transformer for synthetic QNLP sequence modelling. The model maps a fixed context window to a next-token distribution, but replaces classical attention and feed-forward sublayers with variational quantum encoder blocks, connector circuits, decoder blocks and a direct density-matrix readout. The term transformer is therefore used architecturally. The model retains heads, encoder integration, decoder conditioning and autoregressive prediction, but implements these stages with quantum circuits. Related work has also explored quantum natural language generation on nearterm devices [12]. Here we study a complementary architecture-driven setting based on autoregressive variational quantum circuits. We evaluate the architecture on two controlled grammar tasks over a four-token vocabulary: a deterministic cycle, AAA BBB CCC, and a lexicographic language of nondecreasing length-three words. These tasks are deliberately small. Their purpose is to isolate whether the architecture can learn explicit sequence structure under conditions where grammatical validity is measurable. Compared with a compact classical transformer baseline, the quantum variants learn nontrivial grammar regularities and can solve the deterministic task in individual runs, but remain less accurate and less stable overall. Thus this work does not claim quantum advantage. Its contributions are: (i) a concrete variational encoder–connector–decoder architecture for autoregressive QNLP experiments, (ii) parameter-efficient NISQ-scale variants with small active circuit width and (iii) a controlled evaluation using token-level and grammar-level metrics.
2
Background and Motivation
Transformers and QNLP. A transformer maps token representations to contextdependent features and, in autoregressive use, to next-token probabilities [1]. We use this interface as a design pattern rather than reproducing scaled dot-product attention directly. Parallel heads produce intermediate representations, an encoder integrates them, a decoder conditions generation on the encoder state and the final state is measured as a token distribution. This connects to QNLP, where linguistic composition has been studied through tensor-network or circuit-like representations [6,7], practical near-term executions [13] and software pipelines for mapping language to circuits [14]. Recent work has also explored quantum attention and transformer-like models [9,10,11]. Variational circuits under NISQ constraints. Variational quantum circuits encode data into quantum states, transform them with trainable gates and optimize parameters from measurement-derived outputs [3,5]. They are a natural model class for NISQ studies, but they impose architectural pressure by requiring small circuit width and depth together with a readout simple enough for repeated optimization [2]. These constraints motivate the choices made below, including small active registers, shallow strongly entangling layers, repeated reduction to two-qubit states and a direct measurement-based output rather than
Variational Quantum Transformer Architecture
3
a large learned classical head. The aim is not to replace classical transformers at scale, but to define a compact architecture whose sequence-modelling behaviour can be studied under near-term assumptions.
3
A NISQ-Compatible Quantum Transformer
The proposed model is an autoregressive hybrid quantum sequence model. Given a context window xt−c:t−1 , it estimates pΘ (xt = k | xt−c:t−1 ),
k ∈ V,
(1)
where Θ denotes the trainable circuit parameters. In the experiments, V = {A, B, C, space} and c = 4. The output vocabulary therefore matches the four computational-basis probabilities of a two-qubit readout state. A detailed circuit diagram is provided in Appendix A. This section summarizes the architectural data flow.
3.1
Input Encoding
Following the view of quantum models as feature-space methods [4], tokens are mapped to integer identifiers 0, . . . , |V| − 1 and angle-encoded into a four-qubit register by applying one rotation per context position, |ψx ⟩ =
c O j=1
RX
π 2
xj |0⟩,
xj ∈ {0, 1, 2, 3}.
(2)
The wire index represents the token position, so no separate positional encoding is used. For the four-token vocabulary and four-token context, this yields 44 distinguishable context encodings using four data qubits. We write the corresponding density matrix as ρx .
3.2
Encoder, Connector and Decoder
The encoder consists of E blocks. Each block contains two parallel variational head circuits followed by an integration circuit. In the first encoder layer, each head receives ρx on the four data qubits and appends two ancilla qubits. Later layers receive the two-qubit state produced by the previous encoder block. Each head applies V strongly entangling layers, implemented with parameterized rotations and CNOT entanglement in PennyLane [8], and returns a two-qubit reduced state. The integration circuit combines the two head states with the same variational template and again reduces the result to two qubits. After E blocks, the encoder state ρE is therefore a compact two-qubit representation of the complete context.
4
J. Hager et al.
The decoder consists of D blocks and conditions generation on ρE . In each decoder layer, a connector circuit couples the encoder state to a two-qubit decoder ancilla state. The connector applies CNOT gates from each encoder-output qubit to each decoder-ancilla qubit, followed by one trainable RY rotation on each decoder ancilla. The resulting state is combined with the original context state ρx and processed by a variational decoder circuit over the two decoder ancilla qubits and four data qubits. The decoder output is again reduced to two qubits and passed to the next decoder layer. This design provides a residual-style path from the original context to every decoder layer while keeping the intermediate representation small.
3.3
Measurement Readout and Variants
After the final decoder layer, the model obtains a two-qubit density matrix σD . Its diagonal entries are used directly as next-token probabilities, [diag(σD )]k , pΘ (xt = k | xt−c:t−1 ) = P|V|−1 j=0 [diag(σD )]j
k ∈ {0, 1, 2, 3}.
(3)
In exact arithmetic the denominator is one. The implementation clamps and renormalizes probabilities for numerical stability. Generation uses greedy autoregressive decoding by appending the most likely token to the context window. All variants use two heads, four data qubits, two-qubit encoder and decoder outputs and a maximum active circuit width of six qubits. Variant names have the form q_eE_dD_vV , indicating encoder depth, decoder depth and variational depth. For the fixed register sizes used here, the number of trainable parameters is P (E, D, V ) = 48 + 36(E − 1) V + D(18V + 2), (4) which counts encoder heads, integration circuits, decoder circuits and connector rotations. The evaluated variants are q_e2_d2_v1, q_e2_d2_v2, q_e2_d2_v3 and q_e3_d3_v2. Their parameter counts are reported in Table 1.
4
Evaluation on Synthetic QNLP Grammars
We evaluate the architecture on controlled grammar tasks rather than opendomain language. This keeps the setting small enough for repeated quantum simulation while making structural correctness explicit and measurable. Tasks. All experiments use three letter tokens, A, B and C, plus a separator token, space. Texts are sequences of three-character words separated by space and tokenized at the character level. The deterministic grammar is the periodic word cycle Gdet = (AAA, BBB, CCC)∗ , (5)
Variational Quantum Transformer Architecture
5
which tests whether the model can learn both word-internal repetition and longer-range phase structure. The lexicographic grammar contains all lengththree words with nondecreasing characters, Glex = {abc ∈ Σ 3 | a ≤ b ≤ c}.
(6)
This task admits many valid continuations, so exact token agreement with a sampled test sequence is stricter than grammatical validity. Training protocol and baseline. For each grammar, the training text contains 25 words and the test text contains 40 words. Generation starts from the first four test tokens and continues greedily for 100 tokens. We report means and standard deviations over seeds 17, 23 and 42. The same seed controls model initialization and training randomness. For the lexicographic task it also controls the sampled training text, with the held-out test text sampled from the following seed. Quantum variants are trained for 20 epochs with Adam and learning rate 10−2 using negative log-likelihood loss. The classical baseline is a small encoder–decoder transformer with dmodel = 4, two encoder layers, two decoder layers, two attention heads, feed-forward dimension 2, dropout 0.1 resulting in 700 trainable parameters. It is trained for 800 epochs with Adam and learning rate 10−3 . The longer schedule reflects the lower cost of classical batched training and is intended to provide a strong sanity-check baseline rather than a parameter-matched competitor. Metrics. We report token accuracy on the generated continuation and a grammar score measuring structural validity. Token accuracy compares generated tokens with the held-out test sequence after the initial context. The grammar score averages four rule-based components like valid-token rate, separator-position rate, word-length validity and either deterministic-cycle validity for Gdet or lexicographic-word validity for Glex . The score is a diagnostic complement to token accuracy. In particular, a lexicographic output such as repeated AAA can be formally valid while still having low diversity and low agreement with the sampled target sequence.
5
Results
Table 1 reports aggregate performance over the three seeds. Loss values are included for completeness but are most meaningful within a model family because the quantum and classical models use different training schedules. The token and grammar metrics are evaluated under the same autoregressive generation protocol. Deterministic grammar. The classical transformer solves the deterministic task in all runs, reaching perfect token accuracy and grammar score. The quantum variants learn weaker but nontrivial structure. The smallest model, q_e2_d2_v1, obtains low token accuracy but a higher grammar score, indicating valid symbols and partial separator alignment without reliable recovery of the full cycle. Increasing variational depth improves performance. q_e2_d2_v3 gives
6
J. Hager et al.
Table 1. Performance of the quantum transformer variants and the classical transformer baseline on the synthetic grammar tasks. Values are mean ± standard deviation over three seeds. Task
Model
Params
Loss ↓
Token acc. ↑
Grammar score ↑
Det. Det. Det. Det. Det.
Classical q_e2_d2_v1 q_e2_d2_v2 q_e2_d2_v3 q_e3_d3_v2
700 124 244 364 354
0.049±0.020 1.041±0.033 0.857±0.056 0.672±0.085 0.926±0.081
1.000±0.000 0.257±0.045 0.567±0.379 0.557±0.384 0.490±0.207
1.000±0.000 0.561±0.029 0.694±0.271 0.755±0.213 0.646±0.096
Lex. Lex. Lex. Lex. Lex.
Classical q_e2_d2_v1 q_e2_d2_v2 q_e2_d2_v3 q_e3_d3_v2
700 124 244 364 354
0.583±0.031 1.298±0.052 1.156±0.048 1.078±0.042 1.154±0.033
0.557±0.080 0.267±0.029 0.337±0.083 0.280±0.104 0.380±0.156
1.000±0.000 0.446±0.079 0.532±0.136 0.654±0.188 0.828±0.299
the strongest deterministic quantum result with grammar score 0.755 ± 0.213. The large standard deviations are significant. The V ∈ {2, 3} variants solve the deterministic continuation perfectly in one seed but fail to do so consistently. Thus the architecture is expressive enough to represent the rule, but the current optimization procedure does not reliably find that solution. Lexicographic grammar. The lexicographic task separates exact sequence prediction from grammatical validity. The classical baseline has moderate token accuracy, 0.557 ± 0.080, but perfect grammar score, meaning that it generates valid lexicographic words without reproducing the exact sampled test sequence. Quantum grammar scores increase with circuit expressivity, reaching 0.828 ± 0.299 for q_e3_d3_v2, which also obtains the best quantum token accuracy on this task. However, high lexicographic validity can arise from degenerate outputs such as repeated valid words. The result therefore shows that the architecture can learn formal word constraints, but also that nondeterministic grammars require diversity- or distribution-sensitive evaluation beyond validity alone. Interpretation. Across both tasks, the main positive finding is that the quantum transformer is trainable end-to-end and can encode grammar regularities with fewer trainable parameters than the compact classical baseline. The strongest quantum variants use between 354 and 364 parameters, compared with 700 for the baseline. This is not evidence of quantum advantage, since the classical model is clearly more accurate and stable. Rather, the results show that the encoder–connector–decoder circuit design is a viable architecture for controlled autoregressive QNLP experiments, while revealing the optimization instability that future work must address.
6
Discussion and Limitations
The experiments support this work’s central architectural claim, namely that an autoregressive sequence model can be built from variational quantum encoder, connector and decoder blocks while retaining the external next-token interface
Variational Quantum Transformer Architecture Deterministic
7
Lexicographic
1.0
Grammar score
0.8 0.6 0.4 0.2 0.0
Classical
q_e2_d2_v1
q_e2_d2_v2 Model
q_e2_d2_v3
q_e3_d3_v2
Fig. 1. Grammar-level validity of the classical baseline and quantum transformer variants on the deterministic and lexicographic tasks. Bars show mean scores over three seeds, error bars show one standard deviation and black points indicate the individual seed outcomes.
of a transformer-style model. The positive result is not superiority over a classical transformer, but feasibility. Individual quantum runs solve the deterministic grammar and the lexicographic task shows that the model can learn word-level validity constraints that token accuracy alone does not capture. The comparison between variants suggests that, at the tested scale, increasing variational depth inside each block is more useful than simply adding another encoder–decoder layer. However, the large standard deviations show that training remains sensitive to initialization and optimization. Good solutions appear to exist within the architecture, but the current training procedure does not find them reliably. This is consistent with broader difficulties in variational quantum optimization and should temper any interpretation of the aggregate results. Several limitations remain. The tasks are synthetic and do not test semantic representation, compositional meaning, or natural-language generalization. The experiments use exact simulation rather than shot-based or noisy hardware execution. Finally, the direct two-qubit readout ties the present vocabulary size to four tokens. Larger vocabularies will require more readout qubits, a hybrid output map, hierarchical decoding, or another scalable readout scheme. Future work should therefore study richer encodings, more stable optimization, diversitysensitive metrics for nondeterministic grammars and execution under realistic NISQ noise and sampling constraints.
7
Conclusion
We introduced a NISQ-compatible quantum transformer architecture for synthetic QNLP sequence modelling. The model preserves autoregressive next-token
8
J. Hager et al.
prediction while replacing classical sequence-processing blocks with variational quantum encoder heads, connector circuits, decoder blocks and a direct twoqubit measurement readout. On deterministic and lexicographic grammar tasks, the architecture is trainable end-to-end and learns nontrivial structure, including perfect deterministic generation in individual runs and high lexicographic validity in the strongest variant. The classical baseline remains more accurate and stable, so the results should be read as evidence for a viable architecture, not as a quantum advantage claim. Future work should extend the model to larger vocabularies and contexts, improve optimization stability and evaluate the architecture under shot-based and noisy quantum execution. AI Assistance Disclosure. The authors used OpenAI ChatGPT for language editing. All experimental design, code, results, interpretation and final responsibility remain with the authors. Disclosure of Interests. The authors have no competing interests to declare that are relevant to the content of this article.
References 1. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Advances in Neural Information Processing Systems 30, pp. 5998–6008 (2017) 2. Preskill, J.: Quantum computing in the NISQ era and beyond. Quantum 2, 79 (2018). https://doi.org/10.22331/q-2018-08-06-79 3. Cerezo, M., Arrasmith, A., Babbush, R., Benjamin, S.C., Endo, S., Fujii, K., McClean, J.R., Mitarai, K., Yuan, X., Cincio, L., Coles, P.J.: Variational quantum algorithms. Nature Reviews Physics 3, 625–644 (2021). https://doi.org/10.1038/ s42254-021-00348-9 4. Schuld, M., Killoran, N.: Quantum machine learning in feature Hilbert spaces. Physical Review Letters 122, 040504 (2019). https://doi.org/10.1103/ PhysRevLett.122.040504 5. Benedetti, M., Lloyd, E., Sack, S., Fiorentini, M.: Parameterized quantum circuits as machine learning models. Quantum Science and Technology 4(4), 043001 (2019). https://doi.org/10.1088/2058-9565/ab4eb5 6. Coecke, B., Sadrzadeh, M., Clark, S.: Mathematical foundations for a compositional distributional model of meaning. Linguistic Analysis 36(1–4), 345–384 (2010) 7. Meichanetzidis, K., Gogioso, S., de Felice, G., Chiappori, N., Toumi, A., Coecke, B.: Quantum natural language processing on near-term quantum computers. arXiv preprint arXiv:2005.04147 (2020) 8. Bergholm, V., Izaac, J., Schuld, M., Gogolin, C., Ahmed, S., Ajith, V., Alam, M.S., Alonso-Linaje, G., AkashNarayanan, B., Asadi, A., et al.: PennyLane: automatic differentiation of hybrid quantum-classical computations. arXiv preprint arXiv:1811.04968 (2018) 9. Li, G., Zhao, X., Wang, X.: Quantum self-attention neural networks for text classification. Science China Information Sciences 67, 142501 (2024). https://doi. org/10.1007/s11432-023-3879-7
Variational Quantum Transformer Architecture
9
10. Guo, N., Yu, Z., Choi, M., Han, Y., Agrawal, A., Nakaji, K., Aspuru-Guzik, A., Rebentrost, P.: Quantum Transformer: Accelerating model inference via quantum linear algebra. arXiv preprint arXiv:2402.16714 (2024) 11. Khatri, N., Matos, G., Coopmans, L., Clark, S.: Quixer: A Quantum Transformer Model. arXiv preprint arXiv:2406.04305 (2024) 12. Karamlou, A., Pfaffhauser, M., Wootton, J.: Quantum natural language generation on near-term devices. arXiv preprint arXiv:2211.00727 (2022) 13. Lorenz, R., Pearson, A., Meichanetzidis, K., Kartsaklis, D., Coecke, B.: QNLP in practice: Running compositional models of meaning on a quantum computer. arXiv preprint arXiv:2102.12846 (2021) 14. Kartsaklis, D., Fan, I., Yeung, R., Pearson, A., Lorenz, R., Toumi, A., de Felice, G., Meichanetzidis, K., Clark, S., Coecke, B.: lambeq: An efficient high-level Python library for quantum NLP. arXiv preprint arXiv:2110.04236 (2021)
10
J. Hager et al.
A
Detailed Circuit Diagram
The following diagram gives the full circuit-level layout of the architecture instance used to illustrate the model structure. It is moved to the appendix to keep the main paper focused on the architectural definition and experimental results. emb
head
enc
head
1
enc
|0⟩
Rx (θ1 )
R(α11 , β11 , γ11 )
R(δ11 , ϵ11 , ζ11 )
R(α13 , β13 , γ13 )
R(δ12 , ϵ21 , ζ12 )
|0⟩
Rx (θ2 )
R(α21 , β21 , γ21 )
R(δ21 , ϵ12 , ζ21 )
R(α23 , β23 , γ23 )
R(δ22 , ϵ22 , ζ22 )
|0⟩
Rx (θ3 )
|0⟩
Rx (θ4 ) |0⟩ |0⟩
con
2
3
con
R(α31 , β31 , γ31 )
|0⟩
R(α41 , β41 , γ41 )
|0⟩
R(α51 , β51 , γ51 )
|0⟩
|0⟩
R(α33 , β33 , γ33 )
|0⟩
|0⟩
Ry (λ11 )
R(ρ11 , σ11 , ϕ11 )
Ry (λ21 )
R(ρ21 , σ12 , ϕ21 )
R(α61 , β61 , γ61 )
|0⟩
|0⟩
R(α43 , β43 , γ61 )
|0⟩
|0⟩
Ry (λ12 )
R(ρ12 , σ21 , ϕ12 )
Ry (λ22 )
R(ρ22 , σ22 , ϕ22 )
emb
dec
head
dec
head
|0⟩
Rx (θ1 )
R(α12 , β12 , γ12 )
R(δ31 , ϵ13 , ζ31 )
|0⟩
R(α14 , β14 , γ14 )
R(δ32 , ϵ23 , ζ32 )
|0⟩
|0⟩
Rx (θ2 )
R(α22 , β22 , γ22 )
R(δ41 , ϵ14 , ζ41 )
|0⟩
R(α24 , β24 , γ24 )
R(δ42 , ϵ24 , ζ42 )
|0⟩
|0⟩
Rx (θ3 )
R(α32 , β32 , γ32 )
|0⟩
|0⟩
Rx (θ1 )
R(ρ13 , σ31 , ϕ13 )
|0⟩
|0⟩
Rx (θ1 )
R(ρ23 , σ32 , ϕ23 )
|0⟩
|0⟩
Rx (θ4 )
R(α42 , β42 , γ42 )
|0⟩
|0⟩
Rx (θ2 )
R(ρ14 , σ41 , ϕ14 )
|0⟩
|0⟩
Rx (θ2 )
R(ρ24 , σ42 , ϕ24 )
|0⟩
|0⟩
R(α52 , β52 , γ52 )
|0⟩
|0⟩
R(α34 , β44 , γ54 )
|0⟩
|0⟩
Rx (θ3 )
R(ρ15 , σ51 , ϕ15 )
|0⟩
|0⟩
Rx (θ3 )
R(ρ25 , σ52 , ϕ25 )
|0⟩
|0⟩
R(α62 , β62 , γ62 )
|0⟩
|0⟩
R(α44 , β44 , γ44 )
|0⟩
|0⟩
Rx (θ4 )
R(ρ16 , σ61 , ϕ16 )
|0⟩
|0⟩
Rx (θ4 )
R(ρ26 , σ62 , ϕ26 )
|0⟩
emb
emb
Fig. 2. Detailed circuit diagram of the two-encoder, two-decoder quantum transformer instance. Embedding circuits are shown in red, individual heads in blue, encoder integration circuits in yellow, connector blocks in green and decoder blocks in orange.