ConceptioArchivearXiv CS
arXiv CSopen access

Learning to Concatenate Quantum Codes

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Learning to Concatenate Quantum Codes Nico Meyer∗† , Christopher Mutschler∗‡ , Dominik Seuß∗§ , Andreas Maier† , and Daniel D. Scherer∗ ∗ Fraunhofer IIS, Fraunhofer Institute for Integrated Circuits IIS, Nuremberg, Germany † Pattern Recognition Lab, Friedrich-Alexander-University Erlangen-Nuremberg, Erlangen, Germany ‡ Machine Learning and Positioning Systems Lab, University of Technology Nuremberg (UTN), Nuremberg, Germany

|0⟩

|0⟩

|0⟩

inner inner encoding

|ψ⟩

outer encoding

Abstract—Concatenating quantum error correction codes scales error correction capability by driving logical error rates down double-exponentially across levels. However, the noise structure shifts under concatenation, making it hard to choose an optimal code sequence. We automate this choice by estimating the effective noise channel after each level and selecting the next code accordingly. In particular, we use learning-based methods to tailor small, non-additive encoders when the noise exhibits sufficient structure, then switch to standard codes once the noise is nearly uniform. In simulations, this level-wise adaptation achieves a target logical error rate with far fewer qubits than concatenating stabilizer codes alone–reducing qubit counts by up to two orders of magnitude for strongly structured noise. Therefore, this hybrid, learning-based strategy offers a promising tool for early fault-tolerant quantum computing. Index Terms—quantum computing, machine learning, quantum error correction, code concatenation, variational algorithm

outer encoding

arXiv:2604.14931v1 [quant-ph] 16 Apr 2026

§ Center for Artificial Intelligence (CAIRO), Technical University of Applied Sciences Würzburg-Schweinfurt, Germany

outer

Fig. 1: Schematic of noise-aware code concatenation. A logical state |ψ⟩ is recursively encoded by an outer and inner code. Because the effective noise channel changes across levels, we estimate its structure after each concatenation step and use machine-learning methods to tailor the next-level code. This enhances per-level noise suppression and reduces the qubit overhead required to reach a target logical error rate.

I. I NTRODUCTION Quantum error correction (QEC) is essential for reliable quantum computation in the presence of noise and decoherence. Code concatenation by recursively encoding logical states using nested codes achieves doubly exponential suppression of errors [1], [2], assuming beyond-threshold physical components. Using heterogeneous codes across levels can reduce qubit overhead [3], assuming simple stationary noise structure. Realistic noise typically is structured, e.g., dephasing-dominant or otherwise anisotropic [4], [5]. Furthermore, the effective noise channel is reshaped across concatenation levels [1], [2], [6], motivating a level-wise, noise-tailored concatenation strategy. Machine learning (ML) can be utilized to discover new QEC codes with specific properties, in particular using techniques from reinforcement learning [7], [8] and variational methods [9], [10]. In particular, there also exist ML techniques for tailoring encodings to noise structures [11]–[13]. One approach to variational quantum error correction (VarQEC) [14] learns non-additive, measurement-free codes [15] by minimizing the information loss under given noise structures, enabling approximate QEC [16] with fewer qubits than standard stabilizer codes [13], [17]. Yet, how to scale these codes by leveragThe research was supported by the German Federal Ministry of Research, Technology and Space, funding program Quantum Systems, via the project Q-GeneSys, grant number 13N17389. The research is also part of the Munich Quantum Valley (MQV), which is supported by the Bavarian state government with funds from the Hightech Agenda Bayern Plus. Correspondence to: [email protected]

ing the tailored encoders systematically within concatenation is an open research question. We address this gap by introducing learning to concatenate: a pipeline that alternates between (i) estimating the effective single-qubit logical channel produced by the current level and (ii) tailoring the next-level code to the uncovered noise structures. This procedure is sketched in Fig. 1. For the first step, we develop a fidelity-only, two-design estimator to recover the Pauli-Liouville diagonal of the logical channel without full process tomography [18]–[20]. We then tailor small VarQEC patches [13] when structure is exploitable and switch to standard fixed-distance stabilizer codes once the logical channel is effectively depolarizing. Empirical analyses of this procedure confirms prior findings that anisotropic noise converges rapidly to an isotropic channel under concatenation with nonCSS codes [6], [21]. Moreover, we identify noise structures that yield a 4-fold to over 100-fold reduction in overhead for achieving a target error rate, compared to concatenating standard codes. The remainder of this paper is organized as follows: In Sec. II, we summarize the necessary background and prior work on quantum error correction, variational codes, and code concatenation. The procedure for estimating the structure of the effective noise channels between concatenation levels, and the learning and concatenation procedure for the tailored codes is outlined in Sec. III. The empirical analysis of the proposed pipeline for different noise structures is to be found

in Sec. IV. Finally, in Sec. V, we provide a summary, discuss open questions and future work, and position the proposed concept within the context of early fault-tolerant quantum computing [22]. II. P RELIMINARIES AND P RIOR W ORK Noise in open quantum systems can be described using the notion of density matrices ρ, which describe the mixed states as a statistical ensemble of pure n-qubit states. The non-unitary process of noise is described by completely positive trace preserving (CPTP) maps, Ptypically specified using the Kraus representation N (ρ) = k Ek ρEk† , where the Ek P are Kraus operators with k Ek† Ek = I [23]. The most common example of such a noise channel is symmetric depolarizing noise, where all error types (i.e., bitflip, phaseflip, and combinations thereof) are equally likely, described in its single-qubit version by p p p Ndep (ρ) = (1 − p)ρ + XρX † + Y ρY † + ZρZ † , (1) 3 3 3 where p is the overall depolarizing probability. In this work, we target more general noise channels, subsumed under arbitrary single-qubit Pauli noise NPauli (ρ) = (1 − p)ρ + pX XρX † + pY Y ρY † + pZ ZρZ † , (2) where the overall noise strength is given by p = pX +pY +pZ . In particular, we focus on three noise channels: (1) asymmetric depolarizing noise Nadep as investigated in [11], [13], with an asymmetry factor c = 0.5, resulting in ppX = ppY = 0.07 and pZ p = 0.86; (2) standard bitflip noise Nbit , with pX = p and pY = pZ = 0; (3) strictly correlated Pauli X and Z noise, i.e. a Pauli Y -flip Nyflip , with pY = p and pX = pZ = 0. While addressing non-Pauli channels like amplitude damping noise is conceptually possible [13], we leave such considerations for future work. Throughout this manuscript, we assume that errors for multi-qubit systems act independently on each qubit, giving an overall noise model for an n-qubit system as N ⊗n = N n j=1 Nj . Related literature suggests that our analysis should extend to correlated errors [13], which is, however, out of scope for this work. A. Quantum Error Correction Quantum error correction is the primarily pursued concept to protect quantum states against noise channels, by encoding the logical state into a collection of physical qubits [24], [25]. Typically, we refer to QEC codes C by their parameters ((n, K, d)).

(3)

Hereby, n indicates the number of physical qubits the logical K-dimensional state is encoded into, with particular relevance of codes where k = log2 K is a whole number of logical qubits. The code distance d indicates that up to d − 1 arbitrary errors can be detected, and ⌊ d−1 2 ⌋ arbitrary errors can be corrected. When the code does not have a provable code distance, or it is unknown, we simply write ((n, K)). For the

special and frequently considered case of stabilizer (i.e. additive) codes [26], one typically switches to the notation [[n, k, d]].

(4)

Again, n indicates the number of physical qubits, k the number of logical qubits with K = 2k , and d the code distance. The procedure of QEC can be roughly separated into three consecutive steps, with the first one being the encoding of the logical state into multiple physical qubits as   ⊗n−k † ρL = Uenc ρ ⊗ |0⟩ ⟨0| Uenc . (5) The noise-affected logical state ρ̃L = N (ρL ) then undergoes potentially multiple rounds of recovery     ⊗r † ρ̂L = Trr Urec ρ̃L ⊗ |0⟩ ⟨0| Urec , (6) where r is a fresh register of ancilla qubits for every recovery cycle. As we elaborate in Sec. II-B, the QEC procedure in this work is measurement-free [15], i.e., avoids the typical stabilizer measurements and conditional correction operations [25]. Finally, before measurement, the logical state is decoded as  † ρ̂ = Trn−k Uenc ρ̂L Uenc . (7) This gives rise to the overall error correction procedure R, composed from encoding following Eq. (5), potentially multiple rounds of recovery according to Eq. (6), and final decoding as in Eq. (7). The objective of this procedure, associated with an QEC code C, is to keep the effect of noise under control, i.e. (R ◦ N )(ρ) ∝ ρ. B. Variational Quantum Error Correction Recent work on ML-enhanced QEC has shown that the fit of a QEC encoding for a specific noise channel can be quantified by the worst-case distinguishability loss D(N ) = max ∆T (ρ, σ; N ), ρ,σ

(8)

where ∆T (ρ, σ; N ) = T (ρ, σ)−T (N (ρL ), N (σL )) is the lost trace distance for state pairs ρ, σ [13], [17]. The intuition is that this measure, based on the trace distance, quantifies the loss of information under encoding and error channel, which should be kept as minimal as possible to allow for successful error correction. This connection has been formally and empirically supported by showing that a low distinguishability loss guarantees the existence of a high-fidelity recovery operation. Furthermore, the quality of a sophisticated recovery operation can be quantified by the worst-case fidelity loss F(N ) = 1 − min F (ρ, ρ̂), ρ

(9)

with ρ̂ as defined in Eqs. (5) to (7) [9]. To discover encodings Uenc and recovery operations Urec that are desirable following Eqs. (8) and (9), one can establish a machine learning procedure [13]: In order to make the operations tunable, they are instantiated as variational quantum circuits [27] with trainable parameters Θ and Φ, respectively. Concretely, we employ a so-called randomized entangling ansatz (REA) [13] for both Uenc (Θ) and Urec (Φ)

ρ inner Cl+1 outer Cl

outer Cl

|0⟩

Encl

Nl⊗n

Recl

Enc†l

Nl+1 (ρ)

|0⟩

outer Cl

Fig. 3: Pipeline to analyze the effective channel introduced by level l of a concatenated code. An input state ρ from a unitary two-design undergoes encoding, noise, a single round of recovery, and decoding. Parameters pX , pY , pZ of the Pauli channel in Eq. (2) are (least-squares) fitted to estimate Nl+1 .

level l + 1 qubit: logical for Cl+1

level l qubits: physical for Cl+1 , logical for Cl level l − 1 qubits: physical for Cl

Fig. 2: (inspired by Fig. 9.1 of [30]) A logical qubit at level l +1 of the concatenation is encoded into nl+1 physical qubits of level l using the inner code Cl+1 with code parameters ((nl+1 , 2)). These level-l qubits are furthermore logical qubits of the outer code Cl with code parameters ((nl , 2)), which individually are encoded into nl physical qubits at level l − 1. This construction produces a concatenated code with parameters ((nl+1 nl , 2)), where qubits at level l = 0 are the physical qubits on actual the hardware. To ensure fault-tolerance, the respective encodings are typically applied bottom-to-top, as also indicated in Fig. 1.

even multiple quantum error correction codes [31]. This is achieved by encoding the logical state using an inner code, for which its physical qubits are the logical qubits of multiple instances of an outer code. This procedure is sketched for inner codes Cl+1 with hyperparameters ((nl+1 , 2)) and outer codes Cl with ((nl , 2)) at arbitrary concatenation levels l in Fig. 2, where w.l.o.g. we restrict to code patches with only single logical qubits for the remainder of this work [3]. Given encoders Encl+1 for Cl+1 and Encl for Cl , the con⊗n catenated code is therefore produced using Encl+1 ◦Encl l+1 , as depicted in Fig. 1. It is easy to see that the concatenated code has code parameters

throughout this work, which consists of an initial parameterized single-qubit layer, followed by randomly-placed parameterized two-qubit operations. Due to this instantiation of encoding and recovery with parameterized circuits, the procedure is also referred to as variational quantum error correction (VarQEC). We note that the acronym “VarQEC” was originally introduced for a variational procedure whose objective is derived from the Knill-Laflamme conditions [14]. However, this is not the notion we employ in this work. Given a noise channel N , the encoding can be tailored towards the structure of the noise by updating towards

((nl+1 nl , 2)),

min D(N ; Θ). Θ

(10)

After the encoding has been established, we can proceed to train the measurement-free [15] recovery operation by min F(N ; Φ), Φ

(11)

where we restrict to single rounds of recovery. In practice, due to instabilities caused by the extrema operations within the loss function [28], the worst-case formulation is replaced by an average-case proxy. Furthermore, to make the evaluation more efficient, a two-design approximation [18], [29] is employed. This procedure can be used to discover non-additive QEC codes that reduce the loss of information under structured noise using less physical qubits than established stabilizer codes [13]. To further scale the correction capabilities of these VarQEC codes, we envision code concatenation to be an important cornerstone. C. Code Concatenation The underlying idea of code concatenation is to improve the logical error rate by recursively encoding the state into two or

(12)

for which it is known that d ≥ dl+1 dl , where dl+1 and dl are the code distances of Cl+1 and Cl , respectively [30]. For stabilizer codes with parameters [[nl+1 , 1, dl+1 ]] and [[nl , 1, dl ]] this simplifies to [[nl+1 nl , 1, dl+1 dl ]]

(13)

for the concatenated code. It is well known that, for physical error rates below the threshold and with fully fault-tolerant gadgets, code concatenation suppresses logical error rates doubly exponentially in the concatenation level [1], [2]. III. C ONCATENATING VARIATIONAL C ODES To scale the correction capabilities of the VarQEC codes, we propose to concatenate instances that are tailored for the noise structures at every concatenation level. In particular, this requires first estimating the effective single-qubit noise channel induced by the code at the current level of the concatenation, which we discuss in Sec. III-A. Consecutively, we then demonstrate how to tailor the encoding-recovery pair for the next level to this noise structure in Sec. III-B. A. Analyzing Noise Channels under Concatenation To analyze the noise suppression under code concatenation, we require a procedure that reconstructs the effect of a level-l code on the noise strength and structure. For this, we estimate the effective single-qubit channel Nl+1 by fitting a Pauli channel to fidelities computed over a unitary two-design, where each input state ρ evolves governed by Eqs. (5) to (7). This setup is sketched in Fig. 3, and follows the standard encodenoise-recover-decode mapping used to analyze concatenated

codes, based on the insight that encoding and decoding are isometric [1], [2]. Our approach avoids full process tomography by only recovering the diagonal part of the channel, i.e., the diagonal of its Pauli-Liouville matrix [19]. In this work, we are only concerned with unital noise channels, but the concept could be extended to non-unital channels like amplitude damping by incorporating Pauli-twirling techniques [32]. Concretely, we fit a least-squares estimate of the singlequbit channel in the Bloch/Pauli-Liouville representation [20], but restrict the model to a Pauli channel and drive the regression with fidelities from a unitary two-design [29]. To the best of our knowledge, this fidelity-only fitting technique has not been described in literature. Let R = diag(ηX , ηY , ηZ ) be the Bloch matrix of a single-qubit Pauli channel and let ri = [rX , rY , rZ ]i be the Bloch vector for input state ρi from the two-design S = {|±X⟩ , |±Y ⟩ , |±Z⟩}. With prediction 1 1 T 2 2 2 2 (1 + ri Rri ) = 2 (1 + ηX rX + ηY rY + ηZ rZ )i and target bi := 2F (ρi , ρˆi ) − 1 evaluated following Fig. 3, we can set up the least-squares formulation 2

argminη ∥Aη − b∥ .

level 1

Z

Z

Y

Y

Y

X

X

X

level 2

level 3

Fig. 4: Shift of noise structure under concatenation of (i) a ((5, 2)) VarQEC code tailored to Nyflip noise at level 1, resulting in Pauli noise featuring only disjunct bit- and phaseflips at level 2; (ii) another ((5, 2)) VarQEC code tailored to this noise structure, resulting in almost uniform depolarizing noise at level 3. For the next level, as the noise is mostly unstructured, we continue with concatenating a standard [[5, 1, 3]] code. For noise suppression behaviour see Fig. 5, and for details on the noise channels see Tab. 1.

(14)

2 2 ]i ∈ R|S|×3 and b = [bi ]i ∈ R|S| . , rY2 , rZ Hereby, A = [rX

Using the six cardinal two-design states described above, the solution to Eq. (14) is given by [⟩) + F (|−P ⟩ , |−P [⟩) − 1, ηP = F (|+P ⟩ , |+P

Z

(15)

where P ∈ {X, Y, Z} and F is the measured fidelity. The probabilities appearing in the Kraus representation from Eq. (2) follow as pX = 14 (1 − ηY − ηZ + ηX ), and accordingly for pY and pZ . In case the effective channel Nl+1 is exactly represented by single-qubit Pauli noise, we are done. Otherwise, the solutions following Eq. (15) might not satisfy the properties of a CPTP map, in particular violate pP ≥ 0 for all P , and pX + pY + pZ ≤ 1. In such cases, we can recover the closest physically valid Pauli channel by mapping to the probability simplex of pX , pY , pZ with standard sorting-based approaches [33]. B. Tailoring Codes for Concatenation Level For the first level of the concatenation, we straightforwardly use the VarQEC approach described in Sec. II-B to tailor a code C1 , consisting of encoding Enc1 and recovery Rec1 , to the initial noise channel N1 . Afterwards, the tomography method from Sec. III-A is used to estimate the effective noise channel after this code instance, and therefore the input for the next level of the concatenation, denoted as N2 . In the next cycle, a code C2 is tailored to this noise, resulting in another effective channel N3 after encoding and correction. This procedure is iteratively repeated by tailoring a code Cl to noise Nl , producing the effective channel Nl+1 , until a desired target noise suppression rate is achieved. It is crucial to tailor the codes to every level of the concatenation separately, as not only the noise strength, but also the noise structure changes after each concatenation instance [6]. Using different codes for each concatenation level has been proven effective in reducing the overall qubit count, even when

restricting to stabilizer codes [3]. In particular, it has been observed that under non-CSS codes, anisotropic noise quickly changes to effectively isotropic channels under concatenation, i.e., converges towards symmetric depolarizing noise as given in Eq. (1) [6]. All codes we consider in this work are nonCSS, including the [[3, 1, 1]] bitflip and [[5, 1, 3]] perfect stabilizer codes [21], [24], as well as all VarQEC codes, as they are non-additive. This behavior is expected to be different for CSS codes like e.g. the [[7, 1, 3]] Steane or the [[9, 1, 3]] Shor code [24], [25], where anisotropic noise typically remains anisotropic under concatenation [6]. However, for the sake of this paper, the [[5, 1, 3]] code poses the hardest baseline, as it corrects for an arbitrary single-qubit error with the provably smallest number of physical qubits, which keeps the resource overhead analyzed in Sec. IV minimal. To ensure fault tolerance in practice, one needs to encode using logical operations within the codespace of the previous concatenation level, which typically leads to still choosing the [[7, 1, 3]] code over [[5, 1, 3]]. However, such considerations, also for the VarQEC codes, are out of the scope of this manuscript and will be investigated in future work [34]. IV. E MPIRICAL S ETUP AND E VALUATION To empirically analyze the behaviour of a noise channel under code concatenation, we implement the noise tomography procedure described in Sec. III-A in python, relying mainly on the pennylane [35] and qiskit-torch-module [36] libraries. For tailoring the VarQEC codes to the specific noise structures, we employed an open-source implementation of the VarQEC approach [13]. Training is conducted using the quasiNewton L-BFGS optimizer [37] on REA ansätze [13] with n · (n + 1) blocks for the encoding unitary. While training VarQEC codes for every concatenation level from scratch is possible, we observed that using warm-start initialization [38]

100 10

−1

concatenation 0 1 2

Nyflip

[[5,1,3]] perfect

5

6

3

10−3

105-fold

10−4 10−5 10

−6

Nbit

worst-case fidelity loss

100 10

−1

[[3,1,1]] bitflip

10−5 10−6

Nadep

10−1 10−2

4.5-fold

10−3 −4

10−5 10−6

pX/p

Nyflip

0

0.10

0.00

1.00

0.00

[[5, 1, 3]]

1 2 3 4 5 6

0.082 0.056 0.027 0.0068 0.00046 0.0000020

0.50 0.37 1/3 1/3 1/3 1/3

0.00 0.24 1/3 1/3 1/3 1/3

0.50 0.37 1/3 1/3 1/3 1/3

((5, 2)) ((5, 2)) [[5, 1, 3]]

1 2 3

0.0085 0.00060 0.0000037

0.50 0.33 1/3

0.00 0.33 1/3

0.50 0.34 1/3

code

level

strength p

pX/p

Nbit

0

0.10

1.00

0.00

0.00

[[3, 1, 1]]

1 2 3 4

0.028 0.0023 0.000016 0.0000001

1.00 1.00 1.00 1.00

0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00

[[5, 1, 3]]

1 2 3 4 5 6

0.082 0.055 0.027 0.0068 0.00045 0.0000020

0.00 0.26 1/3 1/3 1/3 1/3

0.50 0.37 1/3 1/3 1/3 1/3

0.50 0.37 1/3 1/3 1/3 1/3

((3, 2)) ((5, 2)) [[5, 1, 3]] [[5, 1, 3]]

1 2 3 4

0.028 0.0057 0.00033 0.000011

0.50 0.27 1/3 1/3

0.50 0.30 1/3 1/3

0.00 0.43 1/3 1/3

((5, 2)) ((5, 2)) [[5, 1, 3]]

1 2 3

0.0086 0.00065 0.0000043

0.00 0.20 1/3

0.50 0.35 1/3

0.50 0.45 1/3

code

level

strength p

pX/p

Nadep

0

0.10

0.07

0.07

0.86

[[5, 1, 3]]

1 2 3 4 5 6

0.081 0.054 0.026 0.0062 0.00038 0.0000014

0.44 0.35 1/3 1/3 1/3 1/3

0.44 0.35 1/3 1/3 1/3 1/3

0.12 0.30 1/3 1/3 1/3 1/3

((4, 2)) [[5, 1, 3]] [[5, 1, 3]] [[5, 1, 3]] [[5, 1, 3]]

1 2 3 4 5

0.062 0.035 0.011 0.0012 0.000014

0.36 1/3 1/3 1/3 1/3

0.32 1/3 1/3 1/3 1/3

0.32 1/3 1/3 1/3 1/3

((5, 2)) [[5, 1, 3]] [[5, 1, 3]] [[5, 1, 3]] [[5, 1, 3]]

1 2 3 4 5

0.057 0.030 0.0081 0.00066 0.0000044

0.34 1/3 1/3 1/3 1/3

0.34 1/3 1/3 1/3 1/3

0.32 1/3 1/3 1/3 1/3

90-fold

10−4

10

strength p

((3,2)) VarQEC

10−2 10−3

level

((5,2)) VarQEC

4

10−2

code

((4,2)) VarQEC 100

101

102

103

104

physical qubits Fig. 5: Noise suppression under code concatenation. The colors and markers indicate which code was concatenated at which level, with VarQEC codes targeted to the specific noise structures. As most channels converge to uniform depolarizing noise under concatenation (see Sec. III-B), we concatenate with the standard [[5, 1, 3]] code, once the structure is insufficient for tailoring, at which point we also show the overhead reduction compared to concatenating only standard stabilizer codes from the beginning. Details on the noise at every level of the concatenation are to be found in Tab. 1.

proportion pY /p pZ/p

proportion pY /p pZ/p

proportion pY /p pZ/p

Tab. 1: Change of noise structure under code concatenation. To unify the notation, we note the total noise strength p and the X, Y , and Z contributions of single-qubit Pauli noise, as defined in Eq. (2). We list the noise strength and structure change after each concatenation level with the QEC codes specified in the leftmost column. This structure change is visualized for initially Nyflip noise in Fig. 4, and the noise suppression under concatenation is shown in Fig. 5.

with parameters from the previous level speeds up the convergence significantly. For our experiments, we simulate the encoding and recovery operations themselves to be noise-free, i.e. we assume the existence of fault-tolerant gadgets for realizing encoding and correcting. Using this setup, in Fig. 5 we show the analysis of three noise channels with initial noise strength p = 0.1. For each, we compare fully stabilizer-code concatenations with hybrid concatenations that use tailored VarQEC codes at the outer and stabilizer codes at the inner levels. Fig. 5 shows the resulting suppression by plotting the worst-case fidelity loss (Eq. (9)) versus the number of physical qubits n following Eqs. (12) and (13). Noise structures at each level are listed in Tab. 1. For initial Pauli Y -flip Nyflip noise (top plot of Fig. 5), we compare concatenating ((5, 2)) VarQEC codes at level 1 and 2 and a standard [[5, 1, 3]] code at level 3 to using the [[5, 1, 3]] codes at every level. In the former case, we switch to the standard code at level 3 because the effective noise is nearly uniformly depolarizing with ppX ≈ ppY ≈ ppZ ≈ 31 (see also Fig. 4), leaving no exploitable structure for the VarQEC procedure. For comparable noise suppression, the hybrid strategy reduces resources by about 105-fold, from n ≈ 2625 to n = 25 (i.e., 52 after two levels) physical qubits. For initial bitflip Nbit noise (middle plot of Fig. 5), concatenating tailored 5-qubit codes yield a substantial 90-fold overhead reduction. We also test 3-qubit codes at level 1: the [[3, 1, 1]] bitflip code and ((3, 2)) VarQEC codes. Both achieve the same total noise suppression after the first layer, but the variational code alters the noise structure to Pauli noise with pX pY 1 p ≈ p ≈ 2 , whereas the stabilizer code maintains an effective single-qubit bitflip channel. This, however, requires concatenating the ((3, 2)) with a higher-qubit ((5, 2)) code at level 2, to achieve beyond break-even noise suppression, increasing overall resource requirements. Thus, stabilizer codes tailored to noise structure can still outperform machine-learned nonadditive codes in the concatenated setup. However, this also points to a natural future extension of the concept: instead of minimizing the worst-case fidelity loss, the objective could be to enforce desired channel structure after correction, enabling not only suppression but finer-grained noise control. Lastly, for asymmetric depolarizing noise Nadep (bottom plot of Fig. 5), a single instance of a tailored ((5, 2)) VarQEC code already makes the effective noise channel almost uniform. Therefore, from the second level onwards, we concatenate with standard stabilizer codes. The resulting overhead reduction is smaller, but still significant at 4.5-fold. Using a ((4, 2)) VarQEC code at the initial level yields a similar picture (see inset). Overall, across considered noise structures, combining variational and standard codes reduces overhead by one to two orders of magnitude compared with stabilizer-only concatenations. In Fig. 5, both axes are logarithmic, while concatenation levels are shown on a linear scale. In all cases, we observe the theoretically predicted asymptotically super-exponential suppression of noise under concatenation [1], [2], with the discussed significant practical scaling differences.

V. D ISCUSSION AND O UTLOOK This work presents a pipeline that concatenates non-additive and additive codes, tailoring each level to the effective noise structure. To enable this, we introduced an efficient fidelityonly, two-design estimator that recovers the Pauli-diagonal part of the encode-noise-recover-decode channel. Based upon this, a learning-based procedure is employed to construct variational quantum error correction (VarQEC) codes [13] for the concatenation levels where noise is sufficiently structured, and we resorted to standard stabilizer codes in regimes of nearuniform noise. Across different initial noise channels, layerwise tailoring yields overhead reductions compared to concatenating just stabilizer codes by one to two orders of magnitude, demonstrating the theoretically predicted double-exponentially noise suppression. We restricted our analysis to Pauli noise channels, but emphasize that extending the noise tomography procedure using Pauli twirling allows considering also non-unital channels like amplitude damping noise [13]. Similarly, we focused on codes with a single logical qubit per patch, but conceptually the proposed framework also extends to higher-rate codes [3]. Currently, the major missing piece is the (early) fault-tolerant encoding under code concatenation, which must be performed in the codespace of the lower concatenation levels. In future work, this is to be addressed by the extension of the VarQEC learning procedure with a co-design approach that simultaneously tailors encodings and ensures the existence of low-depth logical operations [34]. Furthermore, one promising future research direction is the modification of the loss function to enforce desired structural properties of the effective noise channel, as currently the advantage of the VarQEC codes is limited due to fast convergence towards isotropic noise under concatenation. In conclusion, concatenating level-wise noise-tailored codes allows for substantial overhead reduction in regimes of structured noise. Extending to higher-rate codes and guaranteeing the native support of fault-tolerant gadgets can evolve the concept to a practical alternative for early fault-tolerant quantum computing. DATA AVAILABILITY The error-correcting codes tailored to the noise structures under consideration and their concatenation levels, as well as the analysis protocol for estimating the noise structure, are available at https://github.com/nicomeyer96/learning-toconcatenate. Additional information and data are available upon reasonable request.

R EFERENCES [1] E. Knill, R. Laflamme, and W. H. Zurek, “Resilient Quantum Computation,” Science, vol. 279, no. 5349, pp. 342–345, 1998. [2] P. Aliferis, D. Gottesman, and J. Preskill, “Quantum accuracy threshold for concatenated distance-3 codes,” Quantum Inf. Comput., vol. 6, no. 2, pp. 97–165, 2005. [3] S. Yoshida, S. Tamiya, and H. Yamasaki, “Concatenate codes, save qubits,” npj Quantum Inf., vol. 11, no. 1, p. 88, 2025. [4] A. Erhard, J. J. Wallman, L. Postler, M. Meth, R. Stricker, E. A. Martinez, P. Schindler, T. Monz, J. Emerson, and R. Blatt, “Characterizing large-scale quantum computers via cycle benchmarking,” Nature Commun., vol. 10, no. 1, p. 5347, 2019. [5] P. Krantz, M. Kjaergaard, F. Yan, T. P. Orlando, S. Gustavsson, and W. D. Oliver, “A quantum engineer’s guide to superconducting qubits,” Appl. Phys. Rev., vol. 6, no. 2, p. 021318, 2019. [6] L. Huang, X. Wu, and T. Zhou, “Robustness of the concatenated quantum error-correction protocol against noise for channels affected by fluctuation,” Phys. Rev. A, vol. 100, p. 042321, 2019. [7] T. Fösel, P. Tighineanu, T. Weiss, and F. Marquardt, “Reinforcement Learning with Neural Networks for Quantum Feedback,” Phys. Rev. X, vol. 8, no. 3, p. 031084, 2018. [8] H. P. Nautrup, N. Delfosse, V. Dunjko, H. J. Briegel, and N. Friis, “Optimizing Quantum Error Correction Codes with Reinforcement Learning,” Quantum, vol. 3, p. 215, 2019. [9] P. D. Johnson, J. Romero, J. Olson, Y. Cao, and A. Aspuru-Guzik, “QVECTOR: an algorithm for device-tailored quantum error correction,” arXiv:1711.02249, 2017. [10] D. F. Locher, L. Cardarelli, and M. Müller, “Quantum Error Correction with Quantum Autoencoders,” Quantum, vol. 7, p. 942, 2023. [11] J. Ollé, R. Zen, M. Puviani, and F. Marquardt, “Simultaneous discovery of quantum error correction codes and encoders with a noise-aware reinforcement learning agent,” npj Quantum Inf., vol. 10, no. 1, pp. 1–17, 2024. [12] J. Ollé, O. M. Yevtushenko, and F. Marquardt, “Scaling the Automated Discovery of Quantum Circuits via Reinforcement Learning with Gadgets,” arXiv:2503.11638, 2025. [13] N. Meyer, C. Mutschler, A. Maier, and D. D. Scherer, “Learning Encodings by Maximizing State Distinguishability: Variational Quantum Error Correction,” arXiv:2506.11552, 2025. [14] C. Cao, C. Zhang, Z. Wu, M. Grassl, and B. Zeng, “Quantum variational learning for quantum error-correcting codes,” Quantum, vol. 6, p. 828, 2022. [15] S. Heußen, D. F. Locher, and M. Müller, “Measurement-Free FaultTolerant Quantum Error Correction in Near-Term Devices,” PRX Quantum, vol. 5, no. 1, p. 010333, 2024. [16] B. Schumacher and M. D. Westmoreland, “Approximate Quantum Error Correction,” Quantum Inf. Process., vol. 1, pp. 5–12, 2002. [17] N. Meyer, C. Mutschler, A. Maier, and D. D. Scherer, “Variational Quantum Error Correction,” in IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 2, 2025, pp. 456–457. [18] A. Ambainis and J. Emerson, “Quantum t-designs: t-wise Independence in the Quantum World,” in IEEE Conference on Computational Complexity (CCC’07), vol. 1, 2007, pp. 129–140. [19] J. Watrous, The Theory of Quantum Information. Cambridge University Press, 2018. [20] D. Greenbaum, “Introduction to Quantum Gate Set Tomography,” arXiv:1509.02921, 2015. [21] R. Laflamme, C. Miquel, J. P. Paz, and W. H. Zurek, “Perfect Quantum Error Correcting Code,” Phys. Rev. Lett., vol. 77, no. 1, p. 198, 1996. [22] A. Katabarwa, K. Gratsea, A. Caesura, and P. D. Johnson, “Early FaultTolerant Quantum Computing,” PRX Quantum, vol. 5, no. 2, p. 020101, 2024. [23] K. Kraus, “General state changes in quantum theory,” Ann. Phys., vol. 64, no. 2, pp. 311–335, 1971. [24] P. W. Shor, “Scheme for reducing decoherence in quantum computer memory,” Phys. Rev. A, vol. 52, no. 4, p. R2493, 1995. [25] A. M. Steane, “Error Correcting Codes in Quantum Theory,” Phys. Rev. Lett., vol. 77, no. 5, p. 793, 1996. [26] D. Gottesman, “Stabilizer Codes and Quantum Error Correction,” Ph.D. dissertation, California Institute of Technology, 1997. [27] M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio et al., “Variational

quantum algorithms,” Nature Rev. Phys., vol. 3, no. 9, pp. 625–644, 2021. [28] T. Hastie, The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer, 2009. [29] C. Dankert, R. Cleve, J. Emerson, and E. Livine, “Exact and approximate unitary 2-designs and their application to fidelity estimation,” Phys. Rev. A, vol. 80, no. 1, p. 012304, 2009. [30] D. Gottesman, “Surviving as a Quantum Computer in a Classical World,” 2024, lecture notes. [Online]. Available: https://www.cs.umd. edu/class/spring2024/cmsc858G/QECCbook-2024-ch1-15.pdf [31] E. Knill and R. Laflamme, “Concatenated Quantum Codes,” arXiv:quant-ph/9608012, 1996. [32] M. R. Geller and Z. Zhou, “Efficient error models for fault-tolerant architectures and the pauli twirling approximation,” Phys. Rev. A, vol. 88, p. 012314, 2013. [33] J. Duchi, S. Shalev-Shwartz, Y. Singer, and T. Chandra, “Efficient projections onto the l 1-ball for learning in high dimensions,” in International Conference on Machine Learning (ICML), 2008, pp. 272–279. [34] N. Meyer, C. Mutschler, D. Seuß, A. Maier, and D. D. Scherer, “Learning Logical Operations for Arbitrary Quantum Error Correction Codes,” 2026, manuscript in preparation. [35] V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, S. Ahmed, V. Ajith, M. S. Alam, G. Alonso-Linaje, B. AkashNarayanan, A. Asadi et al., “PennyLane: Automatic differentiation of hybrid quantum-classical computations,” arXiv:1811.04968, 2018. [36] N. Meyer, C. Ufrecht, M. Periyasamy, A. Plinge, C. Mutschler, D. D. Scherer, and A. Maier, “Qiskit-Torch-Module: Fast Prototyping of Quantum Neural Networks,” in IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1, 2024, pp. 817–823. [37] D. C. Liu and J. Nocedal, “On the limited memory BFGS method for large scale optimization,” Math. Program., vol. 45, no. 1, pp. 503–528, 1989. [38] N. Meyer, J. Murauer, A. Popov, C. Ufrecht, A. Plinge, C. Mutschler, and D. D. Scherer, “Warm-Start Variational Quantum Policy Iteration,” in IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1, 2024, pp. 1458–1466.

Record · ID 19031 · SHA-256 df606e236c19be36
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.