Transversal Fanout for Fault Tolerant Distributed Quantum Computing: Analysis and Application Seng W. Loke
arXiv:2609.08233v1 [quant-ph] 8 Sep 2026
School of Information Technology, Deakin University, Burwood, VIC 3125, Australia. [email protected]
Abstract—We study a resource-efficient approach for implementing logical fanout operations in fault-tolerant distributed quantum computing using transversal operations on quantum error-correcting code blocks. Logical fanout, comprising multiple controlled-NOT operations from a common control qubit to target qubits located at remote nodes, is an important primitive for distributed quantum computation but can require substantial non-local communication when implemented directly between encoded blocks. We exploit the structure of encoded blocks and the availability of transversal logical operations to construct distributed fanout circuits that reduce the required non-local operations while preserving the logical action of the fanout operation. The construction is developed for encoded quantum information and illustrated using Bivariate-Bicycle (BB)-code blocks. We analyze the resulting physical gate, entanglement, circuit-depth, and ancilla requirements. The approach provides a systematic method for implementing large logical fanout operations across distributed error-corrected quantum processors. Also, we study a distributed implementation of the global gate GCZ involving logical qubits (encoded using BB-code blocks), exploiting the concurrency in transversal distributed fanouts. Index Terms—distributed quantum computing, tansversal gate gadgets, fault-tolerant quantum computing, multipartite entanglement, distributed fanout gates
I. I NTRODUCTION Distributed or modular quantum computing has been touted as a means of scaling up quantum computation, beyond monolithic platform approaches [2], [3], [5], [11].1 Much work has focused on the use of Bell pairs for realizing distributed or non-local (or inter-module) CNOT gates to connect disparate QPUs (Quantum Processing Units), as such gates together with single qubit gates can provide a universal gate set for quantum computations. There have indeed been many experimental realizations of such connectivity, see e.g., [15]. Towards fault-tolerant distributed quantum computing, a range of encodings for quantum error correction have been proposed and are being implemented, and even considered for distributed quantum computing including surface codes (e.g., [4], [16], [18]), Bivariate-Bicycle (BB) codes (e.g., transversal non-local CNOTs [21] and inter-module operations [26]), and floquet codes [22]. Previous work advocated the use of fan-out operations, with one control qubit for multiple target qubits, amenable to 1 For example, see [10], https://spectrum.ieee.org/quantum-computers, https://newsroom.ibm.com/2025-11-20-ibm-and-cisco-announce-plans-to-b uild-a-network-of-large-scale,-fault-tolerant-quantum-computers. See also https://quantnet.lbl.gov and https://www.ox.ac.uk/news/2025-02-06-first-distr ibuted-quantum-algorithm-brings-quantum-supercomputers-closer.
Dc-U c
≡ t
U
Fig. 1: Distributed controlled-U (i.e., dCNOT is when U=X (⊕) as shown) between nodes QC and QC ′ (the computation qubits are c, the control qubit, and t, the target qubit), and the wavy line illustrates a Bell pair involving the communication qubits from the two nodes both initially |0⟩. The right hand side shows the notation we will use for a distributed controlledU operation.
efficient implementation in some types of quantum hardware such as trapped-ion quantum computers, enabling reduced depth quantum computations [6], [9]. Distributed fan-out using distributed GHZ states have been proposed in [25], and investigated in [12], [13]. This paper discusses a transversal implementation of a distributed fanout with qubits in a Bivariate-Bicycle (BB) encoding, and also explore their use for realizing distributed versions of global quantum gates (or GM S(θ), short for Global Mølmer–Sørensen multi-qubit gates [20]). In the rest of this paper, we first briefly review distributed fan-out operations in §2, and then look at distributed or nonlocal transversal fanout operations in §3. In §4, we provide a simulation analysis of distributed fanouts, and then explore applications of the transversal fanout for distributed versions of GCZ global gates in §5. We conclude with avenues for future work in §6. II. B RIEF R EVIEW OF D ISTRIBUTED FAN -O UT O PERATIONS We first consider the distributed CNOT gate (or dCNOT, in short, a.k.a. non-local CNOT), which is needed as a resource for each dCNOT gate, as shown in Figure 1. By a distributed fanout operation, we refer to an operation where the same control qubit is used for multiple target qubits, on different nodes. This can arise from the structure of a quantum circuit itself, or from a control-U operation where the unitary U spans multiple qubits and is decomposed
into operations executed on qubits distributed on different nodes [14], [24], [25]. For example, the quantum circuit such as: fan-out c t1 t2 t3
c
and each of the m blocks, Ti , encodes k logical qubits (l) {t̄i | l ∈ {1, ..., k}}. Suppose we wanted to perform the following operation: F AN OU T (C; T1 , ..., Tm ) O (l) F AN OU T L (c̄(l) ; t̄1 , ..., t̄(l) := m) l∈{1,...,k}
U1
t1
U1
t2
U2
t3
U3
≡
U2 U3
where F AN OU T L denotes the logical fanout acting on the corresponding l-th encoded logical qubit, 1 ≤ l ≤ k. We do this by performing n independent physical fanout operations, one for each coordinate of the code: O F AN OU T (ci ; t1i , ..., tmi ) i∈{1,...,n}
i.e., a multitarget control-U , where U1 acts on qubit t1, U2 acts on qubit t2, and U3 acts on qubit t3, each qubit on a different node, will result in the distributed operation as shown in Figure 2; note that the wavy line in the figure connecting the four (communication) qubits a0 , a1 , a2 , a3 represents a distributed 4-qubit GHZ state of the form √12 (|0000⟩ + |1111⟩)a0 a1 a2 a3 , over four nodes A′ , A, B and C. In the rest of the paper, we particularly focus on the case where Ui = X, unless otherwise noted. III. T RANSVERSAL N ON -L OCAL FANOUT We outline why a non-local fanout operation can be implemented transversally. A transversal operation can be defined as follows. Let there be m code blocks, each consisting of n physical qubits. A quantum Nn operation U is transversal if it can be expressed as U = i=1 Ui , where each Ui acts only on the i-th physical qubits of the code blocks and does not couple two qubits within the same code block. From [21], we see that a logical CNOT between two logical qubits is transversal, when the logical qubits are encoded using the Bivariate-Bicycle (BB) qLDPC code; multiple logical CNOT operations are performed simultaneously. Because the logical fanout decomposes into CNOTs between the same code blocks, and because each of those logical CNOTs is transversal, their product can be reordered into a set of independent physical fanout operations, one per physical coordinate; a logical fanout gadget is transversal if it can be composed using transversal logical CNOT gadgets. This comes from [8], where “a product of transversal gate gadgets is a transversal gate gadget”, and since, a logical fanout is a product of (transversal) CNOT operations, by definition. Using the BB encoding [[n, k, d]] of logical qubits, with operations on a block of n qubits, we effectively perform k independent logical fanout operations simultaneously. More specifically, using the [[n, k, d]] BB encoding, where the logical CNOT is transversal, i.e., between two blocks C with qubits ci and T with qubits ti , denoted by CN OT (C; T ) = Πi∈{1,...,n} CN OT (ci ; ti ), consider one control block C = {c1 , ..., cn } with n qubits, and m target blocks Ti , i ∈ {1, ..., m}, each with qubits Ti = {ti1 , ..., tin }, where block C encodes k logical qubits {c̄(l) | l ∈ {1, ..., k}}
where F AN OU T denotes a physical fanout operation, each physical qubit ci fans out as control to the corresponding physical qubits in every target block.2 Note that the above applies to non-local CNOTs and non-local FANOUTs, e.g., with one block per node (or k logical qubits per node). A. Comparison with Pairwise Distributed CNOTs We can compare resources required for a GHZ-based fanout as shown in Figure 2 with an equivalent (in terms of effect) sequence of CNOTs; suppose one control (block on one node) simultaneously controls m − 1 remote targets (i.e., m nodes in total, one target block per target node, for simplicity), with comparison shown in Table I. The GHZ based approach could lead to a reduction in sources of errors, and reduced depth (which could also reduce decoherence due to reduced qubit wait times) and so, in lower error rates. However, it relies on 2 We explain why this is so, as follows:
Πi∈{1,...,n} F AN OU T (ci ; t1i , ..., tmi ) =Πi∈{1,...,n} Πj∈{1,...,m} CN OT (ci ; tji ) (from definition of F AN OU T ) =Πj∈{1,...,m} Πi∈{1,...,n} CN OT (ci ; tji ) (the required CN OT s commute) =Πj∈{1,...,m} CN OT (C; Tj ) (from transversal CN OT s between the control block C and the target block Tj ) =F AN OU T (C; T1 , ..., Tm ) From another perspective, since a transversal CNOT between code blocks acts on k target logical qubits simultaneously, a transversal fanout among code blocks acts on k groups of target logical qubits simultaneously: F AN OU T (C; T1 , ..., Tm ) =Πj∈{1,...,m} CN OT (C; Tj ) (l)
=Πj∈{1,...,m} Πl∈{1,...,k} CN OT L (c̄(l) ; t̄j ) (l)
=Πl∈{1,...,k} Πj∈{1,...,m} CN OT L (c̄(l) ; t̄j ) (the required CN OT L s commute) (l)
(l)
=Πl∈{1,...,k} F AN OU T L (c̄(l) ; t̄1 , ..., t̄m ) (from definition of F AN OU T L ) O (l) (l) = F AN OU T L (c̄(l) ; t̄1 , ..., t̄m ) l∈{1,...,k}
Z m1 ⊕m2 ⊕m3
c
c
A’ a0
m0
a1
X m0
Dfan-out H
c
m1
A U1
t1
t1
U1
t2
U2
t3
U3
≡ a2
X
m0
H
m2
B U2
t2 a3
X m0
H
m3
C U3
t3
Fig. 2: Distributed fanout operation with single control qubit (on A′ ) for multiple target qubits (one on A, one on B and one on C) - all target qubits on different nodes from the control qubit. The right hand side shows the notation we will use for distributed fan-out throughout the paper. Note that mi ∈ {0, 1}. being able to generate distributed GHZ states with low error rates. We show later via simulation that this can work based on particular noise models. With a distributed transversal fanout operation among blocks, using BB block encoding [[n, k, d]], with n physical qubits, say for one block per node, over m nodes, we would have n non-local physical F AN OU T s, i.e., needing n m-qubit physical GHZ states. For distributed GHZ states, a block of n communication qubits are required on each node; this is similar for the distributed CNOTs approach with the communications qubits initialised and reused for each CNOT on the control node. Table I summarises the comparison. The table treats an m-party GHZ as an available entanglement resource; the cost of generating that GHZ is architecture-dependent and is not included; similarly, with the Bell pairs. TABLE I: Physical resources (involving blocks of n qubits) for distributed GHZ-based F AN OU T and (equivalent) distributed CN OT s using [[n, k, d]]BB, 1 control block and m − 1 target blocks, over m nodes (one block per node). Resource Shared entanglement Local CNOTs Hadamards Measurements Classical bits sent Communication qubits
Distributed CN OT s n(m − 1) Bell pairs 2n(m − 1) n(m − 1) 2n(m − 1) 2n(m − 1) nm
GHZ F AN OU T n m-party GHZ states nm n(m − 1) nm 2n(m − 1) nm
IV. S IMULATION S TUDY OF N ON -L OCAL FANOUTS A fanout implements multiple CNOT interactions and may therefore expose the circuit to more error locations than a single CNOT; however, the GHZ-based implementation can
reduce the number of non-local operations relative to a sequential realization. Given a physical error rate (PER) p on qubits, and inter-node ebit noise pebit , and noise in GHZ states pghz , one can compare the LERs of a fanout operation using GHZ states and an equivalent sequence of CNOTs. Also, a fanout operation uses distributed GHZ states, each of which may be constructed using a linear number of distributed CNOTs, or one shot via other mechanisms (e.g., [1], [19]) - both of which might have their own error sources and possibly yielding different LERs. Transversality establishes that the logical fanout can be implemented without coupling two physical qubits within the same code block. The circuitlevel fault tolerance of the complete implementation, however, also depends on the syndrome-extraction circuit and decoder. Consequently, the simulations below evaluate the LER of the complete noisy circuit rather than assuming that the code distance is preserved automatically. Here, we do comparisons, first with a four node fanout operation, with one [[36, 4, 6]]BB block per node, and then with eight nodes. Our simulation study is based on extending the Transversal Multiple Code Block Simulator (TMCBS) library code from [21] (created to study non-local transversal CNOTs) to simulate non-local fault-tolerant transversal fanout operations. TMCBS uses Stim [7], and the syndrome extraction schedule used in the fanout simulation is given in Table II. We utlised ChatGPT-5.6 Luna as a tool to generate Python implementations of the different versions of the fanout operations using the TMCBS library (and the Stim3 library), in an interactive prompt and revise process, with some manual fine-tuning. We use the decoder Tesseract, like in [21]. 3 https://github.com/quantumlib/Stim
TABLE II: Syndrome-extraction schedule (using BB[[36, 4, 6]]) used for the different distributed non-local gate implementations. Each initial extraction consists of one first-pass round followed by three repeated rounds. For 3cnot-seq, an additional extraction round is performed between successive non-local CNOT operations; this isolates successive logical operations and provide the same fault-tolerance boundary used by the TMCBS non-local-CNOT protocol. No such intermediate extraction is required in the GHZ-fanout circuit because the fanout is treated as one logical operation. DEPOLARIZE2 in Stim is a two-qubit depolarizing noise channel, with identity with probability Q 1−pebit , and pebit , otherwise, i.e. P r(II) = 1−pebit , P r(P ) = pebit 15 , for P ∈ {IX, IY, IZ, XI, . . . , ZZ}, and EGHZ (pghz ) = q D1 (q, pghz /2), where the implementation applies independent D1 (DEPOLARIZE1) channels to the GHZ qubits (where DEPOLARIZE1 is Stim’s single-qubit depolarizing channel) giving identity with probability 1 − pghz /2 and (pghz /2)/3 for X, Y and Z error, in this case. Circuit 1cnot 3cnot-seq 3cnotghz-fanout 1shotghz-fanout
Before non-local operation
Between operations
After operation
Non-local noise model
4 rounds 4 rounds 4 rounds 4 rounds
— 1 round after each non-local CNOT — —
3 rounds 3 rounds 3 rounds 3 rounds
DEPOLARIZE2 per dCNOT DEPOLARIZE2 per dCNOT DEPOLARIZE2 per dCNOT EGHZ
Figure 3 shows the graph of LERs vs PERs for three simulated realizations of a distributed (or non-local) fanout, one using a one shot approach to create distributed GHZ states (1shotghz-fanout), one using GHZ states constructed using non-local CNOTs (3cnotghz-fanout), and one is a sequence of (non-local) CNOTS, equivalent to the fanout (3cnot-seq). We construct (blocks of) GHZ states for 1shotghz-fanout and 3cnotghz-fanout via a transversal protocol: among corresponding qubits in the communication qubit blocks on each node, perform the Hadamard H on one “root” (control) qubit, and then perform non-local CNOT between the “root” and each of the other qubits. For the 1shotghz-fanout, we simulate this by constructing the GHZ state via the same protocol but counting the entire process as taking one simulation step. The respective noise models are shown in Table II. For each data point we also show error bars corresponding to 95% confidence intervals. In the first simulation, we used p = pebit = pghz , representing physical error rate p for a single qubit, the error rate for an ebit and a parameter representing the physical error rate in a GHZ state (note that pghz is not the probability of error in the GHZ state but a parameter to add errors to the state in the simulation as used in Table II), and repeated runs, and then calculating the LER, i.e. the probability that at least one logical error occurs anywhere in the complete operation. Indeed, fanout using the 1-shot GHZ states approach generally provides the lowest LER in the low-PER regime, although the ordering of the two GHZ-based implementations can cross at intermediate error rates. The table shows the cross-over when LER < P ER, when the PER p ≈ 5 × 10−4 and below. A corresponding set of results is shown in Figure 4 when the physical error rate in the entanglement (distributed Bell pairs and the GHZ states) is 10 times that of the PER on qubits, i.e. 10p = pebit = pghz ; this is similar to the case of p = pebit = pghz , but with slightly larger differences between the fanout using 1-shot GHZ versus the other two approaches to fanout - higher entanglement costs makes the one shot GHZ approach relatively more efficient, as expected. We also performed similar comparisons with distributed fanout operations over eight nodes, shown in Figure 5 with the values shown in the tables. As expected, LER is higher
p 1 × 10−2 5 × 10−3 1 × 10−3 5 × 10−4 1 × 10−4
1shotghz-fanout 1.00 × 100 4.88 × 10−1 1.71 × 10−3 1.98 × 10−4 8.60 × 10−7
3cnotghz-fanout 9.92 × 10−1 4.77 × 10−1 2.07 × 10−3 1.54 × 10−4 1.19 × 10−6
3cnot-seq 1.00 × 100 5.90 × 10−1 3.50 × 10−3 1.98 × 10−4 2.37 × 10−6
Fig. 3: LER vs PER for a four node distributed fanout with p = pebit = pghz , comparing three different distributed fanout implementations - one using a one shot approach to create distributed GHZ states (1shotghz-fanout), one using GHZ states constructed using a sequence of neighbouring non-local CNOTs (3cnotghz-fanout), and one is a sequence of CNOTS (using only Bell pairs and not using GHZ states) equivalent to the fanout (3cnot-seq), as well as a single distributed CNOT (1cnot) for comparison. The table shows the plotted values.
for 8 nodes, compared to 4 nodes, and the 1-shot GHZ fanout has lower LER. The improvement factor using the 1-shot GHZ fanout compared to using just Bell pairs (a sequence of CNOTs) for the 8 block (or 8 nodes) distributed fanout is shown in Figure 6 - up to around a factor of 2.3 can be achieved in this case with low p. Overall, even with the [[36, 4, 6]], for all cases, a LER in the order of ≈ 10−6 , below
p 1 × 10−2 5 × 10−3 1 × 10−3 5 × 10−4 1 × 10−4
1shotghz-fanout 1.00 × 100 5.98 × 10−1 3.18 × 10−3 3.72 × 10−4 3.00 × 10−6
3cnotghz-fanout 1.00 × 100 7.42 × 10−1 4.23 × 10−3 4.85 × 10−4 3.19 × 10−6
3cnot-seq 1.00 × 100 7.23 × 10−1 4.58 × 10−3 5.39 × 10−4 4.69 × 10−6
Fig. 4: LER vs PER for a four node distributed fanout with 10p = pebit = pghz , comparing three different distributed fanout implementations - one using a one shot approach to create distributed GHZ states (1shotghz-fanout), one using GHZ states constructed using a sequence of neighbouring nonlocal CNOTs (3cnotghz-fanout), and one is a sequence of CNOTS (using only Bell pairs and not using GHZ states) equivalent to the fanout (3cnot-seq), as well as a single distributed CNOT (1cnot) for comparison. The table shows the plotted values. PER, is achieved for p = 10−4 . We studied the circuit-level distance of the non-local transversal fanout circuits using both a sequence of nonlocal CNOTs and using a 1-shot GHZ state fanout approach via the STIM undetectable-error search heuristic (i.e., search_for_undetectable_logical_errors), like in [21]. Code distance refers to the minimum weight of a Pauli operator implementing a logical operation; given a circuit circ, dcirc refers to the minimum number of errors in the logical circuit (across time and space), causing an unwanted logical operation on an observable (i.e., a logical qubit). Table III shows the results; the search found no undetectable logical error of fault weight below 6, i.e., dcirc = 6 for both 4 and 8 node fanouts. V. D ISTRIBUTED GCZ - AN A PPLICATION OF T RANSVERSAL N ON -L OCAL FANOUTS We have seen that each transversal fanout operation performs k logical fanout operations concurrently since the n qubits of each block encodes k logical qubits. This can be advantageous for operations involving multiple fanouts, such as the so-called global gates - this section illustrates this.
p 1shot-fanout CNOT-seq fanout 1 × 10−2 1.00 × 100 1.00 × 100 −1 −3 8.09 × 10 9.77 × 10−1 5 × 10 −3 −3 5.66 × 10 9.16 × 10−3 1 × 10 −4 −4 5 × 10 5.75 × 10 1.08 × 10−3 −6 −4 3.53 × 10 8.12 × 10−6 1 × 10 p 1shot-GHZ-fanout CNOT-seq fanout 1 × 10−2 1.00 × 100 1.00 × 100 −3 −1 5 × 10 9.18 × 10 9.60 × 10−1 −2 −3 1.06 × 10 2.34 × 10−2 1 × 10 −4 −3 5 × 10 1.00 × 10 1.74 × 10−3 −4 −6 1 × 10 8.41 × 10 1.44 × 10−5 Fig. 5: LER vs PER for an eight node (one data block per node) distributed fanout with p = pebit = pghz (top table) and 10p = pebit = pghz (bottom table), comparing two different distributed fanout implementations - one using a one shot approach to create distributed GHZ states (1shotghz) for the fanout, and one is a sequence of CNOTS (using only Bell pairs) equivalent to the fanout (CNOT-seq). TABLE III: Circuit-level distances for BB-code distributed fanout circuits using a sequence of (non-local) CNOTs and using 1-shot distributed GHZ state fanout, using [[36, 4, 6]]BB encoding for the logical qubits, one block per node - the results are the same for both 4 and 8 nodes; included is dcirc for the non-local CNOT, for comparison, from [21]. BB Code [[36, 4, 6]] [[36, 4, 6]] [[36, 4, 6]]
d 6 6 6
Circuit Non-local CNOT CNOT sequence fanout 1-shot GHZ fanout
dcirc 6 6 6
Global gates are efficient for some quantum architectures such as trapped ion qubits where entire arrays of qubits can be targeted and pairwise qubit-qubit interactions can be naturally realized, which can be efficient for certain applications [17], [23]. A global MS gate or operation (GMS) over a set of qubits S is defined as follows: GM SS (θ) = Y θ θ X Xi Xj ) = exp(−i Xi Xj ) exp(−i 2 2 i,j∈S, i<j
i,j∈S, i<j
p 1 × 10−2 5 × 10−3 1 × 10−3 5 × 10−4 1 × 10−4
pebit = pghz = p 1.00 1.21 1.62 1.88 2.30
pebit = pghz = 10p 1.00 1.05 2.21 1.74 1.71
Fig. 6: LER improvement for 10p = pebit = pghz and p = pebit = pghz . LER improvement factor of the 1-shot GHZ fanout relative to the sequential CNOT fanout for 8 blocks. The improvement factor is defined as LERCNOT-seq /LER1shotghz .
where exp(−i θ2 Xi Xj ) is called the “local” MS gate acting on qubits i and j, and can be viewed operationally as: exp(−i θ2 Xi Xj ) = (Hi ⊗ Hj )CN (Ii # ⊗ " OTi→j −i θ2 0 e (RZ (θ))j )CN OTi→j (Hi ⊗Hj ), with RZ (θ) = . θ 0 ei 2 (Note that the local MS gate (or LMS gate, for short) is synmmetrical, i.e. exp(−i θ2 Xi Xj )=exp(−i θ2 Xj Xi ).) GCZ gates are equivalent (up to Clifford gates) to GM S(π/2) gates [23]. An illustration of a GCZ16 operation involving 16 (logical) qubits over four nodes, four (logical) qubits per node, is shown in Figure 7, where there are 120 CZ operations, of which 16(16 − 1)/2 − 4 ∗ 4(4 − 1)/2 = 96 CZs are inter-node CZs. Figure 8 shows the first 54 CZs between the (logical) qubits over the four nodes. Figure 9 shows how the first 54 CZs between the logical qubits shown in Figure 8 can be implemented in a distributed fashion using ancilla blocks on each of the nodes containing the target blocks. We note that the local CZs (on the node containing the control block) can commute with the preceding non-local (inter-node) CZs (left hand side), and so, can be moved and batched prior to all the inter-node CZs (right hand side); • the control logical qubits from the control block q1 , q2 , q3 , q4 are fanned-out to the ancilla qubits on the target nodes - A1 , A2 and A3 are the ancilla logical qubits (or qubit blocks) on target nodes T1 , T2 and T3 , respectively - all initialised to 0 L ; • each ancilla qubit is fanned-out to the internal target •
qubits, within each node. With reordering and uncomputation added to Figure 9, we have the circuit in Figure 10. Let C = {q1 , q2 , q3 , q4 }, and Ti = {q4i+1 , q4i+2 , q4i+3 , q4i+4 }, for i = 1, 2, 3, with ancilla register Ai = {aTi ,1 , aTi ,2 , aTi ,3 , aTi ,4 } on the same node as Ti . Stage 1: intra-C CZs; this isQeffectively a local GCZ operation on node C: GCZC = 1≤j<k≤4 CZqj ,qk . Stage 2: remote CZ block for each node Ti : note the concurrent distributed fanouts among the logical qubits on different nodes, the four red boxes in Figure 10, which can be realized using the transversal fanout between the control block N and the ancilla blocks, i.e. F AN OU T (C; A1 , A2 , A3 ) = l∈{1,...,4} F AN OU T L (ql ; aT1 ,l , aT2 ,l , aT3 ,l ). We denote by dF AN OU T the distributed transversal fanout among blocks on different nodes, i.e. dF AN OU T (C; A1 , A2 , A3 ) = N l∈{1,...,4} dF AN OU T L (ql ; aT1 ,l , aT2 ,l , aT3 ,l ) Hence, in terms of fanouts, a distributed version of the 54 CZs, denoted by dU54 , can be written as a product of local and distributed operations (rightmost operations execute first): dU54 = dF AN OU T (C; A1 , A2 , A3 ) Z
· Π3i=1 Π4j=1 F AN OU T L (aTi ,j ; q4i+1 , q4i+2 , q4i+3 , q4i+4 ) · dF AN OU T (C; A1 , A2 , A3 ) · GCZC Also, note the set of local fanout operations (not CNOTs, but Z CZs), denoted by F AN OU T L , on each node commute (and so, can be done concurrently even if depicted in sequence in † Figure 10). Note that dF AN OU T = dF AN OU T . Coming back to Figure 7, each block of four “triangles” can be realized in a similar way, the second block of four using a distributed transversal fanout between a control block in node T1 and nodes T2 and T3 , and the last four triangles involve only the nodes T1 and T2 , and so needs only (distributed) transversal CNOTs. In general, with a BB encoding [[n, k, d]], for a N (logical) qubit GCZN with k logical qubits per node (one block per node), over m nodes, i.e. N = km, we require 2(m−1) distributed transversal F AN OU T operations (including transversal CNOTs between pairs of nodes), a factor of two due to uncomputation, and local GCZ and local fanout operations. Also, n physical m′ -qubit GHZ states (for m′ ≤ m) are required for each of the distributed transversal fanouts among distributed code blocks involving m′ blocks (i.e., for the concurrent k logical fanouts among logical qubits over m′ nodes). Note that a distributed fanout with one control block and one target block is effectively a non-local CNOT among blocks and instead of GHZ states, we use Bell pairs. Hence, for a GCZN =km , one [[n, k, d]]BB block per node, O(nm) physical GHZ states (over a varying number of nodes) are required, or for nk = c for some constant c, O(ckm) = O(km) = O(N ) physical GHZ states are required. VI. C ONCLUSION AND F UTURE W ORK We demonstrated a distributed transversal fanout for BBencoded logical qubits. We note that its advantages rely on the ability to form distributed GHZ states efficiently and with low enough noise - our simulation demonstrated the advantage of
Node 1: q1 q2 q3 q4 Node 2: q5 q6 q7 q8 Node 3: q9 q10 q11 q12 Node 4: q13 q14 q15 q16
Fig. 7: Global gate operation (GCZ16 ) over 4 nodes, with 120 CZs; GCZ16 =
Node 1: q1 q2 q3 q4 Node 2: q5 q6 q7 q8 Node 3: q9 q10 q11 q12 Node 4: q13 q14 q15 q16
Fig. 8: The first 54 CZs, from Figure 7. one shot fanout, but other noise models for the GHZ states can be investigated. Also, we showed how the concurrent logical fanouts from transversal operations can be used to construct distributed implementations of high-fanout CZ/CNOT circuits, including distributed GMS/GCZ constructions. Future work could further analyse much larger circuits. ACKNOWLEDGEMENTS The author acknowledges the use of ChatGPT-5.6 Luna to assist in (i) generating Python code for the simulation study used in Section IV, (ii) generating the diagram in Figure 7 (the prompt is to generalize from an example of a manually drawn GCZ with a smaller number of CZ gates), and (iii) generating the Python code for plotting the graphs in Figures 3, 4, 5, and 6. The author verified all algorithmic logic and executed the code, and take full responsibility for the results. R EFERENCES [1] E. M. Ainley, A. Agrawal, D. Main, P. Drmota, D. P. Nadlinger, B. C. Nichol, R. Srinivas, and G. Araneda. Multipartite Entanglement for Multi-node Quantum Networks. Jul 2024. arXiv:2408.00149 [quantph]. [2] David Barral, F. Javier Cardama, Guillermo Dı́az-Camacho, Daniel Faı́lde, Iago F. Llovo, Mariamo Mussa-Juane, Jorge Vázquez-Pérez, Juan Villasuso, César Piñeiro, Natalia Costas, Juan C. Pichel, Tomás F. Pena, and Andrés Gómez. Review of distributed quantum computing: From single qpu to high performance quantum computing. Computer Science Review, 57:100747, 2025. [3] Marcello Caleffi, Michele Amoretti, Davide Ferrari, Jessica Illiano, Antonio Manzalini, and Angela Sara Cacciapuoti. Distributed quantum computing: A survey. Computer Networks, 254:110672, 2024. [4] Nitish Kumar Chandra, Reza Nejabati, and Eneet Kaur. Towards the Characterization of Logical Errors in Distributed Lattice Surgery. In 2026 IEEE International Conference on Quantum Computing and Engineering, 7 2026. [5] Ang et al. Arquin: Architectures for multinode superconducting quantum computers. ACM Transactions on Quantum Computing, 5(3), September 2024. [6] Stephen A. Fenner and Rabins Wosti. Implementing the quantum fanout operation with simple pairwise interactions. Quantum Inf. Comput., 23(13&14):1081–1090, 2023.
Q
1≤m<n≤16 CZqm ,qn
[7] Craig Gidney. Stim: a fast stabilizer circuit simulator. Quantum, 5:497, July 2021. [8] D. Gottesman. Surviving as a Quantum Computer in a Classical World. 2026. https://www.cs.umd.edu/ dgottesm/QECCbook-2026.pdf. [9] Peter Høyer and Robert Spalek. Quantum fan-out is powerful. Theory Comput., 1(1):81–103, 2005. [10] Photonic Inc. Distributed Quantum Computing in Silicon. arXiv e-prints, page arXiv:2406.01704, June 2024. [11] Johannes Knörzer, Xiaoyu Liu, Benjamin F. Schaffer, and Jordi Tura. Distributed Quantum Information Processing: A Review of Recent Progress. 10 2025. arXiv:2510.15630 [quant-ph]. [12] Seng W. Loke. Distributed quantum computing with fan-out operations and qudits: the case of distributed global gates. In 2026 International Conference on Quantum Communications, Networking, and Computing (QCNC), pages 134–141, 2026. [13] Seng W. Loke. On distributed quantum computing with distributed fanout operations. In 2026 IEEE International Conference on Quantum Software (QSW), pages 1–6, 2026. [14] S.W. Loke. From Distributed Quantum Computing to Quantum Internet Computing: An Introduction. Wiley, 2023. [15] D. Main, P. Drmota, D. P. Nadlinger, E. M. Ainley, A. Agrawal, B. C. Nichol, R. Srinivas, G. Araneda, and D. M. Lucas. Distributed quantum computing across an optical network link. Nature, 638(8050):383–388, 2025. [16] Frederik K. Marqversen, Gefen Baranes, Maxim Sirotin, and Johannes Borregaard. Fault-tolerant interfaces for modular quantum computing on diverse qubit platforms. Phys. Rev. Res., 8:023097, Apr 2026. [17] Dmitri Maslov and Yunseong Nam. Use of global interactions in efficient quantum circuit constructions. New Journal of Physics, 20(3):033018, Mar 2018. [18] Siddhant Singh, Fenglei Gu, Sébastian de Bone, Eduardo Villaseñor, David Elkouss, and Johannes Borregaard. Modular architectures and entanglement schemes for error-corrected distributed quantum computation. npj Quantum Information, 11(1):146, 2025. [19] Siddhant Singh, Rikiya Kashiwagi, Kazufumi Tanji, Wojciech Roga, Daniel Bhatti, Masahiro Takeoka, and David Elkouss. Fault-tolerant modular quantum computing with surface codes using single-shot emission-based hardware. Phys. Rev. Appl., 26:024005, Aug 2026. [20] Anders Sørensen and Klaus Mølmer. Quantum computation with ions in thermal motion. Phys. Rev. Lett., 82:1971–1974, Mar 1999. [21] John Stack, Ming Wang, and Frank Mueller. Transversal fault tolerant distributed quantum computing operations. Nature Communications, 2026. Published online: 20 July 2026. [22] Evan Sutcliffe, Bhargavi Jonnadula, Claire Le Gall, Alexandra E. Moylett, and Coral M. Westoby. Distributed quantum error correction based on hyperbolic floquet codes. In 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), pages 649–657. IEEE, 2025. [23] John van de Wetering. Constructing quantum circuits with global gates. New Journal of Physics, 23(4):043015, Apr 2021. [24] Shufeng Xu, Ya-Li Mao, Lixin Feng, Hu Chen, Bixiang Guo, Shiting Liu, Zheng-Da Li, and Jingyun Fan. Multiqubit quantum logical gates between distant quantum modules in a network. Phys. Rev. A, 107:L060601, Jun 2023. [25] Anocha Yimsiriwattana and Samuel J. Lomonaco. Generalized GHZ States and Distributed Quantum Computing. Mar 2004. arXiv:quantph/0402148 [quant-ph]. [26] Theodore J. Yoder, Eddie Schoute, Patrick Rall, Emily Pritchett, Jay M. Gambetta, Andrew W. Cross, Malcolm Carroll, and Michael E. Beverland. Tour de gross: A modular quantum computer based on bivariate bicycle codes, 2025.
Fig. 9: Partial distributed implementation of the 54 CZs from Figure 8 using ancilla qubits, showing (i) reordering with the local CZs “batching” and preceding the rest of the inter-node operations; (ii) the four inter-node fanouts (marked in N the four red boxes) can be implemented via the (distributed version of) transversal F AN OU T (C; A1 , A2 , A3 ) = l∈{1,...,4} F AN OU T L (ql ; aT1 ,l , aT2 ,l , aT3 ,l ). Uncomputation of the ancilla blocks is needed and shown in Figure 10.
Fig. 10: Starting with clean ancillas set to |0⟩L , the complete distributed implementation of the 54 logical CZs from Figure N 8; the four inter-node fanouts (marked in the four red boxes) can be implemented via dF AN OU T (C; A1 , A2 , A3 ) = l∈{1,...,4} dF AN OU T L (ql ; aT1 ,l , aT2 ,l , aT3 ,l ).