ConceptioArchivearXiv CS
arXiv CSopen access

A Resource Estimation Model for the Hardware-Software Co-Design of Distributed Quantum Architectures

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

A Resource Estimation Model for the Hardware-Software Co-Design of Distributed Quantum Architectures Raymond P. H. Wu

Chathurika Ranaweera

Sutharshan Rajasegarar

School of Information Technology Deakin University Burwood, Victoria 3125, Australia [email protected]

School of Information Technology Deakin University Burwood, Victoria 3125, Australia [email protected]

School of Information Technology Deakin University Burwood, Victoria 3125, Australia [email protected]

Ria Rushin Joseph

Jinho Choi

Seng W. Loke

arXiv:2607.22998v1 [quant-ph] 25 Jul 2026

School of Information Technology School of Electrical and Mechanical Engineering School of Information Technology Deakin University Adelaide University Deakin University Burwood, Victoria 3125, Australia Adelaide, South Australia 5005, Australia Burwood, Victoria 3125, Australia [email protected] [email protected] [email protected]

Abstract—In distributed quantum computing (DQC), executing monolithic quantum circuits across multiple interconnected quantum processing units (QPUs) requires dedicated communication qubits to generate and distribute entanglement. Because the number of physical qubits within a QPU is finite, a trade-off emerges where allocating more communication qubits increases the capacity of quantum channels for concurrent non-local operations, but reduces the number of computational qubits available for local gate operations. Distributed quantum compilation routinely ignores this channel capacity, while hardware architects lack a method to determine it prior to quantum circuit partitioning. Moreover, scheduling entanglement on demand introduces severe latency, whereas pre-fetching exposes stored pairs to decoherence. We propose an economic order quantity model from perishable inventory theory to optimize the trade-off between entanglement distribution latency and the time cost of decoherence. The resulting estimate is driven by algorithmic demand and physical constraints, offering a dual application for the hardware-software co-design of high-performance DQC: for hardware architects, it gives the optimal allocation of dedicated communication qubits in static heterogeneous architectures; for compiler developers, it gives the optimal number to reserve dynamically in homogeneous architectures. Index Terms—Distributed quantum computing, quantum networks, quantum compilation, entanglement distribution.

I. I NTRODUCTION In distributed quantum computing (DQC) [1], [2], monolithic quantum circuits are partitioned into a set of subcircuits that can be executed concurrently across multiple interconnected quantum processing units (QPUs). Quantum communication between remote QPUs relies on teleportation protocols, which require both classical communication and the generation and distribution of entangled pairs to establish This research was supported by Deakin University and CSIRO’s Data61 through the Next Generation Quantum Graduate Program (118-NGQGP).

quantum channels [3]. Within each QPU, qubits are functionally classified as either computational or communication qubits. A set of computational qubits is dedicated to executing local gate operations, while a separate set of communication qubits is reserved for generating and distributing entangled pairs across the network to enable inter-processor communication [4]. As the number of available QPUs in the quantum network increases, the number of possible quantum channels scales quadratically [2]: for k nodes on a complete graph Kk , there are k(k − 1)/2 channels, and if each channel requires n communication qubits, each QPU needs n(k −1). The number of allocated communication qubits is therefore one of the constraints that limit the capacity of these quantum channels. Because the number of physical qubits is finite, a trade-off emerges: allocating more communication qubits increases the entanglement resources for concurrent non-local operations, but reduces the number of computational qubits available for local gate operations [5]. Optimizing this allocation is critical for the performance of DQC and represents a fundamental open problem [5]. Distributed quantum compilation routinely ignores the capacity of quantum channels, assuming an unconstrained reservoir of communication qubits [5]. Because the exact demand for communication qubits is only known after the qubit assignment schedule is generated from quantum circuit partitioning, compilers may produce physically infeasible schedules. Recent works [6], [7] have acknowledged this limitation and proposed imposing the number of communication qubits as a constraint during quantum circuit partitioning, but working within tight constraints degrades solution quality and hence the performance of DQC if the architecture is designed with insufficient quantum channel capacity. Avoiding such bottlenecks requires co-designing the architecture from the ground up. Yet, because

the communication qubit requirement is obtained only after partitioning, hardware architects and compiler developers lack a method to predetermine this optimal capacity, which is governed simultaneously by algorithmic demand, physical systems, and network configuration. Modular architectures address quantum interconnects in two regimes [8]. In homogeneous architectures, the roles of computational and communication qubits are dynamically reconfigurable by the compiler. For example, trapped-ion qubits are shuttled between memory and interaction regions using dynamic electric fields in quantum charge-coupled devices (QCCDs) [9], [10], and superconducting QPUs utilize an allto-all router reconfigured by external magnetic flux [11]. In contrast, heterogeneous architectures scale by interconnecting QPUs via specialized hardware interfaces, such as trapped-ion QCCDs with photonic interconnects [12] or superconducting QPUs with microwave or optical connections [13]. These interfaces must be permanently engineered into the QPU during manufacturing, requiring hardware architects to predetermine the optimal allocation to prevent interconnect bottlenecks. Regardless of the underlying architecture, a primary compilation challenge is scheduling entanglement generation and distribution. An on-demand strategy generates pairs only when required, introducing severe latency that halts sub-circuit execution. A pre-fetching strategy greedily generates and stores pairs in advance to minimize this latency, but idling exposes them to decoherence, eventually rendering them useless for computation and incurring a time cost from quantum error correction or re-execution. Consequently, there exists an optimal number of communication qubits that minimizes the overall execution time. Recognizing that time is an expensive resource in quantum networks [4], we propose an economic order quantity (EOQ) model [14] from perishable inventory theory [15] to optimize the trade-off between entanglement distribution latency and the time cost of decoherence. II. R ESOURCE E STIMATION M ODEL We treat entangled pairs as perishable inventory and develop an EOQ model [15] to determine the optimal number of ∗ in a QPU. The problem is communication qubits Ncomm formulated as optimizing the trade-off between the latency of generating and distributing entangled pairs and the time cost of decoherence while they idle. Let RQPU be the entanglement demand rate, in pairs per unit time, for a single QPU. For a quantum circuit with N qubits and depth D, distributed across a quantum network with k ≥ 2 QPUs, the total number of entangled pairs E required for execution follows from the qubit assignment schedule after quantum circuit partitioning. Assuming a total execution time of Dtgate , the network demand rate is Rnetwork = E/(Dtgate ), where tgate is the gate time. Since each bipartite entangled pair consumes resources at two nodes, and assuming the workload is distributed evenly across k nodes, the single-QPU demand rate is given by RQPU = 2Rnetwork /k. We define Clat and clat as the entanglement distribution time cost per batch request and per individual pair, respectively. Clat

represents the fixed, one-time network overhead for classical communication and routing setup, whereas clat encapsulates the physical hardware execution time required to generate each successive entangled pair. Distributing a batch of size Ncomm thus requires latency Clat +clat Ncomm . We further define Cdec as the effective time cost of decoherence per entangled pair per unit time. Cdec quantifies the time cost incurred when an idling entangled pair stored in a communication qubit undergoes decoherence. Assuming continuous consumption, a pre-fetched buffer of size Ncomm is depleted linearly down to zero at rate RQPU , so on average Ncomm /2 communication qubits are occupied and exposed to decoherence. The QPU queries the network at frequency RQPU /Ncomm for a new batch of entangled pairs. Although the true consumption is bursty and depends on the specific scheduling of non-local gates, we adopt the timeaveraged rate RQPU as a first-order design metric: hardware allocation and pre-fetch provisioning are static decisions made prior to execution, for which the mean demand over the circuit governs the amortized cost, while instantaneous bursts are absorbed by the buffer itself. The total system time cost is formulated as Ncomm RQPU . (1) + Cdec C(Ncomm ) = (Clat + clat Ncomm ) Ncomm 2 Setting dC(Ncomm )/dNcomm = 0 and solving for Ncomm yields the unconstrained optimum r 2Clat RQPU Ncomm = , (2) Cdec analogous to the result of the EOQ model [14], [15]. This gives an architectural insight: the theoretical optimal capacity of quantum channels is independent of the per-pair cost clat and is driven solely by the demand rate RQPU , batch cost Clat , and decoherence cost Cdec . Assuming Markovian noise dominated by pure dephasing, we model the fidelity of an idling state as F (t) = e−t/T2 [16], where T2 is the coherence time. Since error suppression in quantum error correction requires a physical error rate below a critical threshold [17], maintaining the fidelity above an algorithmic requirement Fth bounds the idling time by tmax = −T2 ln Fth . A buffer of size Ncomm depleted at rate RQPU has maximum storage duration Ncomm /RQPU , imposing Ncomm /RQPU ≤ tmax . Bounding (2) by this usable lifetime, the optimal number of communication qubits is the piecewise minimum ! r 2Clat RQPU ∗ Ncomm = min , −RQPU T2 ln Fth . (3) Cdec This estimate bridges hardware physical constraints and software algorithmic demand: for static heterogeneous architectures, it gives the optimal allocation of dedicated communication qubits, and for dynamic homogeneous architectures, it gives the optimal number for compilers to reserve, with the reserved qubits returned to computational roles upon completion.

A. Cutoff Strategy

30

20 15 10

0

Clat . −T2 ln Fth

(4)

Substituting (4) into (3) gives p  ∗ Ncomm = min −2RQPU T2 ln Fth , −RQPU T2 ln Fth . (5) The optimum then depends entirely on the algorithmic demand RQPU , the fidelity threshold Fth , and the coherence time T2 , and not on Clat : the local QPU architecture does not depend on the global quantum network. This indicates modular decoupling in distributed quantum architectures. III. R ESULTS

∗ Ncomm

5

5

10 15 20 Number of communication qubits Ncomm

25

(a) Time cost as a function of the number of communication qubits Ncomm , illustrating the trade-off between latency and decoherence. ∗ Optimal number of communication qubits Ncomm

Cdec =

Clat RQPU /Ncomm + Cdec Ncomm /2 Clat RQPU /Ncomm Cdec Ncomm /2

25

Time cost

Consuming a decohered pair introduces uncorrectable errors, typically necessitating re-execution of the shot. To avoid this catastrophic time penalty, we assume the compiler employs a proactive cutoff strategy [18] that tracks the idle duration of pre-fetched pairs and discards any exceeding tmax = −T2 ln Fth . On discarding an expired pair, the QPU halts and requests a replacement, incurring a penalty dominated by the fixed latency Clat . Amortizing Clat over the usable lifetime yields an effective decoherence penalty rate

80 RQPU = 5 pairs/µs RQPU = 15 pairs/µs RQPU = 25 pairs/µs RQPU = 35 pairs/µs RQPU = 45 pairs/µs

70 60 50 40 30 20

Table I lists representative order-of-magnitude parameters used to evaluate the model. The entanglement distribution latency depends on the network scale: for a quantum local area network (LAN, ≲ 10 m) it is dominated by hardware overhead, while for a quantum wide area network (WAN, ≳ 100 km) it is dominated by the photon travel time in optical fiber.

∗ (b) Optimal number of communication qubits Ncomm against fidelity threshold Fth for different entanglement demand rates RQPU .

TABLE I

Fig. 1. Estimation of the optimal number of communication qubits for distributed quantum architectures.

10 0 0.5

0.6

0.7 0.8 Fidelity threshold Fth

0.9

1

R EPRESENTATIVE PARAMETERS FOR D ISTRIBUTED Q UANTUM A RCHITECTURES

Parameter

Technology

Gate time tgate Coherence T2 Threshold Fth Distribution Clat

Superconducting (transmon) Superconducting (transmon) Surface code Quantum LAN (≲ 10 m) Quantum WAN (≳ 100 km)

Value ∼10 ns–100 ns [19] ∼100 µs [19], [20] 98.6 %–98.9 % [21] 0.18 µs–6 µs [22], [23] ∼1 ms [24]

A. Superconducting Transmon Qubits We work through the case of superconducting transmon qubits, for which tgate ∼ 100 ns [19] and T2 ∼ 100 µs [19], [20]. The entanglement demand rate RQPU depends on quantum circuit partitioning. For a quantum circuit with N qubits and depth D distributed across k ≥ 2 QPUs, the worst case requires D⌊N/2⌋ pairs. Reference [7] reports an entangled pair fraction (the ratio of required pairs to the worst case D⌊N/2⌋) of 0.023 for QASM-large circuits [25] partitioned over 2, 3 and 4 QPUs. Taking a conservative fraction of 0.1 for N = 100 and D = 1000, distributed across k = 4 QPUs, the worst-case 1000⌊100/2⌋ = 50000 twoqubit gates give Rnetwork = (50 000 × 0.1)/(1000 × 100 ns) =

50 pairs/µs, hence RQPU = 25 pairs/µs. With a surface∗ code target of Fth = 0.99 [21], (5) gives Ncomm = 7.1. A QPU with Nphys = 32 physical qubits then supports this computation with Ncomp = 25 computational and Ncomm = 7 communication qubits, dedicating 22 % of quantum resources to communication. For a quantum LAN with Clat = 1 µs, the cost, plotted in Fig. 1a, is minimized at Ncomm = 7.1. The effect of Fth for different values of RQPU is shown in Fig. 1b. B. Suboptimal Allocation Penalties Here, we address the possibility that the allocation of communication qubits is suboptimal due to hardware constraints. Let Nphys be the number of physical qubits and Ncomp be the number of computational qubits on a single QPU. The maximum possible allocation of communication ∗ qubits is Ncomm,max = Nphys − Ncomp . If Ncomm ≤ Ncomm,max , ∗ the compiler allocates Ncomm communication qubits and the quantum circuit executes at the theoretical optimum. However, ∗ if Ncomm > Ncomm,max , the compiler is forced to limit the allocation to Ncomm = Ncomm,max and the system incurs a time penalty. A quantifiable estimate of this penalty is given by the EOQ model as the difference in the total system time cost ∗ C(Ncomm,max ) − C(Ncomm ).

IV. C HALLENGES AND O UTLOOK Several extensions are needed before deployment. The demand rate RQPU depends on the targeted quantum algorithm and partitioning quality, so benchmarking diverse quantum circuits and partitioning algorithms is required for accurate calibration, together with more experimental data on entanglement distribution latency. Because realistic demand is bursty and stochastic rather than continuous, extending this model with probabilistic error modeling and stochastic demand is a key step toward a dynamic compiler module. Finally, the optimal allocation is further affected by entanglement management strategies: the consumption order of stored pairs (first-infirst-out versus last-in-first-out), and entanglement purification [26], which consumes multiple low-fidelity pairs to yield fewer high-fidelity ones. V. C ONCLUSION We proposed a resource estimation model that resolves the disconnect between quantum hardware design and compilation, optimizing the capacity of quantum channels by minimizing the total time cost of entanglement distribution and decoherence. Applied to superconducting transmon qubits with realistic parameters, the model yields an estimate driven by algorithmic demand and physical constraints. It offers a dual application for the hardware-software co-design of high-performance DQC: the optimal allocation of dedicated communication qubits for hardware architects designing static heterogeneous architectures, and the optimal number for compilers to reserve dynamically in homogeneous architectures. R EFERENCES [1] S. W. Loke, From distributed quantum computing to quantum internet computing: an introduction. John Wiley & Sons, 2023. [2] M. Caleffi, M. Amoretti, D. Ferrari, J. Illiano, A. Manzalini, and A. S. Cacciapuoti, “Distributed quantum computing: A survey,” Computer Networks, vol. 254, p. 110672, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S1389128624005048 [3] A. S. Cacciapuoti, M. Caleffi, R. Van Meter, and L. Hanzo, “When entanglement meets classical communications: Quantum teleportation for the quantum internet,” IEEE Transactions on Communications, vol. 68, no. 6, pp. 3808–3833, 2020. [4] W. Kozlowski, S. Wehner, R. V. Meter, B. Rijsman, A. S. Cacciapuoti, M. Caleffi, and S. Nagayama, “Architectural Principles for a Quantum Internet,” RFC 9340, Mar. 2023. [Online]. Available: https://www.rfc-editor.org/info/rfc9340 [5] D. Ferrari, A. S. Cacciapuoti, M. Amoretti, and M. Caleffi, “Compiler design for distributed quantum computing,” IEEE Transactions on Quantum Engineering, vol. 2, pp. 1–20, 2021. [6] F. Burt, K.-C. Chen, and K. K. Leung, “Generalised circuit partitioning for distributed quantum computing,” in 2024 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 02, 2024, pp. 173–178. [7] ——, “A Multilevel Framework for Partitioning Quantum Circuits,” Quantum, vol. 10, p. 1984, Jan. 2026. [Online]. Available: https://doi.org/10.22331/q-2026-01-22-1984 [8] D. Awschalom, K. K. Berggren, H. Bernien, S. Bhave, L. D. Carr, P. Davids, S. E. Economou, D. Englund, A. Faraon, M. Fejer, S. Guha, M. V. Gustafsson, E. Hu, L. Jiang, J. Kim, B. Korzh, P. Kumar, P. G. Kwiat, M. Lončar, M. D. Lukin, D. A. Miller, C. Monroe, S. W. Nam, P. Narang, J. S. Orcutt, M. G. Raymer, A. H. Safavi-Naeini, M. Spiropulu, K. Srinivasan, S. Sun, J. Vučković, E. Waks, R. Walsworth, A. M. Weiner, and Z. Zhang, “Development of quantum interconnects (QuICs) for next-generation information

technologies,” PRX Quantum, vol. 2, p. 017002, Feb 2021. [Online]. Available: https://link.aps.org/doi/10.1103/PRXQuantum.2.017002 [9] D. Kielpinski, C. Monroe, and D. J. Wineland, “Architecture for a largescale ion-trap quantum computer,” Nature, vol. 417, no. 6890, pp. 709– 711, 2002. [10] J. M. Pino, J. M. Dreiling, C. Figgatt, J. P. Gaebler, S. A. Moses, M. Allman, C. Baldwin, M. Foss-Feig, D. Hayes, K. Mayer et al., “Demonstration of the trapped-ion quantum ccd computer architecture,” Nature, vol. 592, no. 7853, pp. 209–213, 2021. [11] X. Wu, H. Yan, G. Andersson, A. Anferov, M.-H. Chou, C. R. Conner, J. Grebel, Y. J. Joshi, S. Li, J. M. Miller, R. G. Povey, H. Qiao, and A. N. Cleland, “Modular quantum processor with an all-to-all reconfigurable router,” Phys. Rev. X, vol. 14, p. 041030, Nov 2024. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevX.14.041030 [12] C. Monroe, R. Raussendorf, A. Ruthven, K. R. Brown, P. Maunz, L.-M. Duan, and J. Kim, “Large-scale modular quantum-computer architecture with atomic memory and photonic interconnects,” Phys. Rev. A, vol. 89, p. 022317, Feb 2014. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevA.89.022317 [13] S. Bravyi, O. Dial, J. M. Gambetta, D. Gil, and Z. Nazario, “The future of quantum computing with superconducting qubits,” Journal of Applied Physics, vol. 132, no. 16, p. 160902, 10 2022. [Online]. Available: https://doi.org/10.1063/5.0082975 [14] F. W. Harris, “How many parts to make at once,” Operations Research, vol. 38, no. 6, pp. 947–950, 1990. [Online]. Available: https://doi.org/10.1287/opre.38.6.947 [15] S. Nahmias, “Perishable inventory theory: A review,” Operations Research, vol. 30, no. 4, pp. 680–708, 1982. [Online]. Available: https://doi.org/10.1287/opre.30.4.680 [16] M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information: 10th Anniversary Edition. Cambridge University Press, 2010. [17] Google Quantum AI and Collaborators, “Quantum error correction below the surface code threshold,” Nature, vol. 638, no. 8052, pp. 920– 926, 2025. [18] B. Li, T. Coopmans, and D. Elkouss, “Efficient optimization of cutoffs in quantum repeater chains,” IEEE Transactions on Quantum Engineering, vol. 2, pp. 1–15, 2021. [19] J. M. Gambetta, J. M. Chow, and M. Steffen, “Building logical qubits in a superconducting quantum computing system,” npj quantum information, vol. 3, no. 1, p. 2, 2017. [20] I. Siddiqi, “Engineering high-coherence superconducting qubits,” Nature Reviews Materials, vol. 6, no. 10, pp. 875–891, 2021. [21] D. S. Wang, A. G. Fowler, and L. C. L. Hollenberg, “Surface code quantum computing with error rates over 1%,” Phys. Rev. A, vol. 83, p. 020302, Feb 2011. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevA.83.020302 [22] P. Kurpiers, P. Magnard, T. Walter, B. Royer, M. Pechal, J. Heinsoo, Y. Salathé, A. Akin, S. Storz, J.-C. Besse et al., “Deterministic quantum state transfer and remote entanglement using microwave photons,” Nature, vol. 558, no. 7709, pp. 264–267, 2018. [23] C. J. Axline, L. D. Burkhart, W. Pfaff, M. Zhang, K. Chou, P. CampagneIbarcq, P. Reinhold, L. Frunzio, S. Girvin, L. Jiang et al., “On-demand quantum state transfer and entanglement between remote microwave cavity memories,” Nature Physics, vol. 14, no. 7, pp. 705–710, 2018. [24] K. Azuma, S. E. Economou, D. Elkouss, P. Hilaire, L. Jiang, H.-K. Lo, and I. Tzitrin, “Quantum repeaters: From quantum networks to the quantum internet,” Rev. Mod. Phys., vol. 95, p. 045006, Dec 2023. [Online]. Available: https://link.aps.org/doi/10.1103/RevModPhys.95.045006 [25] A. Li, S. Stein, S. Krishnamoorthy, and J. Ang, “QASMBench: A low-level quantum benchmark suite for NISQ evaluation and simulation,” ACM Transactions on Quantum Computing, vol. 4, no. 2, Feb. 2023. [Online]. Available: https://doi.org/10.1145/3550488 [26] C. H. Bennett, G. Brassard, S. Popescu, B. Schumacher, J. A. Smolin, and W. K. Wootters, “Purification of noisy entanglement and faithful teleportation via noisy channels,” Phys. Rev. Lett., vol. 76, pp. 722–725, Jan 1996. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevLett.76.722

Record · ID 405620 · SHA-256 2d6c5c3aaec1e92b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.