Sequential vs. Simultaneous Entanglement Swapping under Optimal Link-Layer Control Priyam Srivastava1,† , Akshat R. Sabavat1 , Siddharth Jain2 , Alan Scheller-Wolf3 , Sridhar Tayur3 , David Tipper1 , Prashant Krishnamurthy1 , Amy Babay1 , Kaushik P. Seshadreesan1,2 1
arXiv:2605.04047v1 [quant-ph] 5 May 2026
Department of Informatics & Networked Systems, School of Computing & Information, University of Pittsburgh, Pittsburgh, PA 15260, USA 2 Department of Physics & Astronomy, University of Pittsburgh, Pittsburgh, PA 15260, USA 3 Tepper School of Business, Carnegie Mellon University, Pittsburgh, PA, USA Emails: [email protected]† , [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected] †
Corresponding author.
Abstract—Connection-less, packet-switched quantum network architectures distribute entanglement across multi-hop paths through sequential entanglement swapping, in which each node acts on purely local state information. The architectural advantages over the connection-oriented alternative—simultaneous SWAP-ASAP—are compelling, but sequential swapping holds partial chains in intermediate buffers between successive swaps, exposing them to memory decoherence in a way simultaneous SWAP-ASAP avoids by design. We present a proof-of-principle study at fixed chain length n = 4 in which each elementary link is governed by a fixed reinforcement-learning policy optimizing the secret-key rate of the six-state protocol, leaving the networklayer protocol as the sole independent variable. Sweeping the network-layer memory coherence time Tcext over four orders of magnitude reveals a clear regime structure governed by the dimensionless ratio Tcext /τ , where τ is the per-link entanglement heralding latency. Simultaneous SWAP-ASAP delivers a constant rate across the full sweep. Sequential swapping, by contrast, collapses to zero end-to-end deliveries below Tcext /τ = 25, and begins recovering at Tcext /τ = 50. It remains limited by the simultaneous rate, which it saturates only at the relaxed end of the sweep. These results suggest that the connection-less penalty is a near-term phenomenon tied to present-day memory coherence rather than a fundamental property of sequential swapping. Index Terms—Quantum Networks, Reinforcement learning, Entanglement distribution, Sequential entanglement swapping
I. I NTRODUCTION Quantum networks are expected to enable distributed applications including device-independent quantum key distribution [1], [2], distributed quantum computation [3], [4], and quantum-enhanced sensing [5], [6]. All of these depend on entanglement distributed between distant nodes, produced by the same recipe: intermediate nodes generate entanglement on shorter elementary links and stitch them together into a longer end-to-end resource via entanglement swapping [7], [8]. The order in which this stitching occurs is the focus of this paper. This work is supported by the U.S. Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program, for support under Award Number DE-SC0026264, and PQI Community Collaboration Awards.
Two paradigms have emerged for the timing of swapping relative to link-level entanglement generation [9]. Simultaneous (or “wait-and-swap”) protocols generate entanglement on every link first and then execute all swaps in parallel [10], [11]. This requires centralized coordination for both route reservation and the synchronized swap trigger. Sequential (or “swap-and-wait”) protocols extend a chain hop-by-hop: as soon as adjacent links produce entanglement they swap, and the partially assembled chain waits for the next link to be ready [12], [13]. Sequential admits a fully distributed, connection-less implementation in which each node acts on purely local state information. This implementation is the basis for the packet-switched quantum network architecture proposed by Bacciottini et al. [9]. The systems-level case for the connection-less architecture is compelling. Route-setup overhead is eliminated, the control plane is simpler, and degradation under multiple flows is graceful. This makes it an attractive target for near-term deployments. Yet the feature that makes sequential swapping locally implementable also makes it vulnerable to memory decoherence in a way simultaneous SWAP-ASAP is not. Partial chains sit in intermediate buffers between successive swaps, exposing them to storage noise that simultaneous SWAPASAP avoids. The relevant research question is therefore not which protocol is better—simultaneous SWAP-ASAP is, for a single flow. The question is rather: in what hardware regime does the connection-less protocol remain viable? We address this question empirically at fixed chain length n = 4. The study is a proof of principle: it establishes the methodology and provides a first empirical anchor, not a fully general regime characterization. Our setup is a twolayer architecture in which at the link layer, each elementary link is governed by a reinforcement-learning policy trained to maximize the per-link six-state QKD secret-key rate, following Yau et al. [14]. Every link continuously runs this fixed policy, deciding on local actions (generate, distill, discard, deliver)
with no representation of which end-to-end flow its output will serve, and delivers entangled pairs into an external memory buffer—called the link buffer, available for network layer protocols to consume. A deterministic, centralized networklayer controller implements either of the sequential and simultaneous SWAP-ASAP protocols. By holding the trained linklayer policy fixed and varying only the network-layer protocol, our setup isolates network-level effects empirically. Note that it is for uniformity that we model both protocols with a centralized network-layer controller that consumes pairs from these buffers and always selects the freshest available pair from each buffer to suppress storage decoherence. We emphasize that for the sequential protocol this is a modeling convenience only: growing entanglement from one end to another in a sequential manner is entirely local and implementable without centralized coordination, as each intermediate node need only know that a partial chain has arrived, e.g., from the left and that a fresh pair is available on the right. On the other hand, the centralized controller is genuinely necessary only for simultaneous SWAP-ASAP, which requires global visibility to coordinate route reservation and the synchronized swap trigger. A natural concern that may arise with analyzing the connection-less sequential protocol in a centrally controlled, reserved-route simulation is whether this faithfully captures the truly connection-less case, in which links have no advance knowledge of any route. The two scenarios are in fact statistically equivalent provided that the path length is fixed, the links are homogeneous, and every link continuously runs its optimal link-layer policy—preemptively delivering entanglement to its link buffer regardless of whether it currently serves an active connection. This last condition is fully consistent with connection-less operation: a link need not know it belongs to a route to generate entanglement opportunistically [9]. Our reserved-route simulation analyzes precisely this setting. We sweep the external memory coherence time—Tcext from a relaxed-coherence reference down to a stressed regime in which it becomes comparable to the per-link heralding latency. We separately disentangle internal link-layer memory coherence time Tcint from the external coherence time with an off-diagonal Tcint × Tcext sweep. This methodology is complementary to existing analytic treatments of repeater chain policies, which derive closed-form expressions for sequential and simultaneous protocols under simplified link-layer models [11]–[13]. Listed below are our main contributions and findings. Main contributions and findings 1) Architectural factorization. The link-layer task admits a single dimensionless operating point at 0.1357– 0.1358 bits/tick, where a tick corresponds to the per-link entanglement heralding latency, i.e., the time incurred per attempt at remote, heralded entanglement generation across an elementary link, and the off-diagonal sweep shows chain-level outcomes under either protocol depend only on Tcext . Any network-level performance
difference is therefore attributable to the choice of the network layer protocol alone. 2) Coherence-time regime structure. Sequential collapses to zero end-to-end entanglement deliveries below Tcext /τ = 25, begins recovering at Tcext /τ = 50, and saturates at the simultaneous rate by the relaxedcoherence reference at Tcext ∈ {0.5, 2} s (Tcext /τ ≈ 10,000–80,000), where the two protocols agree to within 0.4%. Simultaneous SWAP-ASAP delivers a constant rate across the full stressed sweep. The crossover from substantial gap to equivalence is bracketed by these two regimes but not directly localized in the unmeasured intermediate range. 3) Mechanism. Sequential’s pipelined assembly requires partial chains to survive multiple ticks in chain buffers, while simultaneous SWAP-ASAP consumes link pairs in the tick they are pushed and holds no intermediate chain storage. The collapse threshold coincides with the regime in which the per-pair cutoff (described in Section II-C) falls below one tick. Read together, these results suggest that the connectionless penalty is a near-term phenomenon tied to present-day memory coherence rather than a fundamental property of sequential swapping. We frame this as a hypothesis consistent with the present data rather than a general claim, given the n = 4 scope of the study. In the hardware regime accessible today, the penalty is real and can be substantial, and protocol selection should therefore be guided by the operating Tcext /τ ratio rather than by topology alone. The remainder of the paper is organized as follows. Section II describes the two-layer simulation framework, including the WN2M2 link-layer agent and both network-layer protocols. Section III reports the link-layer dimensional invariance, the regime structure under the Tcext sweep, the chain-buffer dwell-time diagnostic, and the off-diagonal sweep. Section IV interprets these results and locates their scope. Section V concludes with the limitations and natural extensions of the present study.
II. M ETHODS We compare the two network-layer protocols introduced above, sequential swapping (swap-and-wait) and simultaneous SWAP-ASAP (wait-and-swap), under a controlled methodology in which the link layer is held fixed and the networklayer protocol is the sole independent variable. Section II-A describes the two-layer system model and underlying physics shared between both layers. Sections II-B and II-C describe the link layer (with its WN2M2 reinforcement-learning agent) and the network layer (with both protocol variants), respectively. WN2M2 denotes two nodes with two memories each and a Werner state generated upon successful entanglement. Section II-D defines the per-pair efficiency metric and evaluation setup.
A. System Model and Physics We simulate an n-link quantum network chain consisting of n + 1 nodes (two end nodes and n − 1 intermediate switches) connected by elementary links along a single, pre-selected endto-end path. Route selection itself is not part of our study. The system is organized into two layers separated by a buffer interface (Fig. 1).
Physics of stored states: Werner states stored for time ∆t in memory of coherence time Tc undergo depolarizing decay according to [15]: 1 1 −2∆t/Tc D(F, ∆t, Tc ) = + F − e . (1) 4 4 We apply (1) with Tc = Tcint inside the link agent’s internal memory and with Tc = Tcext in the external buffer tier. Two Werner states with fidelities F1 and F2 combine via a twirled BSM swap [7] to produce a state with fidelity (1 − F1 )(1 − F2 ) . (2) 3 Local operations (gates, measurements, memory readout) are assumed instantaneous and noiseless. Fswap (F1 , F2 ) = F1 F2 +
B. Link Layer
Fig. 1. Two-layer architecture. Top: per-node link-layer environment. The WN2M2 agent operates on two internal memory slots (coherence Tcint ); its CONSUME action delivers a Werner pair to the external link buffer (coherence Tcext ). Bottom: the n = 4 chain under each of the two network layer protocols. Sequential holds partial chains in n − 1 chain buffers (highlighted); simultaneous SWAP-ASAP holds none.
Memory hierarchy: Each switch carries two physically distinct types of quantum memory per link. Internal communication memories, with coherence time Tcint , are the shortcoherence registers used by the link-layer agent for active operations: heralded entanglement generation, distillation attempts, and intermediate storage during a single training episode. An external storage memory, with coherence time Tcext , holds pairs the link-layer agent has finished operating on and released to the network layer. This external memory is realized as the link buffer, a per-link deque of capacity B = 20 storing tuples (Fdel , tdel ). The sequential protocol (Section II-C) additionally maintains n − 1 chain buffers C1 , . . . , Cn−1 , each of capacity B, in this same external tier. Layer separation as the experimental control: A deterministic network-layer protocol assembles end-to-end (E2E) entangled pairs by performing n − 1 Bell-state measurement (BSM) swaps on entries drawn from the link buffers and, in the sequential case, from the chain buffers. All buffers use freshest-first selection: a pop_freshest operation returns the entry with the largest delivery time, suppressing storage decoherence. The buffer capacity B = 20 is chosen large enough that the protocol never operates capacity-limited. Pairs whose age exceeds a per-pair cutoff tcut (defined in Section II-C) are discarded at each tick. The network layer protocol is the only component that differs between the two schemes we compare.
The link layer comprises one independent reinforcementlearning agent per elementary link. Each agent runs the WN2M2 policy of Yau et al. [14], which performs heralded entanglement generation, distillation via the DEJMPS protocol [16], and local memory management within its two internal memory slots. When the agent’s policy outputs its terminal CONSUME action, the distilled Werner pair is transferred from internal memory into the external link buffer, after which a new episode begins immediately. By construction, every delivered pair satisfies Fdel ≥ F0 , where F0 is chosen per chain length so that the end-to-end fidelity after n − 1 swaps would remain above the six-state QKD threshold Fmin = 0.81 [17], [18] (e.g., F0 = 0.94 for n = 4). Agent state and actions: Each link agent is trained independently on a single elementary link via REINFORCE [19] within the WN2M2 framework [14]. The agent observes a state s = (F1 , F2 , p, t), where F1 and F2 are the Werner-state fidelities of the two internal memory slots (decohering at Tcint ), p ∈ (0, 1] encodes residual uncertainty about slot contents (with p = 1 when both pairs are fully heralded and p < 1 when a recent distillation outcome remains pending), and t is the elapsed episode time. From this state, the agent selects one of four actions: WAIT (attempt heralded entanglement generation), DISCARD (drop the lower-fidelity slot), PURIFY (apply DEJMPS distillation across both occupied slots), or CONSUME (deliver the current pair as (Fdel , tdel ) from internal to external memory, terminating the episode). Physically infeasible actions are blocked by a hard action mask applied before the softmax. Training procedure: The policy is a two-hidden-layer MLP (4 → 64 → 64 → 4) with masked softmax output, trained over batches of 10,000 episodes per iteration via REINFORCE with the Adam optimizer [20]. Each episode produces two return streams: GF (terminal delivery fidelity) and GT (time-to-go), combined through the gradient of the SKR utility uSKR defined in (3). When the delivered fidelity falls below the six-state SKR-positive threshold Fdel ≈ 0.811 (where the unclamped 1 − H(F ) formula crosses zero), we substitute a bootstrap gradient with ∂u/∂JT = 0. In this
regime the unclamped formula is monotonically increasing in episode length, which would otherwise reward stalling rather than fast delivery. The true partial derivatives ∂u/∂JF and ∂u/∂JT are applied above this threshold. Trained policy bank: We trained ten policies, one per (L, Tcint ) configuration with L ∈ {5, 10} km and dimensionless ratio Tcint /τ ∈ {5, 10, 25, 50, 100} (where τ = L/cfiber is the per-link heralding latency), all at F0 = 0.94. The trained policies are reused without modification across all multi-hop configurations. We characterize their dimensional invariance empirically in Section III-A.
C. Network Layer Controller and Protocols The full multi-hop architecture is shown in Fig. 1. Each elementary link runs its WN2M2 agent independently and delivers pairs into its own link buffer, while the network-layer controller draws from these buffers (and, in the sequential case, from chain buffers) to assemble end-to-end pairs. Each elementary link has its own per-attempt latency τℓ = Lℓ /cfiber . The global simulation clock advances at τmin = minℓ τℓ , and each link agent steps at multiples of its own τℓ . End-toend fidelity is computed by recursive application of (2), with intermediate states decohered via (1) over any storage interval between deliveries and swaps. The per-pair cutoff tcut is derived from the fidelity budget. For a delivery fidelity Fdel , tcut is the time at which depolarizing decay (1) would drive fidelity below the required floor Freq (n), defined as the level at which the end-to-end fidelity after n − 1 swaps would still meet Fmin . The closed-form expression for tcut and the associated buffer-tier definitions of Freq are given in Appendix A. Sequential swapping (swap-and-wait): The sequential swapping protocol maintains the n − 1 chain buffers C1 , . . . , Cn−1 , where Ci stores partial chains spanning i + 1 links. At each tick, the controller extends every existing chain by one link where possible, drawing the freshest available pair from the corresponding link buffer and performing a BSM swap. New length-2 chains are seeded from B1 , and chain entries exceeding tcut are expired. This pipelined design supports multiple in-flight chains, so a single tick can produce several E2E deliverys when buffers are well-stocked. Full pseudocode is given in Appendix A. Simultaneous SWAP-ASAP (wait-and-swap): In the case of the simultaneous protocol, the controller waits until every link buffer contains at least one valid pair, then pops one pair from each, decoheres them all to the current time at Tcext , and combines them via a balanced binary swap tree (recursive application of (2)) to produce a single E2E pair. In contrast to sequential, this protocol maintains no intermediate chain storage between ticks, performing all n − 1 swaps in ⌈log2 n⌉ rounds rather than n − 1. Full pseudocode is given in Appendix A.
D. Evaluation Methodology We evaluate the two network-layer protocols using the perpair efficiency metric uSKR =
max{0, 1 − H(F̄E2E )} , ⟨∆tdeliver ⟩
(3)
where F̄E2E is the mean E2E fidelity across deliverys in a trial, ⟨∆tdeliver ⟩ = Tlast /N is the mean inter-delivery interval at the application boundary (with N the delivery count and Tlast the simulation time of the last delivery), and H(F ) = is the Werner-state Shannon −F log2 F − (1 − F ) log2 1−F 3 entropy used in the six-state SKR formula of Yau et al. [14]. The metric is symmetric across both network-layer protocols by construction: it depends only on the mean fidelity and mean inter-delivery interval, both defined identically for sequential and simultaneous SWAP-ASAP deliverys. The following physical constants are fixed throughout: fiber attenuation length Latt = 22 km, speed of light in fiber cfiber = 200,000 km/s, coupling/loss factor K = 0.9, and application threshold Fmin = 0.81. Per-link generation probability is pgen = K exp(−L/Latt ). We focus on chains of length n = 4 and report two complementary sweeps. The matched-coherence sweep sets Tcint = Tcext across Tcext /τ ∈ {5, 10, 25, 50, 100} for both L ∈ {5, 10} km symmetric topologies and four bottleneckposition configurations (one L = 10 km link at each of positions p ∈ {1, 2, 3, 4} in an otherwise L = 5 km chain). The off-diagonal sweep varies Tcint and Tcext independently across all 5 × 5 combinations of {250, 500, 1250, 2500, 5000} µs at L = 10 km. A separate relaxed-coherence reference at Tcext ∈ {0.5, 2} s is reported in Appendix B. Each configuration is run for Ntrials = 200 independent trials of Tsim = 5 s simulated wall-clock time per trial. The simulation engine is implemented in Python, using NumPy for the physics, PyTorch for the WN2M2 policy networks, and Gymnasium for the per-link environment. Sweeps are dispatched as SLURM array jobs on the Pittsburgh CRC cluster. Each trial uses an independent random seed. Results are reported as means with 95% confidence intervals computed from the per-trial distribution where applicable, and via pooled estimators across delivery events otherwise. III. R ESULTS We evaluate sequential swapping and simultaneous SWAPASAP over an n = 4 chain across two complementary sweeps. The first holds the link-layer policy fixed and samples the external coherence time Tcext at two clusters: a relaxedcoherence regime Tcext ∈ {0.5, 2.0} s (Appendix B) and a stressed-coherence regime Tcext ∈ {125, . . . , 5000} µs in which Tcext is comparable to the per-link heralding latency τ . The intermediate range is left to future work. The second is an off-diagonal sweep in which Tcint and Tcext are varied independently. We report performance using the symmetric per-pair efficiency uSKR defined in (3).
A. Link-Layer Dimensional Invariance (L, Tcint )
We trained ten WN2M2 link policies, one per configuration with L ∈ {5, 10} km and dimensionless ratio Tcint /τ ∈ {5, 10, 25, 50, 100}. All policies were trained at the same delivery-fidelity target F0 = 0.94, the floor required for n = 4 against the six-state threshold Freq . Each policy was evaluated on its native configuration. Delivered SKR is reported in dimensionless units of bits per heralding tick.
a constant 165.0 bps at L = 5 and 16.1 bps at L = 10 across the entire sweep. Sequential emits zero end-to-end pairs at Tcext ≤ 625 µs for L = 5 and at Tcext ≤ 1250 µs for L = 10. It recovers partially at higher Tcext , reaching 8.7 bps at Tcext = 5000 µs for L = 10. At both link lengths, sequential is non-delivering up to Tcext /τ = 25 and delivers at Tcext /τ = 50. Bottleneck topologies: Figure 4 shows the same comparison for chains in which one L = 10 km link occupies a single position p ∈ {1, 2, 3, 4} within an otherwise L = 5 km chain. Simultaneous SWAP-ASAP delivers 56–91 bps depending on p, again invariant in Tcext . Sequential emits zero pairs at Tcext = 250 µs at every position. It recovers partially at higher Tcext , with the gap to simultaneous SWAP-ASAP at the Tcext = 2500 µs slice ranging from +157% at p = 2 to +475% at p = 4.
Fig. 2. Each bar is one trained WN2M2 policy evaluated on its training configuration, in bits per heralding tick. Ten policies span L ∈ {5, 10} km and Tcint /τ ∈ {5, 10, 25, 50, 100}.
Figure 2 shows that all ten policies converge to the same dimensionless operating point within four-decimal precision: 0.1357–0.1358 bits/tick. The mean delivery fidelity is F̄del = 0.9575 and the mean inter-delivery interval is 7.36τ in tick units. The convergence holds across both link lengths and across a 20× range in Tcint /τ , including the most stressed case Tcint = 5τ . B. Sequential Collapse Below the Coherence Threshold We sweep Tcext from 5τ to 100τ at both L = 5 km and L = 10 km, with the link-layer policy fixed at the operating point of Section III-A. A separate reference sweep at Tcext ∈ {0.5, 2} s is reported in Appendix B. The figures below report the microsecond-regime sweep.
Fig. 4. uSKR vs. Tcext for chains with one L = 10 km link at position p ∈ {1, 2, 3, 4} in an otherwise L = 5 km chain. Filled bars: simultaneous SWAP-ASAP. Hatched bars: sequential.
C. Chain-Buffer Storage Diagnostic We instrument the simulator with a mean_chain_storage diagnostic that records, for each emitted end-to-end pair, the cumulative time its chain spent in chain buffers between successive swaps—the chainbuffer dwell time. By construction this quantity is identically zero for simultaneous SWAP-ASAP, which maintains no chain buffers, and nonzero for sequential, which maintains n − 1 chain buffers C1 , . . . , Cn−1 that hold partial chains between successive swaps.
Fig. 3. uSKR vs. Tcext for the symmetric topologies [5, 5, 5, 5] km and [10, 10, 10, 10] km. Filled bars: simultaneous SWAP-ASAP. Hatched bars: sequential.
Fig. 5. Left: mean chain-buffer dwell time of growing sequential pairs vs. Tcext . By construction, simultaneous SWAP-ASAP has no chain buffers. Right: end-to-end delivery rate for both network-layer protocols vs. Tcext .
Symmetric topologies: Figure 3 shows uSKR for the two symmetric topologies [5, 5, 5, 5] km and [10, 10, 10, 10] km across the Tcext sweep. Simultaneous SWAP-ASAP delivers
Figure 5 (left) shows the mean chain-buffer dwell time for emitted sequential pairs across the Tcext sweep. The dwell time is approximately zero for Tcext ≤ 1250 µs and rises to ∼50 µs
at Tcext = 2500 µs and ∼120 µs at Tcext = 5000 µs (both at L = 10). The right panel shows that simultaneous SWAPASAP’s delivery rate is constant in Tcext , while sequential’s rate drops to zero in the same regime where dwell time is approximately zero, and recovers only at the highest Tcext values tested. D. Off-Diagonal Tcint × Tcext Sweep We sweep Tcint and Tcext independently across all 5 × 5 = 25 combinations of {250, 500, 1250, 2500, 5000} µs, evaluated under the L = 10 km symmetric topology. Each cell uses the trained policy matched to its Tcint value from the bank of Section III-A.
Fig. 6. Off-diagonal Tcint × Tcext sweep at L = 10 km, B = 20. (a) Simultaneous SWAP-ASAP. (b) Sequential. Both panels are flat in Tcint (rows) to four decimals.
Figure 6 shows the uSKR heatmap for both network-layer protocols. Rows are flat to four decimals: at fixed Tcext , varying Tcint across all five trained policies produces identical uSKR values for both protocols. Simultaneous SWAP-ASAP delivers 16.1 bps at every cell. Sequential’s output is determined entirely by the column index (Tcext ): 0.0 bps for Tcext ≤ 1250 µs, 5.5 bps at Tcext = 2500 µs, and 8.7 bps at Tcext = 5000 µs, independent of Tcint . IV. D ISCUSSION The link-layer invariance of Fig. 2 and the Tcint -blindness of Fig. 6 together support the two-layer architecture as a genuine empirical decomposition rather than a modeling convenience. The link-layer task is characterized by a single dimensionless operating point at fixed delivery-fidelity target F0 . The offdiagonal sweep (Fig. 6) confirms that the chain-level outcome under either protocol depends only on Tcext , not on which Tcint the link policy was trained at. Any chain-level performance difference between sequential and simultaneous SWAP-ASAP must therefore originate at the network layer and depend only on Tcext . This factorization isolates network-layer effects in the comparisons of Section III-B. Figs. 3 and 4 show a qualitative asymmetry: simultaneous SWAP-ASAP is invariant while sequential collapses to zero below a threshold. This asymmetry traces to a structural difference between the two network-layer protocols’ state machines. Sequential’s pipelined design requires partial chains to survive multiple ticks of waiting in chain buffers while downstream link deliveries arrive. Simultaneous SWAP-ASAP’s single-tick
collapse of the entire swap tree requires only that each link buffer hold at least one valid pair simultaneously, which can be the pair pushed on the same tick. The chain-buffer dwell time diagnostic of Fig. 5 confirms this as the operative mechanism: the regime in which sequential’s delivery rate collapses is precisely the regime in which the dwell time becomes a nontrivial fraction of Tcext . The mechanism can be made quantitative through the perpair cutoff of Section II-C. A delivered link pair stored in a link buffer ages at rate Tcext and is discarded once its fidelity falls to the level Freq (n) at which the end-to-end fidelity after n−1 swaps would still meet Fmin . For a fresh link pair delivered at fidelity F0 = 0.9575 against Freq (n = 4) = 0.9472, the perpair cutoff at Tcext = 1250 µs is approximately 9.2 µs, less than one L = 10 km tick (τ = 50 µs). When the link-buffer cutoff falls below τ , link pairs expire faster than the controller can use them. Sequential’s pipelined assembly cannot then reliably pair a partial chain with a non-expired link pair on the next tick. Simultaneous SWAP-ASAP, which consumes link pairs in the same tick they are pushed and holds no chain storage by construction, is mechanically immune to this failure mode. The mechanism implies a regime-based interpretation of the sequential-vs-simultaneous comparison. Beyond the collapse threshold reported in Section III-B, sequential remains substantially below simultaneous SWAP-ASAP at the next two dimensionless ratios tested (Tcext /τ = 50 and 100), where it delivers 14–54% of the simultaneous SWAP-ASAP rate depending on L. At the relaxed end, the equivalence reference at Tcext ∈ {0.5, 2} s (Appendix B) places both protocols within 0.4% of each other across all topologies tested. In dimensionless units these reference values correspond to Tcext /τ in the range 10,000–80,000 depending on L, well above any plausible crossover. We emphasize that the intermediate range Tcext /τ ∈ (100, 10,000) was not swept in this study. The precise location of the crossover from substantial gap to equivalence is therefore not directly measured, and quantitative localization is left to future work. These results are consistent with a picture in which the connection-less penalty is a near-term phenomenon tied to the limited coherence times of present-day quantum memories rather than a fundamental property of sequential swapping itself. We present this as a regime-based interpretation supported by the present data rather than as a general claim, given the n = 4 scope of the study. As Tcext /τ moves into the equivalence regime established by our reference, sequential’s systems-level advantages should become accessible at progressively lower performance cost. In the present hardware regime, however, the penalty is real and can be substantial: simultaneous SWAP-ASAP’s centralized coordination delivers a viable end-to-end rate in conditions where sequential’s pipeline cannot sustain itself. Protocol selection should therefore be guided by the operating Tcext /τ ratio. The regime structure is governed entirely by Tcext , not by internal communication memory or by the link-layer policy,
suggesting that improvements in the coherence of the networklayer storage tier are the productive direction for closing the connection-less gap. V. C ONCLUSION AND O UTLOOK We have presented a proof-of-principle study of the regime in which connection-less sequential entanglement swapping remains operationally viable. The methodology holds the link layer fixed through a single trained reinforcement-learning policy and varies only the network-layer protocol. The principal empirical result is a coherence-time structure in the dimensionless ratio Tcext /τ : at both link lengths tested, sequential is non-delivering up to Tcext /τ = 25, begins recovering at Tcext /τ = 50, and saturates at the simultaneous rate by the relaxed-coherence reference at Tcext ∈ {0.5, 2} s (Tcext /τ ≈ 10,000–80,000), where the two protocols agree to within 0.4%. Simultaneous SWAP-ASAP delivers a constant rate across the full sweep. The chain-buffer cutoff mechanism explains both sequential’s collapse below the boundary and why memory-side mitigations cannot remove it: the per-pair cutoff falls below one tick before the chain-assembly window does, so link pairs expire faster than the controller can pipeline them. Together with the off-diagonal factorization, this regime structure identifies external memory coherence as the dominant design surface for closing the connection-less penalty within the scope studied. Internal communication memory does not enter, and the link-layer policy does not enter. This narrowness is a useful diagnostic for hardware roadmaps: the same Tcext /τ ratio that separates the two regimes in our simulations is the figure of merit a hardware platform would need to surpass before the systems-level advantages of connectionless, packet-switched operation become accessible without performance compromise. Our relaxed-coherence reference at Tcext ∈ {0.5, 2} s sits at Tcext /τ of 10,000–80,000, comfortably inside the equivalence regime. The stressed regime we report below Tcext /τ ≈ 25 is what near-term quantum-memory platforms in the millisecond-coherence range face today on tens-of-kilometer fiber spans. Several limitations bound the present scope and motivate direct extensions. The regime characterization is performed at n = 4. Sequential’s chain-assembly window scales linearly in n while simultaneous SWAP-ASAP’s collapse remains a single tick. The threshold Tcext /τ should therefore grow with chain length, and quantitative confirmation at larger n would localize the scaling. The link-layer policies were trained at a single delivery-fidelity target F0 = 0.94. The dimensional invariance of Section III-A is established at this fixed F0 and not across it. Larger n would require retraining at a correspondingly larger F0 , and the fidelity margin’s effect on the per-pair cutoff would itself shift the regime boundary. The crossover from substantial gap to equivalence is bracketed by our two measurement regimes but not directly localized; sweeping the intermediate Tcext /τ range would identify the precise location of the transition. A classical-communication budget would refine the practical comparison beyond the idealized synchronization
assumed here, in particular by accounting for the higher signaling-round count for sequential (n−1 versus ⌈log2 n⌉ for simultaneous SWAP-ASAP). Finally, multi-flow scenarios— the setting in which the connection-less architecture’s graceful behavior under contention is most directly relevant—are where the systems-level case for sequential swapping should ultimately be made. A single-flow comparison cannot exhibit the contention dynamics that motivate the architecture in the first place. We view the present study as a first empirical anchor for that broader program, locating the hardware regime in which distributed, packet-switched quantum networks become operationally viable in the single-flow setting we examine. ACKNOWLEDGMENT KPS thanks the U.S. Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program, for support under Award Number DE-SC0026264. KPS, PK, ASW, and ST thank the PQI Community Collaboration Awards. KPS thanks Don Towsley for insightful discussions. The authors used Anthropic’s Claude AI model for language refinement and presentation improvement of the manuscript. R EFERENCES [1] A. K. Ekert, “Quantum cryptography based on Bell’s theorem,” Physical Review Letters, vol. 67, pp. 661–663, 1991. [2] C. H. Bennett and G. Brassard, “Quantum cryptography: Public key distribution and coin tossing,” Theoretical Computer Science, vol. 560, pp. 7–11, 2014. [3] J. I. Cirac, A. Ekert, S. F. Huelga, and C. Macchiavello, “Distributed quantum computation over noisy channels,” Physical Review A, vol. 59, no. 6, p. 4249, 1999. [4] L. Jiang, J. M. Taylor, A. S. Sørensen, and M. D. Lukin, “Distributed quantum computation based on small quantum registers,” Physical Review A, 2007. [5] D. Gottesman, T. Jennewein, and S. Croke, “Longer-baseline telescopes using quantum repeaters,” Physical Review Letters, vol. 109, no. 7, 2012. [6] P. Kómár, E. M. Kessler, M. Bishof, L. Jiang, A. S. Sørensen, J. Ye, and M. D. Lukin, “A quantum network of clocks,” Nature Physics, vol. 10, no. 8, pp. 582–587, 2014. [7] H.-J. Briegel, W. Dür, J. I. Cirac, and P. Zoller, “Quantum repeaters: the role of imperfect local operations in quantum communication,” Physical Review Letters, vol. 81, no. 26, p. 5932, 1998. [8] K. Azuma, S. E. Economou, D. Elkouss, P. Hilaire, L. Jiang, H.-K. Lo, and I. Tzitrin, “Quantum repeaters: From quantum networks to the quantum internet,” Reviews of Modern Physics, vol. 95, no. 4, 2023. [9] L. Bacciottini, A. Chandra, M. G. de Andrade, N. K. Panigrahy, S. Pouryousef, N. S. V. Rao, E. Van Milligen, G. Vardoyan, and D. Towsley, “A packet-switched architecture for two-way quantum networks,” in CLEO 2025, Technical Digest Series. Optica Publishing Group, 2025, paper JPS200 149. [10] S. Haldar et al., “Fast and reliable entanglement distribution with quantum repeaters: Principles for improving protocols using reinforcement learning,” Physical Review Applied, vol. 21, p. 024041, 2024. [11] A. G. Iñesta, G. Vardoyan, L. Scavuzzo, and S. Wehner, “Optimal entanglement distribution policies in homogeneous repeater chains with cutoffs,” npj Quantum Information, vol. 9, p. 46, 2023. [12] M. G. de Andrade, E. A. Van Milligen, L. Bacciottini, A. Chandra, S. Pouryousef, N. K. Panigrahy, G. Vardoyan, and D. Towsley, “On the analysis of quantum repeater chains with sequential swaps,” arXiv preprint arXiv:2405.18252, 2024. [13] S. Pouryousef, H. Shapourian, and D. Towsley, “Analysis of asynchronous protocols for entanglement distribution in quantum networks,” arXiv preprint arXiv:2405.02406, 2024.
[14] G. X. Yau, A. Burushkina, F. Ferreira da Silva, S. Maji, P. S. Thomas, and G. Vardoyan, “Reinforcement learning for quantum network control with application-driven objectives,” arXiv preprint, vol. arXiv:2509.10634, 2025, arXiv:2509.10634 [quant-ph]. [15] M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information. Cambridge University Press, 2010. [16] D. Deutsch, A. Ekert, R. Jozsa, C. Macchiavello, S. Popescu, and A. Sanpera, “Quantum privacy amplification and the security of quantum cryptography over noisy channels,” Physical Review Letters, vol. 77, no. 13, pp. 2818–2821, 1996. [17] H.-K. Lo, “Proof of unconditional security of six-state quantum key distribution scheme,” Quantum Information & Computation, vol. 1, no. 2, pp. 81–94, 2001. [18] D. Bruß, “Optimal eavesdropping in quantum cryptography with six states,” Physical Review Letters, vol. 81, no. 14, pp. 3018–3021, 1998. [19] R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, vol. 8, pp. 229–256, 1992. [20] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations (ICLR), 2015.
A PPENDIX A N ETWORK C ONTROLLER A LGORITHMS This appendix gives the simulation main loop and both network-layer protocols in enough detail to support reimplementation. Notation follows Section II. Per-pair cutoffs: Each entry stores its own cutoff, computed at PUSH time: Freq (m) − 0.25 Tc tcut (F ; m, Tc ) = − 2 log , F − 0.25 1/m
where Freq (m) = 0.25 + 0.75 [(4Fmin − 1)/3] . Link buffers Bℓ use m = n, requiring each delivered link pair to survive long enough to absorb the full n-link fidelity budget. Chain buffers Ci use m = 1: the buffered chain is required only to remain above Fmin at the buffer tier, with the live decoherence-and-swap calculation at consumption time accounting for any further fidelity loss. Both buffer tiers age at Tcext . Entries with tcut ≤ 0 at push time (i.e., F ≤ Freq (m) already at delivery) are rejected immediately rather than buffered. DISCARD E XPIRED(t) drops every entry whose age exceeds its own stored tcut . Buffer entry tuples: Link-buffer entries are (Fℓ , tℓ , tcut ). Chain-buffer entries are (Fch , tsw , told , tcut ), where tsw is the time of the most recent swap that produced the chain and told is the delivery time of its oldest contributing link pair. The two timestamps play distinct roles: tsw is the reference for further decoherence (the chain has already absorbed the swap’s noise at that moment and ages thereafter), while told is the reference for cutoff expiry. Diagnostic fields used only for the figures of Section III are omitted here. POP F RESHEST semantics: For link buffers, POP F RESH EST returns the most-recently-pushed entry (LIFO), which has the largest tℓ since pushes are in delivery-time order. For chain buffers, POP F RESHEST returns the entry with the largest told . Algorithm: Two-layer simulation — main loop and controller Input: n, {Lℓ }, Tsim , controller C
Initialize link agents, link buffers B1 , . . . , Bn (and chain buffers C1 , . . . , Cn−1 if C is sequential) τℓ ← Lℓ /cfiber ; τmin ← minℓ τℓ t ← 0; E ← [ ] while t < Tsim do for each ℓ whose next τℓ tick has been reached: step agent ℓ Bℓ .DISCARD E XPIRED(t) for all ℓ C.STEP(t, B1 , . . . , Bn ); append deliverys to E t ← t + τmin end while return E
Sequential Protocol STEP(t, B1 , . . . , Bn ) for i = 1, . . . , n − 1 do Ci .DISCARD E XPIRED(t) end for for ℓ = 2, . . . , n do while Cℓ−1 and Bℓ both non-empty do (Fch , tsw , told , ·) ← Cℓ−1 .POP F RESHEST() if t − told > entry’s tcut then continue ▷ chain expired end if (Fℓ , tℓ , ·) ← Bℓ .POP F RESHEST() if t − tℓ > entry’s tcut then re-push the original chain entry to Cℓ−1 continue end if ′ Fch ← D(Fch , t − tsw , Tcext ) Fℓ′ ← D(Fℓ , t − tℓ , Tcext ) ′ , Fℓ′ ) Fnew ← Fswap (Fch ′ told ← min(told , tℓ ) if ℓ = n then emit (Fnew , t − t′old ) else ▷ new tsw = t Cℓ .PUSH(Fnew , t, t′old ) end if end while end for ▷ seed C1 from non-expired link-1 pairs: while B1 non-empty do (F1 , t1 , ·) ← B1 .POP F RESHEST() if t − t1 ≤ entry’s tcut then C1 .PUSH(F1 , t1 , t1 ) ▷ tsw = told = t1 end if end while
Simultaneous SWAP-ASAP Protocol STEP(t, B1 , . . . , Bn ) if any Bℓ is empty then return end if F ← [ ]; T ← [ ] for ℓ = 1, . . . , n do pop and discard expired entries from Bℓ until a non-expired pair (Fℓ , tℓ , ·) is found, or Bℓ becomes empty if Bℓ is empty then return end if F.APPEND(D(Fℓ , t − tℓ , Tcext )) T .APPEND(tℓ ) end for emit S IM . SWAP-ASAP(F), t − min T S IM . SWAP-ASAP([F1 , . . . , Fm ]):
if m = 1 then return F1 end if mid ← ⌊m/2⌋ FL ← S IM . SWAP-ASAP([F1 , . . . , Fmid ]) FR ← S IM . SWAP-ASAP([Fmid+1 , . . . , Fm ]) return Fswap (FL , FR )
A PPENDIX B E QUIVALENCE AT R ELAXED C OHERENCE Section III-B of the main text reports that sequential swapping and simultaneous SWAP-ASAP are statistically equivalent at relaxed coherence, Tcext ∈ {0.5, 2.0} s, across all topologies tested. We report the underlying data here. Figure 7 shows uSKR for the two symmetric topologies [5, 5, 5, 5] km and [10, 10, 10, 10] km at Tcext = 2 s. Sequential delivers 1164 bps and simultaneous SWAP-ASAP 1165 bps at L = 5. Both deliver 255 bps at L = 10. The relative difference is below 0.1% at both link lengths.
Fig. 8. uSKR vs. bottleneck position for chains containing one L = 10 km link in an otherwise L = 5 km chain. Full color: Tcext = 2 s. Faded: Tcext = 0.5 s. Sequential and simultaneous SWAP-ASAP overlap within statistical noise at every position. Fig. 7. uSKR at Tcext = 2 s for the two symmetric topologies. Filled bars: simultaneous SWAP-ASAP. Hatched bars: sequential.
Figure 8 reports the same comparison for the four bottleneck topologies (one L = 10 km link at each of positions p ∈ {1, 2, 3, 4} in an otherwise L = 5 km chain), at both Tcext = 2 s and Tcext = 0.5 s. Sequential and simultaneous SWAP-ASAP track each other position-by-position to within the 95% confidence interval at every position and both coherence values, with the relative gap remaining below 0.4% everywhere. There is mild position-dependence in absolute efficiency, ranging across p ∈ {1, 2, 3, 4}, but this dependence is shared between the two protocols. A four-fold reduction from Tcext = 2 s to 0.5 s leaves the comparison essentially unchanged. The collapse documented in the main text emerges only at three to four orders of magnitude lower, in the microsecond range where the per-pair cutoff mechanism of Section II-C becomes operative.