Research Article
1
A Wavelength Borrowing Architecture for Optical Data Center Networks - Extended Version A NDREA D ETTI1 , C HIARA L ODOVISI1 , AND S ILVELLO B ETTI1 1 CNIT - University of Rome “Tor Vergata”, Electronic Engineering Dept., Italy
arXiv:2609.04874v1 [cs.NI] 4 Sep 2026
Compiled September 7, 2026
The growth of east-west traffic, along with the cost and power consumption of electronic switching, is motivating the integration of a low-power, high-rate, all-optical layer within the data center network. This paper presents a spine-leaf all-optical architecture in which the default wavelength configuration, one wavelength per source–destination leaf pair, can be reconfigured to accommodate unbalanced traffic demand: wavelengths that are unused or lightly loaded at one leaf are borrowed by another leaf with higher demand. This topology-engineering capability is combined with a traffic-engineering scheme, based on two-hop detouring, enabling the control of wavelength load while limiting the amount of detoured traffic. A key feature of the architecture is that its degree of wavelength reconfigurability is set by a single tunable parameter, the borrowing degree B, ranging from none to full; performance evaluation shows that near-optimal performance is achieved well below the maximum B, saving the complexity and cost of fully reconfigurable solutions. Furthermore, the optical fabric relies on mature, data-center-grade components, namely AWGs, AWGRs, colorless OXCs, and combiners, whose reconfiguration speed makes the architecture deployable at network tiers where traffic demand persists over seconds or longer, e.g., among groups of racks (pods). The architecture is also TDMA-transparent, a property that future work could exploit to refine the borrowing unit below a whole wavelength without changing the optical fabric. http://dx.doi.org/10.1364/ao.XX.XXXXXX
1. INTRODUCTION East-west traffic, from server-to-server or rack-to-rack, dominates volume in modern data centers [1], driven by distributed applications, storage replication, and increasingly by large-scale distributed training and inference workloads that demand high bandwidth between servers and racks. Spine-leaf network architectures [2] are commonly used to support this need. A leaf node is an L3 Ethernet switch: some of its ports serve either a single rack, acting as a Top-of-Rack (ToR) switch, or a group of ToR switches forming a Pod, while the remaining ports connect to a bank of spine L3 switches, providing leaf-to-leaf connectivity. Increasing the leaf-to-leaf bandwidth requires scaling the spine-leaf segment. When feasible, this can be achieved by raising the Ethernet line rate (e.g. from 100 Gbps to 400 Gbps); such an upgrade keeps the fiber plant untouched but requires replacing switches, or at least their optical transceivers, at every leaf and spine switch. Alternatively, capacity can be added by deploying more parallel leaf-to-spine fibers or additional spine switches, but this entails rewiring and possibly a hardware upgrade whenever free ports are unavailable on the switches. To simplify bandwidth scaling and curb the cost and power consumption of spine electronic switches and transceivers,
cloud hyperscalers and academia are therefore exploring the replacement of the electronic spine layer with an all-optical one (see [3, 4] for surveys). Leaf-to-leaf traffic is carried end-to-end over wavelengths routed by the spine layer with no intermediate opto-electronic conversion [4–7]. An optical spine cuts power consumption, since no buffering or electronic processing is needed, and is transparent to data format and rate, so migrating to a newer Ethernet generation only requires upgrading the leaf transceivers, while the optical spine remains untouched. Since optical switching usually does not support buffering, a wavelength used by a leaf to receive traffic can be used at the same time by only one source leaf with no in-network resource contention. This raises an end-to-end wavelength assignment problem that adapts the wavelength provisioning to traffic demand. For finer-grained resource sharing, the same receiving wavelength can be shared among different sources with time division multiple access (TDMA), adding a further time-slot dimension to the optimization problem [8]. The wavelength, and optionally time-slot, assignment strategy can be regarded as a topology engineering problem, which can nonetheless be coupled with a traffic engineering one: how to route traffic on top of the optical topology, possibly accepting intermediate opto-electronic conversion [9]. For instance, if the topology engineering solution assigns no wavelength between
Research Article
leaves j and d connected to the optical spine, traffic from j to d must instead be detoured through an intermediate leaf i, which has wavelengths towards both j and d, thus forming a two-hop j → i → d path 1 . Most spine optical fabrics proposed in the literature, however, are designed for either no reconfigurability [11] or full wavelength-level reconfigurability, where each receiver wavelength can be assigned to any source, and/or rely on wavelengthselective switches (WSS) or in-network optical signal processing [4, 9, 12]. Optical signal processing remains laboratory-grade technology, while WSS devices are currently expensive, complex, and limited in port count, compared to simpler devices, with no optical processing, such as the colorless Optical Cross-Connect (OxC) or the Arrayed Waveguide Grating Router (AWGR)— technology already deployed in real data centers [5, 11] or experimentally demonstrated at scale [6]. In this paper we propose a modular architecture for an optical spine layer that closes all three gaps: it dispenses with optical signal processing, avoiding laboratory-grade technology; it dispenses with WSS, avoiding their cost, complexity, and port-count limits; and it replaces the all-or-nothing reconfigurability choice with a reconfiguration capability, and hence system cost/complexity, that can be scaled out gradually by adding optical devices only where needed. The baseline configuration provides a single default wavelength between any leaf pair. To adapt wavelength assignment to traffic demand, a leaf with idle capacity toward a given destination—the donor—can lend its default wavelength to another leaf that needs additional capacity toward that same destination—the borrower. Any residual traffic that the donor still has toward the destination is optimally split, via traffic engineering, across two-hop detours through other leaves. The complexity-reconfigurability tradeoff of the architecture is governed by a single borrowing degree B: the number of donors a leaf can concurrently borrow from, and symmetrically the number of borrowers a leaf can concurrently lend to, is at most B − 1. This parameter determines the wavelength reconfiguration capability of the architecture, along with its complexity and cost. A fully reconfigurable architecture, where any receiver wavelength can be assigned to any source, would require B = L, where L is the number of leaf nodes. However, we show that full reconfigurability is not necessary to achieve the desired performance objective: the borrowing architecture can instead be tuned in a cost-adaptive way. Overall, the contributions of this paper are threefold: • We propose an optical spine-leaf architecture whose reconfigurability level, hence cost and complexity, can be tuned to fit data center needs, relying only on commercially mature technology—AWGR, AWG multiplexers/demultiplexers, passive optical combiners, and a colorless OxC. The architecture is also TDMA transparent, supporting a subsequent introduction of time domain for finer resource allocation. Its configuration is driven by an SDN controller targeting traffic demand that persists over timescales of seconds or longer, positioning the architecture among slow, coarsegrained optical switching solutions rather than fast, perpacket ones. 1 Orthogonal to both topology and traffic engineering is a third optimization axis, placement engineering, which we do not consider in this paper. Rather than adapting the network to the traffic, it acts upstream, at the traffic-source level, e.g., through traffic-aware virtual machine placement, to reduce the load that the network must carry in the first place [10].
2
Spine Optical Switching Fabric
SDN Controller
LEAF node 1
To R
To R
LEAF node L
To R
To R
Rack Pod 1
To R
To R
Rack Pod L
Fig. 1. Architecture of the proposed wavelength borrowing data
center • We model the architecture limits as a set of optical and electronic mixed-integer linear-programming (MILP) constraints, paving the way for any related topology/traffic engineering optimizations. • Among the many possible ones, we focus on a specific optimization objective: keeping the load of every wavelength below a given threshold while allowing two-hop detoured traffic, but of minimum necessary volume. We propose a greedy heuristic that jointly selects the wavelengthborrowing configuration (topology engineering) and the detouring fractions (traffic engineering) to this end. The remainder of the paper is organized as follows. Section 2 describes the proposed wavelength-borrowing architecture in detail. Section 3 formulates the wavelength assignment and twohop detouring problem. Section 4 presents the greedy heuristic. Section 5 uses a Python simulator to compare the proposed architecture and topology/traffic engineering algorithm against simple solutions representative of a static, non-borrowing optical core, with and without traffic detouring. Finally, Section 6 discusses related work.
2. ARCHITECTURE DESCRIPTION A. Overview
As shown in Fig. 1, the proposed architecture comprises L rack groups/pods, each connected to a leaf node through a Top-ofRack (ToR) switch over standard Ethernet. Inter-leaf traffic is carried over all-optical circuits using Wavelength Division Multiplexing (WDM) with W wavelengths, and, optionally, Time Division Multiple Access (TDMA) [8]. A spine optical switching fabric routes these circuits, while an SDN controller manages their configuration [13]. The reconfiguration timescale is primarily limited by the switching time of a colorless OxC in the spine fabric, and thus falls in the range of tens of milliseconds for MEMS-based OxC [11]. This positions the proposed solution at a coarser granularity than packet switching, targeting traffic demand that persists over timescales of seconds or longer [5]. Under default operation, each leaf node uses one default wavelength per destination, yielding a fully balanced allocation of optical resources across all leaves. The key innovation
Research Article
is a dynamic wavelength borrowing mechanism: a leaf with idle capacity—the donor—lends its unused default wavelengths to other leaves—the borrowers—increasing each borrower’s instantaneous bandwidth toward a specific destination. For example, leaf 1 and leaf L each have a default wavelength, λ2 and λ1 respectively, to reach leaf 2. When leaf 1 has little or no traffic toward leaf 2, it can lend λ2 to leaf L, expanding the latter’s available wavelengths toward leaf 2 to {λ1 , λ2 }. The degree of reconfigurability is governed by a parameter B termed the borrowing degree. Specifically, the number of leaves from which a leaf may concurrently borrow or lend resources is at most B − 1. Increasing B improves wavelength allocation flexibility at the cost of additional optical hardware. Optionally enabling TDMA reduces borrowing granularity from a full wavelength to individual time slots, enabling finer-grained matching of traffic demand at the expense of increased hardware complexity. The architecture thus offers a tunable trade-off between hardware complexity/cost and resource allocation flexibility. Fig. 2 shows the optical components implementing the transmitting (left) and receiving (right) functionalities of the leaves, together with the interconnecting spine optical fabric. Each block is described in the following subsections, covering components and wiring first, followed by data transfer operations. For simplicity, the description focuses on the case where the number of leaves equals the number of wavelengths, i.e., L = W. Appendix I in [14] extends the discussion to the more general case. B. Components and Wiring B.1. Leaf Nodes
For data transmission, each leaf node contains B copies of a transmission module, each composed by a bank of W fixed-wavelength lasers (λ1 , . . . , λW ) feeding a W ×1 AWG multiplexer whose output fiber connects to the spine optical fabric. The first module of a leaf i (green in the figure) is the default one: its default fiber connects directly to the i-th combiner, bypassing the spine OxC, and its W wavelengths are the default ones toward the remote leaves, one per leaf, switched off and lent to borrowing leaves as needed. The remaining B − 1 borrowing modules (light-red in the figure) each connect to the spine OxC via a borrowing fiber and transmit over one or more wavelengths borrowed from a single donor leaf; since wavelengths from different donors require different modules, at most B − 1 donors can be used concurrently. Any laser within a borrowing module can be activated on demand, according to the wavelength assignment strategy implemented by the SDN controller 2 . Finally, an ingress SDN-controlled load balancer routes outgoing traffic to the buffers drained by the lasers of the different transmission modules, following a specific traffic engineering strategy. Transmitting operations.
For data reception, each leaf node is connected to the spine optical fabric via a single input fiber carrying W wavelengths. An AWG demultiplexer separates each wavelength onto a dedicated fiber; the transported bit stream is then converted to the electronic domain by a dedicated WDM receiver for subsequent packet forwarding, either to the final rack or to the next-hop leaf in case of detouring. Reception operations.
2 To reduce the number of lasers of a borrowing TX module, a limited set of tunable lasers may be used and connected opportunistically to AWGs through a local OxC configured by the SDN controller. This would impose an additional optical constraint on the wavelength assignment problem
3
Table 1. Wavelength mapping per output port for a 4 × 4 cyclic
AWGR Output port
Input wavelength and port (λn,i )
1
λ1,1 , λ4,2 , λ3,3 , λ2,4
2
λ2,1 , λ1,2 , λ4,3 , λ3,4
3
λ3,1 , λ2,2 , λ1,3 , λ4,4
4
λ4,1 , λ3,2 , λ2,3 , λ1,4
B.2. Spine Optical Switching Fabric
The spine optical switching fabric consists of three elements: a colorless Optical Cross-Connect (OxC), a bank of combiners, and a cyclic Arrayed Waveguide Grating Router (AWGR). The AWGR is a fully passive component that routes each wavelength λn arriving at input port i to a deterministic output port d according to the cyclic routing rule: Cyclic AWGR.
d = (n − i ) mod W + 1,
(1)
Tab. 1 illustrates such cyclic routing for W = 4, where λn,i denotes wavelength λn arriving on input fiber i. Each AWGR output port is connected to a specific destination leaf: output port d is connected to leaf d, and each AWGR input port is connected to a dedicated combiner. The AWGR can be realized as a single device or replaced by Sato’s cascaded small cyclic AWG architecture [15], which synthesizes a KM ×KM equivalent switch from M copies of K ×K AWGs and K copies of M × M AWGs, with K and M mutually coprime integers. A multi-stage Thin-CLOS wavelength-routing fabric built from smaller AWGRs offers an alternative, experimentally demonstrated scale-out solution [16]. Combiner i, connected to AWGR port i, merges optical signals from B fibers: i) the default fiber from AWG1 of node i, and ii) a group of B − 1 fibers arriving from the OxC, each carrying the wavelengths of a distinct borrowing fiber. The combiner size B limits a donor node to serving at most B − 1 borrowers simultaneously3 . The wavelength sets carried by the B input fibers of any combiner are guaranteed to be disjoint by the wavelength assignment strategy. Accordingly, the combiner can be implemented as a passive coupler, resulting in a simple, fully passive design that is transparent to TDMA operation, but incurring an intrinsic optical loss of 10 log10 ( B) dB, which must be compensated by a shared optical amplifier at the combiner output 4 . Combining stage.
The OxC has size ( B − 1)W × ( B − 1)W and routes the wavelengths of borrowing fibers to combiners as configured by the SDN controller. The OxC performs purely spatial switching with no wavelength awareness. Currently, MEMS technology is a valuable choice for the OxC implementation, as it can achieve the required port count (e.g., on the order of hundreds) with acceptable insertion loss and switching time [11]. Colorless OxC.
3 The number of borrowing TX modules per node and the number of borrowing input fibers per combiner are both equal to B − 1. Relaxing this equality by introducing two distinct parameters—BTX for the TX module count and BCOMB for the combiner size—yields an asymmetric resource relocation constraint: a borrower may use at most BTX − 1 donors, while a donor may lend resources to at most BCOMB − 1 borrowers. 4 Another possible implementation uses a WSS in combiner mode with nearzero combining loss, but requires active control coordinated with the OxC and incurs higher cost. Furthermore, when TDMA is enabled, the WSS must support time-slot-level reconfigurability, which may pose technological challenges.
Research Article
4
LEAF 1 (TX) default TX module
AWG1 MUX Wx1
...
default fiber
from racks
SDN Routing and Load Balancing
borrowing TX modules
LEAF 1 (RX)
Spine Optical Switching Fabric AWG2 MUX Wx1
...
Combiner 1
to racks
AWG DEMUX 1xW
port 1 AWGB MUX
port B-1
B-1 borrowing fibers
Colorless OxC
B-1 borrowing fibers
port 1
SDN
Cyclic AWGR
(B-1)W x (B-1)W
...
LEAF W (RX)
WxW
LEAF W (TX) default TX module
port W
port (W-1)(B-1)+1 1xW
AWG1 MUX Wx1
...
to racks
AWG DEMUX
port W(B-1)
...
Combiner W
from racks
SDN Routing and Load Balancing
borrowing TX modules
AWG2 MUX Wx1
...
AWGB MUX
Fig. 2. Schematic of the proposed wavelength borrowing architecture.
C. Leaf-to-Leaf Data Transfer C.1. Fully-Balanced Configuration
Under a fully-balanced traffic pattern, no wavelength borrowing takes place, and each leaf uses only its default TX module to simultaneously reach all W destinations, one default wavelength per destination. For instance, in Fig. 3, leaf 1 and leaf W use their default wavelengths λ2 and λ1 to reach leaf 2, respectively5 . C.2. Unbalanced Configuration without TDMA
During an unbalanced traffic configuration, a leaf i can require extra bandwidth toward a destination d, while another leaf j has its default wavelength toward d idle or lightly loaded. Two approaches can handle this traffic variation. The first is an electronic-only approach based on traffic detouring [5]: the excess traffic from the overloaded leaf i toward d is rerouted over a two-hop path through leaf j (or more than one), processed electronically there, and then forwarded on that leaf’s unloaded default wavelength to d. The second is an optical–electronic hybrid approach based on wavelength borrowing and detouring: the underloaded leaf j lends its default wavelength toward d to i; the borrower leaf i activates the laser of the borrowed wavelength on a borrowing TX module dedicated to wavelengths 5 Note that in the general configuration of Fig. 2 a leaf has a default wavelength toward itself that carries no traffic and can therefore always be borrowed
borrowed from donor j, and the OxC routes the corresponding borrowing fiber to the combiner of donor leaf j. Any residual traffic from donor leaf j toward d is detoured through other leaves that still have an active default wavelength to d. For instance, in Fig. 4, leaf 1 lends its default wavelength λ2 to leaf W. Leaf W uses one of its borrowing TX modules to transmit data on the borrowed wavelength λ2 . The OxC routes λ2 from the borrowing TX module of leaf W to the donor combiner 1, and the AWGR then routes λ2 to leaf 2. Consequently, leaf W can use two wavelengths to reach destination leaf 2. Specifically, the borrowing operation is managed by the SDN controller as follows: 1. The SDN controller detects that the default wavelength λb of a leaf j toward destination leaf d is underutilized, and that leaf i is congesting its wavelengths toward the same destination. 2. The controller checks the feasibility of donor j lending λb to borrower i and evaluates its potential benefit with respect to a specific optimization objective. 3. When borrowing is feasible and convenient: (a) the controller activates the λb laser on the borrowing TX module of i dedicated to donor j, configures
Research Article
5
LEAF 1 (TX)
3. WAVELENGTH ASSIGNMENT AND TRAFFIC DETOURING
comb. 1
default TX module
borrowing TX module n. b
OxC
LEAF 2 (RX) AWGR
LEAF W (TX) default TX module
borrowing TX module n. b
comb. W
The wavelength borrowing architecture can be dynamically controlled to achieve different optimization goals. In this paper, we focus on the non-TDMA case and consider the base architecture in Fig. 2 with a number of leaves equal to the number of wavelengths, i.e., L = W. The resource allocation problem for TDMA-based solutions, as well as the architectural extensions in Appendix I in [14], are left for future work. The following subsections first derive the MILP constraints defining the feasible region for any optimization problem within the wavelength borrowing framework with two-hop detouring, and then present our specific optimization problem.
Fig. 3. Fully balanced configuration, no borrowing. Leaf 1 and
leaf W reach leaf 2 with their default wavelengths λ2 and λ1 , respectively.
LEAF 1 (TX)
comb. 1
default TX module
borrowing TX module n. b
OxC
LEAF 2 (RX) AWGR
LEAF W (TX) default TX module
borrowing TX module n. b
A. Variables and Constraints E denote the end-to-end traffic generated by racks served by Let Ai,d leaf i and directed to racks of leaf d, normalized to the bitrate of E = 1 means a traffic bitrate equal to the one wavelength, i.e., Ai,d wavelength one; we collect these entries into the traffic matrix E }. The borrowing configuration is represented by the A E = { Ai,d binary matrix b = {bi,j,d }, where an entry equal to 1 indicates that leaf i borrows the default wavelength used by leaf j to reach destination leaf d. The traffic detouring configuration is represented by the matrix w = {w j,i,d }, whose entries denote the fraction of end-to-end traffic A Ej,d that is electronically detoured through node i.
We define the following integer variables related to the optical architecture in Fig. 2: ! Wavelength assignment constraints.
comb. W
Fig. 4. Unbalanced configuration with wavelength borrowing. Leaf 1 lends the default wavelength λ2 to leaf W. The resulting optical capacity from leaf W to leaf 2 comprises λ2 and λ1 .
xi,j = min 1, ∑ bi,j,d
(2)
di,d = 1 − ∑ b j,i,d
(3)
ci,d = di,d + ∑ bi,j,d
(4)
d
j
the OxC to route the related output borrowing fiber to combiner j, and switches off λb on the default TX module of donor j. Destination leaf d now receives wavelength λb from leaf i rather than leaf j; (b) the controller reconfigures the load balancer of leaf i to distribute traffic i → d across default and borrowed wavelengths, and the load balancer of leaf j to detour traffic to d only through intermediate leaves providing a two-hop paths from j to d. 4. If wavelength borrowing is not feasible or not convenient, the SDN controller may still reduce the load on overloaded leaf i by detouring its excess traffic through underloaded leaves providing a two-hop path from i to d. C.3. Unbalanced Configuration with TDMA
The borrowing architecture is TDMA-transparent, requiring changes only to the transmitting lasers and WDM receivers. With TDMA, the SDN controller performs the same operations as before, but the donor’s and borrower’s λb lasers can now remain simultaneously active, transmitting in different time slots whose allocation the controller sets to best match traffic demand. On the receiving side, WDM receivers must operate in burst mode, recovering clock synchronization slot by slot and incurring a preamble overhead per time slot [6].
j
where xi,j is a binary variable indicating that leaf i borrows at least one wavelength from leaf j; di,d indicates that leaf i retains its default wavelength towards destination leaf d; and ci,d denotes the total number of wavelengths available from leaf i to destination leaf d, collected into the capacity matrix c = {ci,d }. The hardware limit of the architecture imposes that a wavelength assignment resulting from the borrowing configuration b is feasible only if it satisfies the following constraints:
∑ xi,j < B
∀j
(5)
∑ xi,j < B
∀i
(6)
∑ bi,j,d ≤ 1
∀ j, d
(7)
∑ bj,k,d = 0
∀ j, d : ∑ bi,j,d = 1
i
j
i
k
ri,j = min 1, ci,j ,
(8)
i
ri,d + ∑ ri,j r j,d > 0
∀i, d : i ̸= d
(9)
j
The constraints Eq. (5)–Eq. (6) reflect the limit of B − 1 TX modules and combiner ports available per leaf for wavelength borrowing; constraint Eq. (7) ensures that the default wavelength of leaf j towards destination d is borrowed by at most one leaf;
Research Article
6
constraint Eq. (8) ensures that a leaf j lending its default wavelength towards destination d may not simultaneously borrow any wavelength towards the same destination. Finally, ri,j indicates whether leaf i has at least one wavelength (default or borrowed) towards leaf j, and Eq. (9) enforces that every ordered source–destination pair of distinct leaves (i, d) remains connected within at most two hops: either i has a direct wavelength to d (ri,d = 1), or there exists at least one intermediate leaf j with ri,j = 1 and r j,d = 1, thereby guaranteeing full leaf-to-leaf connectivity in at most two hops. A traffic detouring solution w is subject to the following constraints: Traffic detouring constraints.
∑ w j,i,d ≤ 1
∀ j, d
(10)
H0 E Ai,d = 1 − ∑ wi,j,d Ai,d
(16)
H1 E Ai,d = ∑ wi,d,j Ai,j
(17)
j
j
H2 Ai,d = ∑ w j,i,d A Ej,d
∀ j, i, d
(11)
w j,i,d = 0
∀ j, d, i : ci,d = 0 or c j,i = 0
(12)
H0 A Ej,d = A j,d + AD j,d
∀ j, d
(13)
Constraint Eq. (10) ensures that the total detoured fraction of any source’s traffic does not exceed unity; constraint Eq. (11) requires non-negative detouring fractions; constraint Eq. (12) restricts detoured traffic to be forwarded only through leaves with available two-hop connectivity; and constraint Eq. (13) is the traffic conservation condition for any pair ( j, d), requiring that H0 and the detoured traffic the non-detoured (direct) traffic A j,d E AD j,d together equal the total end-to-end traffic A j,d .
B. Optimization objective
We formulate a single objective: minimize the total electronically detoured traffic T D while ensuring that the traffic load ρi,d of wavelengths between any source–destination pair (i, d) is below a given threshold ρth . For instance, ρth = 0.9 implies that the average traffic offered to wavelengths of any pair (i, d) is lower than 90% of the maximum wavelengths’ bitrate. The optimization thus determines the borrowing configuration b, which is binary, and the detouring fractions w, which are real-valued, minimizing T D under the wavelength assignment and detouring constraints, together with the load cap condition Eq. (15). The resulting problem is an MILP. min b, w
T D = ∑ w j,i,d A Ej,d
(14)
j,i,d
s.t. Eq. (5)–Eq. (8), Eq. (10)–Eq. (12), ρi,d ≤ ρth
∀i, d
(15)
The load ρi,d is the ratio of offered traffic to wavelength capacity of a pair (i, d) and can be computed as follows. The ci,d wavelengths between a pair (i, d) support three types of traffic: H0 : the portion of end-to-end traffic A E forwarded • direct Ai,d i,d by leaf i to destination d without detouring; H1 : the portion of end-to-end traffic from i to • local-detoured Ai,d any destination j ̸= d detoured via d (first-hop detouring); H2 : the portion of traffic from any leaf • remote-detoured Ai,d j ̸= i detoured via i to reach leaf d (second-hop detouring). This traffic is subject to a possible packet loss rate Pj,i on the ( j, i ) wavelengths (first-hop loss)6 . 6 For the loss rate P, we consider a fluidic model in which the loss volume is simply equal to the amount of traffic exceeding the optical capacity ci,d
1 − Pj,i
(18)
j
Pi,d =
H0 + A H1 + A H2 − c , 0 max Ai,d i,d i,d i,d H0 + A H1 + A H2 Ai,d i,d i,d
(19)
The resulting load on source–destination pair (i, d) is:
i
w j,i,d ≥ 0
ρi,d =
H0 + A H1 + A H2 Ai,d i,d i,d
ρi,d = 0,
ci,d
,
∀i, d : ci,d ≥ 1,
∀i, d : ci,d = 0.
(20) (21)
These entries are collected into the load matrix ρ = {ρi,d }.
4. HEURISTIC RESOURCE ALLOCATION The joint optimization problem formulated in Section 3 couples a combinatorial selection of the borrowing variables b with a continuous allocation of the detouring fractions w, and is therefore NP-hard; exact methods become computationally intractable for network sizes of practical interest. For this reason we developed a heuristic algorithm that decomposes the problem into three phases, each of which relies on the same traffic engineering algorithm, called two-hop water-filling (2HWF), described first. As this paper aims to provide initial results on the complexityreconfigurability tradeoff enabled by the borrowing architecture, we leave a formal analysis of the heuristic’s computational complexity and optimality gap to future work. We simply note that, for the largest scenario considered – 64 leaves and B = 16 – a raw Python implementation running on 2019 i9 Intel Macbook completed in approximately 90 s, and that modern CPU hardware together with a compiled-language implementation can be expected to substantially reduce this processing time. A. Two-hop water filling
Water-filling is a well-known algorithm that distributes an amount of “water” among a set of connected “recipients”, minimizing the maximum final level of water among the recipients [17]. We used a variation of this policy to evaluate the best detouring fractions w for a fixed borrowing b configuration and resulting capacities c. In our context, the water amount is the traffic A D j,d to detour from j to d, i.e., the end-to-end traffic A Ej,d left after subtracting H0 . Coherently with our optimization objecthe direct traffic A j,d
tive, the direct traffic is the maximum portion of A Ej,d that keeps the pair’s load within ρth , thus maximizing traffic served directly – and hence minimizing traffic to detour – while respecting the load constraint. Specifically, E H0 AD j,d = A j,d − A j,d ,
H0 A j,d = min A Ej,d , ρth c j,d
(22)
The recipients are the set of possible two-hop paths ( j, i, d) from j to d, whose normalized level of contained water is the two-hop load ρ j,i,d defined as: ρ j,i,d = max(ρ j,i , ρi,d )
(23)
Research Article
7
Algorithm 1. Two-Hops Water-Filling (2HWF) Detouring
Require: b, A E , ρth Ensure: w, ρ, ∆ρmax , T D 1: Compute ci,d from b with Eq. (4) ∀i, d H0 2: Compute A D j,d and A j,d with Eq. (22) ∀ j, d H0 3: ρold j,d = A j,d /c j,d if c j,d > 0, else 0 ∀ j, d 4: w j,i,d = 0 ∀ j, i, d two-hop path n.1
two-hop path n.2
two-hop path n.3
5: S = {( j, d ) : A D j,d > 0}
Fig. 5. Two-hop water filling
8: 9:
The two-hop water-filling algorithm (2HWF) computes the entire detouring matrix w by sequentially distributing, for each pair ( j, d), the traffic to detour A D j,d across the available two-hop paths, minimizing the resulting increase in the maximum twohop load and thereby keeping the system as far as possible from the load constraint in Eq. (15). Fig. 5 illustrates the underlying idea of the algorithm for a single pair ( j, d) with three possible two-hop paths ( j, i, d). For each path, the two boxes show the load ρ j,i and ρi,d of first and second hop before and after the injection of detoured traffic, marked as “old” and “new”, respectively. Starting from the old (pre-detouring) load of each pair, the algorithm searches for the minimum new two-hop load level θ such that raising the paths’ two-hop loads to θ absorbs exactly the traffic to detour A D j,d . In the figure, the solution detours traffic only along paths 1 and 2: path 3 is already more loaded than θ, so the water cannot fill that recipient7 . Formally, for a pair ( j, d) for which A D j,d > 0, Eq. (24) defines the set R j,d of intermediate leaves i providing two-hop connectivity j → i → d. Using a level θ, the path ( j, i, d) absorbs an amount of traffic Fθ,j,i,d , given by Eq. (25), and the two-hop paths in R j,d together absorb a total traffic Fθ,j,d , given by Eq. (26) (see Appendix II in [14]). The final level θ is then found by solving Fθ,j,d = A D j,d , as in Eq. (27), after which the detouring fraction w j,i,d follows directly from Eq. (28). n o R j,d = i : i ̸= j, d, c j,i > 0 and ci,d > 0 (24) old Fθ,j,i,d = min max 0, (θ − ρold j,i ) c j,i , max 0, ( θ − ρi,d ) ci,d (25) Fθ,j,d =
∑ Fθ,j,i,d
(26)
i ∈ R j,d
θ : Fθ,j,d = A D j,d w j,i,d =
Fθ,j,i,d A Ej,d
▷ leaf pairs with detouring traffic
6: Sort S by A D j,d in decreasing order 7: for all ( j, d ) ∈ S, in sorted order do
(27) (28)
Algorithm 1 summarizes the overall procedure to distribute the whole traffic to detour, i.e., for every pair ( j, d) with A D j,d > 0. The algorithm also returns the maximum overload parameter ∆ρmax , representing the maximum difference greater than zero between any load ρ j,d and the load threshold ρth . 7 Because pairs can have a different number of wavelengths, the same traffic amount can impact their load ρ differently, as for the two links of the same path in the figure.
10: 11: 12: 13:
Compute R j,d from Eq. (24) if R j,d == ∅ then continue Find water level θ solving Eq. (27) for all i ∈ R j,d do Compute Fθ,j,i,d from Eq. (25) old ρnew j,i = ( ρ j,i c j,i + Fθ,j,i,d ) /c j,i
▷ new first-hop load
old 14: ρnew ▷ new second-hop load i,d = ( ρi,d ci,d + Fθ,j,i,d ) /ci,d 15: w j,i,d = Fθ,j,i,d /A Ej,d new ∀ j, i 16: ρold j,i = ρ j,i 17: ∆ρmax = max(0, maxi,d ( ρnew i,d − ρth )) D E 18: T = ∑ j,i,d w j,i,d A j,d 19: return w, ρ, ∆ρmax , T D
B. Greedy wavelength assignment and traffic detouring
Algorithm 2 presents the whole heuristic algorithm we use to compute the borrowing b and detouring w configurations. The algorithm is organized in three phases. Starting from the no-borrowing state bi,j,d = 0, the algorithm computes the initial detouring fractions w and the corresponding load matrix ρ, maximum overload ∆ρmax and T D via 2HWF. Phase 1: Initial water filling
The algorithm then iterates as follows. At each round, it identifies candidate borrowing triples (i, j, d) such that: there is traffic to detour on the pair (i, d); leaf i is not already a donor towards d (ci,d ≥ 1); and leaf j retains only its default wavelength towards d (c j,d == 1) and can therefore D − A H0 – the traflend it. Candidates are ranked by the score Ai,d j,d fic currently detoured by i towards d minus the direct traffic that leaf j would have to detour after lending its default wavelength towards d – which estimates the maximum achievable reduction in detouring traffic. Candidates are then tested in ranked order. If activating bi,j,d = 1 would violate any of the optical constraints Eq. (5)– Eq. (9), the candidate is discarded and permanently excluded from future rounds as inserted in the skip set K. Otherwise, the activation is applied tentatively, and the resulting values ′ w′ , ∆ρmax , and T ′ D are computed by 2HWF. The borrowing bi,j,d = 1 is confirmed, and the round restarts from the first step, if it strictly reduces ∆ρmax , or leaves ∆ρmax unchanged while strictly reducing T D . Otherwise, the triple is marked as rejected and permanently excluded from future rounds, and the next candidate in the ranking is tried. This acceptance rule first drives ∆ρmax toward zero, thereby satisfying the load cap in Eq. (15), and then reduces the detoured traffic volume T D . The phase terminates when a round examines Phase 2: Greedy borrowing
Research Article
Algorithm 2. Greedy borrowing and 2HWF detouring
Require: A E , B, ρth Ensure: b, w 1: bi,j,d = 0 ∀i, j, d 2: w, ρ, ∆ρmax , T D ← 2HWF ( b ) ▷ Phase 1: initial water filling 3: K = ∅ ▷ Set of (i, j, d) triples infeasible or unprofitable 4: repeat ▷ Phase 2: greedy borrowing 5: updated = false D > 0, c 6: build borrowing candidates C = {(i, j, d ) : Ai,d i,d > 0, c j,d == 1, (i, j, d) ∈ / K} D − A H0 in decreasing order 7: sort C by score Ai,d j,d 8: for all (i, j, d ) ∈ C , in sorted order do 9: if activating bi,j,d = 1 violates Eq. (5)–Eq. (9) then 10: K ← K ∪ (i, j, d); continue ′ 11: b′ = b with bi,j,d =1 ′ w′ , ρ′ , ∆ρmax , T ′ D ← 2HWF with b′ ′ ′ if ∆ρmax < ∆ρmax or (∆ρmax == ∆ρmax and T ′ D < T D ) then ′ 14: b = b′ ; w = w′ ; ∆ρmax = ∆ρmax ; T D = T ′D 15: updated = true; break 16: else 17: K = K ∪ (i, j, d) 18: until not updated 19: for all j such that bi,j,j == 0 ∀i do ▷ Phase 3: self-wavelength refinement 20: F j ← {i : activating bi,j,j = 1 satisfies Eq. (5)–Eq. (9)} 21: if F j ̸ = ∅ then 22: i⋆ ← arg maxi∈F j ρi,j 23: bi⋆ ,j,j ← 1 24: w, ρ, ∆ρmax , T D ← 2HWF with b 25: return b, w 12: 13:
every possible borrowing candidate without any useful update of the borrowing state. Since A Ej,j = 0, a leaf j’s self-directed default wavelength never carries traffic, so leaving it unborrowed (bi,j,j = 0 ∀i) after the greedy borrowing phase wastes capacity. For every such leaf j, the algorithm assigns this idle wavelength to the feasible leaf i⋆ (w.r.t. Eq. (5)–Eq. (9)) with the highest current load ρi,j , and recomputes w via 2HWF. Phase 3: Self-wavelength refinement
5. PERFORMANCE EVALUATION To assess the performance of the proposed borrowing architecture, we developed a Python simulator implementing the heuristic in Algorithm 2 and compared it against three baselines, all built on the same static optical core, consisting of a cyclic AWGR alone with one wavelength per pair (i, d). In practice, this corresponds to the architecture in Fig. 2 using only default TX modules, directly connected to the AWGR with no OxC, which we refer to as AWGR-only. The three baselines differ in the detouring capability as follows: • AWGR-only: no detouring. • AWGR-only with uniform detouring: traffic is uniformly detoured (i.e., w j,i,d = 1/(W − 1)) regardless of actual demand. This schedule-less approach resembles [6], though we use a dedicated laser per wavelength operating in parallel, rather than a single laser retuned at packet timescale.
8
• AWGR-only with two-hop water-filling detouring (B=1): uniform detouring is replaced by 2HWF, which routes according to actual demand. Since it is completely equivalent, from a networking perspective, to the borrowing architecture with B = 1, its performance is reported as that of the B = 1 case. Since the last two baselines share the same optical network, comparing them isolates only the effect of traffic-engineering solutions. The normalized end-to-end traffic A E is modeled as a Lognormal distribution with coefficient of variation cv: higher cv means higher traffic variability among pairs (i, d) at the same average value. This model simply lets us show the effectiveness of the borrowing architecture and greedy algorithm as the degree of traffic imbalance varies 8 . Fig. 6 shows A E values for W = 32 leaves, average 0.65, varying cv. At cv = 0 all pairs exchange the same traffic equal to 0.65; as cv increases, traffic imbalance among pairs increases. Fig. 7 shows the performance obtained by varying the borrowing degree B for W = 32 leaves (and wavelengths) with a load cap ρth = 0.9. Fig. 7a shows the detouring rate, i.e., the detoured traffic volume T D normalized to the total end-to-end E . AWGR-only has no load balancing functionality traffic ∑ Ai,d and hence no detoured traffic. The uniform detouring strategy used in [6], also known as Valiant load balancing [18], equalizes traffic among all pairs by having every node distribute traffic completely at random to an arbitrarily chosen intermediate node, which then redirects each packet to its actual final destination. This reshaping fits the uniform topology of a static AWGR, removing the need for reconfigurable optics. However, in our opinion it has the significant drawback that a packet traverses the optical domain twice with probability (W − 2)/(W − 1), asymptotically doubling the wavelength load regardless of traffic pattern (e.g., cv), which can lead to wavelength overload and packet loss. In contrast, the borrowing architecture with 2HWF traffic engineering reconfigures the optical domain to minimize detouring traffic, reducing load overhead due to double-crossing of the optical core and packet loss at the cost of higher, but configurable, complexity9 . These observations are confirmed by Fig. 7a. The detouring rate of AWGR-only with uniform load balancing is close to 30/31 ≈ 97%, i.e., almost all traffic is detoured. The borrowing architecture’s detouring rate is much lower and decreases as B increases, since a higher borrowing degree allows more extensive optical reconfiguration. However, for a fixed optical reconfiguration capability B, increasing the imbalance factor cv requires more detouring to satisfy the load constraint. Results for B = 1 are representative of the “AWGR-only with two-hop water-filling detouring”. Accordingly, Fig. 7a shows that 2HWF alone, without any optical reconfiguration capability, already reduces detoured traffic significantly, making it worth considering as a stand-alone traffic-engineering solution for a static, full-mesh core; further improvement requires the borrowing modules (B > 1). Fig. 7b shows that the number of borrowed wavelengths is non-decreasing in B and cv, confirming that the heuristic algo8 Although not shown here due to space constraints, other traffic characterizations, such as the gravity model [5], provide the same comparative conclusions, since performance gaps among the considered solutions mostly depend on the degree of traffic imbalance, rather than on the generative model. 9 A lower detouring rate also reduces delay, since less traffic traverses two hops. We do not report delay performance, as it would require strong assumptions on the traffic model or a packet-level simulator [19]; we instead use a simple fluidic model.
9
4 2 0 0 10
30
20
20 30
Leaf ID
6 4 2 0 0 10 30
Leaf ID Leaf ID
6 4 2 0 0
(a)
10
20
20
Leaf ID
Leaf ID
30 20 30
10 40
cv = 1
8
30
20
10 40
cv = 0.5
8
i,d
i,d
6
End-to-end Traffic (AE )
cv = 0
8
End-to-end Traffic (AE )
i,d
End-to-end Traffic (AE )
Research Article
(b)
10 40
Leaf ID
(c)
Fig. 6. End-to-end traffic A E variation for increasing values of the lognormal coefficient of variation cv with average 0.65.
rithm exploits higher reconfiguration capability (B) to reduce traffic detouring when it is most needed, i.e., as imbalance (cv) increases. We observe that an architecture with B = W is fully optically reconfigurable, since each wavelength can be assigned to any source–destination pair. Yet Fig. 7a and Fig. 7b show that performance is already close to optimal at B = 8 ≪ 32, well before full reconfigurability: the borrowing architecture’s partial reconfigurability is therefore enough to achieve near-optimal performance, letting the system save on the complexity and cost of a fully reconfigurable solution. Fig. 7c shows the wavelengths allocated to each pair (i, d) for B = 8: its similarity to the traffic pattern in Fig. 6c confirms that wavelength borrowing is carried out by the heuristic in proportion to demand. Figs. 7d and 7e show the load ρ(i, d) of each pair for AWGRonly with uniform detouring and borrowing with 2HWF, respectively (the AWGR-only values coincide with A E in Fig. 6c). Packet loss occurs whenever ρ(i, d) > 1, which happens for many pairs in both AWGR-only and AWGR-only with uniform detouring. Fig. 7 confirms that the borrowing architecture with B = 8 keeps every pair’s load below ρth = 0.9, resulting in no packet loss and respecting load cap. Fig. 7f shows the packet loss rate. For cv ≤ 1, the loaddoubling drawback of uniform detouring makes its loss performance even worse than AWGR-only without detouring. 2HWF alone (B = 1) avoids packet loss up to cv = 1, but not at cv = 2, where the traffic imbalance is severe enough to require optical reconfigurability; borrowing architecture with B > 2 ensures zero packet loss across the whole cv range tested. Fig. 8 shows an interesting scale-out behavior: as the number of leaves W increases, fewer borrowing modules per leaf are needed to keep every pair’s load below target and avoid loss. B = 2 suffices for larger networks (W = 32, 64), while smaller networks need higher B, since a given B offers more donors – and hence more borrowing opportunities – as the network grows. Thus, scaling out the data center does not require the total number of borrowing components to grow linearly. Note, however, that the number of wavelengths W must still scale linearly with the number of leaves to support full connectivity. We conclude this section by discussing a preliminary power budget of the architecture. For the AWG, we consider an insertion loss of 5 dB, counted twice since the end-to-end path traverses both an AWG mux and an AWG demux (10 dB total);
for the OxC, 2 dB; and for the AWGR, we consider a loss varying with W, namely 4, 5, and 6 dB for W = 16, 32, 64, respectively. For the combiner, we consider 10 log10 ( B=8) dB, since a small B (e.g., B = 8) already achieves near-optimal performance 10 . We further consider a loss margin of 2 dB, accounting for connector loss and similar contributions. The resulting power loss is approximately 27, 28, and 29 dB for W = 16, 32, 64, respectively. Regarding the transmission power and receiver sensitivity of Ethernet transceivers, we consider those of the 400/800GBASEDR8 [21], namely 4 dBm for the TX power and −5.9 dBm for the sensitivity, yielding an approximate sustainable power loss of 10 dB. Consequently, DWDM-ready amplifiers in the 15–25 dB range, such as those provided by EDFAs, are required along the end-to-end path, e.g., after each coupler.
6. RELATED WORKS The use of all-optical switching in data centers is attracting growing interest, including from major cloud operators such as Google [5, 11] and Microsoft [6]. Proposed architectures differ substantially in the network tier (server-level, rack-level, etc.) at which the optical fabric is deployed, which drives the traffic characteristics to be handled, the required switching speed, and the optical components. The closer to the server the optical core is placed, the finer and more dynamic the traffic it must serve, imposing stringent switching-speed and component requirements; conversely, higher aggregation levels allow slower, more mature, cost-effective technologies, since traffic variability is smoothed by aggregation. Proposed solutions also differ in the maturity and cost of the required optical components. Architectures relying on actively reconfigurable colored fabrics—such as wavelength-selective switches (WSS) or WDM-aware OxCs with SOA gate arrays—or on optical signal processing for in-network forwarding demand components that remain expensive and available only at limited port counts, hindering near-term deployment. By contrast, colorless MEMS-OxCs performing pure spatial switching and passive AWGRs—whose wavelength routing is fixed and requires no active control—are commercially available at datacenter-relevant 10 AWG: below 3.7 dB (40-ch 100 GHz) and up to 5.5 dB (80-ch 50 GHz, W = 64) for commercial DWDM modules https://edgeoptic.com, https://www.hilinktech.com/ aawg/50ghz-aawg-dwdm-mux-demux-80ch.html; we use 5 dB across all W. OxC: below 2 dB for 136×136 MEMS switches [5, 11]. AWGR: consistent with https: //lumilaserchip.com/product-category/awgr/awgr-module/ and [20]. EDFA: commercial C-band pre-amplifiers reach up to 25 dB gain https://www.optilab.com/products/ c-band-pre-amp-edfa-module-14-dbm-25-db-gain.
10
)
Research Article
0.4
0.2
0
per-pair n. of wavelengths (c
0.6
cv = 0 cv = 0.5 cv = 1 cv = 2
300 250 200 150 100 50
0
5
10
0
15
0
5 10 Borrowing degree (B )
Borrowing degree (B )
(a)
10
30
20
20 30
Leaf ID
2
0 0 30
20
20 10 40
Leaf ID
(c)
2
AWGR-only AWGR-only unif. detouring Borrowing (B=1) Borrowing (B=2) Borrowing (B>2)
0.4
i,d
;=1
pair load (;
1
10 40
0.5
)
)
0 0
Borrowing B = 8, cv = 1
i,d
pair load (;
5
(b)
AWGR-only uniform detouring, cv = 1
Leaf ID
Borrowing B=8, cv=1
10
15
; = 0.9 th
1
0 0
Leaf ID
Leaf ID
30
20
20 10 40
(d)
Leaf ID (e)
Packet loss rate
0.8
i,d
350 AWGR-only, Borrowing (cv = 0) AWGR-only uniform detouring Borrowing (cv = 0.5) Borrowing (cv = 1) Borrowing (cv = 2)
Borrowed wavelengths
Traffic Deturing Rate
1
0.3
0.2
0.1
0
0
0.5
1
2
cv
(f)
Fig. 7. Performance varying the borrowing degree B with ρth = 0.9 in a system with W = 32 leaves
i,d
Maximum pair load (max ; )
1.2 W=8 W=16 W=32 W=64
1.15 1.1 1.05 1 ;th=0.9
0.95 0.9 0.85 0.8
0
5
10
15
Borrowing degree (B )
Fig. 8. Maximum pair load for cv = 1 versus number of leaves
(W) scale today, making them the more practical short-term choice. Open research platforms such as OpenOptics [22] aim to lower the barrier to experimenting with this new architectures. Optical Core at the Server and Rack Level
At the finest granularity, the optical core’s endpoints are directly servers or racks, whose traffic bursts must be handled with virtually no buffering. Traffic here is highly dynamic: flows are short-lived, demand changes on timescales of tens to hundreds of nanoseconds, and efficient utilization requires reconfiguration at packet or sub-packet granularity. Sirius [6] is a prominent example. It proposes a flat, all-optical
network in which a single passive layer of AWGRs replaces the entire electrical switching hierarchy above the ToR. ToR uplinks use custom tunable laser chips that encode the destination as a wavelength on a time-slot basis, enabling end-to-end reconfiguration in under 1 ns, with all-to-all connectivity via a cyclic round-robin TDMA schedule that avoids wavelength contention. To accommodate arbitrary traffic atop this uniform bandwidth topology, every packet is randomly detoured through an intermediate rack [18], creating a fully balanced demand at the cost of a two-hop path, an intermediate O/E/O conversion, and neardoubling of the optical load, since almost all packets traverse wavelengths twice. The AWGR core is passive, with no active reconfiguration or optical signal processing – routing follows purely from the transmitter’s wavelength choice – but the architecture requires custom photonic integrated circuits for the nanosecond tunable laser, and its lack of reconfigurability forces uniform two-hop detouring regardless of actual traffic skew. OPSquare [23] and its multi-level extension HFOS [24] push optical switching to the ToR level using fast colored (WDM-aware) optical packet switches with nanosecond-scale reconfiguration. Each switch combines AWGs with SOA-based 1× F broadcastand-select gate arrays, forwarding packets within a group of F ToRs by extracting an in-band RF-tone optical label to control the SOA gates – genuine optical signal processing. Since no prior scheduling is performed, contention causes ALOHA-like packet loss, recovered via ACK/NACK retransmission from electrical buffers. HFOS scales this to multiple parallel switch levels, reaching tens of thousands of servers under the same colored, processing-based paradigm. Both remain technologically demanding – label processors are research-grade and available only at small port counts – and suffer non-negligible loss and
Research Article
low throughput under load due to lack of contention avoidance. ROTOS [12] extends OPSquare [23] with a reconfigurable ToR switch that dynamically reallocates WDM transceivers and a colored WSS between intra- and inter-cluster traffic under SDN control, steering wavelengths via WSS according to the observed traffic ratio. It remains a packet-level solution using the same colored, SOA-based, label-processing switches as OPSquare, and additionally requires a per-ToR WSS, adding cost and a scalability constraint proportional to the node count N. PULSE [25] builds a packet-level all-optical network around x2 passive N × N star couplers (N servers/rack, x racks). Each server has x transceivers, one per rack, each with B banks of continuously-on tunable DS-DBR lasers covering W = N wavelengths [26]; SOA-gated laser outputs open only during the reserved time slot, and a star coupler broadcasts to the destination rack. Wavelength/slot assignment is precomputed by per-rack schedulers to avoid contention. Its scalability is limited by starcoupler splitting loss, which grows as 10 log10 ( N ) dB with rack size and requires SOA amplification at every transceiver, plus x nanosecond-scale tunable transceivers per server – a costly, complex requirement for off-the-shelf hardware. Optical Switching at the Aggregation Level
A complementary class of architectures places the optical fabric at a higher level, interconnecting aggregation points such as leaf nodes or racks (Pods). Here traffic is considerably smoother – demand evolves on timescales of milliseconds to seconds – enabling slower but more mature and cost-effective switching technologies: colorless MEMS-based OxC, actively reconfigurable colored WSS, and passive AWGR. Google’s Jupiter [5] is the most prominent industrial deployment in this class, replacing the electrical spine layer with a datacenter network interconnection layer (DCNI) built on colorless MEMS-based OxCs that connect aggregation blocks via pure spatial switching, with no wavelength awareness or optical signal processing. Reconfiguration is driven by traffic engineering on timescales of seconds to minutes, matching the millisecond switching time of MEMS-OxC. Lightwave Fabrics [11] deploys the same colorless-OxC approach at even larger scale, across multiple datacenter buildings. In [9], the authors propose an architecture based on the Hyper-FleX-LION fabric [27], operating at rack level but with an aggregation-like reconfiguration paradigm. It combines a passive AWGR with actively reconfigurable colored 1× N WSSs at each rack’s TX/RX, steering wavelengths through the AWGR or directly between rack pairs; routing decisions are made by the SDN controller on a slow timescale, requiring no optical signal processing. Interconnecting N racks requires a 1× N WSS per rack per direction, i.e., 2N 2 WSSs total – highly flexible, but the WSS’s cost and commercially available port count (N ≈ 20) impose a scalability boundary under current market conditions, with millisecond-scale reconfiguration. The large-scale fast optical circuit switch of [15] demonstrates colorless MEMS-based OxC at hundreds of ports with acceptable loss and millisecond switching, confirming the viability of this technology class; a multi-stage Thin-CLOS AWGR fabric [16] offers an alternative, passive route to the same scale. Reconfiguring an OxC-based fabric over time introduces a connection defragmentation problem, studied in [28]. Ring-based aggregation fabrics from grouped ROADMs with shared amplification have similarly been proposed [29], and [8] explores sub-wavelength TDMA resource allocation at the optical layer, a technique also optionally supported by the proposed architecture.
11
The Proposed Architecture
The proposed wavelength-borrowing architecture targets the same aggregation-level design space as Jupiter [5], HyperFleX-LIONS [9], and ROTOS [12], but pursues a distinctive complexity-reconfigurability trade-off grounded in technological maturity. Its central optical fabric relies exclusively on a passive AWGR and a colorless MEMS-based OxC performing pure spatial switching — both well-established, commercially available technologies — with no actively reconfigurable colored components and no optical signal processing required at any node. This contrasts with OPSquare and HFOS, which require colored WDM-aware OxCs with SOA gate arrays and optical label processors, and with Hyper-FleX-LIONS and ROTOS, which rely on per-node WSSs. In place of the WSS, the proposed architecture uses a simple passive combiner, imposing no wavelength-awareness requirement and remaining transparent to TDMA. Unlike ROTOS and Hyper-FleX-LIONS, moreover, activating or releasing a borrowed wavelength never requires reconfiguring the destination leaf, which keeps receiving on the same AWGR output port throughout. Resource allocation flexibility is instead controlled through the single borrowing degree B, providing a tunable trade-off between hardware cost and reconfiguration flexibility: Sec. 5 shows that near-optimal performance is already achieved at a small B regardless of network scale, and that the number of borrowing lines required per leaf can even decrease as the data center grows, in contrast with architectures such as Hyper-FleXLIONS whose per-node WSS count scales with the network size. Reconfiguration operates at the millisecond timescale of MEMS-based OxC switches, targeting traffic demands that persist over seconds or longer [5], consistent with aggregation-level deployment. Finally, this paper couples the architecture with a formal MILP formulation of the joint wavelength-assignment and traffic-detouring problem – reusable for other optimization goals – and a greedy heuristic, built around the 2HWF trafficengineering subroutine, that solves it at practical computational cost and can also serve as a stand-alone detouring solution for other optical fabrics. Appendix III of [14] reports a comparative table of the considered architectures.
7. CONCLUSIONS This paper presented a wavelength-borrowing architecture for spine-leaf optical data center networks, in which idle or lightly loaded wavelengths at one leaf are dynamically reallocated to a leaf with higher demand. The borrowing degree B exposes the complexity/reconfigurability tradeoff as a single tunable parameter, from a static core (B = 1) to a fully reconfigurable fabric (B = W). Combined with the proposed 2HWF traffic-engineering heuristic, results show that a moderate, scaleindependent borrowing degree, e.g., B = 8, already achieves near-optimal performance, and that fewer borrowing lines per leaf are needed as the network scales out, since larger networks offer more donor leaves and, thereby, more optimization opportunities for the same value of B. This spares the network from the cost of full reconfigurability without a performance penalty, using only mature, data-center-grade optical components. Being also TDMA-transparent, the architecture leaves room for finergrained, sub-wavelength borrowing in future evolutions with no change to the optical fabric.
Research Article
REFERENCES 1.
2.
3. 4.
5.
6.
7.
8.
9.
10.
11.
12.
13. 14.
15.
16.
17. 18. 19.
S. Kandula, S. Sengupta, A. Greenberg, P. Patel, and R. Chaiken, “The Nature of Data Center Traffic: Measurements & Analysis,” in Proceedings of the 9th ACM SIGCOMM Conference on Internet Measurement (IMC ’09), (ACM, New York, NY, USA, 2009), pp. 202–208. M. Al-Fares, A. Loukissas, and A. Vahdat, “A Scalable, Commodity Data Center Network Architecture,” in Proceedings of the ACM SIGCOMM 2008 Conference on Data Communication (SIGCOMM ’08), (ACM, New York, NY, USA, 2008), pp. 63–74. C. Kachris and I. Tomkos, “A Survey on Optical Interconnects for Data Centers,” IEEE Commun. Surv. & Tutorials 14, 1021–1036 (2012). P. A. Baziana, “Optical data center networking: A comprehensive review on traffic, switching, bandwidth allocation, and challenges,” IEEE Access 12, 186413–186444 (2024). L. Poutievski, O. Mashayekhi, J. Ong, A. Singh, M. Tariq, R. Wang, J. Zhang, V. Beauregard, P. Conner, S. Gribble et al., “Jupiter evolving: transforming google’s datacenter network via optical circuit switches and software-defined networking,” in Proceedings of the ACM SIGCOMM 2022 Conference, (2022), pp. 66–85. H. Ballani, P. Costa, R. Behrendt, D. Cletheroe, I. Haller, K. Jozwik, F. Karinou, S. Lange, K. Shi, B. Thomsen, and H. Williams, “Sirius: A Flat Datacenter Network with Nanosecond Optical Switching,” in Proceedings of the Annual Conference of the ACM Special Interest Group on Data Communication on the Applications, Technologies, Architectures, and Protocols for Computer Communication (SIGCOMM ’20), (ACM, New York, NY, USA, 2020), pp. 782–797. G. Patronas, N. Terzenidis, P. Kashinkunti, E. Zahavi, D. Syrivelis, L. Capps, Z.-A. Wertheimer, N. Argyris, A. Fevgas, C. Thompson, A. Ganor, J. Bernauer, E. Mentovich, and P. Bakopoulos, “Optical Switching for Data Centers and Advanced Computing Systems,” J. Opt. Commun. Netw. 17, A87–A90 (2025). K. Christodoulopoulos, K. Kontodimas, L. Dembeck, and E. Varvarigos, “Slotted optical datacenter networks with sub-wavelength resource allocation,” in 2019 Optical Fiber Communications Conference and Exhibition (OFC), (IEEE, 2019), pp. 1–3. H. Yang and Z. Zhu, “Traffic-aware configuration of all-optical data center networks based on hyper-flex-lion,” IEEE/ACM Transactions on Netw. 32, 2675–2688 (2024). X. Meng, V. Pappas, and L. Zhang, “Improving the Scalability of Data Center Networks with Traffic-Aware Virtual Machine Placement,” in Proceedings of the IEEE INFOCOM 2010, (IEEE, 2010), pp. 1–9. H. Liu, R. Urata, K. Yasumura, X. Zhou, R. Bannon, J. Berger, P. Dashti, N. Jouppi, C. Lam, S. Li, E. Mao, D. Nelson, G. Papen, M. Tariq, and A. Vahdat, “Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning Systems,” in Proceedings of the ACM SIGCOMM 2023 Conference (ACM SIGCOMM ’23), (ACM, New York, NY, USA, 2023), pp. 499–515. X. Xue, F. Yan, K. Prifti, F. Wang, B. Pan, X. Guo, S. Zhang, and N. Calabretta, “ROTOS: A Reconfigurable and Cost-Effective Architecture for High-Performance Optical Data Center Networks,” J. Light. Technol. 38, 3484–3495 (2020). Open ROADM MSA, “Open ROADM MSA Device Model White Paper,” White paper, version 13.1 (2024). Last accessed July 2026. A. Detti, C. Lodovisi, and S. Betti, “A Wavelength Borrowing Architecture for Optical Data Center Networks - Extended Version,” (2026). Last accessed Sept 2026. K.-i. Sato, “Realization and Application of Large-Scale Fast Optical Circuit Switch for Data Center Networking,” J. Light. Technol. 36, 1411– 1419 (2018). R. Proietti, X. Xiao, K. Zhang, G. Liu, H. Lu, P. Fotouhi, J. Messig, Jr., and S. J. B. Yoo, “Experimental Demonstration of a 64-Port Wavelength Routing Thin-CLOS System for Data Center Switching Architectures,” J. Opt. Commun. Netw. 10, B49–B57 (2018). D. Bertsekas, “R. gallager data networks,” Pretice-Hall Int. (1992). L. G. Valiant, “A scheme for fast parallel communication,” SIAM journal on computing 11, 350–361 (1982). W. Chen, Y. Tian, and X. Zhang, “Acceltor: Accelerating tcp for circuit/packet hybrid data centers with packet scheduling,” IEEE Transac-
12
tions on Netw. (2025). 20. S. Kamei, M. Ishii, M. Itoh, T. Shibata, Y. Inoue, and T. Kitagawa, “64× 64-channel uniform-loss and cyclic-frequency arrayed-waveguide grating router module,” Electron. Lett. 39, 83–84 (2003). 21. IEEE, “IEEE Standard for Ethernet–Amendment 9: Media Access Control Parameters for 800 Gb/s and Physical Layers and Management Parameters for 400 Gb/s and 800 Gb/s Operation,” (2024). Amendment to IEEE Std 802.3-2022. 22. Y. Lei, F. De Marchi, J. Li, R. Joshi, S.-T. Wang, X. Chen, B. Chandrasekaran, and Y. Xia, “OpenOptics: Enabling Open Research and Implementation of Optical Data Center Networks,” in Proceedings of the 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI ’26), (USENIX Association, Renton, WA, USA, 2026). 23. W. Miao, F. Yan, and N. Calabretta, “Towards petabit/s all-optical flat data center networks based on wdm optical cross-connect switches with flow control,” J. Light. Technol. 34, 4066–4075 (2016). 24. E. Khani, S. Hessabi, S. Koohi, F. Yan, and N. Calabretta, “Hfos l: hyper scale fast optical switch-based data center network with l-level sub-network,” Telecommun. Syst. 80, 397–411 (2022). 25. J. L. Benjamin, T. Gerard, P. Bayvel, and G. Zervas, “Pulse: Scalable sub-µs wdm-tdm circuit switched data center network,” in 45th European Conference on Optical Communication (ECOC 2019), (IET, 2019), pp. 1–4. 26. A. J. Ward, D. J. Robbins, G. Busico, E. Barton, L. Ponnampalam, J. P. Duck, N. D. Whitbread, P. J. Williams, D. C. Reid, A. C. Carter et al., “Widely tunable ds-dbr laser with monolithically integrated soa: Design and performance,” IEEE J. selected topics quantum electronics 11, 149–156 (2005). 27. G. Liu, R. Proietti, M. Fariborz, P. Fotouhi, X. Xiao, and S. B. Yoo, “Architecture and performance studies of 3d-hyper-flex-lion for reconfigurable all-to-all hpc networks,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, (IEEE, 2020), pp. 1–16. 28. X. Dong, X. Chen, and Z. Zhu, “On the risk-aware connection defragmentation in ocs-based data-center networks,” IEEE Transactions on Netw. Serv. Manag. (2025). 29. L. Zhao, W. Hu, and X. Zhang, “Architecture and Performance of Grouped ROADM Rings with Shared Optical Amplifier and Grouped Add/Drop Ports for Hybrid Data Center Network,” Opt. Switch. Netw. 23, 1–4 (2017). 30. IEEE 802.3dj Working Group, “FEC baseline proposal for 200Gb/s per Lane IM-DD Optical PMDs,” https://www.ieee802.org/3/dj/public (2023). Last accessed June 2023.
Research Article
13
Combiner group 1
LEAF (1,1) (TX)
LEAF 1 (TX) 1
LEAF d (RX)
1
P
P AWG DEMUX
default TX module
AWGR
B1 an Br d oa Se dc le as ct t
P
G default fibers
Spine Optical Switching Fabric (1,1) LEAF (1,1) (RX)
B-1 borrowing fibers (B-1) 1xG2 1x4
1
Spine Optical Switching Fabric (1,G)
LEAF (W,1) (RX)
1xP
borrowing TX module n. 2
P
Colorless OxC
Broadcast and Select
To central fabrics
1
borrowing TX module n. B
Spine Optical Switching Fabric (G,1)
LEAF (1,G) (RX)
(B-1) 2 1xG 1x4
LEAF (W,G) (RX) LEAF (W,G) (TX)
Fig. 9. Capacity scaling with P parallel AWGR-routed layers
APPENDIX I: ARCHITECTURAL EXTENSION A. Small Data Center
The foregoing description assumed that the number of leaf nodes L equals the number of wavelengths W. If the required number of leaves L is smaller than W, the spine optical fabric remains unchanged, and the default wavelengths of the W − L absent leaves can simply be borrowed by the existing leaves. B. Large Data Center
Hyperscale data centers can require very large bisection bandwidth and large numbers of nodes to interconnect. In the proposed architecture, the maximum number of leaves is W, and the bidirectional bisection bandwidth in the balanced configuration is W 2 S/2, where S is the per-wavelength bitrate. For instance, with W = 64 and S = 400 Gbit/s, the resulting bisection bandwidth is about 0.8 Pbit/s [11, 15, 30]. If this bisection bandwidth is insufficient but the number of leaf nodes W is adequate, the architecture can be layered as shown in Fig. 9, where only the components related to transmitting leaf 1 and receiving leaf d are depicted. Specifically, the architecture provides P parallel AWGR-routed layers carrying both default and borrowed wavelengths. The colorless OxC is shared across all P layers, and its size does not grow with P, thereby removing potential scale-out limitations imposed by the unavailability of large OxC switches. The resulting bisection bandwidth scales by a factor of P. In this scaling scheme, each leaf i has the usual B − 1 borrowing fibers connected to the central OxC, but P parallel default TX modules, resulting in P output default fibers each carrying W default wavelengths toward the W remote leaves. The P default fibers are connected to group i of P parallel combiners, whose output fibers are in turn connected to P parallel AWGRs. Output ports d of these AWGRs are connected to P parallel AWG demultiplexers and WDM receivers at destination leaf d. The colorless OxC remains unchanged and routes the B − 1 borrowing fibers per node to the combiner groups. Routing within each group to a specific combiner is managed by a 1× P optical switch, e.g., implemented with broadcast-and-select technology [12], configured by the SDN controller. In this way, a borrowing fiber can opportunistically access resources from any AWGR layer. Bandwidth scaling
When more than W leaf nodes are required, the architecture can be extended as shown in Fig. 10. The nodes are partitioned into G groups of at most W nodes each. Communication within the same group and between different groups is handled by G2 distinct spine optical switching fabrics, one Node scaling
Spine Optical Switching Fabric (G,G)
Fig. 10. Node scaling with cross-connecting optical fabrics. Leaf (i, j) denotes leaf i of group j; central fabric (t, r ) connects TX leaves of group t to RX leaves of group r.
per ordered source–destination group pair. For example, central fabric (1, 1) interconnects TX leaves of group 1 with RX leaves of group 1, while central fabric (1, G ) cross-connects TX leaves of group 1 with RX leaves of group G. Fig. 10 depicts a representative subset of leaves and interconnections: transmitting leaf 1 of group 1, identified as leaf (1, 1) (TX); central fabric (1, G ) connecting group 1 to group G; and receiving leaf 1 of group G, identified as leaf (1, G ) (RX). Each TX leaf has G distinct default fibers, each sourced by a dedicated default TX module. The g-th default fiber carries W wavelengths destined for the RX leaves of group g and is connected to the spine fabric serving that source–destination group pair. In addition, each TX leaf outputs B − 1 borrowing fibers, which carry wavelengths lent by donor leaves. Since a borrowed wavelength may interconnect leaves belonging to any pair of groups, each borrowing fiber must be steerable to the appropriate spine fabric. This is accomplished by a dedicated 1× G2 broadcast-and-select switch per borrowing fiber, configured by the SDN controller11 . Finally, bisection bandwidth can be further increased by combining both scaling approaches.
APPENDIX II: DERIVATION OF THE TWO-HOP ABSORBED TRAFFIC This appendix motivates Eq. (25), the amount of traffic Fθ,j,i,d that a two-hop path ( j, i, d) absorbs at water level θ. old new old Let ∆ j,i = ρnew j,i − ρ j,i and ∆i,d = ρi,d − ρi,d denote the load increments induced on the two hops by the detoured traffic. Since the same physical traffic Fθ,j,i,d traverses both hops of the path, Fθ,j,i,d = ∆ j,i c j,i = ∆i,d ci,d . (29) For a single link taken in isolation, the largest amount of traffic it could absorb without exceeding the target level θ is obtained by setting ρnew = θ, i.e., j,i
Fmax = (θ − ρold j,i ) c j,i ,
i,d old Fmax = (θ − ρi,d ) ci,d .
(30)
Because Fθ,j,i,d in Eq. (29) is the same quantity on both hops – not the same ρ increment – it cannot exceed either per-link 11 The fan-out of these switches can be reduced by restricting the set of central fabrics from which each borrowing fiber may borrow resources.
Research Article
limit in Eq. (30). The more constrained of the two hops therefore saturates first and bounds the absorbed traffic: j,i i,d Fθ,j,i,d = min Fmax , Fmax . (31) The hop attaining the minimum in Eq. (31) reaches ρnew = θ exactly, while the other hop remains below θ, since it is bounded away from saturation by construction. Hence new max(ρnew j,i , ρi,d ) = θ, consistently with the two-hop load definition ρ j,i,d = max(ρ j,i , ρi,d ). Finally, if a hop is already loaded past θ (i.e., ρold > θ), its term in Eq. (30) is negative and must be clamped to zero, since a link cannot absorb a negative amount of traffic; this is the case of path 3 in Fig. 5, which is already more loaded than θ and therefore cannot be filled. Applying this clamp to Eq. (31) yields Eq. (25).
APPENDIX III: COMPARISON OF RELATED ARCHITECTURES
14
Server-level Passive star coupler
Fast tunable WDM + SOA/AWG bank Burst-mode fast tunable WDM + SOA/AWG bank or coherent RX WDM+TDMA
High—Per-packet λ+slot alloc.
No Star coupler: commercial, low cost. Fast DS-DBR + SOA bank: laboratory/early comm., medium-high cost.
Rack/Server-level
Passive AWGRs
Fast tunable WDM (<1 ns)
Burst-mode fixed-λ WDM + phasecaching
WDM+TDMA
Low—Static cyclic WDM/TDMA schedule.
No
Passive AWGR: commercial, low cost. Fast tunable laser: laboratory, high cost wrt fixed laser.
Network tier
Key switching tech.
TX technology
RX technology
Optical plexing
Resource allocation flexibility
Optical processing
Opt. tech. maturity & cost (key switch elem.)
Multi-
PULSE [25]
Sirius [6]
Property
[23], RO-
AWG + SOA gates: commercial, std. cost. Label processor: laboratory, small port count. WSS (ROTOS): commercial, med.-high. Overall: laboratory-grade, medium-high cost.
Yes
High/Medium— ALOHA-like on predefined λ (OPSquare/HFOS); SDN/WSS-driven λ reallocation (ROTOS).
WDM
Burst-mode fixed-λ WDM
Fixed-wavelength WDM
AWG + SOA gates + WSS (ROTOS only)
Rack-level
OPSquare HFOS [24], TOS [12]
Passive AWGR: commercial, low cost. Colored WSS: commercial, med.-high. Overall: medium cost.
No
Medium—λ reconfiguration.
WDM
Fixed-λ WDM
Fixed-wavelength WDM
Passive AWGR + colored WSS
Rack-level
Hyper-FleXLIONS [9]
MEMS
3D MEMS OXC: commercial, med. cost (320×320, ms switching). WDM TRX (CWDM4): commercial, low cost.
No
Low—fiber-level reconfiguration
Not mandatory (whole fibers switched)
Not mandatory (whole fibers switched)
Not mandatory (whole fibers switched)
Colorless OXC
Rack-group-level
Jupiter [5], Lightwave Fabrics [11]
MEMS OXC: commercial, med. cost. Passive AWGR + combiners: commercial, low cost. Fixed-λ WDM TRX: commercial, low cost. Overall: lowest cost among reconfigurable WDM designs.
No
Tunable, static-tohigh—λ reconfiguration via the borrowing degree B, from a static core (B = 1) to full reconfigurability (B = W)
WDM (+TDMA)
Fixed-λ WDM; burstmode if TDMA enabled
Fixed-wavelength WDM
Colorless MEMS OXC + passive AWGR + Combiners
Rack-group-level
This work
label extraction, SOA gating). “Rack group” denotes a cluster of racks served by a common aggregation point (leaf node, aggregation block, or cluster switch). Opt. tech. maturity & cost rates the key switching technology as commercial (off-the-shelf, volume pricing) or laboratory (custom/prototype, limited availability).
Table 2. Comparison of all-optical datacenter network architectures. “Optical processing” indicates in-network forwarding decisions made in the optical domain (e.g., optical
Research Article 15