arXiv:2609.14253v1 [cs.NI] 13 Sep 2026
Breaking the Duplex Barrier: Lane-Granularity OCS Scheduling for LLM Training Bangbo Liang
Yupeng Chen
Sicheng Zhao
[email protected] Hunan University China
[email protected] Hunan University China
[email protected] Hunan University China
Peihao Huang∗
Di Yang
Bohua Xu
[email protected] Hunan University China
[email protected] China Unicom Software Research Institute China
[email protected] China Unicom Research Institute China
Bin Yang
Shizhen Zhao
Guo Chen∗
[email protected] China Unicom Research Institute China
[email protected] Shanghai Jiao Tong University China
[email protected] Hunan University China
Abstract
1
An optical circuit switch (OCS) can reconfigure physical connectivity to match the predictable communication schedules of large language model (LLM) training. Although each OCS light path is physically simplex, existing demand-aware OCS schedulers allocate capacity in duplex-port pairs, forcing equal bandwidth in both directions and stranding capacity under asymmetric node-pair traffic. This paper presents LACE, the first offline OCS schedule compiler that independently allocates transmit (TX) and receive (RX) lanes for LLM training. Without changing the selected collective algorithms, operation order, or rank placement, LACE reconstructs directed node-level demand, jointly determines which consecutive operations share a configuration and how many simplex circuits serve each direction, and realizes these allocations as physical lane bindings and optical paths under per-node lane-inventory and multi-OCS fabric constraints. Software acknowledgments carry feedback over independently provisioned return paths, while coordinated link configuration and recovery verify each configuration before communication resumes. On a separate three-server testbed using fixed topologies and matched perport rate limits, LACE’s asymmetric connectivity achieves 1.80× speedup for communication replay and 1.27× for GPT2 training over a symmetric-topology baseline. At larger scale, simulations of LLaMA-3.1 70B and 405B schedules with sixteen 400-Gb/s ports per server show that LACE achieves 1.21–2.04× communication speedup over the latest duplex OCS scheduler.
Large-scale LLM training generates enormous inter-node traffic, making network bandwidth a first-order determinant of training performance and infrastructure cost. For training jobs with a known communication schedule, operators can reconfigure the physical fabric to match communication demand [41], rather than relying solely on a static topology provisioned for worst-case demand. An optical circuit switch (OCS) fabric, consisting of one or more OCS devices, provides this flexibility: it steers light between fibers to establish dedicated paths without intermediate optical-to-electrical conversion. Commercial MEMS OCSs can change these optical connections on millisecond time scales [11, 30]. Long studied in the data-center literature [2, 6, 12, 27, 28, 40], OCSs now underpin production infrastructure, including Google’s Jupiter network [36, 39] and TPU v4 supercomputer [21], and motivate LLM-specific designs that adapt connectivity to training communication [22, 24, 25, 41]. For all their diversity, existing demand-aware OCS systems use the duplex port as the basic unit of topology scheduling. Each node is either a server or an electrical packet switch (EPS) and connects to the OCS fabric through one or more ports, each comprising one transmit (Tx) lane and one receive (Rx) lane carried on separate fibers. The topology scheduler computes a matching over ports: connecting port 𝑖 to port 𝑗 always instantiates both directions of the circuit at once [21, 22, 25, 36, 39, 41]. Nothing in the optics requires this. The mirrors steer each simplex light path independently, and a “duplex circuit” is merely two simplex cross-connects that happen to be coordinated; schedulers could, in principle, issue them separately. They do not, because of what sits at the other end of the fiber: default link-management mechanisms are designed for a port whose Tx and Rx lanes reach the same peer, and prior measurements show that simplex
Keywords: Optical Circuit Switching, Distributed Machine Learning, Network Architecture ∗ Corresponding authors.
1
Introduction
Liang et al.
reconfiguration can trigger link training and failure handling at NICs and switches, causing unintended link drops[44]. In our testbed, however, ConnectX-6 ports operate with different Tx and Rx peers once autonegotiation is disabled and the line rate is fixed. Duplex-port binding therefore stems from conventional endpoint link management rather than a physical constraint of the OCS fabric. That convention was benign when traffic was roughly symmetric. LLM traffic is not. The bytes a node sends and receives diverge sharply between the two directions: pipelineparallel stages push activations downstream and pull gradients upstream in disjoint time windows[16, 31]; a selected DP ring can send bulk data to one neighbor and receive it from another. A duplex scheduler, however, can only provision capacity in symmetric pairs: matching node A to node B buys A one lane toward B and one lane back from B, whether or not B has anything to send. The cold direction of each circuit idles while the hot direction saturates—and no matter how cleverly the matching is computed [21, 22, 24, 25, 41], symmetry itself is never up for negotiation. This paper argues that the symmetry constraint is not only costly but unnecessary—once we are precise about where bidirectional connectivity is actually required. In the scheduling model, the ports attached to the OCS fabric are grouped by node: a GPU server contributes its NIC ports, while an EPS contributes its uplink ports. Workloads specify communication between nodes; they do not require the Tx and Rx lanes of each port to connect to the same peer. A node with k ports therefore contributes k Tx and k Rx lanes that can be allocated independently across peers according to directional demand. However, building a practical lane-level OCS scheduler presents three challenges. First, nodes must remain operational when a port’s TX and RX lanes connect to different peers, and transport feedback must reach the sender. Second, node-pair allocations compete for shared TX and RX inventories; in a multi-OCS fabric, they must also fit the available optical paths. Finally, directional demand changes across operations, while every configuration change incurs optical switching and node recovery delay. The scheduler must jointly decide which operations share a configuration and how their lanes are allocated. In this paper, we present LACE: Lane-level Circuit Scheduling. LACE treats TX and RX lanes as independently schedulable simplex resources and compiles an ordered training communication schedule offline. Its Demand Profiler constructs directed node-level demand. The Segment Planner chooses configuration boundaries and fractional simplex circuit allocations, and the Circuit Realizer converts these targets into integer counts within the lane inventories. The Lane Binder assigns physical lanes and optical paths, retaining connections where possible and repairing allocations
when internal paths are constrained. Software acknowledgments provide return feedback across ports, while coordinated link configuration and recovery verify connections before transmission resumes. Together, these mechanisms match directional demand while accounting for the cost of changing configurations (§3). We implement LACE’s scheduler and prototype node-side communication and recovery mechanisms using commodity OCS hardware. The code will be made available upon formal publication. With sixteen 400Gbps ports per server, simulations of LLaMA-3.1 70B and 405B schedules show 1.21–2.04× communication speedup over ACTINA. Our three-server fixed-topology testbed achieves 1.80× communication-replay and 1.27× GPT-2 training speedups over ACTINA. Separate experiments measure software-ACK overhead and verify configuration recovery; simulation evaluates switch-node aggregation and circuit mapping through a multi-OCS Clos. In summary, we make the following contributions: • We quantify directional demand between nodes in LLM training and show why fixed symmetric or asymmetric circuit allocations mismatch changing DP/PP traffic (§2). • We design LACE to jointly plan shared configurations and fractional lane allocations, then realize integer simplex circuits and physical paths within node and fabric constraints (§3.3–3.5). • We describe software acknowledgments and coordinated link configuration and recovery for executing lane-level connections, and prototype software acknowledgments and verify configuration changes on commodity server NICs and an OCS (§3.6). • We evaluate scheduling gains, module contributions, and multi-OCS mapping in simulation, together with communication replay, training, and software-ACK performance on a hardware testbed (§4).
2
Background and Motivation
2.1
OCS Fabrics and Terminology
OCS hardware. An optical circuit switch (OCS) is a layer-0 device: it forwards light, not packets. Commercial 3D-MEMS OCSs provide P ports in a non-blocking crossbar; an array of micro-mirrors steers the beam arriving at any input to any output, establishing a transparent, full-bandwidth light path agnostic to the modulation and bit rate carried on it [5, 18]. Two hardware facts matter for this paper. First, every cross-connect is physically simplex: it carries light in exactly one direction, and what is normally called a bidirectional "circuit" is in fact two independent cross-connects installed by two coordinated mirror settings. Second, reconfiguring the mirror array takes on the order of tens of milliseconds [7, 11, 30], so the fabric operates in epochs: a scheduler computes 2
LACE: Lane-Granularity OCS Scheduling
(a) Duplex scheduler
(b) Simplex scheduler
Figure 1. Duplex and simplex scheduling with the same four TX lanes and four RX lanes per node. Each panel shows the OCS connections (left) and the resulting directed logical links (right). Duplex matching assigns two simplex circuits to each direction of every node pair. Lane-level matching assigns three along 𝐴 → 𝐵 →𝐶 →𝐴 and one along the reverse cycle. Table 1. Optical fabric terminology and notation.
a topology, the switch installs it, and the topology remains fixed until the next epoch. Ports and lanes. Each OCS port terminates a fiber pair—one fiber carries light into the switch, the other carries light out—attached to a duplex transceiver at the endpoint. A port therefore bundles exactly two simplex resources: a transmit (TX) lane and a receive (RX) lane, and a P-port switch exposes 2P independently steerable lanes. (Throughout this paper, "lane" denotes one direction of a port’s fiber pair, not a SerDes or wavelength lane inside a transceiver.) We call a light path from one TX lane to one RX lane a simplex circuit, and the conventional pair of simplex circuits between two ports a duplex circuit. Nodes and fabric boundary. A node is either an OCSfacing EPS, aggregating its attached servers’ traffic [36, 39], or a directly attached server [24, 41, 42]. Its 𝑘𝑛 OCS-facing ports provide 𝑘𝑛 TX and 𝑘𝑛 RX lanes. The OCS fabric comprises the optical switches and connecting fibers, with actual wiring and path constraints retained in §3.5. Figure 11 in Appendix E illustrates the two deployment mappings. These deployment mappings do not imply independent TX/RX hardware support (§5). Scheduling model and the duplex convention. Demand-aware OCS scheduling proceeds epoch by epoch. The scheduler takes as input a directed, node-level demand matrix 𝐷 (𝑘 ) = [𝑑𝑖 𝑗 ], where 𝑑𝑖𝑘𝑗 is the byte volume at which node i wishes to send to node j in communication operation 𝑘, and outputs a topology subject to port-count and reconfiguration constraints. All demand-aware OCS fabrics we are aware of express the topology as a duplex matching: an undirected matching over ports in which connecting port p to port q installs both the 𝑝 → 𝑞 and 𝑞 → 𝑝 cross-connects atomically [8, 14, 21, 24, 38, 41, 42]. Lane-granularity scheduling (§3) instead computes two directed matchings—one from TX lanes to RX lanes in each direction—while preserving node-level duplex: every node keeps at least one active TX lane and one active RX lane.
Term / symbol
Meaning
Node 𝑛, 𝑘𝑛 Port; 𝑃
Server or EPS with 𝑘𝑛 attached duplex ports. Duplex endpoint interface; total attached port count. TX / RX lane; 𝑐 Fixed transmit / receive resource; lane rate. Simplex circuit One TX-to-RX light path in use. Duplex circuit Reciprocal pair of simplex circuits between ports. 𝑋𝑢𝑣 Number of circuits on directed logical link 𝑢 → 𝑣. 𝐷 (𝑘 ) Directed node-pair byte demands of operation 𝑘. Configuration Set of simultaneous TX-to-RX lane bindings. Segment Fixed-configuration interval / consecutive operations assigned to it. Reconfiguration de- Exposed time to change configuration and lay restore usable communication.
Table 1 summarizes the terminology and notation used throughout the paper. 2.2
Directional Demand Exists Between Nodes
Large-model training distributes computation across ranks, i.e., the participants in a parallel execution. Tensor parallelism (TP) partitions tensor computations across ranks; data parallelism (DP) maintains model replicas and exchanges their updates through collectives such as AllReduce, ReduceScatter, and AllGather; pipeline parallelism (PP) places successive model stages on different ranks and exchanges activations and gradients between them. In the deployments considered in this paper, each TP group fits entirely within one server. TP communication therefore never traverses the optical fabric, whether a node is a server or an EPS. The inter-node demand considered here comes from the DP and PP transfers. One collective communication seems to be symmetric: each node may transmit and receive equal total volumes. 3
Liang et al. LLaMA-3.1
Qwen2.5
Gemma-2
OLMo-2
Mistral
BLOOM
TP1 / PP4 / DP128
78.8
94.2
97.9
80.1
84.5
92.3
94.4
62.2
82.6
91.7
86.6
90.6
95.9
76.4
81.6
98.9
TP4 / PP8 / DP16
42.9
76.8
90.5
44.9
52.4
70.9
77.3
25.0
49.0
69.2
56.7
66.1
82.7
39.6
47.3
95.0
TP8 / PP16 / DP4
12.3
38.1
63.9
13.2
17.0
31.2
38.8
5.8
15.2
29.5
19.6
26.6
47.2
10.9
14.4
77.9
8B
70B
405B
7B
14B
32B
72B
2B
9B
27B
7B
13B
32B
7B
Nemo 12B
176B
100 75 50 25 0
Figure 2. Traffic share on strongly skewed (𝜌𝑢𝑣 > 4) node pairs across 48 complete mixed schedules at world size 512.
2.3
0 16.7 40.0 75.0 133 250 600
600 500 400 300 200
PP backward 600 250 133 75.0 40.0 16.7 0
100
DP ReduceScatter 4.70 18.5 40.7 75.0 133 248 595 DP AllGather 4.70 18.5 40.7 75.0 133 248 595 PP forward
Weighted slowdown (%)
However, when viewed as node pairs, the same exchange can be asymmetric. In a selected ring 𝐴 → 𝐵 → 𝐶 → 𝐴, A sends bulk data to B but receives bulk data from C. Equal total sends and receives at A do not require equal demand on 𝐴 → 𝐵 and 𝐵 →𝐴. Figure 2 examines how much training traffic exhibits this pairwise imbalance. We analyze 48 graph-derived mixed training schedules across 16 dense models [1, 3, 9, 13, 20, 29, 35, 43], each using 512 GPUs organized into 64 eight-GPU servers. For each schedule, we aggregate transfers from all ranks by ordered node pair, exclude intra-node traffic, and include the calibrated forward and reverse control overheads. Let 𝐷𝑢𝑣 and 𝐷 𝑣𝑢 denote the resulting byte volumes in the two directions over the complete schedule. We identify a pair as strongly skewed when 𝜌𝑢𝑣 = max{𝐷𝑢𝑣 , 𝐷 𝑣𝑢 }/min{𝐷𝑢𝑣 , 𝐷 𝑣𝑢 } > 4. Each heatmap cell reports the fraction of inter-server bytes carried by strongly skewed pairs, counting both directions. It ranges from 5.8% to 98.9%, with a median of 65.0%. Thus, in the evaluated schedules, much of the traffic crosses node pairs whose directional demands differ substantially. For switch nodes, the same calculation aggregates across all servers attached to each switch. Aggregation can absorb transfers within a switch or combine traffic in opposite directions, so server-level skew does not automatically carry over to switch pairs. Grouping four consecutive servers per switch in the same workloads gives a median strongly skewed traffic share of 47.7%; grouping eight gives 0% over the complete schedules. The latter does not imply exact symmetry or rule out skew within individual phases.
PP complete 300 100 33.3 0 33.3 100 300 Mixed complete 58.7 35.7 43.0 66.4 118 223 541
7/1 6/2 5/3 4/4 3/5 2/6 1/7 Signed static split: u v / v u
0
Figure 3. Traffic-weighted slowdown of fixed directional capacity splits relative to the best pair-local split in each communication window. This analysis isolates directional allocation and does not enforce shared node-level lane inventories
RX lane from a TX lane at another node [5, 17, 18]. For example, a port at node A can transmit to B and receive from C. Across its ports, A can therefore allocate more TX lanes to B and more RX lanes to C, while maintaining nodelevel duplex. This capability applies to both server nodes and switch nodes: the OCS configures optical paths between their lanes, regardless of whether those lanes terminate on NICs or EPS uplinks. Each node retains its fixed inventory of 𝑘𝑛 TX lanes and 𝑘𝑛 RX lanes; lane-level matching changes their peers, not their transmit or receive roles. Consider three nodes, each with four ports. Suppose the directions 𝐴 → 𝐵, 𝐵 → 𝐶, and 𝐶 → 𝐴 each carry 3𝑉 bytes, while each reverse direction carries 𝑉 . A duplex matching can allocate two simplex circuits to each direction of every node pair. Lane-level matching can instead allocate three along the forward cycle and one along the reverse cycle. Both configurations use four TX lanes and four RX lanes at every node. At a per-lane rate of 𝑐 bytes per second, assuming concurrent transfers with no other bottleneck, the transfer time falls from 3𝑉 /(2𝑐) to 𝑉 /𝑐. The improvement comes from matching the same optical resources to the directional demand.
OCS Hardware Supports Lane-Level Matching
The directional demand above can be served by assigning different numbers of simplex circuits to opposite directions of a node pair. Under duplex matching, connecting port 𝑝 at node 𝑢 to port 𝑞 at node 𝑣 reserves both 𝑝 TX → 𝑞 RX and 𝑞 TX → 𝑝 RX . With equal-rate lanes, this enforces 𝑋𝑢𝑣 = 𝑋 𝑣𝑢 , coupling the capacity of the two directions even when their demands differ. The separate TX and RX lanes provide a way to remove this coupling. When these lanes attach to independently configurable OCS interfaces, the switch can connect a port’s TX lane to an RX lane at one node while connecting its 4
LACE: Lane-Granularity OCS Scheduling
Optical support alone does not ensure usable connections. Helios explored unidirectional circuits [12], and measurements show that simplex reconfiguration can trigger Ethernet link-failure handling [44]. NICs and EPS uplinks must remain operational with different TX and RX peers, and feedback must reach the sender. Exploiting this capability therefore requires link and feedback support alongside scheduling.
reconfiguration and node recovery delay. Consecutive operations should therefore share a configuration unless changing it saves more transfer time than it costs. The next section presents LACE, which jointly chooses segment boundaries and lane allocations to minimize communication completion time, including reconfiguration delay.
2.4 Asymmetric Connectivity Needs Scheduling The hardware capability above raises a scheduling question: how should lane allocations follow demand? DP collectives and PP transfers change node-pair traffic in both volume and direction, so an allocation that fits one communication window may not fit the next. Figure 3 shows why one fixed allocation is not enough. For each communicating pair, we compare seven allocations of eight simplex circuits between the two directions, from seven in one direction and one in the other (7 : 1) to the reverse (1 : 7). We compare the transfer time under each allocation with that under the best allocation for the same pair and communication window, and average the slowdown weighted by traffic volume. DP ReduceScatter, DP AllGather, and forward PP perform best with 7 : 1, while backward PP performs best with 1 : 7, and complete PP with 4 : 4. No single fixed allocation performs best across all the evaluated communication windows, motivating lane allocations that adapt during training. One possible way to avoid the limitations of fixed allocations is to make communication demand symmetric. Existing collective optimizations, however, do not guarantee this. NCCL selects algorithms and channels for efficient communication [15, 34], but equal total sends and receives at a node do not imply equal traffic in both directions of each node pair. Bidirectional Ring AllReduce and RingBiOdd can balance payload across opposite-direction rings, but require suitable connectivity and collective-specific schedules [23]; this balance does not extend automatically to other communication operations. SCCL and TACCL synthesize topology-aware collective schedules [4, 37], but minimizing communication time does not require pairwise symmetric traffic. AdapCC adapts communication to resource and network changes [45], while ResCCL improves resource use through task scheduling [26]; neither objective guarantees equal opposite-direction traffic between nodes. Thus, even when individual collectives are optimized or balanced, asymmetric demand can remain as DP collectives and PP transfers execute and overlap. Since asymmetric demand remains, lane connections must be chosen jointly across node pairs and time. Each simplex circuit consumes a source TX lane and a destination RX lane, so the pairwise choices in Figure 3 must fit shared node inventories. Changing connections also incurs optical
3
LACE Design
3.1
From Communication Operations to Lane Connections
LACE takes an ordered communication schedule, collective algorithms, rank placement, node port inventories, and the physical OCS topology. It outputs which operations share a configuration and the lane connections used by each configuration. Figure 4 shows four stages. The Demand Profiler aggregates transfers into node-level demand. The Segment Planner groups consecutive operations and computes fractional circuit allocations, balancing transfer and reconfiguration costs. The Circuit Realizer converts them into integer counts, and the Lane Binder assigns physical lanes and optical paths. Nodes execute these connections using software acknowledgments and coordinated link configuration and recovery. 3.2
Constructing Node-Level Demand
The profiler expands collectives into directed rank transfers and maps ranks to server or switch nodes. For each operation, it sums bytes by ordered node pair to form 𝐷 (𝑘 ) , excluding transfers within a node and combining transfers across its ports. It preserves operation order and keeps opposite directions separate. Calibrated reverse control traffic is included to reserve return connections for feedback (§3.6). Appendix A formalizes the demand and resource constraints. 3.3
Planning Segments and Fractional Allocations
The Segment Planner uses fractional allocations to compare candidate segments without solving an integer problem for every candidate. These targets estimate the cost of sharing a configuration; the Circuit Realizer makes them executable afterward. Starting with one segment per operation, the planner merges adjacent segments when sharing costs less than the reconfiguration it avoids. Evaluating a shared allocation. For segment 𝐼 , let 𝑌𝑢𝑣 be the fractional number of simplex circuits from 𝑢 to 𝑣. Each allocation must fit within the source’s TX lanes and the destination’s RX lanes, with positive allocations for all demanded directions. At per-lane rate 𝑐, the estimated segment time is 𝐶 (𝐼, 𝑌 ) =
∑︁ 𝑘 ∈𝐼
(𝑘 ) 𝑑𝑢𝑣 . (𝑘 ) 𝑑𝑢𝑣 >0 𝑐𝑌𝑢𝑣
max
(1)
Transfers within an operation proceed concurrently, so its slowest transfer determines completion. Summing these 5
Liang et al.
Figure 4. LACE converts a training communication schedule into physical lane connections through four stages. Why direct rounding is insufficient. Suppose A sends (13, 8, 4) units to B, C, and D, with five TX lanes, sufficient destination RX lanes, and 𝑐 = 1. The fractional allocation (2.6, 1.6, 0.8) finishes in 5 time units, but rounding gives (3, 2, 1), requiring six lanes. Removing one A-to-B circuit gives (2, 2, 1) and time 6.5; removing one A-to-C circuit instead gives (3, 1, 1) and time 8. Removing A-to-D leaves demand unserved. Thus, repairing rounded allocations requires preserving demand coverage and evaluating transfer time, rather than merely satisfying lane budgets. Constructing and pruning an allocation. The realizer first assigns one circuit to every demanded direction. If this exceeds an inventory, the segment cannot be served directly by one configuration. Otherwise, it adds circuits in order of largest positive gap 𝑌𝑢𝑣 − 𝑋𝑢𝑣 , with a stable node-pair order for ties, checking source TX and destination RX availability. If the tolerance remains unmet and lanes are available, it attempts additions with the largest reduction in segment time. Once the tolerance is met, the realizer tries removing circuits in order of smallest time increase, accepting only removals that preserve coverage and the tolerance. For the example, (2, 2, 1) meets 𝜖 = 30%; any removal would violate coverage or this tolerance. If the search cannot meet the tolerance, it reports the shortfall with the best feasible allocation found. It passes integer counts to the Lane Binder. Appendix C gives the constraints and time changes used to evaluate additions and removals.
times models operations executing sequentially in the selected schedule. Consider demands (8, 2) and (6, 2) on (𝐴 → 𝐵, 𝐴 → 𝐶), with four TX lanes at A, sufficient RX lanes at B and C, and 𝑐 = 1. Separate allocations (3.2, 0.8) and (3, 1) finish in 2.5 and 2 time units. Sharing (3, 1) takes 8/3 + 2 ≈ 4.67, slightly longer than 4.5, but avoids a reconfiguration. Finding a fractional allocation. The allocator normalizes each operation’s demands by its largest entry, sums them by direction, and uses the square roots as initial weights. Larger weights receive more fractional circuits, subject to both lane inventories. It then evaluates the original byte volumes and reweights the slowest or nearly slowest transfers. After a fixed number of rounds, it returns the best allocation found and its estimated time, 𝐶b(𝐼 ). Merging adjacent segments. For adjacent segments 𝐼, 𝐽 , the saving is save(𝐼, 𝐽 ) = 𝐶b(𝐼 ) + 𝐶b(𝐽 ) + 𝛿 − 𝐶b(𝐼 ∪ 𝐽 ),
(2)
where 𝛿 is the reconfiguration delay. In the example, 𝛿 = 0.5 makes merging save 2.5 + 2 + 0.5 − 4.67 ≈ 0.33. If the next operations have demands (2, 8) and (2, 6), they similarly benefit from sharing (1, 3). This produces two segments, [1, 2] | [3, 4], taking 4.67 + 0.5 + 4.67 ≈ 9.83 time units. Combining all four would take 14 under their best shared allocation (2, 2), so the planner retains the two segments. A priority queue selects the largest positive saving. After each merge, only candidates involving the new segment and its neighbors are recomputed; planning stops when no saving remains. Appendix B details the fractional allocator and segment-merging procedure. 3.4
3.5
Binding Circuits Across the OCS Fabric
The Lane Binder assigns each circuit a source TX lane, a destination RX lane, and an optical path. With one nonblocking OCS, it pairs 𝑋𝑢𝑣 distinct TX lanes at 𝑢 with RX lanes at 𝑣, using each lane at most once. Multiple OCSes add internal path constraints. Routing and repair through multiple OCSes. The binder expands circuit counts into individual requests and jointly selects their lanes and paths. Each path reserves its
Realizing Integer Simplex Circuits
The Circuit Realizer converts fractional allocations 𝑌 into integer counts 𝑋 , seeking communication time within 1 + 𝜖 of the planner’s estimate under the node lane inventories. Here, 𝜖 is the allowed increase in communication time. 6
LACE: Lane-Granularity OCS Scheduling
TX/RX lanes, inter-switch fibers, and OCS cross-connects; no two circuits may occupy the same directional resource. Two requests can therefore conflict on an internal fiber even when both nodes have free lanes. The binder first reroutes circuits, including previously assigned paths, while preserving the requested counts. For the evaluated three-stage Clos, the binder represents inter-leaf circuit requests as a bipartite multigraph: source leaves form one side, destination leaves the other, and each requested circuit contributes an edge. An edge color selects a middle plane, with edges sharing a source or destination leaf assigned different colors. This constructs conflict-free paths when enough middle planes are available. When internal resources are constrained, the binder first attempts rerouting, including previously assigned paths, while preserving the requested counts. If routing still fails, it removes an optional circuit with the smallest increase in segment time and retries. A removal must leave every demanded direction served. The binder checks the resulting time against the Realizer’s tolerance and reports any excess. If minimum coverage cannot be routed, it returns the segment for replanning. A heuristic routing failure is an unresolved mapping, not proof of physical infeasibility. Completing connections at active ports. The targeted commodity nodes require both lanes of each enabled port to be connected, potentially to different peers. The binder completes any unconnected lane using available resources. These added circuits need not carry data, but consume lanes and optical paths and are checked together with data-carrying circuits. Unused ports remain disabled. If completion fails, the binder revises lane assignments or returns the configuration for reallocation. Node readiness is verified before transmission (§3.6). Retaining connections across segments. The binder prefers existing bindings and paths. For example, changing (2, 2, 1) to (1, 3, 1) can retain four circuits and redirect only one A-to-B circuit toward C. In a multi-OCS fabric, retaining a circuit also requires retaining its path; paths may move when they block new requests. The output specifies lane bindings and cross-connects for each OCS. Appendix D details port-completion accounting, the Clos path construction, and reuse and repair conditions. 3.6
choice, not a requirement of lane-level scheduling. UCCL similarly uses UC with software reliability for GPU collective communication [46]. Chunks carry transfer IDs and sequence numbers. The receiver records completed chunks and batches cumulative acknowledgments, flushing on a timer or transfer completion. Each NIC owns its queue pairs and registered buffers; host software tracks completion across NICs. Calibrated reverse demand makes the Realizer reserve at least one reverse simplex circuit, so feedback returns directly between communicating nodes, potentially through a different port. An EPS forwards feedback using configured routes. When direct allocation is infeasible, forwarding through intermediate nodes is a fallback, subject to available paths and resources. The sender bounds outstanding data and retransmits after timeout; the receiver suppresses duplicates. A transfer completes only after all chunks are confirmed; source buffers are released only after local send completion and acknowledgment. §4.1 measures software-ACK performance on a single 100 Gbps port; multi-NIC runtime integration remains future work. §5 discusses compatibility with RC. Separate return paths have precedent in RFC 3077’s tunneling over a bidirectional network [10]; LACE instead uses reverse simplex circuits within the OCS fabric. Link configuration and recovery. Simplex reconfiguration can disrupt Ethernet link-failure handling [44]. Disabling autonegotiation addresses parameter negotiation, but not dependence on a working receive signal. The controller, which selects both ends of every circuit, provisions compatible rate and FEC settings from NIC, EPS, and module capabilities. It disables autonegotiation only on ports verified to support fixed settings and retains required link training. The Lane Binder supplies the active-port connections described in §3.5. Reconfiguration can still interrupt receive signals. Nodes drain outstanding transfers over the old connections before the controller installs new circuits and forwarding rules. It checks port status and sends configuration-ID probes; receiving nodes report through the management network. Transmission resumes only after verification. On timeout, the controller attempts bounded port reinitialization, then restoration of the previous configuration if necessary, keeping transmission paused until verification succeeds. Hardware fault detection remains enabled. The Segment Planner charges the exposed draining, switching, and recovery time as reconfiguration delay.
Executing Lane-Level Configurations
Executing lane bindings requires nodes to confirm delivery through potentially different ports and keep ports operational when their TX and RX lanes connect to different nodes. Software acknowledgments. For server nodes, LACE uses RDMA Unreliable Connected (UC) transfers with software acknowledgments. UC retains NIC data movement while placing reliability in software, allowing feedback received by one NIC to confirm data sent by another without sharing hardware queue-pair state. UC is an implementation
4
Evaluation
We evaluate lane-level scheduling through fixed-topology replay and training on a three-server OCS testbed, separate software-ACK and recovery experiments, and large-scale simulations of graph-derived LLaMA-3.1 workloads [9]. 7
ACTINA 30 20
28.03 15.60
10
thread (Figure 6). At 64 MiB, both approach 98 Gbps. The median combined endpoint CPU cost is 0.1759 CPU-s/GiB for RC and 0.1760 for UC with software ACKs. These processlevel costs include polling and do not isolate ACK-processing cycles (Appendix G). Under the Megatron configuration in Appendix H, estimated per-lane DP/PP payloads are 0.95–38.15 MiB, for which our 100 Gbps measurements show modest overhead.
LACE
Epoch time (h)
Communication time (s)
Liang et al.
0
4 3
3.61 2.84
2 1 0
(a) Llama 8B replay
(b) GPT-2 training
4.2
Figure 5. Fixed-topology testbed performance. (a) Median complete replay time over two runs. (b) One GPT-2 training epoch per topology.
Figure 6. Software ACKs: 30 paired trials per size. (a) Mean CCT increase over RC with 95% bootstrap intervals. (b) Median throughput; shading marks estimated per-lane payloads (Appendix H). K/M denote KiB/MiB. 4.1
Simulation Methodology
Workloads and node resources. The main experiments use LLaMA-3.1 70B and 405B Full schedules on 32–512 eightGPU servers. TP=8 stays within each server, so node-level demand contains only inter-server DP and PP transfers. The models use PP=8 and PP=16, respectively, with DP set by server count. The Demand Profiler preserves collective execution, rank placement, and operation order, and aggregates each operation’s rank transfers by server pair. All methods receive the same calibrated forward and reverse control bytes from NCCL/RoCE measurements of our testbed. Each server node has 16 ports: 16 TX lanes and 16 RX lanes, each at 400 Gbps. A sensitivity experiment uses eight 800-Gb/s ports, preserving 6.4 Tbps per direction. Defaults are 𝛿 = 20 ms and 𝜖 = 0.05. The Lane Binder completes both directions of every enabled port and checks lane inventories; additional connections consume resources but do not contribute extra data throughput. Methods and metric. We compare LACE with the duplex OCS scheduler ACTINA [42], using identical demand, ports, line rates, and reconfiguration delay. Symmetric LACE couples opposite-direction circuit counts at LACE’s segment boundaries, isolating directionality without independently optimizing a symmetric schedule. An ideal nonblocking bound, limited only by each node’s aggregate TX/RX rates, provides context. CCT sums operation completion times using the integer simplex circuit counts in Equation 1, plus 𝛿 between different configurations. Transfers within an operation proceed concurrently; operations follow the selected order. Identical adjacent configurations merge before charging reconfiguration. The model excludes computation, compute/communication overlap, packet dynamics, and host-side transport processing limits, including software-ACK service capacity. Simulated CCT gains therefore measure communication scheduling, not training speedup.
Physical Testbed Evaluation
Setup. Three GPU servers connect to an 8 × 8 Polatis OCS, each through two 100 Gbps NIC ports with separate TX and RX fibers. We compare LACE with ACTINA [42], a state-ofthe-art OCS scheduling approach for distributed AI training. ACTINA and LACE compute their topologies offline: a symmetric triangle and a directed double ring, respectively. Each remains fixed throughout replay or training. Both methods use TCP with the same striping policy. Appendix G gives the workload and transport details. Replay and training. We replay the complete mixed DP/PP schedule of a LLaMA-3.1 8B workload with TP=4, PP=2, DP=3 and eight microbatches. Its 67 slots represent 24 logical GPUs. LACE achieves a 1.80× replay speedup (Figure 5). Although bidirectional 1F1B PP is 49.9% slower, the DP improvement reduces total replay time by 44.3%. We also fine-tune GPT-2 Small on WikiText-103 with one GPU per server, identical initial weights and data order. Over 4,846 optimizer steps, training speeds up by 1.27×. Held-out losses are 2.8169 and 2.8166, with no skipped updates. Software acknowledgments. We compare RDMA RC with UC plus cumulative software acknowledgments using the same port for data and feedback. Both use host buffers and one busy-polling thread per host, with no additional ACK
4.3
Communication Completion Time
Lane-level scheduling improves directional communication. At 16 ports per server, LACE speeds up ACTINA by 1.86–1.89× for 70B and 1.21–2.04× for 405B (Figure 8). Relative to the same-boundary symmetric ablation, gains are 1.57–1.67× for 70B. For 405B, the gain is only 1.04× at 32 servers, where DP=2 makes DP transfers symmetric; it 8
1.014
0.932
0.807
0.774 0.603
0.511
0.5
0.0
LACE
Sym. Static Fixed Per-op Prop. Greedy
1.0
1.000
1.037
1.053
0.5
0.0
Cont. Without With target prune prune
(a) Module replacement
Mean physical circuits
1.000
1.0
Relative comm. time
Relative speed
LACE: Lane-Granularity OCS Scheduling
16000 22.0%
8000
0
19.3% 13.9% 13.5% 8.9%10.1%
32
64 128 256 512 1024
(b) Realization
Servers
(c) Pruning
Figure 7. Module ablation (70B Full, 128 servers). (a) Speed normalized to LACE. (b) Communication time normalized to the fractional target. (c) Speed and mean physical circuits normalized to no pruning.
1.5
LACE
ACTINA
Symmetric LACE
Ideal bound
(a) 70B: 16×400 Gbps
(b) 70B: 8×800 Gbps
(c) 405B: 16×400 Gbps
(d) 405B: 8×800 Gbps
32
32
fixed-length partition, or one configuration per operation reduces speed by 22.57%, 6.83%, and 48.88%, respectively. The best fixed length is 50 among {2, 5, 10, 20, 50, 100}. Allocating fractional circuits in proportion to aggregate demand within the same segments reduces speed by 19.27%. These controlled replacements show why both changing demands and the cost of changing configurations matter. Pruning saves circuits with a small time penalty. At 128 servers, communication time exceeds the fractional target by 3.72% without pruning and 5.29% with it (Figure 7(b)). Pruning reduces mean physical simplex circuits per configuration from 1,928 to 1,668 (13.49%), including port completion, at a 1.36% speed loss. Circuit savings rise from 8.9% at 32 servers to 22.0% at 1024 servers (Figure 7(c)). Greedy matches LACE under the same pruning setting, demonstrating the pruning tradeoff rather than an additional gain from performance augmentation.
1.0
CCT (s)
0.5 0.0 8 6 4 2 0 128
512
128
512
Servers
Figure 8. Full-schedule CCT at equal aggregate bandwidth: 6.4 Tbps per server per direction.
4.5
Switch-node aggregation. Figure 10 shows that LACE reduces CCT relative to symmetric allocation in both deployments. Server-node deployments show larger absolute time savings, while switch-node aggregation lowers the OCS-side CCT baseline. LACE retains substantial relative gains after aggregation, showing that independent directional allocation remains beneficial when switches serve as fabric nodes. Appendix E examines how servers per EPS and uplink provisioning affect gains when server-access costs are included. Multi-OCS routing and repair. We map the main configurations onto a three-stage OCS Clos with 16 servers per leaf, 256 access TX lanes and 256 access RX lanes per leaf, and 256 middle planes. Bipartite edge coloring constructs paths for all 368 segments across the 20 LACE and Symmetric LACE configurations. Both data and port-completion circuits satisfy lane and internal-link exclusivity. CCT is unchanged under the assumption that the stages reconfigure in parallel with one fabric-level 𝛿. Appendix D gives the edge-coloring construction and the resource checks applied during repair.
reaches 1.53–1.58× at 64–512 servers. Thus the benefit depends on the directional demand remaining in the selected schedule. At 512 servers, LACE completes the 70B and 405B schedules in 0.726 and 3.569 s, respectively. More ports improve allocation at the same total rate. As shown in Figure 8, using eight 800 Gbps ports instead of sixteen 400 Gbps ports raises LACE CCT by 6.33–11.19% for 70B and 3.66–3.83% for 405B at 64–512 servers. More ports permit smaller circuit-allocation steps and more simultaneous peers. The 405B, 32-server case is an exception: eight ports are 0.05% faster, reflecting the heuristic integer allocation rather than a strictly monotonic guarantee. 4.4
Deployment on Optical Fabrics
Module Ablation
Directionality, segmentation, and allocation each matter. Figure 7 (a) uses 70B Full at 128 servers. Coupling opposite directions retains 60.27% of LACE’s speed. Replacing adaptive segmentation with one static configuration, the best 9
1 0 256
Route only +Repair 192
128
64
0 256
60 40 20 0
1
10
100
LACE (70B)
5 0
0
0.1 (e) Tolerance ε
128
64
0.2
0.4
70B
1
64 128 256 512
405B
10
4.5 3.0 1.5 0.0
32
Server's number
32
64 128 256 512
100
800
Figure 10. OCS-side CCT under server- and switch-node deployments.
400
5
0
1
10
Discussion
Dependence on the training schedule. LACE exploits directional demand remaining after collective selection and rank placement; it does not require every operation to be asymmetric. Sequence parallelism, collective algorithms, and grouping servers behind an EPS can change both the volume and direction of OCS traffic. These choices must therefore be reflected in the Demand Profiler rather than assumed to preserve the reported gains. LACE currently plans an ordered schedule offline. Changes to that schedule require replanning, and extending the planner to overlapping operations or dynamic expert routing requires modeling their readiness and dependencies. Communication gains may not translate into training speedups under computation overlap. Port inventories and direct connectivity. The main deployment targets are multi-port server nodes and switch nodes, whose ports can provide direct return circuits for feedback. Single-port nodes with different TX/RX peers would need indirect return paths, as in prior OCS forwarding designs [27, 28]. For software ACKs, intermediate hosts or switches would relay feedback along a loop-free path, charging every hop’s traffic and resources; without such a path, the node retains a duplex connection. Port count alone does not determine transport compatibility: both single- and multiport servers must deliver feedback to the state that tracks the transfer, as discussed below. Likewise, sufficient TX and RX lanes do not guarantee routability through an arbitrary multi-OCS fabric. Our routing and repair results cover the evaluated Clos topology, and unresolved mappings require replanning rather than execution of a partial configuration. Endpoint hardware compatibility. Compatibility depends on port capabilities and configuration, rather than NIC generation or whether the endpoint is a NIC, DPU, or EPS. Each must operate with different TX/RX peers under the rate, FEC, and training conditions in §3.6. A valid RX optical signal alone is insufficient: module loss-of-signal handling and Ethernet local/remote faults can affect transmission [19, 44]. Link-local PAUSE/PFC also requires feedback to reach the
100
(f) Delay (ms)
Figure 9. Deployment and sensitivity at 128 servers with 16 × 400 Gbps ports. (a–b) Routing and repair for four 70B segments; speed is relative to 256 planes, and routability does not imply meeting 𝜖. (c–d) Planner segmentation of 70B and 405B Full schedules. (e) 70B CCT increase over 𝜖 = 0 at 𝛿 = 20 ms. (f) 70B CCT versus simulated delay at 𝜖 = 0.05. To test constrained internal connectivity, we reduce middle planes for the four 70B, 128-server segments (Figure 9(a– b)). Route only succeeds for all four through 128 planes, but only three at 96 and two at 64. Repair preserves every demanded direction and routes all four at 98.85% and 80.29% of the 256-plane speed. Two segments meet 𝜖 after either repair, compared with three initially. These results validate routing and repair for the evaluated uniform Clos with leaf-local connections. 4.6
0.8
0.0 10
(d) Delay (ms)
10
405B
100
(c) Delay (ms) 15
Switch: Symmetric
1.2
(b) Middle planes
Avg. length (ops)
(a) Middle planes
192
Switch: LACE
Server: Symmetric
70B
Route only: incomplete from 96
50
Server: LACE
CCT (s)
2
100
CCT (s)
3
Relative speed (%)
4
CCT (ms)
CCT increase (%) Number of segments
Routable segments
Liang et al.
Segmentation and Parameter Sensitivity
Higher reconfiguration cost favors longer segments. At 128 servers, the 70B and 405B Full schedules contain 274 and 1,445 operations. With 𝛿 = 1 ms, the Planner forms 28 and 60 segments; at 20 ms, these fall to four and 33; at 100 ms, to two and four (Figure 9(c–d)). These counts precede integer realization and merging of identical configurations. Tolerance and delay affect different parts of the schedule. For 70B Full at 128 servers, raising 𝜖 from 0.05 to 0.1 increases CCT by 2.92%; the increase from zero to 0.2 is 13.40% (Figure 9(e)). Raising 𝛿 from 1 to 100 ms increases CCT from 0.484 to 0.791 s (63.44%), while the number of segments falls from 28 to two. At 30–100 ms, one configuration change remains, so CCT continues to grow with its cost (Figure 9(f)). Appendix F reports offline scheduling runtime separately from communication time. 10
LACE: Lane-Granularity OCS Scheduling
correct upstream sender. Our testbed validates its installed NIC, module, and firmware combination, not all such platforms. Drain-and-verify checks transition success; it cannot make unsupported hardware compatible. Such ports require compatible duplex bindings or platform changes. Optical switching architectures. 3D-MEMS and piezoelectric beam steering permit lane-level allocation when fibers and controls are independently exposed. Wavelengthselective switches (WSS) and planar-waveguide networks require checking their wavelength, path-conflict, and sharedelement constraints. Tunable lasers and arrayed-waveguide gratings (AWGRs) use wavelength to select destinations [2]; adapting LACE would require wavelength-aware allocation and binding. Hardware that couples both directions, or an API exposing only duplex port pairs, cannot directly express our configurations. Reconfiguration delay. The Segment Planner charges optical switching, link recovery, and verification, favoring longer-lived configurations when recovery is slow. Our prototype observes second-scale recovery (Appendix G), as also reported for commodity link initialization by MixNet [24]. The millisecond-delay simulations assume correspondingly faster recovery; the testbed measures recovery separately from fixed-topology replay and training. Faster recovery would enable shorter communication phases without changing the scheduling model. Compatibility with RC. RC feedback must reach the NIC/QP owning the connection; delivery to another independent NIC on the same server is insufficient. RoCEv2 routing permits different forward/return optical paths [32]. An EPS can forward feedback to the original server NIC without terminating the QP. A conservative integration would reserve reciprocal circuits between the same physical port pair, leaving other lanes independently schedulable. This requires port-binding constraints; symmetric node-level circuit counts alone are insufficient. Alternatively, routed return paths could preserve RC endpoints, subject to routing and transport constraints. Both options require return-path resource accounting and draining RC traffic before reconfiguration. They remain unevaluated. Our UC bulk-transfer mechanism does not transparently implement all RC verbs, including RDMA reads and atomics. Transport scaling and node hardware compatibility. Concurrent 400/800 Gbps ports may expose unmeasured CPU limits. Completion batching, state sharding across cores, and NIC/DPU offload are candidates; ACK batching alone cannot remove per-chunk processing. Future NIC interfaces could expose independent data/feedback paths and a stable transport identity, directing feedback to its reliability context. Hardware could handle cumulative ACKs, retransmission, duplicate suppression, and completion; independent NICs would additionally need explicit state forwarding or sharing. A quiescence interface could coordinate reconfiguration with in-flight transfers. Standardizing and implementing these
capabilities could eliminate host-side ACK processing on supported platforms. LACE would continue to allocate directional capacity and configuration boundaries; software ACKs provide an execution path on current hardware.
6
Related Work
Optical circuit scheduling. ACTINA allocates optical resources among predictable TP, DP, and PP communication domains [42]. MixNet adjusts regional connectivity for dynamic MoE traffic [24], while Opus reconfigures photonic rails at parallelism phase boundaries [8]. RotorNet and Opera instead use traffic-oblivious topology schedules: RotorNet supports two-hop indirect forwarding [28], and Opera combines multi-hop forwarding for latency-sensitive traffic with direct circuits for bulk transfers [27]. LACE targets the directed node-pair demand of a selected training schedule. It jointly chooses which consecutive operations share a configuration and how many simplex circuits serve each direction, subject to each node’s TX and RX lane inventories. Collective communication and topology optimization. NCCL selects collective algorithms and channels [15, 34]; SCCL and TACCL synthesize topology-aware collective schedules [4, 37]; and AdapCC and ResCCL optimize collective execution and scheduling [26, 45]. RingBiOdd constructs bidirectional AllReduce schedules for odd MCM meshes [23]; RECCL adapts communication relationships to optical connectivity [47]; and TopoOpt jointly selects parallelization, topology, and routing [41]. These methods optimize how communication is performed or jointly select its topology. LACE takes the communication operations and rank placement as inputs, preserving their order while allocating simplex circuits and selecting reconfiguration boundaries. Unidirectional connectivity and link operation. Helios explores unidirectional circuits and identifies the bidirectional assumption in Ethernet fault management [12]. Zerwas et al. experimentally examine how optical reconfiguration interacts with Ethernet link-failure handling [44]. RFC 3077 supports unidirectional links through link-layer tunneling over a separate bidirectional network [10]. These precedents establish that unidirectional connectivity and separate return paths are not new in themselves. LACE combines training-aware lane allocation and segmentation with software acknowledgments over reverse simplex circuits and coordinated link configuration and recovery. Optical fabric architectures. Apollo and Jupiter establish production OCS deployment and topology engineering [36, 39]; TPU v4 uses OCS to configure ML supercomputer partitions [21]. LumosCore and InfiniteHBD provide scalable optical architectures for large AI clusters [14, 38]. LACE focuses on scheduling for fabrics exposing independently configurable TX-to-RX connections. Its Lane Binder maps simplex circuit allocations to physical lanes and optical 11
Liang et al.
paths under the deployed fabric’s constraints; our multi-OCS evaluation considers a three-stage Clos.
7
Amy Yang, Angela Fan, et al. 2024. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783. https://arxiv.org/abs/2407.21783 [10] Emmanuel Duros, Walid Dabbous, Hitoshi Izumiyama, Naoki Fujii, and Y. Zhang. 2001. A Link-Layer Tunneling Mechanism for Unidirectional Links. RFC 3077. RFC Editor. doi:10.17487/RFC3077 [11] Nathan Farrington, Yeshaiahu Fainman, Hong Liu, George Papen, and Amin Vahdat. 2011. Hardware Requirements for Optical Circuit Switched Data Center Networks. In Optical Fiber Communication Conference/National Fiber Optic Engineers Conference 2011. Optica Publishing Group, Washington, DC, USA, OTuH3. doi:10.1364/OFC.2011. OTuH3 [12] Nathan Farrington, George Porter, Sivasankar Radhakrishnan, Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fainman, George Papen, and Amin Vahdat. 2010. Helios: A Hybrid Electrical/Optical Switch Architecture for Modular Data Centers. In Proceedings of the ACM SIGCOMM 2010 Conference. Association for Computing Machinery, New York, NY, USA, 339–350. doi:10.1145/1851182.1851223 [13] Gemma Team, Morgane Riviere, et al. 2024. Gemma 2: Improving Open Language Models at a Practical Size. arXiv preprint arXiv:2408.00118. https://arxiv.org/abs/2408.00118 [14] Xinchi Han, Shizhen Zhao, Yongxi Lv, Peirui Cao, Weihao Jiang, Shengkai Lin, and Xinbing Wang. 2024. LumosCore: Highly Scalable LLM Clusters with Optical Interconnect. CoRR. [15] Zhiyi Hu, Siyuan Shen, Tommaso Bonato, Sylvain Jeaugey, Cedell Alexander, Eric Spada, James Dinan, Jeff Hammond, and Torsten Hoefler. 2025. Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and Algorithms. In 2025 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE, Piscataway, NJ, USA, 48–59. [16] Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism. Advances in Neural Information Processing Systems 32 (2019), 10 pages. https://proceedings.neurips.cc/paper/2019/hash/ 093f65e080a295f8076b1c5722a46aa2-Abstract.html [17] HUBER+SUHNER Polatis. 2012. POLATIS Series 6000 Single-Mode All-Optical Matrix Switch. Product webpage. Accessed: 2026-08-01. https://www.polatis.com/polatis-series-6000-optical-matrix-switch192x192-sdn-enabled-industry-leading-performance-lowest-lossswitches.asp [18] HUBER+SUHNER Polatis. 2016. POLATIS Series 7000 SoftwareDefined Optical Circuit Switch. Product webpage. Accessed: 202607-31. https://www.polatis.com/series-7000-384x384-port-softwarecontrolled-optical-circuit-switch-sdn-enabled.asp [19] Intel. 2025. F-Tile Low Latency 100G Ethernet Intel FPGA IP User Guide: Link Fault Signaling Interface. Version 24.3.1, Section 5.1.4. https://www.intel.com/content/www/us/en/docs/ programmable/792946/24-3-1/link-fault-signaling-interface.html [20] Albert Q. Jiang et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825. https://arxiv.org/abs/2310.06825 [21] Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, et al. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings. In Proceedings of the 50th Annual International Symposium on Computer Architecture. Association for Computing Machinery, New York, NY, USA, 1–14. [22] Mehrdad Khani, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, and Eiman Ebrahimi. 2021. SiP-ML: High-Bandwidth Optical Network Interconnects for Machine Learning Training. In Proceedings of the ACM SIGCOMM 2021 Conference. Association for Computing Machinery, 657–675. https://research.nvidia.com/publication/2021-08_sip-ml-
Conclusion
This paper established that pairwise directional skew is a recurring property of training traffic, and that assigning duplex circuits to the two directions of a node pair can leave lanes underutilized. LACE addresses this mismatch through independent TX/RX allocation. It jointly plans configuration boundaries and simplex-circuit allocations offline, then realizes physical connections under node and fabric constraints, without changing the selected collective algorithms, operation order, or rank placement. We separately evaluate software acknowledgments and configuration recovery. With sixteen 400 Gbps ports per server, simulations of LLaMA-3.1 70B and 405B show 1.21–2.04× communication speedup over ACTINA. Our three-server fixed-topology testbed achieves 1.80× communication-replay and 1.27× GPT-2 training speedups over ACTINA.
References [1] Allen Institute for AI. 2025. OLMo 2 32B Model Card. Hugging Face model card. Accessed: 2026-07-31. https://huggingface.co/allenai/ OLMo-2-0325-32B [2] Hitesh Ballani, Paolo Costa, Raphael Behrendt, Daniel Cletheroe, Istvan Haller, Krzysztof Jozwik, Fotini Karinou, Sophie Lange, Kai Shi, Benn Thomsen, and Hugh Williams. 2020. Sirius: A Flat Datacenter Network with Nanosecond Optical Switching. In Proceedings of the ACM SIGCOMM 2020 Conference. Association for Computing Machinery. https://www.microsoft.com/en-us/research/uploads/prod/2020/ 07/sirius-sigcomm20.pdf [3] BigScience Workshop, Teven Le Scao, et al. 2022. BLOOM: A 176BParameter Open-Access Multilingual Language Model. arXiv preprint arXiv:2211.05100. https://arxiv.org/abs/2211.05100 [4] Zixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi, Todd Mytkowicz, Jacob Nelson, and Olli Saarikivi. 2021. Synthesizing Optimal Collective Algorithms. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. Association for Computing Machinery, New York, NY, USA, 62–75. doi:10.1145/3437801.3441620 [5] Calient.AI, Inc. 2026. S320 Optical Circuit Switch. https://www.calient. net/. Accessed September 8, 2026. [6] Kai Chen, Ankit Singla, Atul Singh, Kishore Ramachandran, Lei Xu, Yueping Zhang, Xitao Wen, and Yan Chen. 2012. OSA: An Optical Switching Architecture for Data Center Networks with Unprecedented Flexibility. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12). USENIX Association. https://www.cs. northwestern.edu/~kch670/papers/osa-nsdi.pdf [7] William C. Dickson, Bryan P. Staker, Gene Campbell, and William C. Banyai. 2004. 64×64 3D-MEMS Switch Control System with Robustness to MEMS Resonant Frequency Variation and Pointing Drift. In Optical Fiber Communication Conference. Optica Publishing Group, Washington, DC, USA, ThQ5. doi:10.1109/OFC.2004.1362141 [8] Eric Ding, Barry Lyu, Bhaskar Kataria, and Rachee Singh. 2026. Opus: Photonic Rail-Optimized Fabric in ML Datacenters. In Proceedings of the ACM SIGCOMM 2026 Conference. Association for Computing Machinery, New York, NY, USA, 922–939. doi:10.1145/3789240.3829187 [9] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, 12
LACE: Lane-Granularity OCS Scheduling
[36] Leon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh, Mukarram Tariq, Rui Wang, Jianan Zhang, Virginia Beauregard, Patrick Conner, Steve Gribble, et al. 2022. Jupiter Evolving: Transforming Google’s Datacenter Network via Optical Circuit Switches and Software-Defined Networking. In Proceedings of the ACM SIGCOMM 2022 Conference. Association for Computing Machinery, New York, NY, USA, 66–85. [37] Aashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki, Madan Musuvathi, Todd Mytkowicz, Jacob Nelson, Olli Saarikivi, and Rachee Singh. 2023. TACCL: Guiding Collective Algorithm Synthesis Using Communication Sketches. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Berkeley, CA, USA, 593–612. [38] Chenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng, Yu Zhou, Yimin Jiang, Wenqing Lv, Yelong Xu, Yuanwei Lu, Zhang Chen, et al. 2025. InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domains for LLMs with Optical Circuit Switching Transceivers. In Proceedings of the ACM SIGCOMM 2025 Conference. Association for Computing Machinery, New York, NY, USA, 1–23. [39] Ryohei Urata, Hong Liu, Kevin Yasumura, Erji Mao, Jill Berger, Xiang Zhou, Cedric Lam, Roy Bannon, Darren Hutchinson, Daniel Nelson, et al. 2022. Mission Apollo: Landing Optical Circuit Switching at Datacenter Scale. arXiv preprint arXiv:2208.10041. https://arxiv.org/ abs/2208.10041 [40] Guohui Wang, David G. Andersen, Michael Kaminsky, Konstantina Papagiannaki, T. S. Eugene Ng, Michael Kozuch, and Michael Ryan. 2010. c-Through: Part-time Optics in Data Centers. In Proceedings of the ACM SIGCOMM 2010 Conference. Association for Computing Machinery, 327–338. https://conferences.sigcomm.org/sigcomm/2010/ papers/sigcomm/p327.pdf [41] Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Zhihao Jia, Dheevatsa Mudigere, Ying Zhang, and Anthony Kewitsch. 2023. TopoOpt: Co-Optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Berkeley, CA, USA, 739–767. [42] Zhenguo Wu, Benjamin Klenk, Larry Dennison, and Keren Bergman. 2025. ACTINA: Adapting Circuit-Switching Techniques for AI Networking Architectures. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, Piscataway, NJ, USA, 1211–1222. [43] An Yang et al. 2024. Qwen2.5 Technical Report. arXiv preprint arXiv:2412.15115. https://arxiv.org/abs/2412.15115 [44] Johannes Zerwas, Wolfgang Kellerer, and Andreas Blenk. 2021. What You Need to Know About Optical Circuit Reconfigurations in Datacenter Networks. In 2021 33rd International Teletraffic Congress (ITC 33). IEEE, Piscataway, NJ, USA, 1–9. https://dl.ifip.org/db/conf/itc/ itc2021/1570720746.pdf [45] Xiaoyang Zhao, Zhe Zhang, and Chuan Wu. 2024. AdapCC: Making Collective Communication in Distributed Machine Learning Adaptive. In 2024 IEEE 44th International Conference on Distributed Computing Systems (ICDCS). IEEE, Piscataway, NJ, USA, 25–35. [46] Yang Zhou, Zhongjie Chen, Ziming Mao, ChonLam Lao, Shuo Yang, Pravein Govindan Kannan, Jiaqi Gao, Yilong Zhao, Yongji Wu, Kaichao You, Fengyuan Ren, Zhiying Xu, Costin Raiciu, and Ion Stoica. 2025. An Extensible Software Transport Layer for GPU Networking. arXiv:2504.17307. doi:10.48550/arXiv.2504.17307 [47] Hong Zou, Huaxi Gu, Xiaoshan Yu, Zhuodong Wu, and Yifeng Zhu. 2025. RECCL: Optimizing Collective Algorithms for Reconfigurable Optical Networks. Journal of Optical Communications and Networking 17, 6 (2025), 470–484. doi:10.1364/JOCN.555632
high-bandwidth-optical-network-interconnects-machine-learningtraining [23] Sabuj Laskar, Farabi Mahmud, Pranati Majhi, Abdullah Muzahid, Sungkeun Kim, and Eun Jung Kim. 2024. Enhancing Collective Communication in MCM Accelerators for Deep Learning Training. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, Piscataway, NJ, USA, 832–847. doi:10.1109/ HPCA57654.2024.00069 [24] Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang, Zhenghang Ren, Xinyang Huang, Wenxue Li, Kin Fai Tse, et al. 2025. MixNet: A Runtime-Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training. In Proceedings of the ACM SIGCOMM 2025 Conference. Association for Computing Machinery, New York, NY, USA, 554–574. [25] Hong Liu, Ryohei Urata, Kevin Yasumura, Xiang Zhou, Roy Bannon, Jill Berger, Pedram Dashti, Norm Jouppi, Cedric Lam, Sheng Li, Erji Mao, Daniel Nelson, George Papen, Mukarram Tariq, and Amin Vahdat. 2023. Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning Systems. In Proceedings of the ACM SIGCOMM 2023 Conference. Association for Computing Machinery. doi:10.1145/3603269.3604836 [26] Tongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiamin Cao, Tianshu Wang, Ennan Zhai, and Xingwei Wang. 2025. ResCCL: ResourceEfficient Scheduling for Collective Communication. In Proceedings of the ACM SIGCOMM 2025 Conference. Association for Computing Machinery, New York, NY, USA, 55–70. [27] William M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness, Alex C. Snoeren, and George Porter. 2020. Expanding across Time to Deliver Bandwidth Efficiency and Low Latency. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20). USENIX Association, 1–18. https://www.usenix.org/conference/ nsdi20/presentation/mellette [28] William M. Mellette, Rob McGuinness, Arjun Roy, Alex Forencich, George Papen, Alex C. Snoeren, and George Porter. 2017. RotorNet: A Scalable, Low-complexity, Optical Datacenter Network. In Proceedings of the Conference of the ACM Special Interest Group on Data Communication. Association for Computing Machinery, 267–280. doi:10.1145/3098822.3098838 [29] Mistral AI. 2024. Mistral NeMo 12B Model Card. Online model card. Accessed: 2026-07-31. https://docs.mistral.ai/models/modelcards/mistral-nemo-12b-24-07 [30] Yojiro Mori and Ken-Ichi Sato. 2021. High-Port-Count Optical Circuit Switches for Intra-Datacenter Networks [Invited Tutorial]. Journal of Optical Communications and Networking 13, 8 (2021), D43–D52. doi:10.1364/JOCN.425929 [31] Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, Phillip B. Gibbons, and Matei Zaharia. 2019. PipeDream: Generalized Pipeline Parallelism for DNN Training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles. Association for Computing Machinery, New York, NY, USA, 1–15. [32] NVIDIA. [n. d.]. RDMA over Converged Ethernet (RoCE). MLNX_OFED Software Documentation. Accessed September 10, 2026. https://networking-docs.nvidia.com/mlnxofedswum/24010331/ rdma-over-converged-ethernet-roce [33] NVIDIA Corporation. 2019. Megatron-LM: Training Large Transformer Language Models at Scale. GitHub repository. Accessed: 2026-03-05. https://github.com/NVIDIA/Megatron-LM [34] NVIDIA Corporation. 2024. NVIDIA Collective Communications Library (NCCL) User Guide: Overview. Online documentation. Accessed: 2026-03-05. https://docs.nvidia.com/deeplearning/nccl/userguide/docs/overview.html [35] OLMo Team, Pete Walsh, et al. 2025. 2 OLMo 2 Furious. arXiv preprint arXiv:2501.00656. https://arxiv.org/abs/2501.00656 13
Liang et al.
A
TX and RX inventories 𝑟𝑢TX and 𝑟 𝑣RX , each filling round adds 𝑟𝑢TX𝑤𝑢𝑣 𝑟 𝑣RX𝑤𝑢𝑣 . (7) Δ𝑌𝑢𝑣 = min Í ,Í 𝑗:(𝑢,𝑗 ) ∈ D𝐼 𝑤𝑢 𝑗 𝑖:(𝑖,𝑣) ∈ D𝐼 𝑤 𝑖𝑣
Node-Level Demand and Resource Constraints
Let U be the set of server nodes or switch nodes defined in §2.1. Node 𝑢 contributes 𝑘𝑢 TX lanes and 𝑘𝑢 RX lanes. Let 𝑐 be the per-lane rate in bytes per second; a rate specified in Gbps is converted before computing time. For operation (𝑘 ) 𝑘, 𝐷 (𝑘 ) = [𝑑𝑢𝑣 ] records directed byte demand, including the calibrated transport overhead used in the experiments. Transfers within a node are excluded. For a switch node, this also excludes transfers between servers attached to the same EPS. The profiler preserves operation order and does not symmetrize demand. An optional mask 𝐴𝑢𝑣 records whether a direct optical path can exist from 𝑢 to 𝑣. For the abstract nonblocking OCS, 𝐴𝑢𝑣 = 1 for all distinct nodes. Positive demand on an unavailable direction cannot be served directly. The mask expresses reachability, not simultaneous routability: shared optical-path constraints are checked by the Lane Binder (Appendix D).
B
Source proposals sum to at most the source’s remaining TX inventory; destination proposals do the same for RX. Taking the smaller proposal therefore preserves both constraints. Inventories are updated after each round. The allocator evaluates the original demands and emphasizes transfers close to each operation’s completion time. (𝑘 ) (𝑘 ) For nonzero operations, let 𝛽𝑢𝑣 = [𝑑𝑢𝑣 /(𝑐𝑌𝑢𝑣 )]/𝐿𝑘 (𝑌 ). The next weights are proportional to v u t∑︁ (𝑘 ) 𝑑𝑢𝑣 (𝑘 ) 8 𝑤𝑢𝑣 ← 𝛽𝑢𝑣 . (8) 2 𝑐𝑌𝑢𝑣 𝑘 ∈𝐼 Zero entries contribute zero. The exponent gives greater weight to slow transfers without replacing the max-based evaluation objective. The allocator performs three rounds of weight updates, each with twelve filling rounds, and retains the feasible allocation with the lowest 𝐶 (𝐼, 𝑌 ). If each node has at most one outgoing and one incoming demanded direction, the directions share no TX or RX budgets, and 𝑌𝑢𝑣 = min{𝑘𝑢 , 𝑘 𝑣 } is optimal for the fractional model. Adjacent merging. Starting with one segment per operation, the planner keeps adjacent merge savings from Equation 2 in a priority queue. After accepting a merge, it updates the neighboring candidates and discards stale entries. At termination, no evaluated adjacent merge has positive saving under the heuristic costs; this is not a globally optimal partition. With fixed candidate-solver effort and caching, the process evaluates 𝑂 (𝐾) candidates and uses 𝑂 (𝐾 log 𝐾) queue work for 𝐾 operations. Candidate evaluation also depends on the number of operations and demanded directions in each segment.
Fractional Segment Planning
The allocation 𝑌𝑢𝑣 is a fractional number of simplex circuits, as in §3.3; 𝑐𝑌𝑢𝑣 is its associated rate. The fractional feasible set satisfies ∑︁ ∑︁ 𝑌𝑢𝑣 ≤ 𝑘𝑢 , 𝑌𝑣𝑢 ≤ 𝑘𝑢 , ∀𝑢, (3) 𝑣≠𝑢
𝑣≠𝑢
𝑌𝑢𝑣 ≥ 0,
𝑌𝑢𝑣 = 0
if 𝑢 = 𝑣 or 𝐴𝑢𝑣 = 0.
(4)
Each direction demanded in a segment must have positive allocation. For operation 𝑘, define (𝑘 ) 𝑑𝑢𝑣 , (𝑘 ) 𝑑𝑢𝑣 >0 𝑐𝑌𝑢𝑣
𝐿𝑘 (𝑌 ) = max
𝐶 (𝐼, 𝑌 ) =
∑︁
𝐿𝑘 (𝑌 ).
(5)
𝑘 ∈𝐼
A zero-demand operation has zero optical time, and positive demand with zero allocation has infinite time. This is the sequential-operation model in Equation 1. Continuous reference. For fixed segment 𝐼 , minimizing 𝐶 (𝐼, 𝑌 ) over the fractional feasible set is convex: each reciprocal term is convex on positive allocations, maxima and sums preserve convexity, and the constraints are linear. An exact solution provides a reference for that segment, not a guarantee for LACE’s heuristic or its subsequent integer realization. Physical port completion and shared paths impose further constraints. Weighted filling. Let D𝐼 be the union of demanded directions in 𝐼 . The implementation initializes weights using v u t∑︁ (𝑘 ) 𝑑𝑢𝑣 (0) 𝑤𝑢𝑣 = , (𝑢, 𝑣) ∈ D𝐼 , (6) (𝑘 ) 𝑘 ∈𝐼 max𝑖,𝑗 𝑑𝑖 𝑗
C
Integer Simplex Circuit Realization
For segment 𝐼 , let D𝐼 contain every direction with positive demand in at least one operation, including reverse control traffic. Integer data-carrying circuit counts 𝑋 satisfy ∑︁ ∑︁ 𝑋𝑢𝑣 ≤ 𝑘𝑢 , 𝑋 𝑣𝑢 ≤ 𝑘𝑢 , ∀𝑢, (9) 𝑣≠𝑢
𝑣≠𝑢
𝑋𝑢𝑣 ∈ Z ≥0, 𝑋𝑢𝑣 ≥ 1,
𝑋𝑢𝑣 = 0
if 𝑢 = 𝑣 or 𝐴𝑢𝑣 = 0, (10) (𝑢, 𝑣) ∈ D𝐼 .
(11)
The performance target is 𝐶 (𝐼, 𝑋 ) ≤ (1 + 𝜖)𝐶b(𝐼 ), using Equation 5 with integer 𝑋 . This is a target checked after realization, not a universal approximation bound. Coverage and additions. One circuit per demanded direction is the minimum for direct service. This assignment fits the abstract OCS exactly when every direction is reachable and each node’s outgoing and incoming peer counts fit its lane inventories. It need not be physically realizable before port completion and routing. Further additions follow
with zero-demand operations contributing zero. Small numerical guards prevent division by zero. Given remaining 14
LACE: Lane-Granularity OCS Scheduling
positive fractional gaps 𝑌𝑢𝑣 − 𝑋𝑢𝑣 as described in §3.4. For a feasible one-circuit addition or removal, the time changes are + Δ𝑢𝑣 (𝑋 ) = 𝐶 (𝐼, 𝑋 ) − 𝐶 (𝐼, 𝑋 + 𝐸𝑢𝑣 ),
(12)
− Δ𝑢𝑣 (𝑋 ) = 𝐶 (𝐼, 𝑋 − 𝐸𝑢𝑣 ) − 𝐶 (𝐼, 𝑋 ),
(13)
Resources include TX/RX lanes, inter-OCS fibers, and switch input/output interfaces. Paths must also satisfy each OCS’s cross-connect constraints. Routed connections become usable only after the node-side verification in §3.6. Constructive mapping for the evaluated Clos. Each leaf serves sixteen server nodes through 256 access TX lanes and 256 access RX lanes. Circuits within one leaf use local connections. For inter-leaf requests, construct a bipartite multigraph with one copy of each source leaf on the left and each destination leaf on the right. Each simplex circuit is an edge. With maximum degree Δ, bipartite edge coloring uses Δ colors. Assigning each color to a middle plane ensures that no two circuits share a leaf’s outgoing or incoming connection to that plane. The mapping succeeds when the number of planes is at least Δ and the middle switches support the resulting matchings. Repair and reuse. When routing fails, the binder first tries alternative paths. If circuit counts must change, repair removes optional data circuits while preserving every demanded direction, then recomputes port completion and reroutes the entire 𝐻 . Candidate removals favor smaller increases in 𝐶 (𝐼, 𝑋 ); the evaluated Clos repair also checks that excess internal resource demand does not increase. Recomputing 𝑍 is necessary because removing a data circuit may change which Each accepted removal Í ports need completion. Í decreases 𝑋𝑢𝑣 , so at most 𝑢,𝑣 (𝑋𝑢𝑣 − 1[(𝑢, 𝑣) ∈ D𝐼 ]) removals are possible. Repair may exceed 𝜖; it reports the resulting time separately from routing success. A bounded or heuristic routing failure is an unresolved mapping, not proof of physical infeasibility. For consecutive configurations, the binder retains old connections whose lanes remain in the selected enabled-port sets, then binds remaining requests. The pairwise quantity old, 𝐻 new } bounds potential reuse, but port selection min{𝐻𝑢𝑣 𝑢𝑣 and internal paths can prevent attaining it.
where 𝐸𝑢𝑣 has a single unit entry. The evaluation implementation also attempts performance-guided additions using Δ+ when target alignment leaves the tolerance unmet and lanes remain. The reported 70B ablation does not exercise this additional pass. Pruning and outcomes. After meeting the target, pruning removes the circuit with the smallest Δ − among removals that preserve coverage and the time tolerance. Neither additions nor pruning may exceed a TX or RX inventory. If the search cannot meet the target, it returns its best feasible allocation with the measured excess time; this does not prove that no better integer allocation exists. If minimum direct coverage fails, the segment requires replanning or an explicitly modeled forwarding alternative. Identical adjacent configurations can share their operation lists and avoid a reconfiguration; physical bindings must also be retained for that boundary to incur no switching cost.
D
Port Completion, Routing, and Connection Reuse
Completing enabled ports. The commodity-port requirement in §3.5 applies to both lanes of each enabled port. Let 𝑍𝑢𝑣 count additional simplex circuits installed solely to complete unused directions of these ports. The Lane Binder checks the full configuration 𝐻 = 𝑋 + 𝑍 : ∑︁ ∑︁ 𝐻𝑢𝑣 = 𝐻 𝑣𝑢 = 𝑎𝑢 , 0 ≤ 𝑎𝑢 ≤ 𝑘𝑢 , (14) 𝑣
𝑣
where 𝑎𝑢 is the number of enabled ports at node 𝑢. Each of those ports has exactly one connected TX lane and one connected RX lane; their peers may differ. The binder must find actual lane assignments and paths, not merely satisfy these counts. Additional circuits consume physical resources but contribute no throughput to 𝐶 (𝐼, 𝑋 ).ÍThus the physical Í circuit footprint is 𝑢,𝑣 𝐻𝑢𝑣 , rather than 𝑢,𝑣 𝑋𝑢𝑣 . Peak dataonly TX/RX counts provide lower bounds; they do not by themselves establish a minimum physical port footprint in a constrained fabric. Path constraints. Expand 𝐻 into individual simplex circuit requests L (𝐻 ). For request ℓ, let Pℓ contain candidate optical paths, including their source and destination lanes. Binary choices 𝑧 ℓ𝑝 obey ∑︁ 𝑧 ℓ𝑝 = 1, ℓ ∈ L (𝐻 ), (15)
E
Table 2 reports a separate analysis of earlier 70B/405B traces on 128/512 servers, with 800 Gbps server access, 100 Gbps circuits, 𝛿 = 20 ms, and 𝜖 = 0.05. Placement minimizes crossEPS bytes plus unmatched directional bytes. Each operation costs the larger of its access and optical times, plus reconfiguration between segments; local transfers retain access cost. These are model estimates, not EPS hardware measurements. Aggregation reduces crossing traffic, but its benefit depends on placement and uplink provisioning. Full-rate uplinks can shift the bottleneck to server access, leaving little gain despite residual optical asymmetry; more servers per EPS do not imply monotonically smaller gains. Directconnect fabrics such as TopoOpt [41] avoid EPS aggregation and retain server-level directional demand.
𝑝 ∈ Pℓ
∑︁
∑︁
𝑧 ℓ𝑝 ≤ 1,
each directional resource 𝑒.
Switch-Node Aggregation and CCT
(16)
ℓ 𝑝 ∈ Pℓ :𝑒 ∈𝑝
15
Liang et al.
Table 2. EPS aggregation. 𝑔: servers per EPS; Inter/Excess: percentages of original bytes. Fixed and Full-rate report access-aware speedup over same-boundary symmetric allocation. Fixed provides 800 Gbps per EPS per direction; Full-rate provides 𝑔 × 800 Gbps, matching aggregate server access bandwidth. Inter Excess Fixed Full-rate
70B
4 DP PP Mixed Full 8 DP PP Mixed Full
25.0 14.3 41.7 42.4 12.5 0.0 32.0 32.8
24.6 8.8 32.9 33.2 12.3 0.0 13.7 14.3
1.750 1.000 1.554 1.545 1.750 1.000 1.173 1.179
1.000 1.000 1.038 1.037 1.000 1.000 1.038 1.039
4 DP PP Mixed Full
25.0 20.0 47.4 47.8
24.6 11.1 37.0 37.3
1.750 1.283 1.493 1.498
1.000 1.000 1.040 1.042
8 DP PP Mixed Full
12.5 6.7 31.5 31.6
12.3 3.7 26.6 26.6
1.750 1.000 1.593 1.591
1.000 1.000 1.002 1.002
405B
OCS fabric
OCS 1
OCS 2
Node A
Node B
Node A
Node B
EPS 1
EPS 2
Server 1
Server 2
(a) EPS-attached
64 128 256 512
32
64 128 256 512
Figure 12. Offline scheduling time on one CPU core. Lines show medians and bands min–max over three repetitions; totals include planning, realization, port completion, and the binding passes described above. on segment structure and need not grow monotonically with cluster size. These times are offline planning costs, separate from the simulated CCT.
Physical Testbed Measurement Details
Communication replay and training. Both experiments use OCS topologies on three servers, each with two 100G ports. Replay uses two complete runs per topology, a native three-server mapping, and TCP transfers with real acknowledgments. Per-port rate limits are 10 Gbps for replay and 5 Gbps for training, with the same striping policy across topologies. These limits keep traffic below the host processing ceiling and isolate the effect of which peers the lanes connect to. Training uses one RTX 5000 and two RTX 4000 GPUs, sequence length 1024, microbatch size one, accumulation eight (global batch 24), FP16, and AdamW. Each topology runs one epoch from the same pretrained checkpoint and data order. Wall time includes all optimizer steps and excludes validation and checkpoint saves. Validation uses 96 fixed packed blocks; the maximum paired training-loss difference is 0.002984. Software acknowledgments and CPU cost. The microbenchmark uses a fixed 100G link with data and feedback on the same NIC port and host-memory buffers. Each message size has 30 paired RC/UC trials, with randomized method order and 32 messages per trial. Figure 6 reports the mean of 100(𝑡 UC+ACK /𝑡 RC − 1) with a paired 95% bootstrap interval, and median throughput. Both modes use 64-KiB chunks and a 128-chunk window. UC batches cumulative ACKs every eight chunks, with transfer-boundary flushing and a 10-𝜇s timer. Table 3 reports the median, over 30 trials per mode and size, of the sum of sender and receiver process CPU time divided by application GiB transferred. Both endpoints use one
(b) Server-attached
Figure 11. Node and fabric boundaries in two deployments. Tabs mark OCS-facing ports; optical links represent TX/RX fiber pairs.
F
1
Servers
OCS fabric
OCS 2
(b) 405B
10
32
G OCS 1
8 × 800G
(a) 70B Offline time (s)
Model 𝑔 Mode
16 × 400G
Offline Runtime and Scalability
We measure runtime on one CPU core of a Linux x86-64 server using Python 3.10, with BLAS/OpenMP restricted to one thread. Each of the 20 model/scale/port configurations has one warm-up and three timed repetitions. The measured stages are the Segment Planner, Circuit Realizer, port completion with a deterministic reference binding, and a subsequent binding pass that retains existing connections. The total includes both binding passes. Input loading, serialization, demand profiling, and multi-OCS routing are excluded. Totals are summed within each repetition before taking the median. At 512 servers, measured total time is 7.10 s for 70B and 34.06 s for 405B with 16 ports at 400G. With eight ports at 800G, the times are 2.91 and 10.67 s (Figure 12). More integer circuits increase the search work, but runtime also depends 16
LACE: Lane-Granularity OCS Scheduling
Table 3. Combined endpoint CPU cost on the fixed 100G link. Values are median CPU-s/GiB, including busy polling. Message size 64 KiB 1 MiB 16 MiB 64 MiB
RC
UC + software ACK
0.4697 0.2033 0.1769 0.1759
0.5674 0.2067 0.1772 0.1760
original generated communication schedules. Traffic for a direction is assumed to be evenly striped across ℓ = 16 TX lanes, with rank/channel traffic sharing a sustained stream. Using fewer lanes increases the estimated bytes per lane proportionally. DP Ring rounds. For DP group size 𝐷, use the default target bucket size 𝐵 = max{40,000,000, 1,000,000𝐷 } parameter elements per rank. One logical Ring round for a full bucket sends 𝐵/𝐷 elements per rank. Aggregating eight ranks and striping across ℓ lanes gives 8𝐵𝑏 𝑀DP = bytes per lane per round, (17) 𝐷ℓ where 𝑏 = 2 for BF16 AllGather and 𝑏 = 4 for FP32 ReduceScatter. A complete Ring collective has 𝐷 − 1 rounds; the estimate does not combine those dependent rounds into one message. Tail buckets may be smaller, and actual bucket sizes depend on parameter grouping. PP transfers. With sequence parallelism, each TP rank transfers one eighth of the activation tensor at a PP boundary. For sequence length 𝑆, microbatch size 𝑚, and hidden dimension ℎ, the eight ranks together send 𝑆𝑚ℎ BF16 elements. Hence 2𝑆𝑚ℎ bytes per lane per transfer. (18) 𝑀PP = ℓ For ℎ = 8192 (70B) and ℎ = 16384 (405B), this gives 8 and 16 MiB per lane. A backward activation-gradient transfer has the same volume under the same shape and BF16 precision. The 1F1B order changes when these transfers occur, not their individual size.
busy-polling thread, with median CPU-time/wall-time ratios above 0.996 in all groups. There is no additional ACK thread. These costs include polling and all transport processing; similar values do not imply zero ACK cost or establish spare CPU capacity. The experiment does not measure concurrent multiport processing or GPU-memory transfers. Large transfers still require chunk-level processing: a 64MiB message contains 1,024 chunks. Ignoring wire overhead, retransmissions, and timer/boundary ACKs, a payload rate 𝑅 in bits/s requires 𝑅/(8 · 65,536) chunks/s and one progresstriggered ACK per eight chunks. At 400/800 Gbps this is approximately 0.763/1.526 million chunks/s per port; at 6.4 Tb/s it is 12.207 million chunks/s per node in one direction, for either port configuration. These are processing-demand estimates, not measured CPU costs or predictions of the required core count. Configuration verification and recovery. A separate control experiment performs three triangle-to-ring and three ring-to-triangle transitions with 100G configured, autonegotiation disabled, and the existing FEC configuration preserved. Each transition checks port status and reception of configuration-ID probes on all intended connections. All six pass; controller-observed readiness is 5.79–7.44 s. This interval includes control commands, polling, and probe verification. In one missing-RX fault injection, verification fails, one port-reset attempt does not restore the missing path, and the controller restores and verifies the preceding topology. The observed 40.65-s interval includes two 15-s timeouts and recovery actions. Original topology and port settings are restored after the experiment. The recovery test is unloaded and excludes application draining and RDMA QP recreation. As also observed by MixNet [24], current commodity link initialization can dominate optical switching time. The simulation delay is a separate fabric parameter; deployment uses the full delay exposed to communication, including link recovery and verification.
H
Table 4. Full-bucket DP payload per lane per Ring round, in MiB. Entries show BF16 AllGather / FP32 ReduceScatter with sixteen TX lanes. PP transfers are 8 MiB/lane for 70B and 16 MiB/lane for 405B. Servers 𝑁 32 64 128 256 512
70B (𝐷 = 𝑁 /8)
405B (𝐷 = 𝑁 /16)
9.54 / 19.07 4.77 / 9.54 2.38 / 4.77 1.19 / 2.38 0.95 / 1.91
19.07 / 38.15 9.54 / 19.07 4.77 / 9.54 2.38 / 4.77 1.19 / 2.38
Table 4 gives the range across the evaluated model scales. The 0.95–38.15 MiB band covers these full-bucket DP rounds and the two PP sizes; it is not a percentile interval. These are aggregate payloads per lane, which can consist of multiple NCCL channels and transport chunks.
Per-Lane Training Payload Estimates
The shaded range in Figure 6 is an analytical comparison with the measured message sizes. We use a Megatron configuration [33] with TP=8 within each server, sequence parallelism enabled, context parallelism disabled, sequence length 8,192, and microbatch size one. These assumptions are specific to this estimate; the main scheduling experiments retain their 17