MatchRDMA: A Segmented and Rate-Matched Long-Haul RDMA Scheme for Geo-distributed LLM Training over OTN Jun Dai(1), Xiaorun Wang(1), Xingde Li(1), Zheng Yang(1), Kexiong Fang(1), Zhiqun Gu(1), Hongxiang Wang(1), Yuefeng Ji(1), and Jiawei Zhang(1)* (1) State Key Lab of Information Photonics and Optical Communications, Beijing University of Posts and Telecommunications (BUPT), Beijing, 100876, China. *Corresponding author email: [email protected]
Abstract We propose MatchRDMA, a proactive, segmented, and rate-matched long-haul RDMA scheme for geo-distributed LLM training over OTN. By coordinating source and destination OTN rates, it improves inter-DC throughput by up to 20 compared with conventional RDMA, and reduces destination-OTN buffer occupancy by up to 62.7%. ©2026 The Author(s) Introduction Remote Direct Memory Access (RDMA) has been widely adopted in datacenter and high-performance computing environments due to its low latency, high bandwidth efficiency, and minimal CPU overhead [1-4], and is increasingly important for large language model (LLM) training [5]. Meanwhile, the rapid growth of model parameters is pushing training demand beyond the capacity of a single AI datacenter (AI-DC), making geo-distributed training across multiple AI-DCs increasingly necessary [6-9]. However, extending conventional RDMA over Optical Transport Networks (OTNs) for geo-distributed training introduces significant performance bottlenecks due to the long-haul transmission. As shown in Fig. 1, three bottlenecks emerge: (1) ACK-limited progress. RDMA relies on acknowledgments (ACKs) returned by the receiver to confirm successful reception of earlier data (e.g., RoCE packets), thereby allowing the sender to continue transmitting new data. Over long-haul transmission, this ACK-return loop is significantly prolonged, slowing sender progress and reducing effective throughput. (2) Buffer stress. A longer inter-DC path means that more data remains in flight along the end-to-end transport path. If downstream forwarding is temporarily slowed by a congestion (as shown in Fig. 1), this excess in-flight data imposes strong buffer pressure near the destination OTN node; if the available buffer is insufficient, packet loss and inefficient retransmission become more likely. (3) Feedback imbalance. End-to-end congestion-control schemes designed for short and relatively uniform intra-DC feedback loops become much less effective in inter-DC settings. When inter-DC and intra-DC traffic with different feedback delays jointly contribute to a congestion, rate adaptation slows and bandwidth sharing becomes increasingly unbalanced. Together, these effects reduce transmission efficiency, exacerbate buffer stress at the destination OTN, and complicate efficient coordination between inter-DC and intra-DC traffic. Existing studies have improved long-haul
Leaf
Server
OTN
Spine In-flight data
Congestion
Control signal Limit
Long-haul transmission ①ACK-limited progress
...
New data
③Feedback imbalance
②Buffer stress
...
Earlier data awaiting ACK
Fig. 1: Three bottlenecks for long-haul RDMA over OTN.
RDMA by accelerating ACK return [10], introducing relay-assisted control [11,12], and mitigating unfairness caused by different feedback delays [13-15]. Overall, these approaches are mainly reactive, as they respond only after congestion has already emerged, and are largely designed for the small, irregular, and hard-to-predict traffic patterns that characterize conventional DC workloads. In this paper, we propose MatchRDMA, a proactive, segmented, and rate-matched long-haul RDMA scheme for geo-distributed LLM training over OTN. Leveraging the predictable communication structure of LLM training [16,17], MatchRDMA coordinates inter-DC transmission through segmented control and source-to-destination rate matching. It is realized via sourceOTN budget-gated pseudo-ACK generation and congestion-control proxying, inter-OTN ratebudget signaling, and destination-OTN communication-aware slot-weighted rate estimation. The simulation results demonstrate that, compared with conventional RDMA, MatchRDMA improves inter-DC throughput by up to 20 , reduces destination-OTN buffer occupancy by up to 62.7%. Principles of MatchRDMA Fig. 2(a) compares the workflow of conventional end-to-end RDMA with the segmented OTNassisted MatchRDMA. Unlike conventional RDMA over long-haul, where the OTN acts only as a passive transport pipe, MatchRDMA turns the source and destination OTN nodes into active control elements and decomposes the round-trip
Conventional end-to-end RDMA
RDMA data
ECN mark
Rout’(t)
CNP ACK
CNP ACK
MatchRDMA with segmented OTN Source OTN
ECN mark
Time
Budget Sourceside loop
CNP Inter-OTN ACK Destinationloop side loop
Long-haul
Receiver
(a)
ECN
Opcode
…
RDMA header DQP
…
ECN: Explicit Congestion Notification Opcode: Operation Code DQP: Destination Queue Pair
PSN: Packet Sequence Number
Sender
Intra-DC
Budget
Source OTN
(c)
Destination OTN processing
Output
(1) Slot-level observation and congestion estimation
Estimated future inter-DC rate budget
1
Slot index
CNP frequency
2
3
4
…
6
5
N
…
Low
CNP freq.
…
Moderate
Egress rate.
…
Congestion Low Low Mod. High Mod. Low … level Local egress observation
N-2 N-1
ACK delay
Time
Payload
PSN Extension
ACK/CNP Gen
ACK
Input
Rate
IP UDP Ethernet header header header
Intra-DC
ECN Check
ACK return time
Freq.
Destination OTN
CNP
RDMA NIC
Rout(t)
Destination OTN
RDMA data
(b)
Time Source Sender Intra-DC OTN
RoCE packet
Rin(t)
ECN mark
Delay
RDMA data
Propagation delay D D
BW
ECN mark
High
Time
Low Mod. High
(2) Stable recurrent rate-window identification Stable recurrent rate window
Time
(d)
Higher weight
Jitter-dominated slots Lower weight
Rate
Time
Rate-budget updates
Stable recurrent Jitter-dominated slots rate window
…
Higher weight
Lower weight
To source OTN
(e)
Fig. 2: Principles of MatchRDMA. (a) Comparison of conventional long-haul RDMA and MatchRDMA with segmented OTNassisted control; (b) Reservoir model of destination-OTN buffer stress; (c) Source-OTN control workflow; (d) RoCE packet fields used for OTN-side control; (e) Destination-OTN slot-level rate estimation and rate-budget generation.
time (RTT)-driven control path into three coordinated segments: a source-side loop, an interOTN coordination loop, and a destination-side loop (as shown in Fig. 2(a)). In lossless RDMA transport, sender progress depends on ACK return, while congestion control depends on explicit congestion notification (ECN) marking and congestion notification packet (CNP) feedback. Over a long-haul OTN transmission, both the ACKreturn loop and the ECN/CNP-based control loop are significantly stretched by the long RTT. By splitting this long RTT-driven control path into three segments, MatchRDMA makes sender advancement and congestion regulation more responsive, thereby mitigating ACK-limited progress and feedback imbalance. While segmented control alleviates two bottlenecks, the long-haul distance still induces destination OTN buffer stress. As illustrated in Fig. 2(b), we model the source and destination OTN nodes as two coupled reservoirs linked by a longhaul optical pipe. Let D denote the one-way propagation delay. The arrival process at the destination OTN is then simply the delayed version of the source-OTN output process. If the source injects traffic at a time-varying rate 𝑟 𝑡 while the destination can forward traffic into the receiving AIDC only at rate 𝑟 𝑡 , the minimum runtime buffer required at the destination is governed by the accumulated rate mismatch over the controluncertainty window. This can be expressed as: 𝑟 𝑢 1 sup 𝑑𝑢 𝑟 𝑢 𝐵 where 𝜏 denotes the effective time window during which source injection cannot be instantly adjusted—due to propagation and processing delays. Since LLM training traffic exhibits predictable communication structure, 𝑟 𝑡 can be proactively shaped to follow the expected destination-side rate, directly motivating the proactive, rate-matched principle of MatchRDMA. Guided by these principles, MatchRDMA
realizes proactive, segmented, and rate-matched control by inserting a lightweight RDMA-aware layer at the OTN nodes. As illustrated in Fig. 2(c), the source OTN learns RDMA connection-state information during flow setup and, during transmission, parses ECN marks together with selected RDMA header fields; Fig. 2(d) shows the RDMA-over-Converged Ethernet (RoCE) packet fields used for connection identification and sequence tracking. Based on this RDMA-aware view, the source OTN generates a budget-gated pseudo-ACK and a source-side congestion-control signal. The budget refers to the maximum sustainable injection rate that the destination side can accommodate. This accelerates sender progress while constraining inter-DC injection within the destination-sustainable rate budget, thereby preserving destination-side rate matching and mitigating unfairness caused by the interaction with intra-DC feedback loops. Across the inter-OTN path, the two OTN nodes measure the one-way propagation delay and exchange rate-budget updates, along with concise congestion summaries, over a small high-priority control subchannel. This ensures that source-side release remains matched to the forwarding capability estimated at the destination side. At the destination OTN, MatchRDMA uses returned CNPs for reactive rate tightening, but more importantly constructs communicationaware slot-level observations from ACK/CNP feedback and local egress observations. As illustrated in Fig. 2(e), ACK return time and CNP generation frequency are compared against preset thresholds to assess the congestion level in each slot. By aggregating consecutive slots, the destination OTN identifies stable recurrent rate windows and distinguishes them from short-term jitter-dominated slots. This communication-aware view is then used to estimate the future inter-DC
Leaf x4
...
...
...
4 servers / leaf
4 servers / leaf
Key parameters
Workload
Server-Leaf: 100Gbps
Alibaba AICB [18]
Leaf-Spine: 4*100Gbps
Message size: 1KB-8MB
Spine-OTN: 4*100Gbps
Concurrency: 1-64
Intra-DC delay: 1μs
Mixed inter/intra-DC traffic
80 60 40 20 0
①
1
10
100
1000
Distance (km)
100
Message size = 32K
80 60 40 20 0
②
1
10
100
1000
Distance (km)
Inter-DC throughput (Gbps)
...
DCQCN Pseudo-ACK Themis MatchRDMA
100
Message size = 128K
80 60 40 20 0
③
1
10
100
1000
Distance (km)
(b) 100
10
Message size = 8M DCQCN Pseudo-ACK Themis MatchRDMA
1
0.1
1
10
100
Distance (km)
(a)
(c)
20 15
Message size = 8M DCQCN Pseudo-ACK Themis MatchRDMA
10 5 0
1
10
100
Distance (km) (d)
Normalized average FCT
Leaf x4
...
Message size = 4K
Inter-DC throughput (Gbps)
...
Spine x4
100
Pause time (%)
Spine x4
Inter-DC throughput (Gbps)
AI-DC 2
16 * 100Gbps Distance:1-1000km One-way delay:5μs-5ms
Peak buffer occupancy (MB)
Long-haul OTN link AI-DC 1
15
10
Message size = 8M DCQCN Pseudo-ACK Themis MatchRDMA
5
0
1
10
100
Distance (km) (e)
Fig. 3: (a) Simulated dual-AI-DC leaf-spine-OTN topology; (b) Throughput vs. distance under different message size; (c) Destination OTN runtime buffer occupancy; (d) Pause time ratio; (e) Overall average FCT in the mixed-traffic scenario.
rate budget and generate the corresponding ratebudget updates fed back to the source OTN. For robustness, unstable short-term slots are weighted conservatively, whereas stable recurrent rate windows receive higher weights. In this way, MatchRDMA proactively coordinates source admission with destination-side rate estimation under segmented control. Simulation setup and results We evaluate MatchRDMA using the ns-3 simulator. The simulated topology, shown in Fig. 3(a), consists of two AI-DCs interconnected by 16 bidirectional OTN links, each operating at 100 Gbps. The one-way intra-DC delay is fixed at 1 μs, while the inter-DC distance is varied from 1 km to 1000 km, corresponding to a one-way propagation delay from 5 μs to 5 ms. The AI training workload is generated by the Alibaba AICB benchmark [18], which captures the alternating computation–communication structure of LLM training iterations. In the simulator, each training communication request is modeled as a message, which may be segmented into one or more RoCE packets depending on its size. We sweep the message size from 1 KB to 8 MB and the number of parallel messages (concurrency) from 1 to 64. MatchRDMA is compared with three baselines: a conventional DCQCN-like baseline [1], a pseudoACK baseline by NTT [10], and a THEMIS-like fairness-oriented baseline [14,19]. Fig. 3(b) shows the inter-DC throughput under different message sizes and inter-DC OTN delays. As the delay increases, the DCQCN-like and THEMIS-like baselines suffer severe throughput degradation because sender progress remains constrained by the stretched end-to-end ACK feedback loop. In contrast, both pseudo-ACK baseline and MatchRDMA are much less sensitive to distance, since pseudoACK keeps sender progress active. In addition, throughput increases with message size for all schemes because larger messages make better use of the long-haul path. Compared with the DCQCN-like baseline, MatchRDMA improves
inter-DC throughput by up to 20 . Fig. 3(c) reports the destination-OTN runtime buffer occupancy. The DCQCN-like baseline allows excess inter-DC traffic to accumulate near the destination OTN before downstream constraints are fully reflected, resulting in larger buffer buildup. Fig. 3(d) shows the pause time ratio. Both the THEMIS-like baseline and MatchRDMA achieve lower pause overhead than the DCQCNlike and pseudo-ACK baselines, while MatchRDMA yields the lowest overall pause time ratio because it aligns source release with the destination-estimated rate budget. Compared with the DCQCN-like baseline, MatchRDMA reduces peak runtime buffer occupancy by up to 62.7% and the pause time ratio by up to 94.1%. Fig. 3(e) shows the overall average flow completion time (FCT) under different message sizes in the mixed-traffic scenario. MatchRDMA outperforms the DCQCN-like baseline, reducing the average FCT by 31.5%–43.9%. The advantage becomes more significant for larger message sizes, because larger messages tend to produce more stable traffic patterns and are therefore easier to regulate through slot-level rate estimation. Additional sweeps over traffic jitter and parallel-message concurrency show that MatchRDMA maintains the same performance trend and degrades more gracefully than the baselines. Conclusions We proposed MatchRDMA for geo-distributed LLM training over OTN. It addresses the challenges of long-haul RDMA transmission through source-OTN budget-gated pseudo-ACK generation and congestion-control proxying, inter-OTN rate-budget updates, and destinationOTN communication-aware slot-weighted rate estimation. MatchRDMA improves inter-DC throughput while reducing destination-OTN runtime buffer occupancy and pause time, and remains robust under distance variation, message-size scaling, and parallel-message concurrency.
Acknowledgements This work was supported by the National Key R&D Program of China (No. 2024YFB2908303). References [1] Y. Zhu, H. Eran, D. Firestone, C. Guo, M. Lipshteyn, Y. Liron, J. Padhye, S. Raindel, M. H. Yahia, and M. Zhang, “Congestion Control for Large-Scale RDMA Deployments,” in Proceedings of the 2015 ACM Conference on Special Interest Group on Data Communication (SIGCOMM), London, United Kingdom, 2015, pp. 523– 536, DOI: 10.1145/2785956.2787484. [2] R. Mittal, V. T. Lam, N. Dukkipati, E. Blem, H. Wassel, M. Ghobadi, A. Vahdat, Y. Wang, D. Wetherall, and D. Zats, “TIMELY: RTT-based Congestion Control for the Datacenter,” in Proceedings of the 2015 ACM Conference on Special Interest Group on Data Communication (SIGCOMM), London, United Kingdom, 2015, pp. 537– 550, DOI: 10.1145/2785956.2787510. [3] R. Mittal, A. Shpiner, A. Panda, E. Zahavi, A. Krishnamurthy, S. Ratnasamy, and S. Shenker, “Revisiting Network Support for RDMA,” in Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication(SIGCOMM), Budapest, Hungary, 2018, pp. 313–326, DOI: 10.1145/3230543.3230557. [4] Y. Li, R. Miao, H. H. Liu, Y. Zhuang, F. Feng, L. Tang, Z. Cao, M. Zhang, F. Kelly, M. Alizadeh, and M. Yu, “HPCC: high precision congestion control,” in Proceedings of the 2019 Conference of the ACM Special Interest Group on Data Communication (SIGCOMM), Beijing, China, 2019, pp. 44–58, DOI: 10.1145/3341302.3342085. [5] J. Lu, J. Gao, F. Feng, Z. He, M. Zheng, K. Liu, J. He, B. Liao, S. Xu, K. Sun, Y. Mo, Q. Peng, J. Luo, Q. Li, G. Lu, Z. Wang, J. Dong, K. He, S. Cheng, J. Cao, H. Jiao, P. Zhang, S. Ma, L. Zhu, C. Shi, Y. Zhang, Y. Chen, W. Wang, S. Zhu, X. Li, Q. Wang, J. Liu, C. Wang, W. Lin, E. Zhai, J. Wu, Q. Liu, B. Fu, and D. Cai, “Alibaba Stellar: A New Generation RDMA Network for Cloud AI,” in Proceedings of the ACM SIGCOMM 2025 Conference, Coimbra, Portugal, 2025, pp. 453–466, DOI: 10.1145/3718958.3750539. [6] J. Sun, D. Wang, B. Qi, T. Gao, D. Zhang, W. Chen, and H. Li, “Decentralized Training over 100km Based on Optical Transport Network for Artificial Intelligence,” in Proceedings of 50th European Conference on Optical Communication (ECOC), 2024, pp. 1-3, DOI: 10.1109/ECOC00010.2024.10739621 [7] Y. Liu, A. Zhang, X. Wang, L. Feng, K. Lv, H. Liu, X. Sheng, X. Huo, J. Li, “Field Trial of Multi-Datacenter Distributed Training for LLM Based on Bandwidth Convergence and Two Parallel Strategies over 120km High-reliability 800Gbit/s C+L OTN”, in Proceedings of 50th Optical Fiber Communication Conference (OFC), 2025, pp. 1-3. [8] T. Chen, A. Kubicek, L. Huang, and T. Hoefler, “CrossPipe: Towards Optimal Pipeline Schedules for CrossDatacenter Training,” in Proceedings of the 2025 USENIX Conference on USENIX Annual Technical Conference(ATC 25), Boston, MA, USA, 2025, Art. no. 64, 20 pages. DOI: 10.5555/3768039.3768103. [9] J. Dai, X. Wang, K. Fang, Z. Yang, Y. Ji, and J. Zhang, “GeoPipe: a Geo-distributed LLM Training Framework with enhanced Pipeline Parallelism in a Lossless RDMAenabled Datacenter Optical Transport Network,” in Proceedings of 2025 Asia Communications and Photonics
Conference (ACP), 2025, pp. 1–6, DOI: 10.1109/ACP66871.2025.11350566. [10] J. Ichikawa, H. Masutani, K. Obana, H. Takahashi, and K. Takasugi, “RDMA Acceleration Scheme for Long-Distance Optical Network,” in Proceedings of 2024 IEEE Global Communications Conference (GLOBECOM), Cape Town, South Africa, 2024, pp. 4842–4847, DOI: 10.1109/GLOBECOM52923.2024.10901383. [11] Y. Chen, C. Tian, J. Dong, S. Feng, X. Zhang, C. Liu, P. Yu, N. Xia, W. Dou, and G. Chen, “Swing: Providing Long-Range Lossless RDMA via PFC-Relay,” IEEE Transactions on Parallel and Distributed Systems, vol. 34, no. 1, pp. 63–75, 2023, DOI: 10.1109/TPDS.2022.3215517. [12] M. Long, J. Han, W. Wang, J. Yang, and K. Xue, “LSCC: Link-Segmented Congestion Control for RDMA in CrossDatacenter Networks,” in Proceedings of 2024 IEEE/ACM 32nd International Symposium on Quality of Service (IWQoS), Guangzhou, China, 2024, pp. 1–10, DOI: 10.1109/IWQoS61813.2024.10682909. [13] D. Yan, Y. Liu, S. Zhang, M. Xu, Z. Yang, and B. Fang, “LRCC: Long-haul RDMA congestion control for crossdatacenter networks,” Computer Networks, vol. 273, art. no. 111756, 2025, DOI: 10.1016/j.comnet.2025.111756. [14] Z. Niu, M. Zhang, J. Zhang, R. Xie, Y. Yang, and X. Hu, “THEMIS: Addressing Congestion-Induced Unfairness in Long-Haul RDMA Networks,” in Proceedings of 2025 IEEE 33rd International Conference on Network Protocols (ICNP), 2025, pp. 1–13, DOI: 10.1109/ICNP65844.2025.11192376. [15] T. Bonato, S. Abdous, A. Kabbani, A. Ghalayini, N. Gebara, T. Lam, A. Agarwal, T. Chen, Z. Yu, K. Taranov, M. Elhaddad, D. De Sensi, S. Ghorbani, and T. Hoefler, “Uno: A One-Stop Solution for Inter- and IntraData Center Congestion Control and Reliable Connectivity,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC '25), St. Louis, MO, USA, 2025, pp. 1195– 1210, DOI: 10.1145/3712285.3759884. [16] W. Li, X. Liu, Y. Li, Y. Jin, H. Tian, Z. Zhong, G. Liu, Y. Zhang, and K. Chen, “Understanding Communication Characteristics of Distributed Training,” in Proceedings of the 8th Asia-Pacific Workshop on Networking (APNet '24), Sydney, Australia, 2024, pp. 1–8, DOI: 10.1145/3663408.3663409 [17] Q. Hu, W. Wang, C. Huang, X. Wang, Y. Li, Y. Zhao, Y. Zheng, Y. Tan, and J. Zhang, “Task placement and traffic interleaving for cross-datacenter LLM training over optical networks,” Journal of Optical Communications and Networking, vol. 18, no. 2, pp. 137–149, 2026, DOI: 10.1364/JOCN.579324. [18] Alibaba Cloud, “AICB: Artificial Intelligence Communication Benchmark,” GitHub repository. [Online]. Available: https://github.com/aliyun/aicb. Accessed: 2026. [19] Networked-System-and-Security-Group, “THEMIS,” GitHub repository. [Online]. Available: https://github.com/Networked-System-and-SecurityGroup/Themis. Accessed: 2026.