Conceptio › Archive › arXiv CS
arXiv CSopen access

Weaver: A System for AI-RAN Compute Sharing with Foundation Model Training

Leyang Xue et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

W EAVER: A System for AI-RAN Compute Sharing with Foundation Model Training Leyang Xue1∗ , Tianxin Wang1∗ , Xin Zhe Khooi2∗ , Jiaxun Yang1 , Dheeraj Mahendiran1 , Yufeng Xia1 , Mun Choon Chan2 , Myungjin Lee3 , Mahesh K. Marina1

arXiv:2609.35276v1 [cs.DC] 28 Sep 2026

1 The University of Edinburgh

2 National University of Singapore

Abstract

compute infrastructure between RAN and non-RAN (AI) workloads [60]. As envisioned by the AI-RAN Alliance [1], AI-RAN promises not only reduced operational costs through various efficiency gains but also helps maximize RAN asset utilization and generate new revenue streams (e.g., through value-added services like sensing). Enabling the full potential of AI-RAN does entail upgrading existing mobile network infrastructure at cell sites to embed accelerators with RAN compute hardware that can natively support AI workloads. Recent developments, including the emergence of NVIDIA AI Aerial platforms [43] and Nokia’s partnership with NVIDIA [37], indicate movement in this direction. We believe the shift toward GPU-accelerated RAN infrastructure for AI-native 6G networks presents a compelling and timely opportunity for hosting non-RAN AI workloads on AIRAN compute infrastructure, aligned with the AI-and-RAN aspect of AI-RAN. Looking ahead, potential AI compute distributed across the RAN infrastructure can be enormous. A back-of-the-envelope calculation suggests that the aggregate GPU compute capacity across RAN cell sites could potentially rival the currently deployed global GPU compute capacity.1 Not only that, such compute would typically be only partially utilized for RAN processing. RAN compute usage varies with the traffic it carries while compute infrastructure is typically provisioned for the peak usage (worst case). The spatiotemporal variations inherent to RAN traffic at both micro and macro scales [17, 25, 60, 74, 75] may therefore create plentiful spare compute that can be harnessed by other AI workloads. In this paper, we explore the aforementioned opportunity considering foundation model (FM) training as a representative yet challenging non-RAN AI workload. FMs [5] are models trained at scale on broad data that can then be adapted to a wide range of downstream tasks (e.g., GPT-4 [48], Sta-

The emergence of AI-RAN infrastructure, which equips cell sites with GPU-accelerated hardware, creates an opportunity to colocate non-RAN workloads with primary RAN processing. We explore using this spare capacity for decentralized training of foundation models (FMs), one of the most compute-intensive AI workloads. We present the first characterization of spare GPU capacity in AI-RAN systems at both micro-scale—across transmission slots within a cell site—and macro-scale—across sites. Our analysis finds that 40–85% of GPU capacity is unused; although this capacity is temporally bursty at individual sites, it is spatially complementary across sites. To safely and efficiently harness these resources, we present W EAVER, a system that opportunistically trains FMs alongside latency-critical RAN workloads without degrading RAN performance. W EAVER adopts a RAN-first design: a spare-compute controller integrated into the MAC scheduler uses compute-aware scheduling to smooth RAN GPU demand and exposes more usable spare GPU capacity. A two-level elastic training framework then adapts to dynamic, heterogeneous spare capacity within and across sites. Experiments on an O-RAN-aligned system prototype show that W EAVER creates up to 4.9× more usable spare compute and utilizes up to 83% of the available spare capacity. On a multi-site testbed, W EAVER improves training throughput by 2.1–3.7× over baseline approaches.

1

3 Cisco Research

Introduction

As we head towards 6G, there is a wide recognition that AI for radio access network (RAN) – “AI-for-RAN” – can unlock significant spectral, energy, and operational efficiency gains (e.g., [14, 18, 59, 62, 68]). Building on this, there is growing momentum towards a more holistic integration of AI with the RAN to realize AI-RAN [31], which goes beyond AI-forRAN to also cover “AI-on-RAN” to enable novel edge (AI) services [27, 73], as well as “AI-and-RAN” to share RAN

1 There are currently around 20 million 4G/5G RAN cells globally [50]. Assuming the number of cells remains at least at that level going forward to 6G, deploying a GH200 server (each with about 1 PFLOP capacity) for every 20 cells, as per the benchmarking in [42], yields an estimated total of 1 ZFLOP across all RAN sites. Current deployed GPU capacity globally is around 4 ZFLOPS [15].

* These authors contributed equally to this work.

1

ble Diffusion [57], AlphaFold [28]). Large language models (LLMs) such as GPT-4 are a prominent subclass of foundation models that are trained on language data. FM training requires enormous compute. For example with LLMs, scaling model size, data size, and training duration consistently improves capability, driving rapid growth in computational demand [30]. Growing interest in domain-specific FMs, including for mobile networks [80], further amplifies this demand. This has in turn contributed to a highly centralized, cloud-centric AI ecosystem currently. In view of the above, leveraging the spare compute on naturally decentralized AI-RAN compute infrastructure can serve as an alternative platform for FM training that is complementary to the cloud. FM training over AI-RAN compute infrastructure, however, requires addressing three main challenges: 1. We need to clearly understand the scale and nature of spare compute likely to be available on an AI-RAN compute (GPU-accelerated RAN) infrastructure. Prior work offers limited insight, both in the scope of metrics and spatiotemporal scales considered, as discussed in §3 and §8. 2. Even when spare compute exists, making it usable for nonRAN workloads such as FM training is challenging. This is because RAN traffic is inherently bursty at microscopic scales [17], which in turn manifests as burstiness in RAN related GPU compute usage. Smoothing that burstiness to maximize usable spare GPU compute must be done carefully to ensure RAN performance is unaffected. 3. Efficiently utilizing the available spare GPU compute resource for FM training within and across cell sites in an AI-RAN compute infrastructure necessitates a training strategy that is robust and efficient in the face of resource dynamism and heterogeneity across sites and timescales. To tackle the first challenge, we conduct the first comprehensive characterization study on spare GPU compute capacity in GPU-accelerated RAN infrastructure, both at the micro-scale (RAN-slot-level) within a cell site and at the macro-scale across cell sites (§3), focusing on offloading computationally heavy part of RAN processing on LDPC decoding to the GPU as in prior work [54, 60, 61]. Our study considers NVIDIA DGX Spark as a representative AI-RAN compute platform along with large-scale production mobile network traffic trace. Our study reveals that substantial unused GPU resources (overall, about 40–85%) exist that could be harnessed by non-RAN workloads. We observe that memory bandwidth-limited RAN processing complements computebound FM training. At a macro scale, spare compute mirrors diurnal traffic patterns at individual cell sites, while being spatially diverse across sites. Building on the above outlined spare-compute characterization, we address the second and third challenges with a novel RAN-first system architecture design termed W EAVER (§4), illustrated in Fig. 5. W EAVER enables efficient FM training on GPU-accelerated RAN infrastructure by opportunistically sharing GPU resources with RAN processing while protect-

ing RAN performance. W EAVER employs Green Contexts based dynamic GPU sharing (Table 1). The W EAVER system is made up of two principal components: • First, W EAVER introduces a RAN-centric spare compute controller (§5) embedded inside the RAN stack, in contrast to relying on externally predicting RAN-related GPU compute demand as done by prior works [60]. By integrating with the MAC scheduler, it performs compute-aware scheduling to smooth RAN related GPU usage (despite RAN traffic burstiness) to unlock maximal usable spare GPU resource for non-RAN workloads while enforcing per-slot deadlines to guarantee RAN performance. • Second, W EAVER introduces a tailored two-level elastic FM training framework (§6) to efficiently cope with spatiotemporal variations in spare GPU compute within and across cell sites. At the inter-site level, a centralized coordinator dynamically redistributes training computation and selectively reconfigures model layouts across cell sites only when long-term compute/traffic shifts occur. Within a cell site (i.e., intra-site), general matrix multiplication (GEMM) operations (that dominate FM training) are split into fine-grained and elastic tiles to match the instantaneous available spare compute. A publish–subscribe interface couples these two components, allowing FM training to continuously adapt to the RANapproved spare GPU resource at each cell site. Together, these design choices enable safe, efficient co-location of latencycritical RAN processing and opportunistic FM training. We conduct a comprehensive evaluation of W EAVER spanning both its main components (§7). We build a system prototype that deploys W EAVER’s RAN-centric spare-compute controller on an O-RAN-aligned 5G NR stack with GPUaccelerated LDPC decoding [8], while co-locating FM training computation on the same GPU. Results from evaluating the system under production traffic traces and live end-toend UE transmissions show that W EAVER’s RAN-centric approach reduces fluctuations in RAN related GPU usage by up to 4.9× without degrading RAN performance, while exposing substantially more usable spare compute to the colocated training workload than prior and baseline methods. Then, on an eight-site distributed GPU testbed, W EAVER’s two-level elastic FM training framework achieves 2.1–3.7× higher training throughput than baseline systems while harvesting up to 83% of available spare compute. Finally, via simulations, we show that W EAVER’s training framework effectively scales to hundreds of sites and larger models. The following section provides the essential background before we proceed to describing the main contributions of the paper in the subsequent sections. 2

2

Table 1: Types of GPU sharing in current systems.

Background

Paradigm Mechanism

GPU acceleration for vRAN. Virtualized Radio Access Network (vRAN) moves baseband processing onto generalpurpose platforms [6,17]. The Distributed Unit (DU) hosts the most compute-intensive MAC and PHY functions, with LDPC decoding as the dominant bottleneck [8, 61, 71]. Because LDPC decoding must meet strict millisecond-scale deadlines while maintaining carrier-grade reliability, hardware acceleration is often required in production deployments [8]. While ASICs and FPGAs have traditionally served this role, GPUs are increasingly attractive due to their massive parallelism and programmability [60]. For example, NVIDIA’s SionnaRK [8] accelerates LDPC decoding, while Aerial [39, 40] offloads the entire 5G PHY pipeline to GPUs, demonstrating their viability for compute-intensive vRAN processing. GPU architecture overview. A modern NVIDIA-style GPU consists of tens to hundreds of Streaming Multiprocessors (SMs), each hosting hardware threads and execution pipelines, while all SMs share a high-bandwidth global memory [38]. The GPU programming model exposes these resources via kernels, which are functions launched by the host CPU, each with a specified SM demand. GPU utilization metrics. GPU load can be characterized using three metrics: (i) SM utilization (SMU), the fraction of SMs assigned to a kernel; (ii) Arithmetic compute utilization (ACU), the fraction of peak sustained arithmetic issue rate achieved; and (iii) Global bandwidth utilization (GBU), the fraction of peak global-memory bandwidth achieved. GPU sharing for AI-and-RAN. Prior work has explored GPU sharing for co-locating RAN and AI workloads. YinYangRAN [60] uses multi-process service (MPS) to partition SMs between 5G DU PHY processing and ML inference, while CAORA [63] employs multi-instance GPU (MIG) to create isolated GPU instances. However, both approaches rely on RAN workload prediction for dynamic resource provisioning. Mis-estimation is problematic: overestimation starves co-located workloads, whereas underestimation may violate PHY processing deadlines. They also incur substantial reconfiguration overheads (∼0.3 s for MPS and ∼7 s for MIG), during which the GPU cannot service workloads or processing must fall back to CPUs, making them impractical for realtime-sensitive and mission-critical RAN workloads. More broadly, modern GPUs expose a range of temporal, spatial, and hybrid sharing mechanisms [38] (see Table 1). However, existing GPU-sharing systems are designed for multi-tenant environments where workloads are treated as peers [11, 16, 23, 84]. In contrast, RAN workloads must be unconditionally prioritized and require per-slot (≤1 ms) deadline guarantees. While some systems distinguish latency-sensitive and best-effort tenants [24,64,67], they do not provide the hard isolation and deadline guarantees required for RAN processing. Among existing mechanisms, Green Contexts are unique in combining hardware SM isolation with microsecond-scale

Isolation

Reconfiguration

Temporal

Context switch [16] Preemption [24, 64]

Device exclusive Instruction-level

Restore ∼100 ms Application rebuild –

Spatial

MIG [84] Green Contexts [38]

SM & memory SM

GPU reset CUDA Stream

Mixed

MPS [61] Best-effort SM (no mem. isol.) GPU reset CUDA streams [38, 67] Best effort Static

∼7 s ∼10 µs ∼300 ms ∼1 µs

reconfiguration, making them the only practical substrate for per-slot RAN control, which we will later explore in this paper. Moreover, Green Contexts are natively supported in the NVIDIA AI-RAN software stack [44]. Foundation model training. FM training involves distributed computation over a directed acyclic graph (DAG) of dense matrix multiplications (GEMMs) [10, 65]. Each layer produces a sequence of GEMM operations arising from attention projections, feed-forward transforms, and embedding lookups, whose execution order and synchronization points are determined by the distributed training strategy used (e.g., data parallelism (DP), pipeline parallelism (PP), or their combination). Over 99 % of training FLOPs are GEMMs [10, 65], making FM training compute-bound and capable of saturating the GPU’s arithmetic pipelines rather than its memory bus. We leverage this complementary resource profile (see later in §3.1) to co-locate FM training with memory-bound RAN workloads and exploit otherwise idle compute capacity.

3

Spare Compute Characterization

The adoption of GPUs in vRAN raises a fundamental question: how much spare compute is available, and along which dimensions can it be safely exploited without affecting RAN performance? We study this at two levels: (i) micro-scale (§3.1), examining fine-grained opportunities within slots and across GPU resources (compute, memory bandwidth, and SM occupancy) across a diverse range of RAN workloads; and (ii) macro-scale (§3.2), analyzing whether spare capacity remains consistently available over time and across cell sites. Answering this question requires more than the coarse scalar utilization metrics reported in prior work (e.g., CloudRIC [61]); see §8. At the micro-scale, standard tools (e.g., nvidia-smi) provide only coarse activity indicators and do not expose multi-dimensional GPU utilization [41]. At the macro-scale, traffic traces [52] capture demand dynamics but cannot directly reveal actual GPU utilization. In this work, our focus is on vRAN deployments with GPUaccelerated LDPC decoding, inline with prior work [54, 60, 61], given that it is a widely known compute bottleneck [71].

3.1

Micro-Scale Opportunities

First, to study the micro-scale opportunities across slots within a cell, we use the OpenAirInterface5G-based SionnaRK [8] to profile GPU-accelerated LDPC decoding on an NVIDIA DGX Spark (48 SMs). We collect measurements 3

30

400

20

Zone-I

600

10

200 0

c) Compute utilization ACU (%) 273 250 225 200 175 150 125 100 75 50

0 3 6 9 12 15 18 21 24 27 MCS

10.0

40

7.5

20

5.0

5

0

30 20 10 0

0 3 6 9 12 15 18 21 24 27 MCS

3.2

(b) LDPC Decoding

Compute-Bound (ACU > GBU)

60

Compute-Bound (ACU > GBU)

1.0 0.8 0.6

40

Memory-Bound (GBU > ACU)

Memory-Bound (GBU > ACU)

20 0

20

40 60 GBU (%)

80

0

20

40 60 GBU (%)

80

0.4 0.2

Macro-Scale Opportunities

Next, to assess whether spare capacity remains consistently available over time and across sites, we replay a large-scale production traffic trace from six urban and suburban cell sites [52] using a modified version of the OpenAirInterface (OAI) phy-test tool [34]. For brevity, we report only SMU, as ACU and GBU exhibit similar trends (§3.1). Single-site temporal spare-compute patterns. We observe a clear diurnal pattern in both traffic and SMU across sites. Traffic is generally low and stable overnight (02:00–06:00) and peaks during the day (Fig. 3a). Correspondingly, the available spare compute is highest overnight, while even at the busiest site during peak demand, roughly 85% of SM capacity remains available (Fig. 3b). Spatial cross-site spare-compute patterns. We compare traffic demand and SMU across six sites, revealing three observations. First, highly skewed traffic: Zone-II (suburban) sites (4–6) carry noticeably higher load than Zone-I (urban) sites (1–3), leaving low-traffic sites with significant spare compute even during local peaks (Fig. 3a). Second, staggered daily peaks: sites peak at different times of day, so simultaneous peak load across sites is rare (Fig. 3a). Third, SMU tracks traffic: SMU follows the same spatiotemporal pattern as traffic, showing that spare compute is abundant at individual sites and complementary across sites (Fig. 3b).

norm. TB size

(a) GEMM 80

00h 04h 08h 12h 16h 20h Hour of Day

Figure 3: Spatial traffic patterns and their SMU. 40

2

0

lization (GBU) ranges from 10–60% depending on transport block (TB) size and redundancy (Fig. 1d), indicating that LDPC decoding is memory bandwidth-bound (see detailed breakdown in Appendix A.1) while leaving much of the compute pipeline idle. This creates a natural complementarity with FM training, which is compute-intensive (see Fig. 2).

Figure 1: a) Worst-case latency, b) SMU, c) ACU, and d) GBU of LDPC decoding across different #PRBs and MCS.

ACU (%)

60

4

00h 04h 08h 12h 16h 20h Hour of Day

50

0 3 6 9 12 15 18 21 24 27 MCS

0

3

0

8

4

15.0 12.5

d) Bandwidth utilization GBU(%)

6

b) Average SMU (%)

80

2

6

0 3 6 9 12 15 18 21 24 27 MCS

PRB Count

a) Normalized Traffic (% of global peak) 100 1

b) SM Utilization (%)

Zone-II

PRB Count

a) LDPC Decoding Latency (us) 273 250 225 200 175 150 125 100 75 50

0.0

Figure 2: ACU and GBU of GPU operations for FM training (model sizes: 1.3B to 70B) vs. LDPC decoding. using ulsim to sweep across the full modulation and coding scheme (MCS 0–28) and physical resource block (PRB 0– 273) configuration space. To capture worst-case execution time, the LDPC decoder is fixed at 10 iterations. Slack within a slot. We observe two sources of spare GPU capacity within a slot. First, even the worst-case RAN configuration completes LDPC decoding in ∼700µs, leaving 30–70% of the headroom available given a typical 1-ms deadline [61] (Fig. 1a). Second, GPU resources are underutilized: worstcase SM utilization reaches only ∼35%, for a single DU, and remains below 20% across most operating points (MCS ≤ 20, PRB ≤ 200), leaving over 80% of the available SMs on the GPU idle. Thus, LDPC decoding primarily scales in execution time rather than SM occupancy (Fig. 1b).

Key takeaway: Spare GPU capacity exists at multiple scales. Microscopically, LDPC decoding is memorybound (complementary to compute-bound FM training) and leaves substantial compute resources idle. Macroscopically, this slack persists across time and space (cell sites), making co-location with FM training feasible.

Microscopic (compute and memory) view of slack. Examining spare GPU resources after RAN execution (i.e., LDPC decoding) reveals two key observations. First, RAN is not compute intensive: even under the worst-case configuration, the Arithmetic Compute Utilization (ACU) reaches only ∼12% (Fig. 1c). This stems from the latency-driven nature of RAN processing: LDPC decoding is parallelized across SMs to minimize completion time, resulting in low ACU. Second, RAN is memory-bandwidth intensive: Global Bandwidth Uti-

4

W EAVER System Overview

Building on the characterization study in §3, W EAVER seeks to exploit the substantial spare GPU capacity available in 4

RAN Centric Spare Compute Controller

142

RAN Stack

b) Stable SM Reservation, Avg SM: 16.0

48 32 16 0 20

c) Comp. Throughput

150

30

40 50 RAN Slot Index

60

TFLOPS

SM Count

a) Bursty SM Demand, Avg SM: 16.9

48 32 16 0

Compute-Aware Scheduler

100

50

0

RAN-Controlled SM Reservation

O-RAN dApp(Real-Time) 54

Bursty

Standard MAC Scheduler

GPU Scope

Physical Layer (PHY) [GPU-accel. LDPC Decoding]

Stable

SM Reservation

Figure 4: SM reservation comparison: a) bursty SM demand derived directly from RAN traffic; b) stable SM reservation; c) impact of SM reservation stability on training computational throughput (the higher, the better).

SM Reservation Via CUDA Green Context

Spare Compute Pub-Sub Interface (Spare SM Budget: [Time Slice, SM Count]) Notification Two-Level Elastic FM Training Strategy

vRAN deployments for FM training. Driven by that target, W EAVER is designed with the following goals: 1. Maximize harvestable spare compute. W EAVER must expose as much GPU slack as possible to FM training without compromising RAN performance. 2. Efficiently utilize distributed slack. W EAVER must allow FM training to effectively exploit spare compute distributed across time and cell sites. Challenges and solutions. Achieving these goals is challenging for three reasons. Below, we describe each challenge and the corresponding solution adopted in W EAVER. 1. Only SM allocation is controllable. Our characterization reveals spare capacity across multiple dimensions, including SMU, ACU, and GBU (§3.1). However, current GPUs expose only SM allocation as a practical runtime control knob [38]; ACU and GBU cannot be directly partitioned or allocated. Consequently, any sharing mechanism must expose spare compute through SM allocation; Appendix A.2 shows that capping the RAN’s SM count also frees headroom in all three dimensions. Among available GPU-sharing primitives (§2), Green Contexts are uniquely suited for this purpose, providing hardware-isolated SM partitions with microsecond-scale reconfiguration. We therefore define an SM reservation: the per-slot SM allocation reserved for RAN processing, whose complement is made available to FM training. 2. The SM reservation is highly dynamic. The amount of SMs required by RAN processing is not constant. Bursty RAN traffic [52] causes frequent fluctuations in SM demand (Fig. 4a). While Green Contexts enable fine-grained SM partitioning, exposing these fluctuations directly to FM training results in a rapidly changing compute budget. As shown in Fig. 4, such variability reduces the effective computational capacity available to GEMM workloads, both through frequent SM reconfigurations and the disruption of long-running training kernels. The challenge therefore is to transform a highly dynamic SM demand from the RAN into a stable spare-compute budget without compromising RAN correctness. To address this, W EAVER introduces a RAN-centric spare compute controller (§5) that derives a stable SM reservation for RAN to maximize

Intra-Site Slot-Level Scheduling

Aggregate Compute Capacity Inter-Site Distributed Scheduling Distributed Coordination

Data Reroute Model Reshard

Figure 5: W EAVER system architecture overview. spare compute and minimize reconfiguration overhead, while preserving sufficient headroom for RAN processing to absorb short-term fluctuations in RAN demand. 3. Spare compute is distributed and time-varying. Available slack varies across both time and cell sites (§3.2), creating dynamic and heterogeneous available GPU compute across training workers. Without adaptation, stragglers would limit distributed-training throughput. To address this, W EAVER introduces two-level elastic training (§6), which elastically adjusts training execution both within a site and across sites to match available spare compute. Taken together, these challenges motivate the W EAVER system architecture shown in Fig. 5. W EAVER combines (i) Green-Context-based SM reservation, (ii) a RAN-centric spare-compute controller that stabilizes the SM reservation to maximize usable spare compute, and (iii) a two-level elastic training framework that efficiently utilizes distributed spare compute. In W EAVER, a publish-subscribe (pub-sub) interface is used for (ii) to continually communicate RAN’s SM reservation to (iii). We describe the design of components (ii) and (iii) in detail in the following two sections.

5

RAN-Centric Spare Compute Controller

As discussed in §4, a key challenge is to transform highly dynamic RAN SM demand into a stable spare-compute budget that can be exploited for FM training. Prior approaches for AIand-RAN infer available GPU resources externally through RAN workload prediction [60]. However, RAN SM demand is highly dynamic, and so mis-prediction either risks RAN deadline violations or leaves spare compute underutilized. So, instead of predicting future demand, W EAVER adopts a 5

a) Latency (ms) w/ SM = 8

Scheduler Input PUSCH Info.

Entry point:[Prev. SM Count, Prev. Avg. MCS]

PRB Budgeting (Look-Up Table Query)

#PRB Budget

Standard PUSCH Scheduling Per-UE #PRB Allocation

b

Adaptive SM Count Adjustment Buffer-Backlog Monitoring & History Tracking SM Ramp-Down

SM Hold

SM Ramp-Up

[SM Count, Allocated #PRB, Avg. MCS]

c

Iteration Control #PRB Allowance Check (Based on cur. [SM Count, Avg. MCS])

Loop back Scheduler Output

Success

0.7

162

b) Latency (ms) w/ SM = 16

106

273 0.8

8

8

16 24 32 48 48

1.2 217 8

8

16 16 24 40 40

162

8

8

8

16 16 32 24

106

8

8

8

8

16 16 24

52

8

8

8

8

8

8

16

25

8

8

8

8

8

8

8

0

5 10 15 20 25 27 MCS

0.8 0.7

0.9 0.8

0.7 0.6 0.8

0.5 0.7

5 10 15 20 25 27 MCS

0

c) Min SM Required

1.5

0.7

0.7

0

5 10 15 20 25 27 MCS

0.6

0.3

Figure 7: a) and b): Worst-case LDPC decoding latency with various (PRB, MCS) tuples given an SM count; c): minimum SM demand as a function of (PRB, MCS). the required SM count. We capture this relationship using an offline-generated lookup table (Fig. 7c), derived from worstcase LDPC decoding latency profiling across SM budgets and RAN configurations (e.g., Fig. 7a–b) using ulsim (§3.1). As an example, (MCS 15, 162 PRBs) satisfies the latency deadline with 16 SMs but violates it with 8 SMs (Fig. 7b); the table therefore records 16 as the minimum required SM count (Fig. 7c). The resulting table provides safe operating bounds because it is derived from conservative worst-case profiling, covers the full MCS/PRB configuration space, and is calibrated through end-to-end testing under extreme channel and traffic conditions (Appendix B.1). PRB budgeting. A practical issue is that determining the PRB budget requires both the target SM count and the slot’s average MCS. However, the true average MCS is known only after PRB allocation completes, creating a circular dependency: the scheduler must cap PRBs before allocation, yet the required lookup-table input is finalized only afterwards. To break the MCS-dependency cycle, W EAVER bootstraps scheduling using the previous slot’s average MCS. At the start of each slot, and on each subsequent re-entry iteration, the scheduler pairs an MCS estimate with the target SM count and queries the lookup table to obtain a PRB budget. We define average MCS as the PRB-weighted average of per-UE MCS values. The initial MCS estimate and target SM count are inherited from the previous slot’s scheduling decision; on re-entry, they are updated using the true average MCS and adjusted SM count from the preceding iteration. Standard proportional-fair scheduling [49] then proceeds within the resulting PRB budget, preserving UE fairness while respecting the target SM reservation.

Safeguard

[Final UL PRB Grant, SM Count for slot]

RAN-first approach. Specifically, it introduces a RAN-centric spare-compute controller embedded inside the RAN stack that directly shapes RAN compute demand and exposes the remaining spare capacity to FM training – a non-RAN AI workload we focus on. The MAC scheduler is the natural control point, as its scheduling decisions determine the amount of LDPC decoding work and thus the required SM reservation. The key idea behind the spare-compute controller in W EAVER is computeaware scheduling, which smooths the short-term fluctuations in RAN compute demand while preserving RAN processing deadline and QoS guarantees. Rather than serving every burst immediately, W EAVER absorbs transient traffic spikes in UE buffers and only increases the RAN’s SM reservation when backlog growth approaches a QoS-aware threshold. This effectively rebalances PRB allocation across neighboring slots, transforming bursty slot-level demand into a more stable SM reservation. The resulting spare-compute budget is substantially more useful for FM training. To achieve this, W EAVER implements a stack-agnostic ORAN dApp [12, 53] that wraps the standard-aligned MAC scheduler and augments it with compute awareness (Fig. 5). The resulting per-slot SM reservation is enforced through CUDA Green Contexts, which provide hardware-isolated SM partitions and microsecond-scale reconfiguration.

Compute-Aware Scheduling Pipeline 5.1.2

Fig. 6 depicts the compute-aware scheduling pipeline in a PRB W EAVER. The pipeline comprises three key steps: ⃝ b adaptive SM reservation adjustment, and ⃝ c budgeting, ⃝ iteration control. We elaborate on these steps below. 5.1.1

217

25

Figure 6: Per-slot compute-aware scheduling pipeline.

5.1

0.7

52

SM Ramp-Up Request

[ Tgt. SM Count, Avg. MCS]

PRB Count

a

273

Adaptive SM reservation adjustment

The second issue is deciding when and by how much to adjust the SM reservation. Frequent adjustments expose slotlevel traffic fluctuations directly to co-located workloads (FM training in our case), reducing the effective compute budget available to those workloads. Conversely, too coarse-grained adjustments may waste harvestable compute during low load or under-provision the RAN during sustained traffic increases. So, the scheduler must balance stability with responsiveness. To achieve this, W EAVER continuously monitors RAN backlog and QoS indicators. Rather than reacting to every

PRB budgeting using lookup tables

Lookup table. A key enabling property is that SM reservation for RAN is directly controllable through PRB allocation. Given an MCS index, capping the number of scheduled PRBs bounds the amount of LDPC decoding work and therefore 6

6

transient traffic burst, it absorbs short-lived fluctuations within UE buffers and adjusts the SM reservation only when backlog growth approaches a QoS-aware threshold. Given the uplink PRB scheduling decisions, the scheduler examines per-UE backlogs and the current SM reservation:

The controller in §5 yields a stable SM reservation for the RAN and consequently makes spare-compute budget relatively stable within a cell site. But from a FM training perspective, available spare compute still evolves over time (see Fig. 4b) and is also heterogeneous across cell sites. Static distributed training configurations are therefore inefficient: workers with insufficient spare compute become stragglers, while workers with insufficient work under-utilize available spare compute. Note that we consider distributed FM training using standard PP+DP training strategy as with modern LLM frameworks [9, 78, 79]: the model is partitioned into pipeline stages (PP), each replicated across workers via data parallelism (DP), so the runtime views a job as stage tasks connected by activation transfers and synchronized through data-parallel gradient exchange. To cope with diverse and time varying spare compute across cell sites, W EAVER dynamically distributes the FM training computation across workers adapting with the spare-compute budget exposed by the RAN. Specifically, W EAVER employs a two-level elastic training framework (Fig. 8) that adapts to spare-compute availability both within a site and across sites: a §6.1): harvests capacity across (i) inter-site scheduling (⃝, sites at slower timescales, from sub-second asynchronous inter-site rerouting (AIR) decisions to hourly traffic shifts via model layout rebalancing (MLR); and (ii) intra-site schedulb §6.2): adapts to RAN-slot-level variation locally, ing ( ⃝, and exposes an aggregated capacity view to the inter-site coordinator. This two-level approach is essential because RAN-slot-level fluctuations (<1 ms) are too fast for intersite coordination (>100 ms), while macro-scale shifts are too slow-varying2 to be handled efficiently by intra-site scheduling alone. Note that MLR is also termed as model resharding as in Fig. 8.

• Hold: If all UE backlogs are empty or below a tolerance threshold, the SM reservation stays unchanged. • Ramp-up: If any UE backlog exceeds the pre-defined threshold, the scheduler immediately increases the SM count to drain the UE buffer backlog within the current slot, preserving UE performance. • Ramp-down: If no ramp-up occurs, the scheduler checks whether the current slot could have been served with fewer SMs. A hysteresis counter tracks sustained overprovisioning and triggers a gradual step-down when appropriate. A “fast-exit” path applies when the current slot allocated zero PRBs, allowing prompt down-sizing without waiting for hysteresis, so that spare SMs can be released to co-located workloads. This policy is intentionally asymmetric: ramp-up is immediate to protect RAN performance, whereas ramp-down is conservative to improve stability. Together, the hold and hysteresis mechanisms suppress unnecessary SM reservation fluctuations, while fast-exit avoids prolonged over-provisioning and quickly updates the reservation to release spare SMs.

5.1.3

Two-Level Elastic Training

Iteration control with PRB allowance check

After PRB allocation and SM adjustment, both the average MCS and the target SM reservation may deviate from the values assumed for PRB budgeting. To this end, the scheduler recomputes the PRB allowance (i.e., maximum supported PRB count) by querying the lookup table (Fig. 7c) with these two updated values. If the allocated PRBs are within a close margin below this allowance, the grants are directly emitted. Otherwise, the scheduler loops back to PRB budgeting a ⃝ c with the updated MCS and SM count, repeating steps ⃝– in Fig. 6 until PRB allocation converges and the current-slot UL PRB grant is generated. A small iteration bound (i.e., five) safeguards against non-convergence; if exceeded, the scheduler forces an SM ramp-up to guarantee forward progress.

6.1

Inter-Site Scheduling

Inter-site scheduling must cope with model and optimizer states on the order of 10–100 GB [55] at each cell site over 1– 10 Gbps inter-site links [20,21], so migrating states would take seconds to minutes. Because RAN cell sites connect to the mobile network operator’s core infrastructure in a star topology with site-to-core links that have 10–50× higher bandwidth inter-site links [21, 36], we use a coordinator-centric runtime situated at the operator’s core network infrastructure for lightweight control-plane coordination while minimizing data-plane state transfers between sites. Centralized training state coordinator. The inter-site coordinator tracks whereabouts of model state and in-flight training work, and maintains per-sample-stage commit records. This is complemented by AIR that reroutes micro-batch

By combining all three mechanisms above, our computeaware scheduler determines the final UL PRB grant for each UE while deriving the current-slot SM reservation for the RAN and publishing the remaining SMs as spare. Over time, a stable SM reservation is produced, allocating only the minimum necessary SM count per slot to guarantee RAN performance while maximizing spare-compute opportunity. Algorithm details are provided in Appendix B.2, with parameter settings and their tuning in Appendix B.3.

2 Note that an hour spans millions of RAN slots.

7

result for each sample-stage key. Distinct from the migrationbased approaches (illustrated in Fig. 9b), AIR reacts without waiting for ongoing work to drain and moves only KB–MBscale inputs or activations, not multi-GB model state. Model layout rebalancing (MLR). AIR absorbs transient fluctuations within a fixed model-parallel layout. But when one site remains persistently slower, deferred micro-batches accumulate and rerouting alone cannot remove the structural imbalance. The coordinator therefore tracks the per-trainingstep deferral fraction with a moving average and escalates to MLR only when it stays above a threshold across consecutive training steps (illustrated in Appendix Fig. 13), so this mechanism is applied only for persistent spikes and imbalances. When MLR fires, the pipeline-parallel partition boundary shifts at a training-step barrier so that overloaded sites shed layers to less-loaded sites; only boundary-crossing layers’ parameters and optimizer state move, via a P REPARE /C OPY /C OMMIT sequence whose drain phase alone stalls training (measured breakdown in Appendix Table 13). The full rebalancing workflow and fault-tolerance support appear in Appendix C.1. Correctness. AIR and MLR affect only the execution location and timing of each workload. They do not change which samples contribute to training or how their gradients are aggregated. The training correctness guarantees (including atmost-once fenced commits, weighted aggregation matching a synchronous run, bounded debt) are stated and proved in Appendices C.2–C.4.

Training State Coordinator Distributed Runtime Ops

Gradient Aggregation Manager Model Parameter Store/Index

Hook Layer

Exactly-Once Aggregation Asynchronous Inter-Site Rerouting

Model Resharding

Input-Only Work Stealing

Reroute Threshold

Work Recomputation

Staleness Checks

a Distributed Training Runtime Local Runtime Ops

Worker Pool

RAN Partition

Train Task Queue

SM Reservation Transition

b

Latency & BW Est.

Hook Layer

Task Runner Training Partition

Sub-Task Tiles (Fit to Slot)

Host CPU

GPU

Figure 8: Two-level elastic training architecture: the inter-site coordinator manages AIR rerouting and MLR rebalancing across RAN sites (top), while each intra-site scheduler fits RAN-slot-bounded GEMM tiles into RAN-slot-level spare compute (bottom). RAN Load Spike Site A

MB1

MB2

MB4

MB5

RAN Load Spike MB1

MB2

Reroute

Est. Budget Miss Site B

MB in Queue MB3

MB3

a) Asynchronous Inter-Replica Rerouting

Migration Model Param

MB4

MB5

MB3

b) Common State Migration

Figure 9: An example illustrating the benefit of AIR compared against common state migration (MB denotes Micro-Batch). stage tasks, and MLR that reshapes the model-parallel layout. The coordinator couples to the training runtime through lightweight hooks at existing synchronization points and derives latency signals from task arrivals and completions at site boundaries. More details in Appendix C.1. Asynchronous inter-site rerouting (AIR). AIR seeks to balance load across sites at sub-second timescales by rerouting micro-batch stage tasks (i.e., a set of GEMM tiles) while keeping model state pinned at each site. The underlying principle is move computation work, not model state. At each AIR control interval, the coordinator assigns queued micro-batch stage tasks to equivalent-stage executors (workers) whose predicted completion time fits within the current training-step admission deadline. If an executor can no longer meet that deadline, AIR reroutes a pending micro-batch stage task to another equivalent-stage executor that holds the same model parameters; when no replica can finish before the training step closes, the not-yet-owned task is deferred for admission in the next training step. Only inputs or hidden states move between sites, and gradients are recomputed when necessary. Consider a two-site example in Fig. 9 with two equivalentstage executors in different data-parallel replicas. If a spike in RAN load slows site A while it executes MB2, AIR stops assigning new micro-batch stage tasks to site A and reroutes the next queued micro-batch stage task (MB3 in the example) to site B, while letting MB2 finish on site A. If MB2 is the last micro-batch of the current training step, AIR speculatively reexecutes its stage task on site B and accepts only the first valid

6.2

Intra-Site Scheduling

Next, we explain how W EAVER adapts to available spare compute (SMs) within a single cell site. Here the challenge arises because FM training kernels are inherently long-running, whereas spare compute is exposed at RAN-slot granularity. Since FM training is dominated by GEMM operations (§2), W EAVER achieves intra-site elasticity by controlling their execution granularity. However, individual GEMMs as a whole can be computationally heavy and may need multiple RAN slots, making them poorly matched to the rapidly changing spare SMs exposed by the spare-compute controller. A GEMM that fits within the available SM budget in a slot may no longer fit if the RAN expands its SM reservation in the next slot. To address this, W EAVER decomposes GEMMs into fine-grained GEMM tiles, which become the basic schedulable unit of computation and can be paused, resumed, and migrated with minimal disruption. RAN-slot-bounded GEMM tiles. As per the above, the intra-site scheduler queues the GEMM tiles assigned by the inter-site scheduler and admits them against each RAN slot’s spare SM and bandwidth budgets, so training compute stays bounded and safely co-exists with the RAN workload. A smoothed window of aggregate spare capacity is reported back to the inter-site scheduler (§6.1). As outlined above, 8

Table 2: OTA results (2 UEs): iperf (5 min bidirectional); WebRTC call quality (10 min bidirectional). Test

Metric

iperf

Throughput (Mbps)

12.1

12.2

−0.8

WebRTC Avg. bitrate (kbps) Packet loss (%) Jitter (ms) RTT (ms) Freeze count

2440 0.35 3.4 22.4 0

2460 0.33 3.3 22.8 1

−0.8 +6.1 +3.0 −1.8 —

Table 3: Aggregated application-layer KPIs (3 UEs, 90 s UL UDP via D-ITG w/ real-world channel replay). Metric

W EAVER NoCtrl ∆ (%)

Throughput, sum (Mbps) Delay, mean (ms) Delay, p95 (ms) Jitter, mean (ms) Jitter, p95 (ms) Packet loss (%)

Scen.

Evaluation

We evaluate W EAVER along three dimensions. First, we determine whether its RAN-centric spare-compute controller stabilizes the RAN SM reservation while preserving RAN performance, and whether the resulting stable spare-compute budget improves co-located training throughput (§7.1). Second, we evaluate whether W EAVER’s two-level elastic training framework efficiently harvests heterogeneous and timevarying spare compute across sites, and study its scalability to larger deployments and models (§7.2). Finally, we quantify W EAVER’s runtime overheads (§7.3).

∆ (%)

4.56 19.0 124.3 4.5 38.1 66.7

+0.1 −3.1 −3.3 −3.0 −5.9 −0.03

Method

Mean SM cnt.↓ SM trans.↓ ddl. met↑

N O C TRL Scen. 1 Y IN YANG RAN W EAVER

14.2 34.1 14.4

435 (0.7/s) 96 (0.16/s) 89 (0.1/s)

99.8% – 99.8%

N O C TRL Scen. 2 Y IN YANG RAN W EAVER

8.4 22.5 8.5

744 (1.2/s) 124 (0.21/s) 187 (0.3/s)

99.8% – 99.8%

evaluate interactive traffic over the air (OTA) using a USRP B210 and two 5G phones. Then, we use D-ITG [7] to generate bursty UDP traffic for three OAI RFSIM UEs to measure finegrained network KPIs while varying the channel conditions based on the production O-RAN traffic [19]. More evaluation details are provided in Appendix D.1. The results show that W EAVER preserves application performance. In OTA experiments (Table 2), the iperf3 throughput differs by less than 1% from N O C TRL while WebRTC performance is also comparable. As for the D-ITG workload (Table 3), W EAVER and N O C TRL achieve the same aggregate throughput and packet-loss rate, while mean and p95 delay and jitter remain comparable. Overall, these results show that W EAVER’s compute-aware scheduling does not degrade application performance.

7.1 Effectiveness of RAN-Centric Spare Compute Control We first evaluate W EAVER’s RAN-centric spare compute controller on a live 5G NR stack, focusing on RAN performance, compute stability, and benefits to co-located training. Testbed. Unless otherwise stated, we use the OAI [49] 5G stack with GPU-accelerated LDPC decoding [8] on a 48-SM DGX Spark, using band n78 (3.5 GHz TDD, 40 MHz), for the experiments conducted in this section. Baselines. N O C TRL uses the unmodified OAI MAC scheduler and exposes its raw SM demand without shaping. Both N O C TRL and W EAVER use CUDA Green Contexts to enable co-location with the training workload. Y IN YANG RAN [60] represents MPS-based GPU sharing: we grant it oracle knowledge of future RAN demand and size each one-second partition for the peak requirement within that interval. We emulate MPS for the RAN workload rather than enforcing MPS directly, since MPS reconfiguration requires CPU fallback and disrupts RAN operation on our testbed. 7.1.1

NoCtrl

4.56 18.4 120.2 4.4 35.9 66.7

Table 4: SM reservation stability and HARQ deadline-met rate in two TRACTOR scenarios (4 UEs, 600 s). ↑/↓ means the higher/lower the better.

W EAVER decomposes GEMMs into GEMM tiles, and admits them only when they satisfy two constraints under the current RAN-slot spare-SM budget: (i) predicted completion before the end of the current RAN slot, and (ii) required memory fits within the remaining memory/interconnect headroom. W EAVER further uses persistent kernels [22, 51].

7

W EAVER

Impact on SM reservation stability. Next, we evaluate whether W EAVER can effectively stabilize the RAN SM reservation and thereby expose a more stable spare-compute budget. We replay two four-UE TRACTOR 5G scenarios (production O-RAN traffic) [19] in OAI RFSIM for 600 s, while varying the channel conditions accordingly: static UEs with steady traffic (Scen. 1, Trial2) and walking UEs with dynamic traffic (Scen. 2, Trial3). Table 4 reports the mean RAN SM reservation, number of SM transitions, and HARQ deadline-met rate. Compared with N O C TRL, W EAVER reduces SM transitions by 4.0–4.9×, while maintaining nearly the same mean SM reservation and the same 99.8% deadline-met rate. Y IN YANG RAN achieves comparable transition counts but reserves 2.4–2.7× more SMs because each partition must accommodate the peak demand within its interval. Thus, consistent with our design in §5, W EAVER effectively smooths short-term RAN demand into longer stable SM-reservation intervals, as further substantiated by the timeline in Fig. 10.

Impact on RAN Performance and SM Reservation

Impact on application performance. First, we use iperf3 to measure sustained throughput and WebRTC [4] calls to 9

Table 6: L40S training under recorded RAN SM reservations. ↑/↓ means the higher/lower the better.

a) NoCtrl (SM Transitions: 24, Avg SM: 15.6) 48

SM Count

32

Method Violations ↓ Efficiency ↑ Token/s ↑ TFLOPS ↑

16 0

Y IN YANG RAN N O C TRL W EAVER

b) Weaver (SM Transitions: 10, Avg SM: 16.0)

48

N/A 2.7% 0.6%

32% 67% 83%

1447 3827 4088

54 118 142

32

able spare compute, compared with 67% for N O C TRL and 32% for Y IN YANG RAN. This translates to 4088 tokens/s and 142 TFLOPS, corresponding to 1.07× higher token throughput and 1.20× higher compute throughput than N O C TRL. Interestingly, the higher training throughput also comes with better allocation compliance. An isolation violation occurs when a training kernel still occupies SMs required by the RAN at decoding start. W EAVER reduces the violation rate from 2.7% under N O C TRL to 0.6%, corresponding to 118 versus 26 events, while reducing their aggregate duration by 2.7×. This shows that the more stable RAN reservation not only exposes more usable compute, but also allows the training workload to adapt more cleanly to the RAN’s changing compute demand.

16 0

0

25

50

75

100

125

150

175

200

RAN Slot Index

Figure 10: SM reservation over time. For brevity, we only show the case of a sample period from Scenario 2. Table 5: Useful throughput of 64-GEMM tasks during RAN coexistence (four UEs, 120 s per case). Useful throughput (TFLOPS) Method Y IN YANG RAN N O C TRL W EAVER

7.1.2

Scenario 1

Scenario 2

0.183 0.481 1.398

0.183 0.779 1.374

Impact on Usable Spare Compute

7.2 Effectiveness of Two-Level Elastic Training

We next evaluate whether the more stable SM reservations produced by W EAVER translate into more usable compute for a co-located workload. Controlled compute workload. Using the same RAN setup and TRACTOR traces (§7.1.1), we run the live end-to-end RAN workload, including actual UE transmissions through the OAI stack, while co-locating a synthetic training workload on the same GPU. The workload consists of repeated chains of 64 dependent GEMMs; let M and N denote the output matrix dimensions, while K is the shared reduction dimension in CM×N = AM×K BK×N , and we have M = 20480, N = K = 1024. This allows us to evaluate the impact of RAN control on useful computation while keeping the training-side execution identical across methods. As shown in Table 5, W EAVER sustains approximately 1.4 TFLOPS in both scenarios, achieving 2.91× and 1.76× the useful throughput of N O C TRL, and 7.50–7.62× that of Y IN YANG RAN. This gain follows directly from the more stable reservations while not degrading RAN performance, as observed in §7.1.1: fewer SM transitions provide longer uninterrupted intervals in which the co-located workload can make useful progress. End-to-end FM training workload. We next verify that this benefit carries over to actual FM training. Here, we use a datacenter grade GPU, an NVIDIA L40S, and replay the earlier RAN SM-reservation traces (rescaled to the L40S SM count) to train a OPT-13B [83] configured with 12 layers (with target global batch size B⋆ = 4 and sequence length 512) under the same reservation traces. All settings use an identical model and optimizer state. For Y IN YANG RAN, each reconfiguration checkpoints training state to host memory. As shown in Table 6, W EAVER harvests 83% of the avail-

Having shown that W EAVER exposes more usable spare compute at RAN sites, we next evaluate whether its two-level elastic training framework can efficiently harvest this heterogeneous and time-varying compute across sites. We first evaluate end-to-end training on our eight-site GPU testbed, then isolate its adaptation to dynamic spare compute, and finally study scalability to larger deployments and models. Testbed. Unless otherwise stated, we evaluate W EAVER on an eight-site GPU-backed 5G testbed, with one NVIDIA L40S GPU (48 GB) at each site. The sites are distributed across different cities; measured inter-site bandwidth and latency are reported in Appendix D.2, and implementation details are provided in Appendix E. We replay the SM-reservation traces from §7.1.1, rescaled to the L40S SM count, to reproduce millisecond-scale RAN compute dynamics. Models and datasets. We use the OPT family throughout the training evaluation. Our physical-testbed experiments use OPT-125M–6.7B, while the large-scale simulation in §7.2.3 evaluates OPT-13B–70B. We train OPT on SST-2 [66] with global batch size B⋆ = 32 and sequence length 512. The training scheduler determines the micro-batch size based on the data-parallel (DP) and pipeline-parallel (PP) configuration. Baselines. We compare against three distributed training systems designed for heterogeneous and time-varying resources: (i) DTFM [79], which uses DP+PP with coarsegrained checkpoint-and-restore; (ii) A STEROID [78], which uses hybrid PP and pipeline replay to tolerate stragglers; and (iii) C ONFIDANT [9], which combines local updates with periodic global aggregation to tolerate slow or unavailable sites. In our evaluation, these baselines use MPS for GPU sharing, while W EAVER uses CUDA Green Contexts. 10

DTFM

Asteroid

Confidant

4 3 2 1 0

25m 50m PT-1.3b PT-2.7b PT-6.7b O O O OPT-1 OPT-3

Table 7: OPT-13B training throughput (K tokens/s) as site count increases; 48 GB GPUs.

Weaver

b) Achieved TFLOPS (% of Cluster Total) Achieved TFLOPS (%)

Throughput Norm.

a) Normalized Throughput

30

Rescaled System

128

256

512

128

256

512

10

W EAVER Confidant* Confidant

76.8 47.6 46.6

87.8 47.6 46.6

92.2 47.6 46.6

106.7 58.5 46.5

198.8 77.6 48.6

350.6 87.3 49.7

0

OPT-1

25m T-350m PT-1.3b PT-2.7b PT-6.7b O O O OP

training across sites as spare compute changes. Over a 6-minute trace with 41 capacity-drop events across eight sites, W EAVER sustains 15–32% of peak cluster TFLOPS, compared with 6–11% for C ONFIDANT and 1–2% for DTFM. This shows that W EAVER effectively preserves useful training computation as capacity shifts across sites. Inter-site adaptation behavior. We zoom into the scheduler internals to understand how it reacts to these dynamics. Over the same 6-minute replay, spare compute shifts 80 times at the training-step timescale. The AIR fast-path handles 86.2% of these fluctuations, by rerouting pending micro-batch stage tasks, without invoking model layout rebalancing (MLR), with a median exposed stall of 66 ms (p95: 68 ms). Only 11 events escalate to the MLR slow-path. For these events, the only exposed critical-path cost is the 528 ms PREPARE drain; model-state transfer and layout activation overlap with useful training on unaffected stages. Together, these results validate the two-timescale design in §6.1: AIR absorbs most transient capacity changes, while MLR is invoked only for persistent imbalance. Additional mechanism-level results and parameter sensitivity are reported in Appendix D.4.

Figure 11: Training throughput and achieved TFLOPS as a percentage of peak cluster compute across OPT model sizes. 7.2.1

End-to-End Training Performance

We first evaluate the end-to-end benefit of W EAVER’s twolevel elastic training framework under heterogeneous and time-varying spare compute across multiple RAN sites. We evaluate the full W EAVER stack against each baseline under the same SM-reservation trace replay for 1000 training steps, covering both per-site capacity fluctuations and longer-term inter-site drift. We report token throughput as the end-to-end training metric and achieved compute throughput, normalized by aggregate peak cluster compute throughput (in terms of TFLOPS), as a measure of how effectively available compute is converted into useful training work. Training throughput. From the results in Fig. 11a, we observe that W EAVER consistently achieves higher end-to-end training throughput across the evaluated model sizes. For OPT-1.3B, W EAVER achieves 35K tokens/s, 2.1× that of the strongest baseline, C ONFIDANT (17K tokens/s). The gap widens to 3.7× for OPT-6.7B (6.3K versus 1.7K tokens/s). W EAVER also consistently achieves a higher fraction of peak cluster compute throughput across model sizes (Fig. 11b). Together, the higher token throughput and compute utilization show that W EAVER more effectively converts dynamic spare compute into useful training work. In contrast, the baselines rely on coarser adaptation mechanisms that are less effective under dynamic and heterogeneous resource availability. Training stability and convergence. Beyond mean throughput, W EAVER also maintains stable training under replayed RAN dynamics, with lower step-level throughput variation and fewer throughput drops than the baselines. In our experiments, the training-loss trajectory closely tracks a fixed-cluster synchronous SGD reference, indicating that W EAVER’s dynamic adaptation does not compromise training convergence. We provide more detailed stability and convergence results in Appendix D.3. 7.2.2

Typical

20

7.2.3

Scalability to Larger Deployments and Models

Finally, we evaluate whether W EAVER’s training design continues to scale beyond the size of our physical testbed, both to larger RAN deployments and larger foundation models. Environment settings. We evaluate 128–512-site deployments, all with 48 GB GPUs, under two RAN settings. Rescaled scales our measured testbed topology while preserving its link characteristics. Typical uses a representative RAN star topology with 32 sites per core, 25 Gbps/1 ms site-to-core, 400 Gbps/20 ms core-to-core, and 10 Gbps peer links [21, 36]. Methodology. We use a profile-based simulator built on Morphling [77]. Its inputs combine measured per-layer compute profiles from truncated-depth training runs, inter-site bandwidth and latency, per-link contention, per-site memory constraints, and TRACTOR-derived SM-reservation traces; further details are provided in Appendix D.5. The simulator does not model site failures or stragglers beyond their effect on the available SM reservation. We compare against C ON FIDANT , the strongest baseline in our physical testbed, and C ONFIDANT *, which strengthens it with exhaustive DP+PP configuration search and perfect spare-compute harvesting. This separates W EAVER’s scaling benefits from limitations due to baseline configuration or compute harvesting.

Adaptation to Dynamic Spare Compute

We next isolate how effectively W EAVER’s inter-site scheduler adapts training to dynamic spare compute across sites. To remove differences from intra-site execution, we enable the same intra-site scheduler for all methods and initialize them with the same model-parallel partition. Thus, the remaining performance difference comes from how each method adapts 11

Table 8: Larger-model feasibility and throughput on 512 sites with 48 GB GPUs. OOM = out of memory. Peak mem.

K tokens/s

Model

System

(GB)

Rescaled

Typical

OPT-30B

W EAVER Confidant* Confidant

30.2 57.6 70.0

38.4 OOM OOM

146.1 OOM OOM

W EAVER Confidant* Confidant

40.2 75.2 149.8

15.4 OOM OOM

58.4 OOM OOM

OPT-70B

are provided in Appendix E. Broader deployment considerations, including full-L1 acceleration, applicability to other non-RAN workloads, tenant isolation, and energy/thermal constraints, are discussed in Appendix F.

8

Compute sharing for RAN. We discuss recent GPU-sharing techniques for AI-and-RAN in §2; here, we focus on complementary work on sharing CPU resources in vRAN. Concordia [17] dynamically reallocates CPU cores across RAN and co-located workloads, while other systems use CPU quotas [45] or fine-grained redirection and scheduling [26] to reclaim otherwise idle CPU capacity. These approaches benefit from fine-grained CPU thread preemption and migration, mechanisms that do not directly transfer to GPU execution. W EAVER instead targets GPU-accelerated RAN infrastructure, where resource sharing must be coordinated through SM allocation while preserving per-slot RAN deadlines. Characterizing RAN GPU Utilization. Prior work has profiled the computational behavior of GPU-accelerated RAN processing. CloudRIC [61] studies GPU utilization under different RAN loads, ETHOS [72] characterizes the performance and efficiency of virtualized O-RAN processing, and DecodeX [54] benchmarks LDPC decoding across different compute platforms. These studies primarily focus on aggregate utilization (which is coarse-grained) or individual RAN components. In contrast, our characterization is aimed specifically at understanding the spare GPU capacity available for co-located workloads. We therefore study GPU usage along multiple resource dimensions, i.e., SM occupancy, arithmetic utilization, and memory-bandwidth utilization, at RAN-slot granularity, and further examine how spare compute varies temporally and spatially across cell sites. Resource harvesting. Volunteer-computing systems [2, 3] and harvested cloud VMs [85] aggregate opportunistic idle resources for embarrassingly parallel workloads, typically treating capacity availability coarsely and scavenging slack over much longer timescales. In contrast, spare compute in GPU-accelerated RANs varies across fine-grained temporal, spatial, and resource dimensions, while the primary RAN workload must retain strict performance guarantees. These systems therefore do not address the fine-grained resource adaptation or cross-site coordination required for tightly coupled FM training. Dynamic ML training systems. Existing training systems [9, 13, 79, 82] adapt to heterogeneous or time-varying resources, but are not designed for the RAN setting, where GPU capacity can change at sub-millisecond timescales due to a high-priority, latency-critical co-tenant. Other systems tolerate resource churn through comparatively heavyweight mechanisms, such as checkpoint-and-restart (e.g., Bamboo [69], Mario [33]) or recomputation (e.g., SWARM [58], Asteroid [78]), which are poorly matched to such fine-grained re-

Scaling across sites. Table 7 shows that W EAVER continues to benefit from additional sites despite heterogeneous spare compute and constrained inter-site connectivity. Under Typical, OPT-13B throughput increases from 106.7K to 350.6K tokens/s as the deployment grows from 128 to 512 sites, a 3.29× increase, compared with 1.49× for C ONFIDANT *. At 512 sites, W EAVER achieves 4.0× the throughput of C ONFI DANT * and 7.1× that of C ONFIDANT . Scaling is more constrained under the Rescaled interconnect: W EAVER increases from 76.8K to 92.2K tokens/s, while both baselines remain nearly flat. These results show that W EAVER can effectively exploit large-scale distributed spare compute. Scaling to larger models. We next evaluate whether W EAVER can extend training to models that exceed the sizes feasible on our physical testbed. On 512 sites, W EAVER fits OPT-30B and OPT-70B within the 48 GB per-site memory budget, with peak memory footprints of 30.2 GB and 40.2 GB, respectively (Table 8). In contrast, both C ONFIDANT and C ONFIDANT * exceed the per-site memory budget for both models. Under Typical, W EAVER sustains 146.1K tokens/s for OPT-30B and 58.4K tokens/s for OPT-70B. Thus, W EAVER remains capable of harvesting distributed RAN spare compute as both deployment size and model size increase.

7.3

Related Work

Runtime Overheads

We quantify the runtime overheads introduced by W EAVER across both the RAN and training control paths. W EAVER introduces modest runtime overheads. The introduction of the RAN-centric spare compute controller into the MAC scheduler incurs an overhead of 0.8% (0.13µs) and 13.4% (6.56µs) at the p50 and p99, respectively. W EAVER’s training control plane requires only ∼120 µs of CPU scheduling per training step, accounting for <0.3% of step time, confirming that W EAVER control plane is lightweight. To support rapid adaptation to SM-reservation changes, each worker pre-creates up to 16 Green Context streams, consuming 320 MB of GPU memory (<0.7% of a 48 GB L40S), while rebinding the next GEMM tile to a pre-created Green Context takes less than 10 µs (Appendix E). Among sites, AIR reroutes only KB– MB-scale inputs or activations rather than model state. These overheads are substantially smaller than the multi-GB state transfers and multi-second recovery costs incurred by coarsegrained inter-site scheduling. Further implementation details 12

source dynamics. Moreover, approaches such as per-device AllReduce in DTFM [79], and communication that grows with participating device count in Bamboo [69] and Mario [33], can become costly over constrained inter-site links. In contrast, W EAVER separates adaptation across timescales: fine-grained intra-site scheduling reacts to RAN-slot-level changes, while lightweight work rerouting and infrequent model-layout rebalancing handle cross-site dynamics.

9

State-of-the-art and the road ahead. Computer Networks, 182:107516, 2020. [7] A. Botta, A. Dainotti, and A. Pescapé. A tool for the generation of realistic network workload for emerging networking scenarios. Computer Networks, 56(15):3531– 3547, 2012. [8] S. Cammerer, G. Marcus, T. Zirr, F. A. Aoudia, L. Maggi, J. Hoydis, and A. Keller. Sionna research kit: A GPUaccelerated research platform for AI-RAN, 2025.

Conclusions

[9] Y. Chen, Y. Yan, S. Ge, Y. Qin, Y. Zheng, Q. Yang, S. He, Z. Shi, J. Chen, and Y. Shu. Confidant: Customizing transformer-based LLMs via collaborative training on mobile devices. In ACM MobiCom, pages 483–497. ACM, 2025.

This paper presents W EAVER, a system for safely and efficiently harvesting spare GPU compute in AI-RAN infrastructure for FM training. W EAVER combines a RAN-centric spare-compute controller that stabilizes SM reservations with a two-level elastic training framework that adapts to resource variability across sites and timescales. Our evaluation shows that W EAVER improves FM training throughput while preserving RAN performance.

[10] A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.

Acknowledgements

[11] P. H. Coppock, B. Zhang, E. H. Solomon, V. Kypriotis, L. Yang, B. Sharma, D. Schatzberg, T. C. Mowry, and D. Skarlatos. LithOS: An operating system for efficient machine learning on GPUs. In ACM SOSP, pages 1–17. ACM, 2025.

This work was supported in part by the UKRI/EPSRC grants UKRI860 and UKRI554, and the AI-RAN Alliance Innovation Award.

References

[12] S. D’Oro, M. Polese, L. Bonati, H. Cheng, and T. Melodia. dApps: Distributed Applications for Real-Time Inference and Control in O-RAN, 2022.

[1] AI-RAN Alliance. AI-RAN alliance vision and mission white paper. https://ai-ran.org/wp-content/u ploads/2024/12/AI-RAN_Alliance_Whitepaper .pdf [Accessed: Mar 2026], 2024. [2] D. P. Anderson. BOINC: A platform for volunteer computing. Journal of Grid Computing, 18(1):99–122, 2020.

[13] J. Duan, Z. Song, X. Miao, X. Xi, D. Lin, H. Xu, M. Zhang, and Z. Jia. Parcae: Proactive, liveputoptimized DNN training on preemptible instances. In USENIX NSDI, pages 1121–1139. USENIX Association, 2024.

[3] A. L. Beberg, D. L. Ensign, G. Jayachandran, S. Khaliq, and V. S. Pande. Folding@home: Lessons from eight years of volunteer distributed computing. In IPDPS, pages 1–8. IEEE, 2009.

[14] A. Duttagupta, M. Jabbari, C. Fiandrino, M. Fiore, and J. Widmer. SYMBXRL: Symbolic explainable deep reinforcement learning for mobile networks. In IEEE INFOCOM, pages 1–10. IEEE, 2025.

[4] N. Blum, S. Lachapelle, and H. Alvestrand. WebRTC: real-time communication for the open web platform. Communications of the ACM, 64(8):50–54, 2021.

[15] L. Emberson and D. Owen. The stock of computing power from NVIDIA chips is doubling every 10 months. https://epoch.ai/data-insights/nvidia-chi p-production [Accessed: Mar 2026], 2025.

[5] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. v. Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, et al. On the opportunities and risks of foundation models, 2022.

[16] R. Fan, T. Ren, M. Xie, S. Gao, J. Shu, and Y. Lu. GPREEMPT: GPU preemptive scheduling made general and efficient. In USENIX ATC, 2025. [17] X. Foukas and B. Radunovic. Concordia: teaching the 5G vRAN to share compute. In ACM SIGCOMM, pages 580–596. ACM, 2021.

[6] L. Bonati, M. Polese, S. D’Oro, S. Basagni, and T. Melodia. Open, programmable, and virtualized 5G networks:

13

[29] A. Kalia, N. Lazarev, L. Xue, X. Foukas, B. Radunovic, and F. Y. Yan. Towards energy efficient 5g vran servers. In USENIX NSDI, pages 1205–1219. USENIX Association, 2025.

[18] C. Ge, A. Mahimkar, Z. Ge, R. Fernandez, J. Maniaci, S. Pathak, and M. Shah. Iridescence: Improving configuration tuning in the presence of confounders for 5G NSA networks. Proc. ACM Netw., 3(CoNEXT1):1–22, Mar. 2025.

[30] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models, 2020.

[19] J. Groen, M. Belgiovine, U. Demir, B. Kim, and K. Chowdhury. TRACTOR: Traffic analysis and classification tool for open RAN. In IEEE ICC, pages 4894– 4899. IEEE, 2024.

[31] L. Kundu, X. Lin, R. Gadiyar, J.-F. Lacasse, and S. Chowdhury. AI-RAN: Transforming RAN with AIdriven computing infrastructure. IEEE Communications Magazine, 64(1):168–174, 2026.

[20] GSMA. 5G-era Mobile Network Cost Evolution. http s://www.gsma.com/solutions-and-impact/tech nologies/networks/gsma_resources/5g-era-m obile-network-cost-evolution/ [Accessed: Sep 2026], 2019.

[32] E. Li, L. Zeng, Z. Zhou, and X. Chen. Edge AI: Ondemand accelerating deep neural network inference via edge computing. IEEE Transactions on Wireless Communications, 19(1):447–457, 2020.

[21] GSMA. Mobile Backhaul: An Overview. https://ww w.gsma.com/solutions-and-impact/technologi es/networks/gsma_resources/mobile-backhau l-an-overview/ [Accessed: Sep 2026], 2019.

[33] W. Liu, M. Li, G. Tan, and W. Jia. Mario: Near zerocost activation checkpointing in pipeline parallelism. In PPoPP, pages 197–211. ACM, 2025.

[22] K. Gupta, J. A. Stuart, and J. D. Owens. A study of persistent threads style GPU programming for GPGPU workloads. In InPar, pages 1–14. IEEE, 2012.

[34] F. Mani. Introduction to 5G RAN PHY Simulators in OpenAirInterface. OAI Webinar Series, Chapter 4. http s://openairinterface.org/introduction-to-5 g-ran-phy-simulators-in-openairinterface/ [Accessed: Mar 2026], 2025.

[23] L. Han, Z. Zhou, and Z. Li. Pantheon: Preemptible multi-DNN inference on mobile edge GPUs. In ACM MobiSys, pages 465–478. ACM, 2024.

[35] R. Motwani and P. Raghavan. Randomized Algorithms. Cambridge University Press, 1995.

[24] M. Han, H. Zhang, R. Chen, and H. Chen. Microsecondscale preemption for concurrent GPU-accelerated DNN inferences. In USENIX OSDI, pages 539–558, 2022.

[36] Nokia. A Revealing Look into the Evolution of Mobile Transport (Anyhaul White Paper). https://pages.no kia.com/14265.Mobile.Anyhaul.html [Accessed: Sep 2026], 2020.

[25] S. Hui, H. Wang, Z. Wang, X. Yang, Z. Liu, D. Jin, and Y. Li. Understanding mobile traffic patterns of large scale cellular networks: The shanghai cellular network dataset. Proc. ACM IMWUT, 18(4), 2024.

[37] Nokia. Nokia partners with NVIDIA. https://www. nokia.com/newsroom/nokia-partners-with-nvi dia/ [Accessed: Mar 2026], 2025.

[26] Y. Jia, Y. Zhong, M. Wang, J. Gao, P. Zhang, X. Liu, and X. Jin. Aquifer: Transparent microsecond-scale scheduling for vRAN workloads. IEEE Transactions on Services Computing, 17(6):3171–3184, 2024.

[38] NVIDIA. CUDA C++ Programming Guide, 2024. Version 12.6. https://docs.nvidia.com/cuda/archi ve/12.6.0/cuda-c-programming-guide/.

[27] S. Jin, S. Kim, S. Ha, and K. Lee. End-to-end coordination of RAN and edge server for latency-critical inference serving over cellular networks. Proc. ACM Netw., 3(CoNEXT4):1–23, Nov. 2025.

[39] NVIDIA. NVIDIA Aerial SDK: GPU-accelerated 5G radio access network. https://docs.nvidia.com/ aerial [Accessed: Mar 2026], 2024.

[28] J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, et al. Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873):583–589, 2021.

[40] NVIDIA. NVIDIA cuPHY: GPU-accelerated 5G physical layer. https://docs.nvidia.com/aerial/cud a-accelerated-ran [Accessed: Mar 2026], 2024. [41] NVIDIA. NVIDIA Data Center GPU Manager (DCGM), 2024. https://developer.nvidia.c om/dcgm.

14

[42] NVIDIA. Enhanced DU Performance and Workload Consolidation for 5G/6G with NVIDIA Aerial CUDAAccelerated RAN. https://developer.nvidia.com /blog/enhanced-du-performance-and-workloa d-consolidation-for-5g-6g-with-aerial-cud a-accelerated-ran/ [Accessed: Mar 2026], 2025.

[54] Z. Qi, Y. Yao, Y. Li, C.-H. Tung, J. Zheng, D. Zhuo, and T. Chen. DecodeX: Exploring and Benchmarking of LDPC Decoding Across CPU, GPU, and ASIC Platforms. In Proceedings of the 27th International Workshop on Mobile Computing Systems and Applications, HotMobile ’26, page 127–132, New York, NY, USA, 2026. Association for Computing Machinery.

[43] NVIDIA. NVIDIA AI Aerial. https://developer. nvidia.com/industries/telecommunications/a i-aerial [Accessed: Mar 2026], 2026.

[55] J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD, pages 3505–3506. ACM, 2020.

[44] NVIDIA Corporation. NVIDIA Aerial™ CUDAAccelerated RAN. https://github.com/NVIDI A/aerial-cuda-accelerated-ran [Accessed: Sep 2026], 2025.

[56] J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He. ZeRO-Offload: Democratizing billion-scale model training. In USENIX ATC, 2021.

[45] A. F. Ocampo, M.-R. Fida, J. F. Botero, A. Elmokashfi, and H. Bryhni. Opportunistic CPU sharing in mobile edge computing deploying the cloud-RAN. IEEE Transactions on Network and Service Management, 20(3):2201–2217, 2023.

[57] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10674–10685. IEEE, 2022.

[46] Ookla. Speedtest Global Index. https://www.speedt est.net/global-index [Accessed: Mar 2026], 2024.

[58] M. Ryabinin, T. Dettmers, M. Diskin, and A. Borzunov. Swarm parallelism: Training large models can be surprisingly communication-efficient. In ICML, 2023.

[47] OpenAI. Model optimization. OpenAI API documentation. https://developers.openai.com/api/docs /guides/model-optimization [Accessed: Jul 2026], 2026.

[59] J. X. Salvat Lozano, J. A. Ayala-Romero, A. GarciaSaavedra, and X. Costa-Perez. Kairos: Energy-efficient radio unit control for O-RAN via advanced sleep modes. In IEEE INFOCOM, pages 1–10. IEEE, 2025.

[48] OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, et al. GPT-4 technical report, 2024.

[60] L. L. Schiavo, J. A. Ayala-Romero, A. Garcia-Saavedra, M. Fiore, and X. Costa-Perez. YinYangRAN: Resource multiplexing in GPU-accelerated virtualized RANs. In IEEE INFOCOM, pages 721–730. IEEE, 2024.

[49] OpenAirInterface. OpenAirInterface: 5G Software Alliance. https://openairinterface.org [Accessed: Mar 2026], 2024.

[61] L. L. Schiavo, G. Garcia-Aviles, A. Garcia-Saavedra, M. Gramaglia, M. Fiore, A. Banchs, and X. Costa-Perez. CloudRIC: Open radio access network (O-RAN) virtualization with shared heterogeneous computing. In ACM MobiCom, pages 558–572. ACM, 2024.

[50] OpenCelliD. OpenCelliD – the world’s largest open database of cell towers. https://www.opencellid.o rg/stats.php [Accessed: Mar 2026], 2025. [51] M. Osama, D. Merrill, C. Cecka, M. Garland, and J. D. Owens. Stream-K: Work-centric parallel decomposition for dense matrix-matrix multiplication on the GPU. In PPoPP, pages 429–431. ACM, 2023.

[62] R. Shafin, L. Liu, V. Chandrasekhar, H. Chen, J. Reed, and J. C. Zhang. Artificial intelligence-enabled cellular networks: A critical path to beyond-5G and 6G. IEEE Wireless Communications, 27(2):212–217, 2020.

[52] P. F. Pérez, C. Fiandrino, and J. Widmer. Characterizing and modeling mobile networks user traffic at millisecond level. In ACM WiNTECH, pages 64–71. ACM, 2023.

[63] S. D. A. Shah, M. Hafeez, A. Salama, and S. A. R. Zaidi. Proactive AI-and-RAN workload orchestration in ORAN architectures for 6G networks. IEEE Open Journal of the Communications Society, 6:7939–7954, 2025.

[53] M. Polese, S. D’Oro, A. Lacava, N. Mohamadi, Y. Lee, D. Villa, and C. Dick. Research Report on dApp Architecture and Interfaces. Technical Report nGRGRR-2025-05, O-RAN ALLIANCE, next Generation Research Group (nGRG), Feb. 2026. O-RAN nGRG Contributed Research Report, Version 1.0, released 2026-0211.

[64] W. Shen, M. Han, J. Liu, R. Chen, and H. Chen. XSched: Preemptive scheduling for diverse XPUs. In USENIX OSDI, 2025.

15

[76] T. Xu, L. Xue, Z. Lu, A. Jackson, and L. Mai. MoE-Gen: High-throughput MoE inference on a single GPU with module-based batching, 2025.

[65] M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020.

[77] L. Xue, Y. Xia, E. Mendi, I. Bashir, J. Yang, M. Lee, and M. K. Marina. Morphling: Emulator for distributed machine learning at the edge. In Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services Workshops (MobiSys Workshops), pages 307–314. ACM, 2026.

[66] R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP, pages 1631–1642. ACL, 2013. [67] F. Strati, X. Ma, and A. Klimovic. Orion: Interferenceaware, fine-grained GPU sharing for ML applications. In EuroSys, pages 1075–1092. ACM, 2024.

[78] S. Ye, L. Zeng, X. Chu, G. Xing, and X. Chen. Asteroid: Resource-efficient hybrid pipeline parallelism for collaborative DNN training on heterogeneous edge devices. In ACM MobiCom, pages 312–326. ACM, 2024.

[68] C. Sun, U. Pawar, M. Khoja, X. Foukas, M. K. Marina, and B. Radunovic. SpotLight: Accurate, explainable and efficient anomaly detection for open RAN. In ACM MobiCom, pages 923–937. ACM, 2024.

[79] B. Yuan, Y. He, J. Q. Davis, T. Zhang, T. Dao, B. Chen, P. Liang, C. Re, and C. Zhang. Decentralized training of foundation models in heterogeneous environments. In NeurIPS, 2022.

[69] J. Thorpe, P. Zhao, J. Eyolfson, Y. Qiao, Z. Jia, M. Zhang, R. Netravali, and G. H. Xu. Bamboo: Making preemptible instances resilient for affordable training of large DNNs. In USENIX NSDI, 2023.

[80] T. Zanouda, M. Masoudi, F. G. Gebre, and M. Dohler. Telecom foundation models: Applications, challenges, and future trends, 2024.

[70] V. Volkov. Understanding Latency Hiding on GPUs. PhD thesis, EECS Department, University of California, Berkeley, Aug. 2016.

[81] H. Zhang, A. Cardoza, P. B. Chen, S. Angel, and V. Liu. Fault-tolerant and transactional stateful serverless workflows. In USENIX OSDI, pages 1187–1204, 2020.

[71] C. Wei, A. Kak, N. Choi, and T. Wood. 5gperf: profiling open source 5g ran components under different architectural deployments. In ACM SIGCOMM Workshop on 5G and Beyond Network Measurements, Modeling, and Use Cases, page 43–49, 2022.

[82] S. Zhang, L. Diao, C. Wu, Z. Cao, S. Wang, and W. Lin. Hap: SPMD DNN training on heterogeneous GPU clusters with automated program synthesis. In EuroSys, pages 524–541. ACM, 2024.

[72] Z. Wu, R. Doost-Mohammady, and A. Sabharwal. ETHOS: Demystifying performance, energy, and computational efficiency in virtualized 5G O-RAN networks. In ACM WiNTECH, pages 97–104, 2025.

[83] S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, et al. OPT: Open pre-trained transformer language models, 2022.

[73] D. Xu, A. Zhou, G. Wang, H. Zhang, X. Li, J. Pei, and H. Ma. Tutti: coupling 5G RAN and mobile edge computing for latency-critical video analytics. In ACM MobiCom, pages 729–742. ACM, 2022.

[84] S. Zhang, A. Xu, Q. Chen, H. Zhao, W. Cui, Z. Wang, Y. Li, L. Xiao, and M. Guo. Efficient performance-aware GPU sharing with compatibility and isolation through kernel space interception. In USENIX ATC, 2025.

[74] F. Xu, Y. Li, H. Wang, P. Zhang, and D. Jin. Understanding mobile traffic patterns of large scale cellular towers in urban environment. IEEE/ACM Transactions on Networking, 25(2):1147–1161, 2017.

[85] Y. Zhang, I. Goiri, G. I. Chaudhry, R. Fonseca, S. Elnikety, C. Delimitrou, and R. Bianchini. Faster and cheaper serverless computing on harvested resources. In ACM SOSP, pages 724–739. ACM, 2021.

[75] K. Xu, R. Singh, M. Fiore, M. K. Marina, H. Bilen, M. Usama, H. Benn, and C. Ziemlicki. SpectraGAN: spectrum based generation of city scale spatiotemporal mobile network traffic data. In ACM CoNEXT, pages 243–258. ACM, 2021.

[86] L. Zheng, Z. Li, H. Zhang, Y. Zhuang, Z. Chen, Y. Huang, Y. Wang, Y. Xu, D. Zhuo, E. P. Xing, J. E. Gonzalez, and I. Stoica. Alpa: Automating inter- and intra-operator parallelism for distributed deep learning. In USENIX OSDI, 2022.

16

SMU

ACU

GBU

SM Reservation

(a) Without SM Reservation (SM = 48)

main takeaway is that SM reservation trades a modest latency increase for substantially more stable and shareable headroom. SM reservation creates headroom in all three dimensions. Without an SM reservation (Fig. 12a), LDPC decoding bursts across all 48 SMs for ∼550 µs (SMU ∼100 %, GBU ∼65 %, ACU ∼40 %), after which the GPU sits idle for the remaining ∼450 µs. With a 16-SM reservation (Fig. 12b), the same work is spread over a longer interval: SMU stabilizes at ∼33 % for ∼820 µs, and the freed resources along all three dimensions become available for spatial sharing with co-located workloads. GBU drops because fewer SMs issue concurrent memory requests, spreading the same traffic over a longer window. ACU also drops because fewer SMs are active, but in our measurements the per-SM ACU within the active SM increases: denser thread-block packing gives the warp scheduler more warps to interleave compute with memory stalls [38, 70]. 3× fewer SMs incur only 1.5× slowdown. This headroom comes at a relatively small cost in completion time. Reducing the SM allocation by 3× (from 48 to 16 SMs) increases decoding latency by only 1.5× (∼550 µs to ∼820 µs), because fewer SMs each host more resident warps, improving per-SM occupancy and hiding memory stalls more effectively.

(b) SM Reservation = 16

100

Utilization (%)

80

Spatial Sharing

60 40

16 SMs

20 0 0

200 400 600 800 Time within slot (us)

1000

0

200 400 600 800 Time within slot (us)

1000

Figure 12: Per-slot GPU utilization (SMU, ACU, GBU) without and with SM reservation: the gray region shows headroom for spatial sharing.

A

Spare Compute Characterization

A.1

The LDPC Decoding Is Memory-Bound

We apply a roofline analysis to SionnaRK’s [8] GPUaccelerated LDPC decoding implementation to show that the PHY kernels are decisively memory-bandwidth-bound. The relevant comparison point is the roofline ridge point, η∗ = 1000 · P [TFLOPS]/B [GB/s] in ops/byte, where P is FP16 Tensor Core peak throughput and B is device memory bandwidth. For representative AIRAN accelerators (DGX Spark, GH200, and L40S), this ridge point lies between 246 and 458 ops/byte. We then estimate the arithmetic intensity of a representative 5G NR LDPC configuration (BG1, lifting factor Z=128, Offset Min-Sum, 10 iterations) by modeling the two dominant kernel types at edge granularity (FP16, 2 B per message):

B

This appendix elaborates on §5. It first explains the lookup table that maps RAN configurations to SM demand, then presents the full pseudocode of the compute-aware UL scheduler, and finally summarizes the parameters that govern its responsiveness and stability.

• Check-node (CN) update: 10 ops/edge over 4 B/edge ⇒ ηCN ≈ 2.5 ops/byte.

B.1

• Variable-node (VN) update: 2 ops/edge over 10 B/edge ⇒ ηVN ≈ 0.2 ops/byte.

4.9 × 106 ≈ 0.9 ops/byte, 5.7 × 106

Lookup Table

The lookup table captures the relationship between RAN configuration and SM demand. It is indexed by (MCS, PRB count, SM count) and records whether LDPC decoding meets the RAN-slot deadline for each triple. During scheduling, the controller uses it in two directions: PRB budgeting asks how many PRBs a given SM budget can support, while SM adjustment asks how many SMs the RAN requires for a given PRB demand. These two accesses are exposed as LUT_PRB and LUT_SM, which are different projections of the same table: • LUT_PRB(mcs, sm_count): fixes the MCS index and SM count, then returns the maximum PRB count whose decoding meets the RAN-slot deadline. • LUT_SM(mcs, num_prbs): fixes the MCS index and PRB count, then returns the minimum SM count that meets the RAN-slot deadline. We construct the table in two steps. First, we profile standalone LDPC decoding with the OAI ulsim tool: for each (MCS, PRB, SM count) triple, decoding is benchmarked on

BG1 has 316 non-zero entries in its base graph (E = 316Z = 40,448 edges/block). Aggregating over I=10 iterations yields total traffic of ≈5.7 MB and ≈4.9×106 ops, which gives an overall intensity of ηLDPC ≈

Compute-Aware Scheduler Details

(1)

This is two to three orders of magnitude below the device ridge points (η∗ ∈ [246, 458] ops/byte), confirming that LDPC decoding is decisively memory-bandwidth-bound on these accelerators.

A.2 Effect of SM Reservation on GPU Utilization Fig. 12 compares per-slot utilization on a 48-SM GPU (MCS 15, 162 PRBs) with and without SM reservation. The

17

Table 9: Compute-aware scheduler parameters. Symbol

Description

SM_MIN SM_MAX STEP BACKLOG_THR DECR_THR PRB_TOL MAX_ITER

Minimum SM count (floor) Maximum SM count (all SMs) SM increment/decrement step Per-UE backlog ramp-up threshold Ramp-down patience (slots) PRB slack tolerance in step (c) Pipeline iteration bound

the DGX Spark. Setting the minimum to 8 SMs ensures that even the lightest RAN workload can still be decoded within the RAN-slot deadline. BACKLOG_THR. governs when the scheduler switches from smoothing to draining. We set it to 64 KB, approximately half of a full-bandwidth transport block at moderate MCS. A smaller value triggers ramp-up more eagerly, improving UE latency but causing more SM transitions; a larger value absorbs more traffic bursts at a constant SM tier, reducing transitions at the cost of slightly higher buffer occupancy. DECR_THR. governs how much evidence of overprovisioning is required before releasing SMs. We set the ramp-down patience to 30 consecutive RAN slots (∼15 ms at 30 kHz SCS). The scheduler therefore steps down only after sustained over-provisioning, preventing premature SM release from fragmenting the SM reservation. This value is chosen to exceed the typical burst duration observed in the RAN. PRB_TOL. bounds how much PRB under-allocation is tolerated in iteration control. We set it to 10 PRBs, so the scheduler accepts small slack gaps without another pass. MAX_ITER. bounds the iteration loop. We cap the pipeline at 5 iterations per slot. In practice, it converges in 1–3 iterations in >99% of slots; this bound is only a safeguard against non-convergence under extreme channel conditions or workloads.

Value 8 SMs 48 SMs 8 SMs 64 KB 30 slots 10 PRBs 5

the target GPU with the SM reservation applied, and the P99 latency is checked against a conservative 800 µs RAN-slot decoding budget. Second, we calibrate the table against endto-end OAI runs with rfsim under synthetic extreme channel conditions and workloads, ensuring that the resulting budgets introduce no additional HARQ deadline violations beyond the baseline scheduler. For (MCS, PRB) pairs between profiled sample points, we apply ceiling interpolation in both dimensions so every query remains conservative by construction.

B.2

Pseudocode

Algorithm 1 instantiates the per-slot control loop described in Fig. 6. It takes the standard PUSCH scheduler input together with persistent SM state, then wraps the OAI proportional-fair UL scheduler with three scheduler-specific mechanisms: PRB budgeting, adaptive SM adjustment, and iteration control. The output is a set of UL PRB grants together with the current-slot SM budget, which is retained as prior state for the next slot; its complement is current-slot spare-SM capacity. The control flow consists of four stages: (a) query the lookup table to obtain a provisional PRB budget, (a′ ) run standard PUSCH scheduling within that budget, (b) adjust the target SM count based on observed backlog and recent history, and (c) recompute the PRB allowance under the updated state, looping whenever the current allocation is either above that allowance or more than a small tolerance below it.

B.3

C

Two-Level Elastic Training Details

C.1

Scheduling Mechanism Details

This subsection provides implementation details for the two levels defined in §6: inter-site scheduling and intra-site scheduling. Within the inter-site level, AIR is the fast path for transient fluctuations and MLR is the slow path for persistent imbalance. Coordinator metadata. The coordinator maintains three types of runtime state: (i) executor performance state, recent throughput and latency observations used to predict completion times for micro-batch stage tasks and decide when to trigger load rebalancing; (ii) sample-stage commit state, microbatch and per-sample identifiers, attempt tokens, and commit records that provide at-most-once acceptance even when work is rerouted or speculatively re-executed; and (iii) deferral statistics, the fraction of candidate micro-batches deferred in each training step, which distinguishes transient overload from persistent imbalance and drives the MLR slow-path trigger. Coordinator deployment. In RAN deployments, the coordinator runs at the core to exploit strong site-to-core links; in peer-to-peer clusters, the same logic can be hosted on a worker with replicated metadata storage. Rebalancing protocol. When MLR fires, the slow path shifts the pipeline-parallel partition boundary so that persistently

Parameters

Table 9 lists the tunable parameters used by the scheduler. These parameters control four things: the SM quantization granularity, how aggressively the scheduler ramps up in response to backlog, how cautiously it ramps down after demand subsides, and how tightly iteration control tracks the recomputed PRB allowance. We choose their values through end-toend profiling on the DGX Spark testbed across different traffic scenarios. The tuning process sweeps each parameter while measuring SM reservation stability (transition count), HARQ deadline-met rate, and UE-level KPIs (throughput, delay, and loss), then selects values that improve the reservation stability without degrading RAN performance. SM_MIN and STEP. define the scheduler’s SM control granularity. The SM count is quantized in steps of 8, matching the CUDA Green Context co-scheduling granularity observed on 18

Deferral fraction

0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0

per-step deferral frac.

transient spike (AIR absorbs)

Moving Average sustained overload (Slow Path needed)

at which point MLR adjusts the model-parallel layout. MLR only reacts when overload is sustained, avoiding expensive rebalancing (which requires state migration) for short-lived spikes that AIR can handle on its own.

AIR sufficient 0

20

40

60

80

100

Training step

120

140

C.2

160

Correctness Guarantees

Correctness for W EAVER is defined relative to a reference run that trains the same model with standard synchronous dataparallel stochastic gradient descent (SGD) on a fixed cluster. Dynamic rerouting, deferral, and model-layout rebalancing should not change which samples are counted or how their gradients are combined; they should only change where and when the work runs. At-most-once, version-valid sample-stage acceptance. Each micro-batch stage task carries a metadata tuple that includes its training step t, micro-batch identifier b, sampleidentifier set Ub , stage s, and a fenced attempt token. The coordinator maintains one commit record per sample-stage key (t, u, s) for u ∈ Ub . AIR transfers each micro-batch stage task as a bundle. A normal AIR transfer revokes the old attempt before authorizing a new one; the speculative retry in Fig. 9 may temporarily execute both copies, but an atomic open-to-committed transition accepts only the first valid result for each key. Tokens also carry a task epoch, model version, and layout epoch, so late outputs from revoked attempts or old layouts are rejected rather than applied twice. Thus each sample-stage result is accepted at most once per training step and no stale gradient is applied. Per-sample gradient correctness with varying participation. Because the AIR fast path may carry infeasible microbatch stage tasks into the next training step, replicas can complete different sample counts ng,t in a training step and the effective global batch size Bt = ∑g ng,t can vary over time. W EAVER aggregates gradients using the weighted rule in Eq. (3), which uses the actual sample counts ng,t from each replica instead of assuming fixed sample quotas. This produces the exact per-sample average over the set of samples accepted in training step t, matching what a reference synchronous run would compute on that same sample set. The debt update in Eq. (5) then adjusts future sample quotas when a replica under-contributes for many training steps, which keeps long-run participation balanced without changing the per-training-step estimator. Appendix C.4 gives the convergence argument under this varying effective batch size. Bounded latency and eventual progress from fast- and slow-path inter-site adaptation. AIR derives a per-trainingstep admission deadline from the SM reservation published by RAN Control and admits only micro-batch stage tasks whose predicted finish time does not exceed this deadline. Microbatch stage tasks that would miss the deadline are rerouted when another equivalent-stage executor is feasible and otherwise remain unowned until admission in the next training step. When persistent capacity loss keeps the deferral rate

Figure 13: MLR trigger based on AIR deferrals. ⋆ marks the rebalancing trigger. overloaded stages shed layers to less-loaded sites, rather than migrating the full runtime data path. Only the persistent state for layers that cross the boundary (parameters and optimizer state) is transferred; transient activations are regenerated under the new layout. The P REPARE phase drains work that crosses the changing boundary at a training-step barrier. The subsequent C OPY and C OMMIT phases transfer the moved layers’ state and activate the new layout while work on unaffected stages continues, so the seconds-long transfer does not become a seconds-long exposed stall (Table 13). The complete P REPARE /C OPY /C OMMIT protocol ensures that layout transitions are consistent and do not corrupt gradients. GEMM-tile admission and profiling. To avoid runtime jitter, W EAVER pre-allocates GPU GEMM-tile buffers and stream state during initialization, eliminating online allocator overhead (100–300 µs under load) from the GEMM-tile admission path. Before deployment, W EAVER profiles GEMM performance across tile sizes and spare-SM budgets to build a lookup table that maps tile size × spare-SM budget to predicted completion time, so runtime GEMM-tile admission can enforce the RAN-slot deadline and bandwidth constraints without per-decision profiling. Fault tolerance. W EAVER handles crash-stop site failures through the same mechanisms used for capacity fluctuations. When a site becomes unreachable, the coordinator detects the failure via lease timeout, increments the task epoch for each affected micro-batch stage task, revokes the failed attempts, and reassigns pending micro-batch stage tasks to surviving replicas through AIR’s min-cost assignment (Appendix C.3). Fenced attempts guarantee at-most-once acceptance: any delayed output from the failed site carries a stale task epoch and is rejected at commit (Appendix C.4). For persistent site loss, MLR rebalances the model layout to exclude the unavailable site, transferring the affected partition’s parameters and optimizer state. We do not address Byzantine faults, correlated failures across a majority of sites, or coordinator crash recovery; the coordinator is assumed to be a reliable service, deployable with standard replication techniques if higher availability is required. MLR example. Fig. 13 illustrates how MLR distinguishes transient overload from persistent imbalance. Initially, deferrals remain low and the moving average stays well below the threshold, so the AIR fast path alone is sufficient. Later, deferrals remain high over many consecutive training steps; the moving average climbs above the threshold and stays there, 19

where ntarget is the target per-replica sample count, qbase is g the baseline sample quota from placement, and κ is the debtcorrection gain.

high, MLR adjusts pipeline boundaries and sample quotas until the new layout again satisfies the timing constraint. Thus AIR handles transient inter-site variation, MLR handles persistent inter-site imbalance, and the intra-site scheduler independently absorbs RAN-slot-level changes through GEMM-tile admission. Summary of guarantees. AIR rerouting and deferral change where and when micro-batch stage tasks execute, but not how accepted gradients are formed or which model and layout version they update: weighted aggregation is unbiased when admission, deferral, and rerouting do not depend on sample content; the per-key atomic commit gate prevents duplicate contributions during rerouting or speculative re-execution; the debt update keeps long-run contribution skew bounded for 0 < κ < 1; and during MLR, the barrier-aligned P REPARE /C OPY /C OMMIT protocol binds every accepted completion to one layout epoch. Appendix C.4 gives the corresponding derivations and protocol arguments.

C.3

C.4

This subsection provides the derivations and protocol arguments behind W EAVER’s aggregation, commit, debt, and layout-transition semantics. Unbiased weighted aggregation. For training step t, let St be the set of accepted samples after rerouting and deduplication. For each sample i ∈ St , let gi (wt ) denote its per-sample gradient at synchronized weight wt . Replica g contributes Gg,t = ∑ gi (wt ), F

This subsection defines the latency, aggregation, and debt notation that formalizes the scheduling behavior described in §6. AIR enforces a per-training-step admission deadline defined by the nominal barrier deadline Tbar with bounded elastic slack ∆t ≤ ∆max . Admitting only micro-batch stage tasks that finish by this deadline yields the training-step-latency bound

∑G g=1 ng,t

.

∑g Gg,t 1 = ∑ gi (wt ). |St | i∈ ∑g ng,t St

(7)

Hence it is exactly the sample average over all accepted samples. Under standard random sampling of the global batch and the assumption that admission, deferral, and rerouting decisions are independent of sample content conditioned on system state, E[ḡt | wt ] equals the gradient of the acceptedsample objective. The weighted aggregation rule is therefore unbiased under this content-independent scheduling assumption. At-most-once acceptance under fenced attempts. Each sample-stage key is χ = (t, u, s) with a commit record cχ ∈ {OPEN, COMMITTED} and one or more authorized attempt tokens τa tied to the current task epoch, model version, and layout epoch. A normal ownership transfer increments the task epoch and revokes the old token; speculative re-execution may authorize an additional token in the same task epoch. Every result must first pass token and version validation, then atomically attempt

(2)

where Lt is observed training-step latency, ∆max is the configured maximum elastic slack, and ∆c is the duration of one inter-site AIR control interval. RAN Control exports the SM reservation; the training scheduler derives these timing quantities from that reservation and recent completion-time errors. After rerouting/deferral, replica g may complete ng,t samples at training step t, with summed local gradient contribution Gg,t . The tree-reduce through the hub computes ∑G g=1 Gg,t

(6)

with disjoint partition St = g Sg,t due to at-most-once acceptance. The estimator in Eq. (3) can be rewritten as ḡt =

Lt ≤ Tbar + ∆max + ∆c ,

ng,t = |Sg,t |,

i∈Sg,t

Scheduling Formalism

ḡt =

AIR Correctness Arguments

CAS

cχ : OPEN −−→ COMMITTED.

(3)

(8)

Safety (at-most-once sample-stage acceptance): Suppose two results for the same key χ are accepted. Both must successfully change cχ from OPEN to COMMITTED, but the atomic compare-and-swap (CAS) permits only one such transition. This remains true when two speculative copies have valid tokens; the slower copy observes an already committed record and is discarded. Results from earlier task or layout epochs fail validation before reaching the commit gate. Liveness (progress with failures): If an executor fails, lease timeout triggers reassignment with a higher task epoch. Any delayed old-attempt output is stale and rejected. Therefore, if a healthy executor eventually obtains a valid attempt token

Here ḡt is the global per-sample average over all uniquely accepted samples in training step t, Bt = ∑g ng,t is the realized effective global batch size, and G is the number of participating replicas in that step. To correct persistent contribution skew, W EAVER maintains per-replica debt state dg and updates next-step sample quota qg using  dg (t+1) = dg (t) + ntarget − ng,t , (4)  base qg (t+1) = clip qg + κdg (t+1) . (5)

20

and completes, one result can commit while all later results are rejected. Bounded contribution debt. Define debt dynamics (Eq. (5)) and let realized count be ng,t = ntarget + κdg (t) + ξg,t ,

model-version, and layout compatibility; (ii) the per-key opento-committed CAS allows at most one accepted result even under speculative re-execution; (iii) PREPARE/COPY/COMMIT performs an atomic cutover at training-step barrier t ⋆ , so each accepted completion is bound to one layout epoch r; and (iv) outputs with stale (e, v, τa ) or stale r are rejected, while reclamation is deferred until drain completion.

(9)

where ξg,t captures bounded actuation/scheduling noise, |ξg,t | ≤ Ξ. Then dg (t+1) = (1 − κ)dg (t) − ξg,t .

D

(10)

D.1 RAN-Centric Spare Compute Controller Evaluation

For 0 < κ < 1, this is a stable linear system with bounded disturbance. Recursing gives |dg (t)| ≤ (1 − κ)t |dg (0)| +

Ξ . κ

Evaluation Details

D.1.1

(11)

RAN configuration

The gNB runs 5G NR standalone (SA) on band n78 (3.5 GHz TDD) with 106 PRBs at 30 kHz subcarrier spacing (numerology 1), corresponding to a 40 MHz carrier. The TDD pattern uses a 5 ms periodicity with 7 DL slots, 1 flexible slot, and 2 UL slots (DDDDDDDSUU), yielding a ∼20 % UL duty cycle. All experiments run on an NVIDIA DGX Spark with 48 SMs.

Thus long-run skew is bounded by O(Ξ/κ) and decays geometrically when noise is small. Layout-transition consistency. MLR cutover begins at a training-step barrier t ⋆ , where the pipeline is quiescent at the changing boundary. In the PREPARE phase, W EAVER freezes AIR assignments whose path would cross the moved layers, drains in-flight work under the old layout, and stops admitting new work at the affected stages. In the COPY phase, it transfers parameters and optimizer state Sℓ for the moved layers and verifies checksums without changing attempt metadata; work that does not traverse the changing boundary can continue. In the COMMIT phase, it atomically publishes the new boundary vector, rebuilds equivalent-stage pools {Es }, and increments the monotonic layout epoch r in the coordinator ledger. After COMMIT, any new micro-batch stage task that enters the affected stages carries the updated layout epoch in its metadata ⟨t, b,Ub , s, e, v, τa , r⟩, where e is the task epoch, v is the model version, and r is the layout epoch, and the coordinator accepts a completion for key χ = (t, u, s) for each u ∈ Ub only if its task epoch e, model version v, attempt token τa , and layout epoch r are current. Hence any accepted completion for χ is executed entirely under a single pipeline layout. Extending the fenced-attempt validation to include r preserves at-most-once acceptance across boundary moves. Suppose two distinct results for the same key χ were both accepted around an MLR event. If they carry the same current task, model, and layout versions, they still compete at the single open-to-committed CAS, which accepts only one. If they carry different layout epochs r < r′ , validation accepts only the current layout epoch, so the stale result is rejected before the commit gate. Thus at most one result for each χ is accepted, and no commit contains mixed or stale layout information. Combined with the liveness condition above, this yields eventual exactly-once sample contribution. Protocol invariant summary. Taken together, these arguments establish one end-to-end acceptance invariant per sample-stage key χ = (t, u, s): (i) metadata-gated validation on ⟨t, b,Ub , s, e, v, τa , r⟩ and u ∈ Ub enforces task-attempt,

D.1.2

Channel condition replay

Static AWGN channels represent the primary fidelity gap between RF-simulated and OTA experiments. We close this gap by replaying per-UE uplink channel measurements from the TRACTOR dataset [19] into OAI RFSIM’s channel model. At each trace timestamp, the recorded UL RSSI and SINR are translated into the simulator’s path-loss and noise-power controls: ploss_dBu,k = Pref − RSSIu,k ,

(12)

noise_power_dBu,k = Noff − SINRu,k ,

(13)

where Pref and Noff are fixed calibration offsets. Updates are sent via OAI’s telnet interface and synchronized with traffic generators using Unix wall-clock timestamps, so each UE observes the same time-varying channel trajectory as the original trace. D.1.3

Application traffic replay

TRACTOR application trace replay. We use production O-RAN traffic from the TRACTOR 5G dataset [19], selecting two four-UE scenarios: Scenario 1 (referred to as MultiUE/Trial2/multi4 in the dataset) and Scenario 2 (referred to as Multi-UE/Trial3/multi4 in the dataset). Per-UE traffic profiles and channel conditions are listed in Table 10. For the RANcontrol evaluation, traces are replayed with the original packet timing over a 600 s window; shorter UE traces are looped to ensure all four UEs remain active.

21

Table 12: Inter-site latency matrix (src_rank → dst_rank), in ms. Diagonal entries are not applicable (—). src/dst

rank0

rank1

rank2

rank0 rank1 rank2 rank3 rank4 rank5 rank6 rank7

— 0.8 25.7 26.4 26.6 25.8 23.7 23.5

0.8 — 25.8 26.5 26.6 25.9 23.6 23.5

25.9 25.7 — 28.6 28.8 0.9 25.7 25.6

different pending process. The scheduler exports cumulative overdue events and initial UL transport-block transmissions. We use differences of these counters over the measurement window. The normalized HARQ deadline-met metric is 1 − (overdue events/initial UL TBs); for live co-location we report the underlying event rate per 1000 initial UL TBs. This MAC-level metric is separate from the fraction of LDPC decoder records exceeding a 1 ms processing budget.

rank3 rank4 rank5 rank6 rank7 26.7 26.5 28.8 — 0.9 28.8 26.5 26.6

26.7 26.7 28.8 30.5 — 28.8 26.6 26.7

25.7 26.5 0.8 28.7 28.6 — 25.8 25.7

23.7 23.6 25.7 26.4 26.6 25.6 — 0.2

23.5 23.7 25.6 26.6 26.5 25.7 0.2 —

D.2

Throughput (tokens/s)

Temporal Stability - opt-1.3b 40000 Weaver (CV 0.6%, Drop 0.0%) DTFM (CV 2.8%, Drop 0.8%) Asteroid (CV 13.8%, Drop 0.8%)

30000 20000

Training Evaluation Settings

Here, rank i denotes the worker at site i. Table 11 provides the inter-site bandwidth and Table 12 provides the inter-site latency.

Confidant (CV 7.3%, Drop 0.8%) 80% Mean

10000 0

D.3 0

20

40

60 80 Training Step

100

Convergence semantics. Our training-side experiments focus on throughput and efficiency, but all systems run under strict synchronous semantics: AIR accepts each completion at most once, gradients are aggregated with the unbiased weighted rule over accepted samples, and no stale gradients are applied (K=0), so convergence on the completed-sample objective matches a reference synchronous run (Appendices C.2–C.4). Fig. 14 shows step-level token throughput for all systems over the same replayed RAN SM reservation. Fig. 15 supports the loss-parity argument of Appendix C.4 with an OPT-1.3B training-loss trace under replayed RAN SM reservations: under at-most-once acceptance and weighted aggregation, the loss trajectory tracks a fixed-cluster synchronous reference.

Figure 14: Step-level training throughput under identical replayed RAN SM reservations. Table 10: TRACTOR dataset configurations used in RAN control evaluation. Datasets

UE

Traffic profile

Type

Scenario 1 UE 1 UE 2 UE 3 UE 4

Background traffic YouTube streaming (stationary) Background traffic YouTube streaming (stationary)

mMTC eMBB mMTC eMBB

13.1 dB 2.3 dB 3.3 dB 6.0 dB

Scenario 2 UE 1 UE 2 UE 3 UE 4

Netflix streaming (stationary) YouTube streaming (campus walk) Background traffic Internet browsing (stationary)

eMBB eMBB mMTC eMBB

20.7 dB 5.5 dB 4.2 dB 12.2 dB

Avg. UL SINR

Table 11: Inter-site bandwidth matrix (src_rank → dst_rank), in Gbps. Diagonal entries are not applicable (—).

D.4

src/dst rank0 rank1 rank2 rank3 rank4 rank5 rank6 rank7 rank0 rank1 rank2 rank3 rank4 rank5 rank6 rank7

— 20.5 1.1 1.1 1.1 1.1 1.0 1.0

11.0 — 1.0 1.0 1.0 1.0 0.8 0.9

0.3 0.3 — 0.3 0.3 10.2 0.3 0.3

1.0 1.0 0.9 — 11.9 0.9 0.8 0.8

0.9 1.0 0.9 14.6 — 0.9 0.8 0.9

0.9 0.9 10.6 0.9 1.0 — 0.8 0.8

1.0 1.0 1.0 1.0 1.0 1.0 — 17.9

Inter-Site Scheduling

We now evaluate how the two-level inter-site scheduler in §6.1 behaves under SM reservations derived from replayed RAN traces. We feed a representative 6-minute production traffic trace through our training simulator on OPT-13B across 8 edge sites with NVIDIA L40S GPUs, so any difference comes only from the policies described in §6.1. For the training side, we compare W EAVER against the same DTFM, Asteroid, and Confidant baselines described in §7.2, which all react at slow timescales while sharing the common oracle-MPS substrate. Fast-path AIR absorbs most capacity fluctuations. Fig. 16 summarizes how Asynchronous Inter-site Rerouting behaves when driven by SM reservations derived from real RAN traces. Over the 6-minute window, spare compute fluctuates 80 times at the training-step timescale (2–8 s). Of these events, 86.2% are handled entirely by fast-path AIR rerouting with a median exposed stall of 66 ms (P95: 68 ms), and only 11 events escalate to the slow path. This matches the intent of the AIR design in §6.1: the coordinator moves work, not state, so most transient drops in spare compute are hidden behind normal training progress.

1.0 1.0 1.0 1.0 1.0 1.0 14.5 —

D-ITG stress traffic. For UE-level KPI evaluation, we use DITG [7] to generate synthetic bursty UDP traffic that exercises more extreme conditions than production traces. Each UE runs a steady 800 kbps flow plus a 20 s burst at 3.6 Mbps, so the aggregate peak load (∼13.2 Mbps across three UEs) saturates the uplink. D.1.4

Training Stability and Convergence

120

HARQ deadline-met rate

We instrument the OAI MAC scheduler to record the expected frame and slot for UL HARQ feedback. A waiting HARQ process is counted as overdue when its feedback remains missing beyond the configured grace window; the same check applies when feedback arrives for a 22

Table 13: MLR rebalancing overhead decomposition per event, averaged over 11 triggers in the 6-minute trace. Phase

Total (ms)

Stall (ms)

PREPARE (drain) COPY (transfer) COMMIT (activate)

528.0 4350.6 550.0

528.0 0.0 0.0

Total

5428.6

528.0

than to short spikes. To make the cost of each MLR rebalancing event concrete, Table 13 decomposes the MLR protocol into its phases. These numbers come from the same trace replay used above and guide the choice of rebalancing thresholds. The PREPARE phase is the only component on the critical path, and its 528 ms stall cost is amortized over many training steps between MLR rebalancing events. COPY and COMMIT mostly overlap with useful work on neighboring stages. Taken together with the 11-event count above, this supports the design choice in §6.1 to keep layout changes rare and hide most of their cost behind ongoing execution. Inter-site scheduling parameter sensitivity. Finally, we study how the thresholds that drive AIR and MLR affect throughput and stability under the same replayed SM reservations used above. We sweep the AIR slack threshold θAIR and the MLR trigger threshold θMLR while keeping all other settings fixed (Table 14). Across these settings, throughput varies only slightly while the CoV remains well below 1%, which indicates stable training-step latency. More aggressive MLR triggering increases the number of rebalancing events but does not improve throughput in a meaningful way. This backs up the design in §6.1: a conservative trigger that lets AIR handle most transients and invokes MLR rarely is enough to keep the system both stable and efficient.

Table 14: Threshold-sensitivity analysis for inter-site scheduling under trace replay. Higher throughput and lower coefficient of variation (CoV) are better. Threshold

Throughput

CoV (%) # MLR events

θAIR =0.05 21765.72 tokens/s θAIR =0.1 21837.73 tokens/s θAIR =0.2 21845.08 tokens/s

0.67 0.08 0.01

0 0 0

θMLR =0.3 θMLR =0.4 θMLR =0.5

0.45 0.39 0.13

14 10 1

21813.14 tokens/s 21822.01 tokens/s 21841.99 tokens/s

Training loss

12.5

Fixed-batch synchronized SGD Weaver (varying effective batch size)

12.0 11.5 11.0 0

200

Step

400

600

D.5

Large-Scale Simulator Methodology

Our simulator extends Morphling [77], a measurement-driven emulator for distributed ML that runs unmodified training scripts over fitted device performance models, with inter-site topology modeling and per-site RAN SM-reservation replay. Compute profiles. Because all transformer layers are architecturally identical, we profile per-layer forward and backward latency separately for the embedding, a single transformer block, and the LM head on every testbed node, sweeping micro-batch size and sequence length (5 warm-up and 20 measured iterations per configuration. “Running training with a reduced number of layers” refers to exactly this: we execute truncated-depth models to obtain measured perlayer profiles rather than extrapolate from analytical FLOP counts, then reconstruct the full-model step time as embedding + Nlayers ×block + head under the chosen parallelism layout and pipeline schedule, keeping the target global batch size and sequence length unchanged. Communication. We transmit the actual hidden-state and gradient tensors required by the training framework over the real links and use the measured inter-site bandwidth/latency matrices (Tables 11–12); the simulator additionally models per-link contention when multiple transfers share a path. Memory. The simulator checks per-site feasibility against parameter, optimizer, and activation footprints—the source of the OOM entries in Table 8.

Figure 15: OPT-1.3B training-loss trace under replayed RAN SM reservations vs. a fixed-cluster synchronous reference. The runtime view in Fig. 16 also shows that AIR decisions occur within the current training step rather than waiting for a global barrier. The coordinator reassigns pending micro-batch stage tasks to equivalent-stage executors in other data-parallel replicas, with no GPU reconfiguration, no MPS reset, and no checkpoint replay. As a result, the RAN-slot deadline is not endangered even when site capacity swings quickly, which is the safety property targeted by the inter-site scheduler in §6.1. Slow-path model layout rebalancing is rare and lowoverhead. When the SM reservation drifts in a persistent way, fast-path AIR starts to defer a growing fraction of microbatches. Each deferral is still within the bound enforced by the inter-site policy, but the accumulated signal in the exponentially weighted moving average (EWMA) trigger state zt eventually crosses the MLR trigger threshold and activates slow-path MLR. In Fig. 16, the middle panel tracks zt and the MLR trigger threshold over time. Across the 6-minute trace, the MLR trigger fires 11 times, which confirms that the deferral-based trigger in §6.1 reacts only when imbalance is sustained rather

23

Spare capacity (% of peak SMs)

Site 0

Site 1

100

Site 2

Site 3

Weaver Stability

Site 4

AIR Reroute

Site 5

Site 6

Site 7

MLR Reshard

80 60 40 20 0 0

50

100

150 200 Wall-clock time (s)

250

300

350

Figure 16: Timeline of inter-site scheduling over a 6-minute replay window under per-site SM reservations. RAN dynamics. Each site replays a TRACTOR-derived SMreservation trace drawn independently at random from the trace pool, so no two sites are artificially synchronized and the 128–512-site configurations exercise heterogeneous, uncorrelated capacity dynamics.

E

GPU memory. Memory overhead and tiering. Beyond the pre-created context streams accounted in §7.3, W EAVER manages GPU HBM and host DRAM as a unified dual-tier pool: the RAN working set occupies a pinned, guaranteed HBM partition; training receives an elastic HBM allocation backed by staging buffers for transparent migration to host memory. Unlike host-offloading schemes that optimize training throughput in isolation (e.g., ZeRO-Offload [56]), W EAVER applies a RAN-priority eviction policy: when PHY memory demand spikes, it demotes training tensors to host memory immediately, preserving RAN latency guarantees at the cost of temporarily reduced training bandwidth. CPU offloading of non-GEMM operators. W EAVER offloads non-GEMM training operations—layer normalization, element-wise activation functions, and GEMV—to idle host CPU cores. Because RAN PHY processing is GPU-bound, the host CPU carries mostly lightweight control-plane signaling and has substantial idle capacity. An adaptive routing policy tracks per-pool queue depth and a sliding-window average of completion times, and assigns each operation to the less-loaded resource at submission time. This converts idle CPU cycles into additional GPU headroom for the matrixintensive passes that dominate training compute; the resulting host CPU utilization is accounted in §7.3.

Implementation Details

SM-level partitioning via Green Contexts. W EAVER partitions the GPU’s SM pool using the CUDA Green Context API, which provides driver-level isolation of SM subsets. To keep SM reservation changes off the RAN-slot critical path, each worker pre-creates Green Context streams for all valid SM partitions during initialization. W EAVER maintains up to 16 such streams, each corresponding to a configurable fraction of the GPU’s SMs and consuming approximately 20 MB of GPU memory, for a total overhead of up to 320 MB. The associated tile buffers, stream state, and (tile size, spare-SM budget) → completion time profiles are also prepared in advance. At each slot boundary, the runtime only rebinds the next GEMM tile to the appropriate pre-created context, thereby enforcing the slot-level SM cap while reserving the remaining SMs for training. This design avoids on-demand resource allocation and per-decision profiling from the admission path, reducing context-switching overhead to below 10 µs (see Table 1). Asynchronous dispatch and persistent kernels. Each Green Context is paired with a persistent kernel whose lifetime matches the context, backed by a task queue in GPU global memory. The host enqueues task descriptors (GEMM-tile coordinates, data pointers, output addresses) via mapped zerocopy memory, bypassing the kernel-launch path; a dedicated warp polls the queue head with atomic reads while the remaining warps execute the current tile. W EAVER intercepts memory-copy completion events on the RAN data path and signals training dispatch via CUDA IPC events, so training GEMMs issue only after RAN decoding data is resident in

F

Discussion on Practical Concerns

General applicability. While we demonstrate W EAVER with LDPC decoding – the dominant compute bottleneck for RAN processing – offloaded to the GPU, the key techniques remain applicable when the entire L1 stack is offloaded (that is supported by NVIDIA Aerial [39]); we will explore this once the corresponding hardware is available to us. Beyond FM training, we will also intend to extend W EAVER’s colocation framework to other non-RAN workloads such as ML

24

inference and RAN digital twins. Coordinator capacity and scaling. Assuming a widely available datacenter configuration of 200 Gbps networking and 128 CPU cores for the CPU-only coordinator (e.g., AWS M6in instances), and typical per-site backhaul on the order of hundreds of Mbps [46], a single coordinator sustains roughly 1,000–2,000 sites. At larger scales, W EAVER shards data and replicates model parameters across multiple coordinators— distributing both bandwidth and computation—following fault-tolerant coordination systems such as Beldi [81]. Robustness to result corruption. For third-party or multitenant deployments, W EAVER can verify distributed GEMM results with random-projection checks [35]: for C = AB, sampling random vectors r, s and testing r⊤ (AB)s = (Ar)⊤ (Bs) detects even single-entry corruption with probability 1 − O (2−n ) at O (n) cost per check; because the check reduces to GEMVs, it runs in real time on host CPUs [76]. Energy considerations. Harvesting already-provisioned, already-powered RAN GPUs avoids the embodied and provisioning cost of dedicated training clusters. Decentralized training on such resources can also improve net energy efficiency relative to sub-linearly scaling centralized clus-

ters [32, 55, 65, 86]. At our traffic volumes, communicationside energy is negligible relative to compute energy. Sustained training nonetheless remains bounded by site-specific thermal limits under operator-provisioned power and cooling budgets, and warrants a dedicated energy study [29]. Limitations. (i) Convergence scope: gradient semantics are equivalent to synchronous SGD (Appendix C.4) and a supporting convergence trace is provided (Appendix D.3), but multi-week convergence campaigns on production-scale deployments remain future work. (ii) Single-operator assumption: W EAVER assumes one operator controls both the RAN and the training workload within an operator-managed execution environment, where users provide models and data while the operator controls the training software [47]; stronger isolation for third-party-controlled multi-tenant workloads is future work, and in multi-operator or neutral-host settings W EAVER reduces to passive sharing via Green Context partitioning alone. (iii) Vendor specificity: SM-level partitioning relies on NVIDIA’s Green Context API; porting to accelerators without an equivalent sub-GPU partitioning interface would fall back on weaker software-level isolation.

25

Algorithm 1 Compute-Aware UL Scheduler: per-slot pipeline Require: Scheduler input for standard PUSCH scheduling Require: Persistent SM history state (carried across slots): sm_prev: prev-slot SM count; mcs_prev: prev-slot weighted-avg MCS; slots_since_under: underutilization counter; slots_since_update: slots since last SM change Ensure: PUSCH grants within current-slot SM compute budget; SM state carried to next slot 1: // Entry-point initialisation 2: tgt_sm ← sm_prev; mcs_est ← mcs_prev; iter ← 0; high_demand_seen ← false 3: // Iterative scheduling pipeline (Fig. 6, steps a–c) 4: P IPELINE E NTRY: iter ← iter + 1 5: // Step (a): PRB Budgeting 6: prbs_avail ← LUT_PRB(mcs_est, tgt_sm) ▷ max PRBs decodable within RAN-slot deadline 7: // Step (a′ ): Standard PUSCH Scheduling within prbs_avail 8: Run standard PF scheduling capped by prbs_avail 9: Record prbs_alloc ← ∑u rbSizeu , sum_backlog ← ∑u Bu 10: mcs_curr ← ∑u mcsu · rbSizeu / prbs_alloc 11: // Step (b): Adaptive SM Count Adjustment 12: // Ramp-Up: immediate increase to drain backlog 13: if ∃ u with Bu > BACKLOG_THR and the PRB budget is exhausted or residual backlog  exceeds threshold then 14: est_prbs ← max prbs_alloc + min_rb, est_prbs_from_backlog(sum_backlog) 15: tgt_sm ← LUT_SM(mcs_est, est_prbs); high_demand_seen ← true 16: end if 17: // Fast Exit: immediate down-sizing when slot is idle 18: if not high_demand_seen and prbs_alloc = 0 and tgt_sm > HIGH_TIER_CAP then 19: exit_sm ← LUT_SM(mcs_est, est_prbs_from_backlog(sum_backlog)) 20: if exit_sm ≤ HIGH_TIER_CAP then  21: tgt_sm ← max tgt_sm − STEP, max(exit_sm, SM_MIN) ▷ step down one tier 22: end if 23: end if 24: // Hold or Ramp-Down: hysteresis-based conservative step-down 25: if tgt_sm ̸= sm_prev then 26: slots_since_under ← 0; slots_since_update ← 0 27: else 28: slots_since_update += 1 29: if sm_prev > SM_MIN and LUT_PRB(mcs_curr, sm_prev − STEP) ≥ prbs_alloc then 30: slots_since_under += 1 ▷ one fewer tier would have sufficed 31: else 32: slots_since_under ← 0 33: end if 34: end if 35: if slots_since_under ≥ DECR_THR then  36: desired_sm ← max LUT_SM(mcs_est, est_prbs_from_backlog(sum_backlog)), SM_MIN  37: tgt_sm ← max sm_prev − STEP, desired_sm ; slots_since_under ← 0 38: end if 39: // Step (c): Iteration Control  40: Pmax ← LUT_PRB mcs_curr, tgt_sm ▷ recompute with true MCS and updated SM 41: if prbs_alloc ≤ Pmax and Pmax − prbs_alloc ≤ PRB_TOL then 42: S UCCESS: emit grants 43: else if iter < MAX_ITER then 44: mcs_est ← mcs_curr; goto P IPELINE E NTRY 45: else if prbs_alloc > Pmax then 46: S AFEGUARD: force tgt_sm ← LUT_SM(mcs_curr, prbs_alloc) ▷ ramp-up to guarantee forward progress 47: else 48: C ONSERVATIVE E XIT: emit grants ▷ safe but under-allocated after hitting the iteration bound 49: end if 50: // Carry state to next slot 51: sm_prev ← tgt_sm ▷ retain current-slot SM count as next-slot prior 52: if mcs_curr > 0 then 53: mcs_prev ← mcs_curr 54: end if 55: Output: final UL PRB grants and tgt_sm for this slot

26

Record · ID 1108650 · SHA-256 98f2eb434d74e4e9
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.