Accurate Simulation of Distributed Training Jobs with Network Contention Modeling Yeonho Yoo∗ , Hyunho Lee† , Hyunmok Choi† , Chuck Yoo† , Gyeongsik Yang† ∗ Department of Computer Science and Artificial Intelligence, Dongguk University
arXiv:2609.23278v1 [cs.DC] 20 Sep 2026
† Department of Computer Science and Engineering, Korea University
Abstract—Trace-driven simulation is widely used to evaluate distributed training (DT) jobs in GPU clusters, but existing simulators either ignore network contention or approximate it with a fixed penalty. This misses how scheduling decisions determine which jobs share server network interfaces and inter-server links, thereby changing networking time during training. As a result, our motivating experiments demonstrate that they incur large errors, reaching up to 73.64% mean absolute percentage error (MAPE) in average job completion time (JCT). This paper introduces M O S IM, a GPU-cluster simulator that models DT job execution under dynamic network contention. M O S IM combines GPU-free characterization with network contention model: it obtains each job’s compute time, networking time, and networking volume without GPUs, then uses the current worker assignment to estimate how shared network interfaces affect each job’s iteration time. Our evaluation shows that, compared with existing simulators, M O S IM reduces simulation error for average JCT by up to 3.28×, tail (99th-percentile) JCT by up to 7.79×, and makespan by up to 8.48×, while modeling NIC contention factors with only 8.63% error on average. By avoiding real-GPU profiling, M O S IM also reduces input construction overhead by 44.6×. Index Terms—Distributed training simulation, Network contention modeling, GPU cluster scheduling, Simulation fidelity
I. I NTRODUCTION Simulation has become an essential tool for analyzing the execution of distributed training (DT) jobs and the efficiency of GPU-cluster infrastructure, since large-scale what-if studies on real hardware are costly and disruptive. A useful simulator must capture how the training time of active jobs changes under shared cluster resources. This is challenging because DT jobs repeatedly synchronize gradients over the network, and concurrent jobs compete with one another when their traffic shares the same server NICs or inter-server links. As job scheduling determines which jobs run concurrently and which network resources they share [1], accurately modeling job contention on the network is key to understanding job execution times, cluster efficiency, and resource management decisions in GPU clusters. In a DT job, gradients are synchronized over the network every iteration, and since gradient size scales with model parameter size, networking alone can occupy up to 90% of training time [2]. When concurrent jobs are co-located on the same server or routed through the same inter-server link, their synchronization traffic contends for shared bandwidth, and this contention (not additional computation) is what slows each job down (§II-B). Crucially, this contention is the common case in production, not an artificial corner case. Over 86% of jobs in the Microsoft Philly cluster [3] and 93% of jobs in AcmeTrace
[4] request 8 or fewer GPUs. These small and medium DT jobs are exactly where a scheduler has many choices—keeping a job within a server, spreading it across servers, or running multiple jobs so that their inter-server traffic overlaps—so whether their traffic shares a NIC or an inter-server link is decided at scheduling time, not fixed in advance. So, accurately simulating DT job performance under such contention is challenging but critical. Yet, in our motivating experiments, existing GPU-cluster simulators incur large errors—73.64% mean absolute percentage error (MAPE) in job completion time (JCT). This error arises because existing simulators handle contention in two limited ways. No-contention simulators, such as Tiresias [5] and Gavel [6], use isolated job execution times and do not model contention between concurrent jobs. Static-contention simulators, such as Pollux [7] and Muri [8], approximate contention by applying a fixed factor, e.g., increasing every job’s training time by 10%. Both are scalable but fail to capture the workload- and schedulingdependent nature of network contention. Our experiment shows that, relative to contention-free execution, network contention increases job training time by up to 114%, with an average increase of 33% (§III). Therefore, ignoring it or replacing it with a fixed factor leads to the large errors noted above. This level of error can substantially distort what-if analysis of job execution and cluster efficiency, making the results unreliable for resource-management decisions. To address the inaccuracy, we introduce M O S IM, a GPUcluster simulator that estimates DT job execution time under dynamic network contention. Rather than ignoring contention or applying a fixed factor, M O S IM derives contention directly from each job’s execution characteristics. For each job, M O S IM constructs the quantities it needs (§IV-A)—compute time, networking time, and networking volume—via GPUfree characterization, which obtains them without real-GPU profiling. M O S IM then combines these per-job quantities with scheduling decisions (§IV-B) and applies network contention model that estimates each job’s contention on shared NICs and inter-server links (§IV-C) and calculates the changing training time. As the cluster state evolves, M O S IM continuously updates each job’s networking time, enabling high-fidelity simulation of DT job execution and overall cluster efficiency. We make the following contributions: Quantification of network contention, showing that training time can increase by up to 114% (33% on average) mainly due to networking contention. • Development of M O S IM , a new GPU-cluster simulator •
© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. This paper has been accepted for publication in IEEE MASCOTS 2026.
Switch Contention
Contention GPU 0
Server 0
GPU 1
Job1
Job1 Contention
GPU 4 Job1
Server 1
GPU 5 Job1
Contention
Job 2
Job 2
GPU 2
GPU 3
Job 2 GPU 6
Job 2 GPU 7
representative policies [11]. Bin packing (or MostAllocated in Kubernetes) schedules a job on the fewest servers possible, preferring compact allocation but spanning multiple servers when necessary. Conversely, Load balancing (or LeastAllocated in Kubernetes) spreads a job’s workers across as many servers as possible to balance GPU usage.
C. Existing GPU Cluster Simulators Running controlled scheduling experiments on real GPU clusters is expensive and hard to reproduce, so researchers Fig. 1: Network contention example. often rely on trace-driven simulation to compare scheduling that is network-contention-aware using GPU-free charac- policies [5]–[8]. A cluster-level simulator replays a production trace [3], [4] under different scheduling policies and compares terization and contention-aware network model. the resulting JCT and makespan. Each trace entry specifies • Production trace-driven evaluation showing that M O S IM substantially reduces JCT simulation error compared to a job’s submission time and GPU demand, but such traces typically do not contain the training time of each target existing simulators. workload. The simulator therefore assigns each job a workload II. BACKGROUND from a curated model pool, such as common vision and A. DT Job language models, and fills in its training time by conducting A DT job uses multiple GPU workers to train a model in profiling on jobs, which measure each workload running alone parallel, reducing training time and scaling to larger models or on a fixed number of GPUs. The simulator advances through datasets [9]. In data parallel training [10], each worker holds a discrete events, including job arrivals, GPU allocations, and job model replica and processes a disjoint partition of each mini- completions. At each allocation event, the scheduling policy batch. After the backward pass, workers synchronize gradients decides which queued jobs receive GPUs and how their workers through all-reduce so that they apply identical parameter are scheduled across GPU servers [9]. Existing simulators fall into two categories. No-contention updates. Each iteration consists of 1) compute time TJcomp simulators use each job’s standalone training time without for the forward and backward passes on local GPUs and 2) net adjustment, implicitly assuming that co-running jobs do not networking time TJ for gradient exchange; their sum is the comp iter net affect one another. Tiresias [5] replays fixed job durations iteration time TJ = TJ + TJ . Since a job completes train iter from recorded traces without modeling contention. Gavel [6] WJ iterations, its total training time is TJ = WJ · TJ . profiles each job across heterogeneous GPU types for resource The key metric for evaluating a DT job is JCT, which allocation, but does not model network contention of jobs. includes both waiting time, defined as the time from job In contrast, static-contention simulators approximate netsubmission until GPU allocation, and training time. Makespan is another key metric for evaluating scheduling quality, which work contention by a fixed factor. Pollux [7] and Muri [8], is defined as the time from the first job submission to the for example, increase training time by 10% to account for contention. However, the same ratio is applied regardless of a completion of the last job in a workload. job’s networking volume, other co-running jobs, or the server B. DT Job Scheduling and Contention NICs and inter-server links shared by synchronization traffic. Thus, existing simulators do not capture the workload- and A job scheduler (e.g., Kubernetes [11] and Volcano [12]) allocates GPUs to submitted jobs and decides how their workers scheduling-dependent nature of network contention. are scheduled across GPU servers. This scheduling decision III. M OTIVATION determines which server NICs and inter-server links carry each This section presents our motivating experiments. We first job’s network traffic. When concurrent jobs share the same describe the experiment setup and then analyze the results. NICs or inter-server links, their traffic can contend for shared A. Experiment Setup bandwidth; we refer to these jobs as co-running jobs. Fig. 1 illustrates how scheduling creates such network 1) Baselines: We evaluate Tiresias [5] and Pollux [7] as contention. The green job occupies GPUs 0 and 1 on Server representative no-contention and static-contention simulators, 0, and GPUs 4 and 5 on Server 1, and the blue job occupies respectively. Tiresias uses each job’s standalone training time GPUs 2 and 3 on Server 0, and GPUs 6 and 7 on Server 1. regardless of co-running jobs. Pollux applies a fixed contention Each job generates networking traffic to synchronize gradients ratio 10% to co-running jobs. Other simulators, such as across GPUs. So, intra-server flows (e.g., between GPUs 0–1 Gavel [6] and Muri [8], follow similar designs and show trends and GPUs 2–3) and inter-server flows (e.g., between GPUs 0–4 consistent with Tiresias and Pollux, respectively; we omit their and GPUs 2–6) overlap, thereby causing network contention results due to space constraints. 2) Metrics: We evaluate the fidelity of simulation results by at both the intra- and inter-server levels. Thus, scheduling determines not only which GPUs a job uses, but also which comparing them against ground-truth measurements. We use network resources its all-reduce traffic shares with other jobs. mean absolute percentage error (MAPE) across four metrics: There exist diverse scheduling policies. For example, Ku- average training time, average JCT, 99-th percentile (P99) JCT, bernetes, de facto job orchestration framework, provides two and makespan. The metrics are explained in §II-A.
TABLE II: Fidelity of existing simulators: MAPE (%, ↓: better).
Model(s)
SQuAD [13] iMDB [15] Synthetic
BERT-base [14] GPT-2 [16] Whisper [17]
Batch size 4–32 4–32 4–32
CIFAR-10 [18]
AlexNet [19], DenseNet-40/100 [20], ResNet-44/110 [21]
128– 32768
ImageNet [22]
Inception-v3 [23], ResNet-50 [21], VGG-16 [24], GoogLeNet [25]
64– 2048
Tiresias (No-contention) Pollux (Static-contention) Average training time Average JCT P99 JCT Makespan
Training model
Dataset
Whisper VGG-16 ResNet-50 ResNet-44 ResNet-110 Inception-v3 GPT-2 GoogLeNet DenseNet-40 DenseNet-100 BERT AlexNet
58.71 73.64 67.19 47.51
56.92 72.23 65.31 46.42 Network time
Traning model
TABLE I: Models used in trace.
Whisper VGG-16 ResNet-50 ResNet-44 ResNet-110 Inception-v3 GPT-2 GoogLeNet DenseNet-40 DenseNet-100 BERT AlexNet
Compute time
3) Workloads: We prepare a 6.6-hour DT job trace sampled from the Microsoft Philly trace [3]. Since the public trace does not provide model details, we use 12 representative models 1.2 1.6 2.0 0 20 40 60 80 100 spanning computer vision, natural language processing, and Average of contention fraction (%) Contention factor (×) speech domains, as summarized in Table I. The trace contains (a) Iteration-time changes. (b) Breakdown of changes. 30 jobs requesting 2–8 GPUs, with per-job durations ranging Fig. 2: Root-cause analysis of poor fidelity. from 3 to 424 minutes. All jobs use data-parallel training with ring all-reduce for gradient synchronization. Job submission accumulate over repeated DT iterations (§II-A), we analyze times follow a Poisson process with arrival rate λ = 0.005, how per-iteration time changes by contention. which makes GPUs nearly fully utilized. We use Kubernetes Each experiment runs two 8-GPU jobs concurrently across load balancing, which spreads jobs across nodes with more two servers with fixed worker locations. On each server, the available resources (§II-B). Other policies show similar trends. measured job A uses four GPUs, and the co-running job B 4) Machine: We run the DT job trace on the two existing disjointly uses the other four GPUs. Both jobs perform gradient simulators and compare their outputs against ground-truth synchronization across the two servers, so their inter-server measurements from a physical GPU cluster. The simulators networking shares the same NICs. For each job A, we change run on a separate host equipped with an Intel Xeon E5-2650 the co-running job B among the 12 models in Table I and v4 @ 2.20 GHz CPU, 24 cores, and 64 GB of DDR4 memory; measure the per-iteration time of A. this host is sufficient for simulation and does not become a Then, we define the iteration-time change of job A bottleneck. To obtain ground-truth measurements for MAPE, when co-running with job B as T (A, B)/T (A), where pair solo we run the same trace on a physical testbed provided by Lambda T (A, B) is the mean iteration time of job A while copair Cloud [26]. The testbed consists of two 8-GPU V100 servers, running with B, and T (A) is the mean iteration time of A solo each equipped with eight V100-SXM2-16GB GPUs, two Intel when running alone. A value of 1 indicates no change, whereas Xeon Gold 5220R CPUs, and 440 GiB of DRAM. The two values > 1 indicate that the iteration time of A increases. servers are connected through a 5 Gbps interface. Fig. 2a shows the iteration-time change of each measured job as the co-running job varies across the 12 models. Across B. Simulation Fidelity the models, the iteration-time changes are 1.33× on average Table II shows the simulation errors of existing simulators and reach up to 2.14× (BERT model). Fig. 2b decomposes the against the physical testbed. Tiresias incurs large errors across iteration-time change into the part caused by network contention all metrics: 58.71% for average training time, 73.64% for and compute contention. On average, network time accounts average JCT, 67.19% for P99 JCT, and 47.51% for makespan. for 88.9% of the change, while compute accounts for only These errors arise because Tiresias ignores the training time 11.1%. Even for the model with the largest compute fraction, increase caused by co-running jobs. Once training time is networking still explains 66.4% of the change. incorrectly modeled, the error propagates to GPU release times, The results show that the network and its contention are the queued-job waiting times, and eventually JCT and makespan. dominant sources of changes in iteration time. As per-iteration Pollux reduces the errors only modestly by applying a fixed time errors accumulate into job-level timing metrics such as contention ratio. However, its MAPE remains high: 56.92% training time and JCT, modeling network contention is key for for average training time, 72.23% for average JCT, 65.31% for improving simulation fidelity. P99 JCT, and 46.42% for makespan. This shows that a static IV. D ESIGN contention factor is insufficient to capture network contention. Errors of this magnitude make simulation results unreliable for Fig. 3 shows the workflow of M O S IM. M O S IM takes DT job scheduler evaluation. Next, we show that network contention configurations, cluster configurations, and a job trace as input is the root cause of the significant errors. (§IV-A). First, M O S IM runs simulation input construction (① in Fig. 3), which prepares all required information to conduct C. Cause of Poor Fidelity DT job simulation. GPU cluster simulators take public traces in To identify why existing simulators present significant errors, general; however, typical public traces do not include each job’s we run the following experiments on the physical testbed using details, such as the model type, dataset, per-iteration compute the same 12-workload pool. Because training time and JCT time, network time, or networking volume. Simulation input
DT Job Config.
① Simulation Input Construction
② Discrete Event Simulation
③ Network Contention Model
ARRIVAL COMPLETE
Job GPU-free augmentation characterization
Training time JCT Makespan
…
Cluster Config.
GPU GPU
NIC
GPU GPU
Traffic
Server
DT Job Trace
Fig. 3: Overview of M O S IM. construction of M O S IM fills the missing information from predefined job pools and GPU-free characterization that gets joblevel details without requiring any GPUs (details in §IV-A2). Next, the discrete event simulation (② in Fig. 3) replays the constructed input over time. It processes job arrivals and completions, applies the scheduling policy of the cluster, updates job progress, and schedules each job’s estimated completion event (§IV-B). Depending on the scheduling, the simulation invokes network contention model (explained next) to update each affected job’s iteration time. After running all jobs in the input, the simulator outputs the training time and JCT for each job, as well as the overall makespan. The network contention model (③ in Fig. 3) captures network contention and estimates its impact on iteration time. Given the current worker assignment on GPUs and servers, which the scheduler determines at simulator run time, the model estimates how contention from co-running jobs inflates each job’s network time and thus its iteration time (§IV-C). A. Simulation Input Construction M O S IM takes three inputs: DT job trace, DT job configuration, and cluster configuration. M O S IM runs simulations from a DT job trace. Representative sources of such traces include public production traces from large-scale GPU clusters [3], [4], [27]. We denote the trace as a set of DT jobs, T = {J}, where each job initially carries only the metadata provided by these traces: J = ⟨aJ , gJ , dJ ⟩, including its submission time aJ , the number of requested GPUs gJ , and its tracereported duration dJ . Production traces rarely include detailed training workload information essential for simulation, such as 1) workload description mJ of the deep learning model type, dataset, training framework, and collective communication algorithm, 2) per-iteration compute time TJcomp , 3) networking time TJnet , and 4) networking volume VJnet . To make the traces usable for simulation, M O S IM performs two steps. First, it augments each trace job with the missing mJ , following the common practice in other simulators [5]–[8]. Second, it characterizes the job to obtain the remaining details, TJcomp , TJnet , and VJnet . The first augmentation is achieved using another input, job configuration, which we explain next. The second job characterization is explained in §IV-A2. 1) Job Augmentation: To augment the details, we take DT job configuration as another input which provides the candidate workload descriptions used for the first augmentation step. Since production traces omit mJ , M O S IM takes candidate DT job configurations as separate inputs and uses them as a pool from which the missing mJ is sampled. Each candidate configuration is provided as a YAML file and specifies the job details that form mJ , including the deep learning model type, dataset, training framework, and collective communication
algorithm. If users do not provide custom configurations, M O S IM uses its default DT job configurations, which consist of the models in Table I. For each trace job, M O S IM samples an mJ from this pool and attaches it to the job, yielding J = ⟨aJ , gJ , dJ , mJ ⟩. Another input, cluster configuration, describes the target cluster specifications and scheduling policy used in the simulation. It includes the number of servers, the number of GPUs per server, the per-server network bandwidth, and the scheduling policy. M O S IM supports heterogeneous cluster configurations, allowing users to specify different GPU types, numbers of GPUs per server, and per-server network bandwidths. M O S IM also provides implementations of several scheduling policies and allows users to implement custom policies, which are used in the simulation (§IV-B). 2) GPU-Free Characterization: In addition to job augmentation, replaying a DT job trace requires per-job execution characteristics that public traces omit (§IV-A) and cannot be filled by augmentation. Existing simulators recover these by profiling each job for a few iterations on a separate GPU cluster, measuring its iteration time when the job runs alone [7].1 This approach is costly and hardware-dependent: it occupies real GPUs and must be repeated whenever the GPU type, model pool, or training configuration changes. M O S IM avoids this cost through GPU-free characterization, which replaces real-GPU profiling with a job-level simulator. Job-level simulators have been developed in the GPU architecture domain to model single-job execution. For example, ASTRA-sim [28] and SimAI [29] simulate a DT job at operatorlevel granularity across different GPU architectures. Assuming the job runs alone on the requested number of GPUs (gJ ), they estimate the mean compute time and networking time, which together constitute the iteration time, and expose operator-level network transfer records from which the networking volume can be derived. M O S IM uses such a simulator to generate a single-job profile for each trace job, which then serves as input to the cluster-level simulation. We use ASTRA-sim2 as the job-level backend, as it is known to be accurate in the compute time, networking time, and networking volume of DT jobs. For each job J, M O S IM obtains its per-iteration compute time TJcomp , networking time TJnet , and networking volume VJnet . M O S IM keeps compute time and networking time separate because it applies network contention only to the networking part of an iteration. This is consistent with our observation in §III-C: iteration time changes almost due to networking time (and its contention), while compute time remains unchanged. The networking volume VJnet is the total network traffic generated by all workers of job J during one iteration. Rather than measuring it on real hardware, M O S IM extracts it from the operator-level transfer records, summing the transmitted and received bytes reported by ASTRA-sim’s networking operators across all workers. 1 Since iteration time depends on the job details, simulators profile it after augmenting each trace job with a DT job configuration. 2 M O S IM is not tied to ASTRA-sim; it can use any single-job simulator that reports per-iteration compute time, networking time, and networking volume.
M O S IM also determines each job’s iteration count WJ . Treating the trace-reported duration dJ as the standalone runtime, M O S IM sets WJ = dJ /TJiter . With WJ replacing dJ , each job is now fully described as: J = ⟨aJ , gJ , mJ , WJ , TJcomp , TJnet , VJnet ⟩.
(1)
The constructed input (including job description) is further used to estimate how network contention changes each job’s iteration time, as explained in §IV-C.
Simulation output. Once the last C OMPLETE event has been processed, M O S IM reports per-job and cluster-level metrics. For each job, training time is the time spent executing after GPU allocation, while JCT is the time from job submission to completion, including both the waiting time before GPU allocation and the training time after allocation. At the cluster level, makespan is the time from the first job’s submission to the last job’s completion across the whole trace. C. Networking Contention Model
B. Discrete Event Simulation
The discrete event simulation in the previous subsection uses M O S IM replays the constructed input as a discrete-event the contention factor S (t) to update each job’s iteration time. J simulation that processes, in time order, two events per job: This subsection explains how M O S IM computes this factor an A RRIVE event at its submission time aJ , and a C OMPLETE from the current worker assignment and per-job networking event scheduled at its estimated finish time fˆJ once the job is profile. The network contention model does not determine the scheduled. A job scheduler applies the scheduling policy from worker assignment. Instead, the assignment is produced by the the cluster configuration (§IV-A) to decide which waiting jobs job scheduler at run time during the discrete-event simulation run and where their workers are assigned. M O S IM updates the (§IV-B). This subsection assumes that the current assignment cluster state only when either of these two event types occurs, is given and defines how the contention factor is computed. as these are the only moments when GPU allocation or NIC We denote job J’s worker assignment—the assignment of sharing can change. its NW = gJ workers to servers—by πJ , and the number When a job arrives at aJ , the scheduler assigns it to free of J’s workers on server i by ni . Using the per-job profile GPUs if it fits; otherwise the job joins a waiting queue. When (T comp , T net , V net ) from §IV-A2W together with π , M O S IM J J J J a job completes, it releases its GPUs, and the scheduler then estimates how much each job’s iteration time changes when revisits the waiting queue to assign any job that fits. Either way, its synchronization traffic shares server NICs and inter-server when a job is assigned or completed at simulation time t, the set links with co-running jobs. M O S IM captures this as a per-job of jobs sharing each NIC can change. Thus, M O S IM computes contention factor, derived from each job’s bandwidth demand the contention factor SJ (t) of every affected job using the on every server NIC as follows. network contention model (to be explained in §IV-C). 1) Bandwidth Demand: M O S IM converts each job’s netBased on the contention factor from the networking conworking volume and scheduling decision into a bandwidth tention model, M O S IM models the iteration time of job J at demand on each server NIC, which is used to determine the simulation time t as: contention factor. We assume synchronous data-parallel training 3 τJ (t) = TJcomp + TJnet · SJ (t). (2) with ring all-reduce. In each iteration, every worker of job J sends and receives the same amount of gradient data, so for a Only the networking term is scaled by SJ (t) because network job with networking volume V net spread over NW workers, J contention affects networking time but not compute time the per-worker bandwidth demand is: (§III-C). At each event time tnow , M O S IM first advances every active (6) c = VJnet /NW . job by the iterations it has completed since its previous update In ring all-reduce, the workers form a logical ring, and time tlast . Since the set of running jobs is fixed between two consecutive events, SJ (t) and hence τJ (t) remain constant each worker communicates only with its two neighbors in the ring. Two neighboring workers on the same server over that interval. The completed iterations are: exchange data locally and do not use the NIC, whereas two tnow − tlast ∆NJ = . (3) neighboring workers on different servers communicate via the τJ (tlast ) NIC connecting the servers. NCCL orders the ring so that workers on a given server are placed next to each other [30]; M O S IM then decrements the remaining iterations: as a result, the workers on server i form a continuous block of WJ (tnow ) = max(0, WJ (tlast ) − ∆NJ ) . (4) the ring, and only the two workers at the ends of that block After processing the event and applying the scheduling policy, communicate with workers on other servers. Each of these two M O S IM recomputes SJ (tnow ) for affected jobs and projects boundary links carries the per-worker demand c, so a job whose workers span multiple servers places a demand of 2c on server their new finish times as: i’s NIC. The one exception is a two-worker job (NW = 2) fˆJ = tnow + WJ (tnow ) τJ (tnow ). (5) split across two servers: the ring then reduces to a single link between the two workers, so the demand is c instead of 2c. The corresponding C OMPLETE event is scheduled or updated ˆ to fJ . Advancing time event by event, rather than through 3 This assumption only affects bandwidth-demand in Eq. 7. M O S IM can fixed time ticks, lets M O S IM track contention that changes support other parallelism strategies by replacing this equation with a strategyover each job’s lifetime. specific communication-demand, without changing the rest of the simulator.
Combining these, the inter-server bandwidth demand of job J on server i is: ( min(1, NW − niW ) × c × 2 NW > 2 i Tex = (7) min(1, NW − niW ) × c NW = 2 where niW is the number of J’s workers on server i. The quantity NW −niW counts J’s workers assigned to other servers, so min(1, NW −niW ) is 1 when at least one such worker exists, meaning J’s traffic crosses server i’s boundary, and 0 when i all of J’s workers are on server i. In the latter case Tex =0 and J uses no inter-server NIC bandwidth. 2) Contention Factor: M O S IM then computes how much bandwidth each job actually receives at each server NIC, i treating the NICs independently. Let DJi = Tex be job J’s demand on server i, C the NIC capacity, and J i (t) the set of active jobs sending traffic through that NIC at time t. If the combined demand of these jobs fits within C, each job receives its full demand; if the combined demand exceeds C, the NIC bandwidth is split among them in proportion to their demands. Both cases are captured by: ! C i i AJ (t) = DJ · min 1, P , (8) i k∈J i (t) Dk where the min(1, ·) term equals 1 when the NIC is not oversubscribed and scales every job’s allocation down by the same ratio when it is. The contention factor of J at server i then measures how far its allocation falls short of its demand: SJi (t) =
DJi , i max(AJ (t), MIN_BW)
(9)
where MIN_BW is a small lower bound on the allocation that keeps the ratio finite. The factor equals 1 when J receives all the bandwidth it needs, and rises above 1 as contention reduces J’s allocation, lengthening J’s networking time by the same proportion. Because synchronous all-reduce advances only as fast as the slowest group of workers, a job’s overall contention factor is the largest factor across the servers it occupies: SJ (t) = maxi SJi (t).
(10)
A job with no inter-server traffic does not suffer contention, so SJ (t) = 1. V. E VALUATION We implement M O S IM in Rust and Python (4.4K lines of source code) and use ASTRA-sim [28] for GPU-free characterization. Based on the implementation, we first describe the experiment setup (§V-A) and then present the results. Our implementation is publicly available4 . A. Experiment Setup We extend the setup in §III-A to a larger scale. We reuse the same model pool and evaluation methodology, but scale the workload from one 6.6-hour trace of 30 jobs to a 13-hour trace of 60 jobs, derived from the Philly trace [3]. We also scale the testbed from two to four 8-GPU V100 servers (32 GPUs in 4 https://github.com/OSSS-KU/MoSim
total), keeping the same per-server hardware and inter-server bandwidth as in §III-A. As in §III-A, we use measurements on this physical testbed as the ground truth. Baselines. We compare M O S IM against the same two simulators as §III-A: Tiresias [5], no-contention simulator that replays each job’s standalone time, and Pollux [7], static-contention simulator that inflates every co-running job’s time by 10%. Scheduling policies. We evaluate M O S IM under two cluster scheduling policies: bin packing, which places workers on the fewest servers possible by preferring servers with the fewest free GPUs, and load balancing, which spreads workers across servers by preferring servers with the most free GPUs. Both are provided by Kubernetes [11]. We evaluate the following items: • Simulation fidelity (§V-B): how closely M O S IM reproduces training time, JCT, and makespan measured on the physical testbed. We report 1) aggregate fidelity using MAPE for these metrics; and 2) distribution-level fidelity using Kolmogorov– Smirnov and Wasserstein distances between the simulated and measured, ground-truth JCT distributions. • Fidelity of network contention model (§V-C): compare the contention factor from M O S IM against the ground-truth value measured on the physical testbed in §III-C. • Input construction overhead (§V-D): the resource overhead of constructing each simulator’s inputs for the models. We report the required machine time, monetary cost, and normalized cost, comparing real-GPU profiling in the baselines with GPU-free characterization in M O S IM. B. Simulation Fidelity 1) Aggregate Fidelity: Table III shows the MAPE of training time, JCT (average, median, and P99), and makespan for Tiresias, Pollux, and M O S IM under both bin packing and load balancing scheduling policies. M O S IM achieves the best (lowest) MAPE in all cases, showing the highest accuracy across all metrics and scheduling policies. For training time, M O S IM improves accuracy by 1.93× over Tiresias and 1.79× over Pollux under bin packing, and by 3.33× and 3.21× respectively under load balancing. For average JCT, M O S IM improves accuracy by 3.28× over Tiresias and 2.97× over Pollux under bin packing. Under load balancing, M O S IM achieves 2.47× and 2.36× improvements, respectively. For median JCT, M O S IM achieves improvements of 21.59× over Tiresias under bin packing and 2.77× under load balancing. For P99 JCT, M O S IM improves accuracy by 2.03× over Tiresias and 1.95× over Pollux under bin packing. Under load balancing, M O S IM achieves larger improvements of 7.79× and 7.62× respectively. For makespan, M O S IM improves accuracy by 3.04× over Tiresias and 2.9× over Pollux under bin packing, and by 8.48× and 8.25× respectively under load balancing. 2) Distribution-Level Fidelity: Here, we evaluate how accurately M O S IM captures the overall JCT distribution of the trace. We use two distance metrics between the measured and simulated JCT distributions. First, we measure Kolmogorov– Smirnov (KS) distance, which quantifies the maximum vertical distance between the measured and simulated cumulative distribution functions (CDFs) as maxx |Fmeasured (x) − Fsimulated (x)|.
TABLE III: Simulation fidelity: MAPE (%) across training time, JCT metrics, and makespan (§V-B1, ↓: better).
Method
Bin packing
Load balancing
Average Average Median training time JCT JCT
Average Average Median P99 JCT Makespan training time JCT JCT
Tiresias Pollux M O S IM
38.19 35.30 19.74
36.38 33.00 11.10
34.98 50.42 31.83 48.46 1.62 24.84
53.45 51.08 17.61
TABLE IV: Simulation fidelity: distribution-level fidelity, KS and Wasserstein distance of JCT distribution (§V-B2, ↓: better). Bin packing
56.75 54.72 17.05
53.93 51.42 21.83
P99 JCT Makespan
54.20 65.28 51.86 63.89 19.60 8.38
65.44 63.64 7.72
TABLE V: Input construction overhead: time (hours) on GPUmachine and CPU-machine, cost ($), and normalized cost (§V-D, ↓: better).
Load balancing
Method KS distance Wasserstein (s) KS distance Wasserstein (s)
Method GPU-hours CPU-hours Cost ($) Normalized cost
Tiresias Pollux M O S IM
Tiresias Pollux M O S IM
3
0.32 0.27 0.13
3499 3184 1579
7082 6753 2864
Ground Truth
MoSim (Ours) NIC factor
0.38 0.37 0.22
2
Tiresias Pollux M O S IM
29.69 21.05 8.63
W hi VG sp G 16 R N 50 R N 4 R 4 N 11 0 In cV 3 G PT G 2 oo gL D N 4 D 0 N 10 0 BE RT Al ex
1
Method MAPE (%)
(a) NIC contention factor comparison.
(b) MAPE(%).
Fig. 4: Fidelity of network contention model: NIC contention factor comparision and MAPE (§V-C, MAPE ↓: better).
41.2 41.2 -
2.7
32.55 32.55 0.73
44.6× 44.6× 1×
across the 12 models. From the measurements, we calculate the real contention factor and compare it with the estimate from the network contention model. Fig. 4 shows the results. In Fig. 4a, for each model on the x-axis, we show two whiskers: one black whisker for the ground-truth value and one blue whisker for the value estimated by our network contention model. For each model, there are 12 counterpart models, so the whiskers show the range of all values. We observe that the estimates from M O S IM are highly similar to the ground-truth values. This corresponds to the MAPE of 8.63% for M O S IM across 144 possible experiments, as shown in Fig. 4b. We also calculate the MAPE values of Tiresias and Pollux, which set the fixed contention factor to 1.0 and 1.1, respectively. Compared with Tiresias and Pollux, M O S IM’s MAPE is 3.44× and 2.44× lower, respectively.
This metric captures the largest distributional gap at any point and is useful for detecting bias across quantiles, including the tail region. Second, we measure Wasserstein distance, which quantifies how far the simulated JCT distribution is from the D. Input Construction Overhead measured JCT distribution in terms of magnitude. Unlike KS Table V compares the overhead of each simulator, especially distance, which captures the largest CDF gap, Wasserstein for constructing simulation inputs. We use on-demand cloud distance reflects overall JCT differences across the distribution. prices for the analysis: $6.32/h for the GPU machine, i.e., an Table IV shows that M O S IM consistently outperforms the 8-GPU V100 VM [26], and $0.2688/h for the CPU machine, baselines across both distance metrics. Under bin packing, i.e., a t4g.2xlarge CPU instance in AWS us-east-2. Existing M O S IM achieves a KS distance of 0.13, which is 2.46× better simulators, such as Tiresias and Pollux, run profiling to fill in than Tiresias and 2.08× better than Pollux. It also achieves missing job details, consuming the GPU machine for about 41 a Wasserstein distance of 1579 s, improving over Tiresias by hours. Instead, M O S IM runs its GPU-free characterization on 2.22× and Pollux by 2.02×. These results indicate that the the CPU machine for about 2.7 hours. As a result, M O S IM simulated distribution closely follows the measured distribution reduces the monetary cost of input construction by 44.6×. across all quantiles. Under load balancing, M O S IM obtains a VI. R ELATED W ORK KS distance of 0.22, which is 1.73× better than Tiresias and 1.68× better than Pollux, and a Wasserstein distance of 2864 s, GPU cluster simulators. Existing GPU cluster simulators which improves over Tiresias by 2.47× and Pollux by 2.36×. differ from M O S IM along two criteria: how they model network The results show that M O S IM provides superior distribution- contention and how they obtain per-job characteristics. As level fidelity, accurately capturing not only aggregate statistics discussed in §II-C, prior simulators either ignore contention but also the overall shape and spread of the JCT distribution. (Tiresias [5], Gavel [6]) or apply a fixed penalty ratio (Pollux [7], Muri [8]), and all obtain per-job characteristics through realC. Fidelity of Network Contention Model GPU profiling. M O S IM differs on both: it models contention We next evaluate the fidelity of M O S IM’s network contention dynamically from placement and networking volume, and model. We reuse the pairwise job-running measurements from obtains the same per-job characteristics through GPU-free §III-C, where one job is fixed, and its co-running job varies characterization, removing GPU dependence from simulation.
Single-job simulators. M O S IM relies on a single-job simulator as the backend for its GPU-free characterization (§IV-A2). ASTRA-sim [28], SimAI [29], and Multiverse [31] simulate a single job at operator-level granularity and are complementary to M O S IM, which operates at the cluster level. M O S IM uses one such simulator (ASTRA-sim) to replace per-job profiling and layers its cluster-level contention model on top; any backend that reports per-iteration compute time, networking time, and networking volume can serve in its place. Communication schedulers. A separate line of work mitigates network contention directly. C ASSINI [32] reduces network contention through fine-grained delay scheduling, and Crux [33] assigns flow priorities by GPU intensity. VALO accurately and efficiently splits datacenter traffic of AI workloads across multiple paths [34]. These techniques intervene to reduce contention, whereas M O S IM models and predicts it; the two are independent, and such mitigation could itself be included and modeled within M O S IM. VII. C ONCLUSION This paper introduces M O S IM, a GPU-cluster simulator that models DT jobs under dynamic network contention. M O S IM combines GPU-free characterization with network contention model, obtaining each job’s compute time, networking time, and networking volume without real-GPU profiling, and modeling how shared server NICs and inter-server links change each job’s iteration time with scheduling decisions. Compared with existing simulators, M O S IM reduces simulation error by up to 3.28× for average JCT, 7.79× for P99 JCT, and 8.48× for makespan. It also estimates NIC contention factors with 8.63% MAPE and reduces input construction cost by 44.6×. The results show that network contention modeling is important for faithful GPU-cluster simulation. ACKNOWLEDGMENT Yeonho Yoo, Hyunho Lee, and Hyunmok Choi contributed equally. This research was supported by National Research Foundation of Korea (NRF) grant funded by Korea government (MSIT) (RS-2024-00336564), by Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by Ministry of Science and ICT (MSIT) (RS2026-25518394), by ICT Creative Consilience Program through IITP grant funded by MSIT (IITP-2026-RS-2020-II201819), by IITP under the Artificial Intelligence Convergence Innovation Human Resources Development grant funded by Korea government (MSIT) (IITP-2026-RS-2023-00254592), by ANCHOR through Seoul ANCHOR Center funded by MOE and Seoul Metropolitan Government (2026-ANCHOR-01-003-09), and by computing support from Lambda Cloud. Corresponding authors: Gyeongsik Yang and Chuck Yoo. R EFERENCES [1] Y. Go, C. Shin, M. Kang et al., “Making sense of job preemption for distributed deep learning acceleration,” in Proc. DAC, 2026. [2] D. Narayanan, A. Harlap, A. Phanishayee et al., “PipeDream: Generalized pipeline parallelism for DNN training,” in Proc. SOSP, 2019, pp. 1–15. [3] M. Jeon, S. Venkataraman, A. Phanishayee et al., “Analysis of large-scale multi-tenant GPU clusters for DNN training workloads,” in Proc. ATC, 2019, pp. 947–960.
[4] Q. Hu, Z. Ye, Z. Wang et al., “Characterization of large language model development in the datacenter,” in Proc. NSDI, 2024, pp. 709–729. [5] J. Gu, M. Chowdhury, K. G. Shin et al., “Tiresias: A GPU cluster manager for distributed deep learning,” in Proc. NSDI, 2019, pp. 485–500. [6] D. Narayanan, K. Santhanam, F. Kazhamiaka et al., “HeterogeneityAware cluster scheduling policies for deep learning workloads,” in Proc. OSDI, 2020, pp. 481–498. [7] A. Qiao, S. K. Choe, S. J. Subramanya et al., “Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning,” in Proc. OSDI, 2021, pp. 1–18. [8] Y. Zhao, Y. Liu, Y. Peng et al., “Multi-resource interleaving for deep learning training,” in Proc. SIGCOMM, 2022, pp. 428–440. [9] C. Shin, Y. Go, Y. Yoo et al., “Prediction-based GPU sharing for distributed training,” Future Generation Computer Systems, vol. 181, p. 108413, 2026. [10] H. Wang, H. Tian, J. Chen et al., “Towards Domain-Specific network transport for distributed DNN training,” in Proc. NSDI, 2024, pp. 1421– 1443. [11] The Kubernetes Authors, “kube-scheduler,” https://github.com/kubernetes/ kube-scheduler, 2026, accessed: 2026-05-31. [12] K. Ma, K. Wang et al., “Volcano: A cloud native batch system,” https: //github.com/volcano-sh/volcano, 2025. [13] P. Rajpurkar, J. Zhang, K. Lopyrev et al., “SQuAD: 100,000+ questions for machine comprehension of text,” in Proc. EMNLP, 2016, pp. 2383– 2392. [14] J. Devlin, M. Chang, K. Lee et al., “BERT: pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL, 2019, pp. 4171–4186. [15] A. L. Maas, R. E. Daly, P. T. Pham et al., “Learning word vectors for sentiment analysis,” in Proc. ACL, 2011, pp. 142–150. [16] A. Radford, J. Wu, R. Child et al., “Language models are unsupervised multitask learners,” OpenAI Blog, p. 9, 2019. [17] A. Radford, J. W. Kim, T. Xu et al., “Robust speech recognition via large-scale weak supervision,” in Proc. ICML, 2023, pp. 28 492–28 518. [18] A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” University of Toronto, Tech. Rep., 2009. [19] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Proc. NeurIPS, 2012, pp. 1097–1105. [20] G. Huang, Z. Liu, L. van der Maaten et al., “Densely connected convolutional networks,” in Proc. CVPR, 2017, pp. 2261–2269. [21] K. He, X. Zhang, S. Ren et al., “Deep residual learning for image recognition,” in Proc. CVPR, 2016, pp. 770–778. [22] J. Deng, W. Dong, R. Socher et al., “ImageNet: a large-scale hierarchical image database,” in Proc. CVPR, 2009, pp. 248–255. [23] C. Szegedy, V. Vanhoucke, S. Ioffe et al., “Rethinking the inception architecture for computer vision,” in Proc. CVPR, 2016, pp. 2818–2826. [24] K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. ICLR, 2015. [25] C. Szegedy, W. Liu, Y. Jia et al., “Going deeper with convolutions,” in Proc. CVPR, 2015, pp. 1–9. [26] Lambda Labs, “Lambda Labs AI Cloud Service,” https://lambda.ai/, 2026, accessed: 2026-06-03. [27] Q. Hu, P. Sun, S. Yan et al., “Characterization and prediction of deep learning workloads in large-scale GPU datacenters,” in Proc. SC, 2021, pp. 104:1–104:15. [28] W. Won, T. Heo, S. Rashidi et al., “ASTRA-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale,” in Proc. ISPASS, 2023, pp. 283–294. [29] X. Wang, Q. Li, Y. Xu et al., “SimAI: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision,” in Proc. NSDI, 2025, pp. 541–558. [30] NVIDIA, “NCCL: Optimized primitives for inter-GPU communication,” https://github.com/NVIDIA/nccl, 2026, accessed: 2026-09-01. [31] F. Gui, K. Gao, L. Chen et al., “Accelerating design space exploration for LLM training systems with multi-experiment parallel simulation,” in Proc. NSDI, 2025, pp. 473–488. [32] S. Rajasekaran, M. Ghobadi, and A. Akella, “CASSINI: Network-aware job scheduling in machine learning clusters,” in Proc. NSDI, 2024, pp. 1403–1420. [33] J. Cao, Y. Guan, K. Qian et al., “Crux: GPU-efficient communication scheduling for deep learning training,” in Proc. SIGCOMM, 2024, pp. 1–15. [34] Y. Yoo, G. Yang, C. Shin et al., “Revisiting traffic splitting for software switch in datacenter,” Proc. ACM Meas. Anal. Comput. Syst., vol. 9, no. 2, pp. 39:1–39:26, 2025.