Characterizing Job Power Elasticity for Power-Flexible AI Training
arXiv:2609.11542v1 [cs.AI] 10 Sep 2026
Philip Colangelo, Charles Dawson, Shayan Sengupta, Ayse Coskun, Varun Sivaram Emerald AI
Abstract Large language model (LLM) training is among the fastest-growing sources of electricity demand in modern data centers, and power availability is a primary bottleneck to continued AI infrastructure growth. Making the power consumption of these workloads flexible could unlock additional power for AI growth, limit increases in electricity prices, and improve the utilization of existing grid infrastructure. However, to realize this flexibility, we must first understand how the performance of training workloads changes when GPU power is reduced. This paper presents the first systematic characterization of job power elasticity (the sensitivity of throughput to power reductions) in LLM training. To quantify elasticity, we introduce the Power Flexibility Index (PFI), a normalized metric that quantifies the performance cost of power reductions and provides a control primitive for SLA-aware power flexibility. We collect data from 131 LLM training runs on H200 (plus 24 H200 validation runs and 34 matched H100 runs), including both dense and mixture-of-experts models, pretraining and fine-tuning tasks, and up to 32 GPUs. We find that LLM training jobs exhibit substantial but variable power elasticity, and we identify telemetry signals that predict PFI at runtime. Finally, we demonstrate that PFIaware power allocation maximizes total tokens/second throughput under power constraints. Under a 30% power reduction, PFI-aware power allocation recovers ∼1.5k tokens/s per job, 63% of the performance gap between an equal-weight allocation and an oracle with perfect information. Our results establish power elasticity as a measurable property of training jobs and provide a foundation for power-aware, grid-responsive AI infrastructure.
1
Introduction
Large language model (LLM) training is a rapidly growing source of data center electricity demand. Power availability has become a bottleneck for the deployment of new AI compute capacity, with limited grid capacity causing multi-year interconnection delays for new data centers [12]. Recent analysis of the U.S. power system found that data centers operating as flexible loads, reducing power consumption in a few peak hours, could unlock nearly 100 GW of additional capacity [17]. The promise of such flexible capacity creates a strong incentive for AI training clusters to be powerflexible. To realize these benefits, AI clusters must reduce power consumption when needed while maintaining acceptable training performance, as defined in customer service-level agreements (SLAs). Achieving SLA-aware power flexibility requires understanding the throughput cost of power reductions. Prior studies of training efficiency primarily focus on energy consumption, Model FLOPs Utilization (MFU [2]), tokens-per-joule at nominal operating points [4], or energy-delay tradeoffs for individual jobs [32, 3]. However, these metrics provide limited insight into the quantity most relevant for Preprint.
PFI vs. Power (W) PFI vs. Tensor active PFI vs. SM active PFI vs. Mean DRAM active (%)
→
→
Performance gap vs. oracle (tok/s/job)
PFI vs. NVLink TX
0
−1000 −2000 −3000 Oracle (perfect information) Equal weight PFI-aware
−4000
0
5
10
15
20
25
30
35
40
45
Power reduction (%)
Figure 1: LLM training jobs vary in how quickly performance degrades under power reductions. Relatively
more power flexible (—) jobs maintain throughput despite power reductions, while relatively less flexible (- -) jobs quickly lose performance at lower power. We introduce the Power Flexibility Index (PFI), a new metric for measuring power elasticity, and conduct large-scale, systematic characterization of modern LLM training jobs. We develop a model for predicting PFI online from readily available GPU telemetry, and we demonstrate how PFI-aware power allocation improves total throughput under cluster-wide GPU power constraints and realistic job mixes.
infrastructure control: a job’s throughput sensitivity to sustained reductions in available power for up to several hours during periods of peak grid load. Understanding job-level power elasticity is important because training jobs with similar baseline throughput and power draw can respond very differently to the same GPU power cap. As a result, heuristics for spreading power constraints across jobs can waste throughput that could otherwise be preserved through elasticity-aware power allocation. To enable these strategies, cluster power managers must be able to identify which jobs can absorb power reductions at the lowest performance cost. To this end, we present the first large-scale empirical characterization of the power elasticity of modern LLM training jobs under controlled GPU power modulation. We conduct 131 training runs across multiple architectures, pretraining and fine-tuning tasks, and up to 32 NVIDIA H200 GPUs, sweeping GPU power caps across the full operating envelope. From these measurements, we compute the Power Flexibility Index (PFI), a normalized metric that quantifies the performance cost of power reductions and allows the cluster power manager to rank jobs by relative flexibility. Building on this empirical characterization, we develop telemetry-driven estimators that infer a job’s PFI online from readily-available signals from NVIDIA Data Center GPU Manager (DCGM), allowing cluster power managers to estimate PFI without power-cap sweeps. We apply this estimator to simulated workloads to demonstrate how PFI-aware orchestration can protect relatively inflexible jobs while assigning deeper power reductions to relatively flexible jobs, maximizing tokens/sec under a global power budget. Our contributions are: • We present the first systematic characterization of power elasticity across LLM training jobs, across a range of model architectures, training tasks, and distributed GPU scales. • We introduce the Power Flexibility Index (PFI), a normalized job-level metric that quantifies the throughput cost of power reduction. • We identify runtime telemetry signals that predict PFI, showing that elasticity can be inferred online from GPU monitoring signals that correlate with different flexibility regimes. • We demonstrate that PFI-aware power allocation improves aggregate cluster throughput under cluster-wide GPU power constraints, establishing a practical path toward power-aware, grid-responsive data center operation.
2
Power Elasticity and the Power Flexibility Index
A job’s power elasticity describes the sensitivity of its throughput to sustained reductions in available power. Highly elastic jobs experience rapid throughput degradation under GPU power capping, whereas inelastic jobs preserve throughput. Figure 1 illustrates representative elastic and inelastic power–throughput regimes. 2
Power elasticity is important because it enables throughput-aware power capping, allowing cluster power managers to reduce GPU power consumption while minimizing throughput loss. Despite its importance, two gaps remain in the current literature. First, there is no accepted quantitative definition of power elasticity that permits systematic comparison across jobs, model architectures, and hardware settings. Second, there is limited understanding of the job characteristics that determine elasticity or enable its prediction a priori. To close these gaps, we begin by developing a formal definition of a power flexibility index (PFI)1 . Let (T, P )i be tuples of the average throughput (in tokens/s) and aggregate GPU power (in W) of a job measured under a range of per-GPU power caps P̄i , i = 1, . . . , N .2 Pi is the sum of per-device GPU power readings across all GPUs participating in the job. The realized Pi generally falls below G · P̄i (where G is the GPU count), since the cap is an upper bound and the job need not saturate. We define (Tmax , Pmax ) as the throughput/power pair measured without a power cap (i.e. with P̄N set to thermal design power, TDP). We define the power flexibility index as the ratio of power decrease to throughput decrease, averaged over measured power caps (excluding the uncapped measurement where ∆T = 0): N
PFI =
1 1 X ∆Pi = N i=1 ∆Ti N
N X i=1;Ti ̸=Tmax
1 − Pi /Pmax 1 − Ti /Tmax
(1)
While elasticity describes throughput sensitivity to power loss, PFI inverts this relationship into a control-oriented measure of how efficiently a job can shed power. Consequently, jobs with high elasticity exhibit low PFI, whereas throughput-preserving inelastic jobs exhibit high PFI. PFI is a hardware-dependent property of a job: the formula normalizes for per-platform reference points (Tmax , Pmax ) and so is comparable across models on the same accelerator, but values from different accelerator generations should not be compared directly because they reflect each platform’s specific power-cap mechanism (§5). The metric has a straightforward interpretation: Inelastic (P F I > 1): throughput declines more slowly than power (e.g., −5% power, −2% throughput). Linear (P F I = 1): throughput is proportional to power (e.g., −5% power, −5% throughput).
Elastic (P F I < 1): throughput declines faster than power (e.g., −5% power, −10% throughput).
Comparison to existing metrics. The closest existing metrics fall into two families. ML training efficiency metrics such as MFU [2] and tokens/joule characterize steady-state efficiency at a fixed operating point; they carry no information about how throughput responds to power reduction. Zeus [32] and Perseus [3] consider the energy–time Pareto frontier for individual jobs, but produce perjob optimization outputs rather than a scalar characterization metric comparable across architectures and training tasks. Power systems flexibility metrics quantify load reduction capacity (MW), ramp rate, and response latency [14] but ignore quality-of-service degradation. PFI unifies both perspectives: it is computable from external power measurements without job internals, normalized for crossarchitecture comparison, and explicitly encodes the performance cost of power reduction. Driving questions. Given this definition, three questions motivate the rest of this paper:
What factors influence power elasticity? Does PFI vary across architectures (e.g. dense models vs. mixture-of-experts) or training tasks (e.g. pretraining vs. fine-tuning)? How does training scale affect PFI? Are there underlying mechanisms (e.g. memory bottlenecks) that explain these variations? Can we predict power elasticity in production workloads? The definition of PFI in (1) requires measurements across multiple power caps, which limits its ability to predict job elasticity in production. Predicting PFI from readily available telemetry (e.g. NVIDIA DCGM [18]) would allow ML engineers and data center operators to track power elasticity in real-time. Can we use PFI for real-time control and power-aware workload orchestration? Prior works have demonstrated the ability of GPU clusters to respond to requests from grid operators [31, 5]. Data 1 Note that under this definition a high-PFI job has performance that is relatively inelastic with respect to power reduction and therefore more power-flexible. 2 Throughout this paper, a job refers to the combination of training task, model, GPU count, parallelism strategy, sequence length, and batch size; each job is run at multiple power caps to measure PFI. A workload refers to a collection of jobs.
3
Table 1: Experiment matrix. Each sweep group fixes one full configuration per job and varies only the per-GPU
power cap; PFI is computed within a sweep group. 8-GPU groups sweep power caps in 100 W increments from 200–700 W, 16- and 32-GPU groups sweep 200/500/600/700 W only. EP = expert parallelism degree (pretraining only); PT = pretraining; LoRA = rank-32 LoRA SFT. GC = gradient checkpointing.
Model
Training Task
GPUs
Seq Len
Batch
EP
Compile
GC
gpt-oss-20b Qwen3-30B-A3B Qwen3-32B Llama-3.1-70B
PT, LoRA PT, LoRA PT, LoRA PT, LoRA
8 8, 16 8, 16, 32 8, 16
2048, 4096 4096 2048, 4096, 8192 2048, 4096
2, 8 2, 16 4, 8, 16 2, 16
1, 4, 8 1, 8 — —
Y, N Y Y, N Y, N
Y, N Y, N Y Y
Table 2: Models used in experiments, split into training and held-out validation sets. Active params denote the Model
Architecture
Total
Active
Routing
Training
gpt-oss-20b Qwen3-30B-A3B Qwen3-32B Llama-3.1-70B
MoE decoder MoE decoder Dense decoder Dense decoder
21 B 30 B 32.8 B 70.6 B
≈3.6 B ≈3 B 32.8 B 70.6 B
4 of 32 experts 8 of 128 experts — —
Held out
per-token activation count for sparse Mixture-of-Experts (MoE) models.
DeepSeek-V2-Lite Mistral-Small-24B-Base-2501 Gemma-4-26B-A4B
MoE decoder Dense decoder MoE decoder
15.7 B 24 B 26 B
≈2.4 B 24 B ≈4 B
6+2 of 64 experts — 8 of 128 experts
center operators could use real-time PFI estimates to respond efficiently, allocating power reductions to jobs that can best tolerate power reductions and minimizing disruption to their tenants.
3
Experimental Setup
To characterize power elasticity for representative LLM training jobs, we sweep per-GPU power caps from 200 W to 700 W (TDP), using 100 W increments for 8-GPU jobs and a four-point sweep for 16- and 32-GPU jobs, while pretraining and fine-tuning four open-weight LLMs on clusters of 8/16/32 GPUs, yielding 25 headline sweep groups (131 individual training runs across power caps) plus 4 held-out validation sweeps (24 runs). Complete data tables showing the number of runs for each model are included in Appendix C. All runs execute on Google Kubernetes Engine (GKE) a3-ultragpu-8g nodes (8× NVIDIA H200, 141 GiB HBM3e, 700 W TDP), with multi-node jobs scheduled through Kueue. Clusters are provisioned from a single Terraform module and replicated across six GCP regions; full infrastructure details are given in Appendix B. Throughput (k tok/s)
Cluster power (kW)
We profile four open-weights LLMs span5 8 ning dense and Mixture-of-Experts (MoE) architectures (gpt-oss-20b, Qwen3-30B-A3B, 4 6 Qwen3-32B, and Llama-3.1-70B) under pre3 4 training and rank-32 LoRA fine-tuning (Ta2 bles 1, 2). We chose sequence length 2 and batch size to saturate GPU utilization 1 0 at each cluster size, with gradient check0 5 10 15 20 25 30 35 Elapsed time (min) pointing applied where needed to fit activations in memory and torch.compile enabled where supported. We also profile Figure 2: Aggregate GPU power and training throughthree validation models and exclude their re- put for a representative Q WEN 3-32B pretraining run on sults from feature selection and predictor fit- 32×H200 at 700 W/GPU. ting: DeepSeek-V2-Lite for MoE pretraining, Mistral-Small-24B-Base-2501 for dense pretraining and dense LoRA SFT, and Gemma-4-26B-A4B for MoE LoRA SFT. We enforce per-GPU power caps via nvidia-smi -pl before any CUDA context allocation and verify using DCGM power limit counters. We collect GPU utilization, power, and memory at 5 s intervals via a DCGM sidecar; training throughput and MFU are logged in each training step via a custom callback. We partition each trace into warmup, startup, and training phases; all reported 4
metrics aggregate over the training phase only, using the mean of the top 10% of samples to exclude evals and checkpoints. A representative run is shown in Fig. 2.
4
Results
Three observations emerge from our study: PFI varies substantially across jobs, the variation depends on architecture and training task, and memory- and compute-related telemetry demonstrate strong correlations with PFI. Complete data tables are included in Appendix C. PFI varies substantially across modern LLM training jobs. Fig. 3 shows the raw and normalized power–throughput curves for all runs in our dataset, and Fig. 4 summarizes PFI across cluster sizes, model architectures, and training tasks. Across the cohort, we observe substantial variation in PFI (1.10–2.67), indicating that modern LLM training jobs differ in how much throughput they preserve under GPU power capping. Some jobs exhibit nearly linear power–throughput degradation, while others maintain high throughput under aggressive power reduction. We did not observe any jobs with PFI < 1; however, as we discuss in §6, differences in relative PFI are still important for runtime power orchestration. Architecture and training task shape the observed PFI distribution. As shown in Fig. 4, we observe that dense jobs have the lowest PFI (median 1.27), MoE fine-tuning jobs have the highest PFI (median 2.10), and MoE pretraining jobs fall in between (median 1.43). Applying Dunn’s test for statistical significance of group differences, we find support for MoE fine-tuning as a distinct cluster (corrected p < 0.05) but are unable to distinguish MoE pretraining as a third cluster distinct from dense or MoE fine-tuning. Appendix D.4 provides the results of our statistical group difference testing. Memory and compute intensity-related telemetry show the strongest association with PFI. Fig. 5 shows the relationship between PFI and several metrics gathered through DCGM, including mean SM, DRAM, and tensor pipe activity, NVLink utilization, and selected composite metrics. We observe statistically significant positive correlation between PFI and memory-related signals (mean DRAM activity and DRAM copy product), as well as statistically significant negative correlation with tensor pipe activity. We tested several additional composite metrics, reported in Appendix D.1; however,
MoE
Pretraining
Fine-tuning
60
1.0
T /Tmax (mean)
1.0 0.8
T /Tmax
Throughput (k tokens/s)
Dense
40 20
0.6 0.4
0.8 0.6 0.4
0 10
20
0.25
0.50
Mean cluster power (kW)
0.75
1.00
0.25
0.50
P/Pmax
0.75
1.00
P/Pmax
Figure 3: Power–performance curves for all sweep groups. Left: Aggregate GPU power vs. throughput. Center: Per-sweep normalized power vs. throughput. Right: Mean normalized power vs. throughput for each category (dense vs. MoE, pretraining vs LoRA SFT). Full data are provided in Appendix C. 2.75
PFI = 1 median 1.35
2
2.25
2.00
2.00
1.75
1.0
1.5
2.0
PFI
2.5
1.75
1.50
1.50
1.25
1.25
1.00
0
2.50
2.25
PFI
6 4
2.75
Dense MoE
2.50
PFI
Sweep groups
8
1.00 8
16
Cluster size (GPUs)
32
Dense Pretraining
Dense Fine-tuning
MoE Pretraining
MoE Fine-tuning
Figure 4: Left: PFI distribution across all jobs. Right: PFI by architecture and training task; median shown as horizontal bar.
5
Dense MoE Validation
2.50
2.25
2.25
2.00
2.00
2.00
2.00
1.75
1.75
1.50
1.50
1.50
1.25
1.25
1.25
ρ=-0.27 (n=25)
0.2
0.3 0.4 0.5 Mean DRAM active (%)
1.75 1.50 1.25
ρ=-0.64** (n=25) 0.2 0.4 0.6 Mean tensor active (%)
ρ=-0.13 (n=25) 0
50 100 150 Mean NVLink TX (GB/s)
2.75
Dense MoE Validation
2.5
2.50 2.25
2.00
2.00
PFI
2.25
1.75
1.50
1.25
1.25
ρ=0.54* (n=25)
2.25
1.75
1.50
10 20 30 DRAM copy product
2.50
2.0
PFI
2.50
0.8 Mean SM active (%)
ρ=0.54* (n=25)
PFI
0.6
PFI
1.75
PFI
2.50
2.25 PFI
2.50
2.25 PFI
PFI
2.50
1.5
ρ=0.63** (n=25) 0.4 DRAM to SM ratio
1.75 1.50 1.25
ρ=-0.61** (n=25)
1.0
0.6
2.00
1 2 3 Arithmetic intensity
ρ=-0.67** (n=25) 0.2
0.4 0.6 Tensor to SM ratio
0.8
Figure 5: PFI vs. GPU telemetry metrics, showing the correlation (ordinary least-squares fit and Spearman ρ) between PFI and metrics measured for the uncapped job in each sweep group. DRAM copy product and tensor to SM ratio are derived metrics. ∗ q < 0.05, ∗∗ q < 0.01, where q is the Benjamini–Hochberg FDR-corrected p-value across all metrics tested in Fig. 14 in the appendix.
due to the limited size of our dataset, we limit our analysis in §5 to features with a strong mechanistic link to power flexibility.
5
Discussion
To understand the mechanism behind the different PFI regimes in §4, as well as the correlation between memory activity and PFI, we can examine how the H200 implements power caps. Fig. 6 shows that the H200 caps power by preferentially reducing SM clock frequency while leaving memory clock frequency essentially unchanged: capped GPUs lose compute throughput, but retain memory bandwidth. Memory-bound jobs therefore absorb caps with a smaller throughput loss than computebound jobs. As shown in Fig. 5, MoE models exhibit higher values for memory-related telemetry, suggesting that these signals provide a measurable proxy for regime, separating compute-bound and memory-bound jobs. While extensively studying whether this mechanism generalizes across accelerator architectures is beyond the scope of this work, Appendix D.2 includes a limited validation on H100s, which also preferentially throttle compute speed over memory bandwidth [28], showing a similar pattern.
2.6 2.4
3000
2.2
2500
Actual
Clock frequency (MHz)
R2val = 0.280
2000 1500
2.0 1.8 1.6 1.4
1000 500 200
300
400
500
600
Dense MoE Validation
1.2
SM clock Memory clock
1.5
700
2.0
2.5
Fitted
Power cap (W/GPU)
Figure 7: Predictions of the PFI model fit
Figure 6: Steady-state SM clock and memory
to DRAM copy product, including the validation set.
clock frequency as a function of power cap.
6
Table 3: Predictor ablation for PFI (n = 25). Job metadata includes architecture type, training task, and log2
GPU count. p: # of features. LOO: leave-one-out. LOCO: leave-one-cluster-out. ρ: Spearman rank correlation. LOO Predictors
p
2 Rin
(i) DRAM copy product (ii) Arithmetic intensity (iii): (i) + (ii) (iv) Job metadata (v): (i) + (iv) (vi): All DCGM metrics (vii): All DCGM metrics + (iv)
1 1 2 3 4 13 16
0.675 0.381 0.676 0.519 0.687 0.885 0.942
R
2
RMSE
0.524 0.238 0.462 0.310 0.350 -0.187 -0.010
0.241 0.305 0.256 0.290 0.282 0.381 0.351
LOCO ρ
R
2
RMSE
ρ
0.489 0.482 0.435 0.377 0.386 0.323 0.412
0.246 -1.652 -0.588 -4.106 -2.391 -18.483 -104
0.303 0.569 0.440 0.789 0.643 1.542 3.572
0.515 0.311 0.357 0.188 0.208 0.077 -0.508
As defined in Eq. 1, computing PFI requires power-throughput measurements at a range of power caps, making it difficult to calculate for production jobs. While we find that PFI varies by job type, the power management layer typically does not know which architecture or task is being run on any given hardware node. As a result, the ability to correlate GPU telemetry with task/architecture regime (and thus PFI) is a critical capability for production use. To build this estimator, we use the feature selection and ablation results presented in Table 3. Specifically, we fit a least-squares linear model using two candidate features derived from DCGM data and motivated by the memory-bound mechanism of power flexibility discussed above: DRAM copy product (the product of DRAM activity and memory copy utilization) and arithmetic intensity (the ratio of tensor pipe activity to DRAM activity). We evaluate each model using both leave-oneout (LOO) and leave-one-cluster-out (LOCO, holding out one architecture-task cluster at a time), reporting in-sample and out-of-sample R2 , RMSE, and Spearman ρ. For completeness, we also show the results of fitting a model on all DCGM metrics as well as ground-truth job metadata (architecture, task, and GPU count). We find that a simple linear model fit on DRAM copy product alone (shown in Fig. 7) provides the best performance across both LOO and LOCO, and we observe consistent R2 values between the validation data and LOCO analysis. Consistent with our observation that PFI appears to cluster by job type (Dense vs. MoE, pretraining vs. fine-tuning), the ability of DCGM signals like DRAM copy product to predict PFI is likely because these signals provide both a measurable proxy for job type and a means of rank-ordering the expected relative PFI of different jobs. Appendix D.5 provides further sensitivity analysis for this model. Given the limited number of observations (n = 25 samples of PFI derived from 131 runs), we cannot justify higher-order models (e.g., nonlinear relationships or using more DCGM metrics), but future work may expand to larger sample sizes and different functional forms.
6
Exploiting Power Flexibility Index
PFI is a control-oriented metric: our ultimate goal is to allow the infrastructure layer to maximize total tokens/sec throughput under a global power budget by protecting relatively inflexible jobs and assigning deeper power reductions to relatively flexible jobs. To demonstrate this capability, we construct simulated workloads by randomly sampling collections of jobs in our dataset and apply five different strategies to allocate power caps across these jobs: • Oracle: optimal allocation using perfect knowledge of each job’s throughput-power curve. • Equal weight: allocate power reductions proportionately to each job’s uncapped power. • MoE FT-weighted: based on equal weight but doubles the power reduction allocated to MoE fine-tuning jobs (the highest flexibility cluster). • PFI-aware: adjusts the proportional allocation using PFI, as described below. • Uncalibrated DCGM (DRAM-copy-product): allocate power reductions proportionately to each job’s DRAM copy product, showing the contrast between the calibrated PFI model and raw telemetry. 7
Performance gap vs. oracle (tok/s/job)
Random job mix
Production job mix
0 −1000 −2000 −3000 −4000 −5000 −6000
Oracle (perfect information) Equal weight DRAM-copy-product PFI-aware MoE-FT-weighted 5
10
Oracle (perfect information) Equal weight DRAM-copy-product PFI-aware MoE-FT-weighted 20
30
Power reduction (%)
5
10
20
Power reduction (%)
30
Figure 8: Difference in throughput for four allocation strategies compared against an oracle with perfect information of each job’s power-performance curve (less negative = closer to optimal) across a range of curtailment targets and workload mixes (random = jobs sampled at random from our dataset, production = jobs sampled according to the pretraining/fine-tuning mix from recent work [10]). i Denote the power reduction applied to job i as Pcurtail , the uncapped power and throughput of job i i i as Pmax and Tmax , and the PFI of job i as P F I i , then: i i Equal weight: Pcurtail ∝ Pmax
i PFI-aware: Pcurtail ∝
i Pmax P F Ii i Tmax
(2)
We simulate the performance of these strategies on 500 synthetic workloads in two scenarios, each with 100 jobs. In the “random” scenario, we sample jobs with replacement from our dataset. In the “production” scenario, we sample jobs according to the pretraining/fine-tuning ratio provided in recent work [10]. We estimate PFI using the model from §5, fitted to telemetry measured on the uncapped run of each job, mirroring the data that would be available from uncapped jobs running in production. We then allocate cluster-level GPU power reductions between 0% and 40% across jobs using each strategy and record the corresponding throughput reduction (simulated by interpolating the measured power-throughput curve for each job). Fig. 8 shows that the PFI-aware strategy yields equal or higher throughput than the equal weight strategy across all power reduction levels and both workload mix scenarios. While the random job mix includes a roughly equal mix of pretraining and fine-tuning, the production mix [10] includes a large number of small fine-tuning jobs (which tend to have high PFI) and a small number of large pretraining jobs (which tend to be lower PFI). The PFI-aware orchestration strategy is better able to exploit the differences in power elasticity between these job types, leading to substantially improved performance with the production job mix. Under a 30% power reduction on the production mix, PFI-aware power allocation recovers ∼1.5k tokens/s per job, 63% of the performance gap between an equal-weight allocation and an oracle with perfect information. These results show how PFI can enable power-aware orchestration of LLM training workloads, providing the foundation for data center operators to offer flexible SLAs while preserving token throughput.
7
Related Work
Hardware-level power management. GPU power capping and DVFS have been studied as efficiency knobs for over a decade. Tang et al. [27] sweep core and memory frequency on Pascal and Volta GPUs for convolutional networks and identify workload-specific optimal frequencies. Krzywaniak et al. [13] apply NVML power caps during CNN training on a single V100 or A100 and report 22 to 32% energy savings at small slowdowns. Zhao et al. [34] demonstrate that capping V100 power across a mixed HPC and AI workload can be done with minimal overall performance loss, and Costa et al. [6] provide similar results on an exascale GPU system. For LLM workloads, Patel et al. [19] characterize A100 cluster power for both training and inference and show the ability to increase capacity via power-capping-based oversubscription. Ujeniya et al. [28] compare power capping on H100 and H200 and show that memory bandwidth differences affect energy and performance tradeoff under power caps. The HPC community has explored related ideas on CPU-based systems [21]. Prior work establishes power capping as a viable control knob but does not provide a job-level characterization 8
of the power-performance tradeoff across architectures, cluster scales, and modern LLM training tasks. ML systems optimization. A large body of systems work studies ML throughput or energy optimization under fixed power budgets. The dominant building blocks include tensor and pipeline parallelism (Megatron-LM [25]), sharded optimizer states (ZeRO [24] and FSDP [35]), IO-aware kernels such as FlashAttention [7], and MoE training stacks [15, 8, 11, 9]. Complementary works such as Zeus [32] and Perseus [3] optimize energy–time tradeoffs for training jobs, while ML.ENERGY [4] benchmarks inference energy use and recommends energy-efficient deployment configurations. While these works provide tools for optimizing energy use, our goal in this paper is to provide a characterization of job-level power flexibility that allows a training cluster to dynamically reduce power demand during periods of grid stress (rather than finding an efficient static operating point). Energy-aware and carbon-aware AI orchestration. A related line of work seeks to reduce the energy or carbon footprint of large-scale computing through job scheduling and control. Carbonaware computing frameworks shift flexible workloads in time or location in response to real-time electricity carbon intensity and power availability [23, 1]. DynamoLLM dynamically adjusts instance counts, routing policies, and GPU frequencies to minimize energy and carbon cost while satisfying latency SLAs in LLM inference clusters [26]. By providing a systematic characterization of the power-throughput response of individual training jobs, our work complements these scheduler-level approaches to workload flexibility. Demand response and power-flexible AI infrastructure. Data center demand response shifts computing demand in response to grid conditions [30, 16, 33] or oversubscribes infrastructure through dynamic power provisioning [22]. POLCA [20] oversubscribes inference clusters using statistical headroom, while Wang et al. [29] shape AI cluster power profiles for grid response by varying GPU frequency. At the operator and grid layer, field demonstrations [5, 31], together with EPRI’s DCFlex initiative [14], have shown that GPU clusters can function as dispatchable grid resources. As with the energy-aware orchestration methods discussed above, our work complements these approaches by characterizing the job-level power-throughput response, which enables grid-responsive operation with minimal throughput loss. Positioning of this work. This paper fills an important gap by quantifying the throughput cost of power reductions for specific LLM training jobs. By introducing PFI and presenting the first (to our knowledge) systematic, large-scale characterization of power flexibility in LLM training, our work helps enable power-flexible data center operations by allowing schedulers to optimize power reductions to preserve cluster performance and maintain SLAs. In addition, by demonstrating that we can predict PFI from readily-available DCGM telemetry, we show that our approach is practically implementable and can measurably improve cluster performance on simulated workloads.
8
Limitations
Our characterization is scoped to NVIDIA H200 GPUs on a3-ultragpu-8g nodes, uses nvidia-smi -pl as the sole power-control mechanism, and covers pretraining and rank-32 LoRA SFT runs of 30 min each. The 32-GPU runs include only Qwen3-32B pretraining, since rank-32 LoRA SFT is rarely deployed at this scale and our 32-GPU capacity was constrained; broader scaling across architectures remains a natural extension. PFI is computed from aggregate GPU power only: host CPU/DRAM, NIC, PSU, fan, and rack-level cooling are not measured, and broadening the power signal in Eq. 1 as those streams become available is a direct extension. Each PFI data point represents a single power sweep; a subset of 14 configurations was repeated (35 runs total) and within those repeats steady-state aggregate GPU power and tokens/s agreed with median coefficient of variation 0.4% and 0.5%. The n = 25 PFI sample size still limits crossarchitecture generalization, which we mitigate via held-out validation models. The curtailment simulation in §6 uses the same power-throughput curves used to measure PFI; validation on out-ofsample workloads is left to future work with a larger dataset. Closed-loop scheduler validation of §6 and extensions to other accelerators (including heterogeneous clusters), RLHF, and inference are left to future work. 9
9
Conclusion
While AI data centers can operate as power-aware, grid-responsive loads, doing so while maximizing training throughput requires a principled understanding of how LLM training workloads respond to sustained reductions in available GPU power. In this work, we show that training-job power elasticity can be quantified, estimated at runtime, and exploited for intelligent power-aware workload orchestration. Broader impact Rapid growth in LLM training is placing significant new demands on electric power systems and is increasingly constrained by limited near-term power availability. By enabling data center operators to reduce power consumption when needed while maximizing throughput and minimizing disruption to customer workloads, this work supports a new class of power-flexible AI infrastructure that can improve utilization of constrained grid capacity, support faster AI deployment timelines, and reduce the infrastructure cost of large-scale AI expansion.
Acknowledgments and Disclosure of Funding We would like to acknowledge Fatih Acun and Brian Kulis at Emerald AI for their support and constructive feedback during the review phase.
10
References [1] T. Anderson, A. Belay, M. Chowdhury, A. Cidon, and I. Zhang. Treehouse: A case for carbon-aware datacenter software. SIGENERGY Energy Inform. Rev., 3(3):64–70, Oct. 2023. [2] A. Chowdhery, S. Narang, J. Devlin, et al. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023. arXiv:2204.02311. [3] J.-W. Chung, Y. Gu, I. Jang, L. Meng, N. Bansal, and M. Chowdhury. Reducing energy bloat in large model training. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, SOSP ’24, page 144–159, New York, NY, USA, 2024. Association for Computing Machinery. [4] J.-W. Chung, J. J. Ma, R. Wu, J. Liu, O. J. Kweon, Y. Xia, Z. Wu, and M. Chowdhury. The ml.energy benchmark: Toward automated inference energy measurement and optimization. In Advances in Neural Information Processing Systems, Datasets and Benchmarks Track, volume 38. Curran Associates, Inc., 2025. [5] P. Colangelo, A. K. Coskun, J. Megrue, C. Roberts, S. Sengupta, V. Sivaram, E. Tiao, A. Vijaykar, C. Williams, D. C. Wilson, B. Records, Z. MacFarland, D. Dreiling, N. Morey, A. Ratnayake, and B. Vairamohan. AI data centres as grid-interactive assets. Nature Energy, 11(2):254–261, Feb. 2026. [6] M. T. Costa, A. Georgiadou, J. B. White, B. V. Alvarez, J. Polo, W. Shin, P. O. A. Navaux, B. Messer, and A. F. Lorenzon. Characterizing the impact of gpu power management on an exascale system. In Proceedings of the SC ’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC Workshops ’25, page 1524–1533, New York, NY, USA, 2025. Association for Computing Machinery. [7] T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. In Advances in Neural Information Processing Systems, volume 35, pages 16344–16359. Curran Associates, Inc., 2022. [8] W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022. [9] T. Gale, D. Narayanan, C. Young, and M. Zaharia. Megablocks: Efficient sparse training with mixture-of-experts. In D. Song, M. Carbin, and T. Chen, editors, Proceedings of Machine Learning and Systems, volume 5, pages 288–304. Curan, 2023. [10] Q. Hu, Z. Ye, Z. Wang, G. Wang, M. Zhang, Q. Chen, P. Sun, D. Lin, X. Wang, Y. Luo, Y. Wen, and T. Zhang. Characterization of large language model development in the datacenter. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), pages 709–729, Santa Clara, CA, Apr. 2024. USENIX Association. [11] C. Hwang, W. Cui, Y. Xiong, Z. Yang, Z. Liu, H. Hu, Z. Wang, R. Salas, J. Jose, P. Ram, H. Chau, P. Cheng, F. Yang, M. Yang, and Y. Xiong. Tutel: Adaptive mixture-of-experts at scale. In Proceedings of Machine Learning and Systems, volume 5, pages 269–287. Curan, 2023. [12] International Energy Agency. Key Questions on Energy and AI. Technical report, IEA, Paris, Apr. 2026. [13] A. Krzywaniak, P. Czarnul, and J. Proficz. Gpu power capping for energy-performance tradeoffs in training of deep convolutional neural networks for image recognition. In Computational Science – ICCS 2022, pages 667–681. Springer International Publishing, 2022. [14] E. Lannoye. DCFlex flex MOSAIC framework. Electric Power Research Institute (EPRI), Mar. 2026. [15] D. Lepikhin, H. Lee, Y. Xu, et al. GShard: Scaling giant models with conditional computation and automatic sharding. In ICLR, 2021. arXiv:2006.16668. [16] Z. Liu, M. Lin, A. Wierman, S. Low, and L. L. H. Andrew. Greening geographical load balancing. IEEE/ACM Transactions on Networking, 23(2):657–671, 2015. 11
[17] T. H. Norris, T. Profeta, D. Patino-Echeverri, and A. Cowie-Haskell. Rethinking Load Growth: Assessing the Potential for Integration of Large Flexible Loads in US Power Systems. Technical Report NI R 25-01, Nicholas Institute for Energy, Environment & Sustainability, Duke University, Durham, NC, Feb. 2025. [18] NVIDIA Corporation. Data center GPU manager (DCGM). https://developer.nvidia. com/dcgm. Accessed 2026. [19] P. Patel, E. Choukse, C. Zhang, I. n. Goiri, B. Warrier, N. Mahalingam, and R. Bianchini. Characterizing power management opportunities for llms in the cloud. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, (ASPLOS 24), page 207–222, New York, NY, USA, 2024. Association for Computing Machinery. [20] P. Patel, C. Zhang, E. Choukse, et al. POLCA: Power oversubscription in LLM cloud providers. arXiv:2308.12908, 2024. [21] T. Patki, D. K. Lowenthal, A. Sasidharan, M. Maiterth, B. L. Rountree, M. Schulz, and B. R. de Supinski. Practical resource management in power-constrained, high performance computing. In Proceedings of the 24th International Symposium on High-Performance Parallel and Distributed Computing, HPDC ’15, page 121–132, New York, NY, USA, 2015. Association for Computing Machinery. [22] S. Pelley, D. Meisner, P. Zandevakili, T. F. Wenisch, and J. Underwood. Power routing: dynamic power provisioning in the data center. In Proceedings of the Fifteenth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS XV, page 231–242, New York, NY, USA, 2010. Association for Computing Machinery. [23] A. Radovanović, R. Koningstein, I. Schneider, B. Chen, A. Duarte, B. Roy, D. Xiao, M. Haridasan, P. Hung, N. Care, S. Talukdar, E. Mullen, K. Smith, M. Cottman, and W. Cirne. Carbonaware computing for datacenters. IEEE Transactions on Power Systems, 38(2):1270–1280, 2023. [24] S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16, 2020. [25] M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv:1909.08053, 2019. [26] J. Stojkovic, C. Zhang, Í. Goiri, J. Torrellas, and E. Choukse. Dynamollm: Designing llm inference clusters for performance and energy efficiency. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA), pages 1348–1362, 2025. [27] Z. Tang, Y. Wang, Q. Wang, and X. Chu. The impact of gpu dvfs on the energy and performance of deep learning: an empirical study. In Proceedings of the Tenth ACM International Conference on Future Energy Systems, e-Energy ’19, page 315–325, New York, NY, USA, 2019. Association for Computing Machinery. [28] A. Ujeniya, J. Eitzinger, G. Hager, et al. Architectural trade-offs in the energy-efficient era: A comparative study of power-capping NVIDIA H100 and H200. arXiv:2604.11391, 2026. [29] Y. Wang, Q. Guo, and M. Chen. Providing load flexibility by reshaping power profiles of large language model workloads. Advances in Applied Energy, 19:100232, 2025. [30] A. Wierman, Z. Liu, I. Liu, and H. Mohsenian-Rad. Opportunities and challenges for data center demand response. In International Green Computing Conference, pages 1–10, 2014. [31] C. Williams, P. Colangelo, A. Coskun, et al. Power-flexible AI factories: A UK-first demonstration of grid-responsive AI infrastructure. Whitepaper, Emerald AI, EPRI, National Grid, and Nebius, Mar. 2026. 12
[32] J. You, J.-W. Chung, and M. Chowdhury. Zeus: Understanding and optimizing GPU energy consumption of DNN training. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 119–139, Boston, MA, Apr. 2023. USENIX Association. [33] Y. Zhang, D. C. Wilson, I. C. Paschalidis, and A. K. Coskun. Hpc data center participation in demand response: An adaptive policy with qos assurance. IEEE Transactions on Sustainable Computing, 7(1):157–171, 2022. [34] D. Zhao, S. Samsi, J. McDonald, B. Li, D. Bestor, M. Jones, D. Tiwari, and V. Gadepally. Sustainable supercomputing for ai: Gpu power capping at hpc scale. In Proceedings of the 2023 ACM Symposium on Cloud Computing, SoCC ’23, page 588–596. ACM, Oct. 2023. [35] Y. Zhao, A. Gu, R. Varma, et al. PyTorch FSDP: Experiences on scaling fully sharded data parallel. Proc. VLDB, 2023. arXiv:2304.11277.
13
A
Alternative definitions of the power flexibility index (PFI)
We consider three candidate definitions of PFI and discuss their tradeoffs. Ratio (adopted definition)
Define PFI at each power cap,
1 − Pi /Pmax , (3) 1 − Ti /Tmax PN and take the average over measured caps: P F I = N1 i=1 P F Ii (excluding the uncapped point). This definition is simple to compute and has a direct interpretation as the ratio of fractional power reduction to fractional throughput reduction. Its main limitation is sensitivity to the choice of sweep points: if different jobs are measured at different power caps, the averages are not directly comparable. However, Appendix D.3 shows that changing the sweep grid does not change the PFI results significantly (the relative ordering of PFI is preserved under different sweep grids). P F Ii =
Area between curves Define PFI as the area between the normalized power–throughput curve and the unit-slope diagonal, Z 1 T P dP PFI = − , (4) T P P max max max 0 approximated with the trapezoid rule: PFI ≈
N X 1 1 [(Ti + Ti−1 )(Pi − Pi−1 )] − , 2Tmax Pmax i=1 2
(5)
with boundary condition (P0 , T0 ) = (0, 0). This definition is robust to the choice of sample points, but it is less intuitive than the ratio definition. Constant-elasticity model Fit the log-linear model log T = α log P + log k and define P F I = 1 − α, so that the implied power–throughput relationship is T = kP 1−P F I . Under this parameterization, P F I = 0 corresponds to throughput proportional to power (fully inflexible), and P F I = 1 corresponds to throughput independent of power (fully flexible). The model has a single degree of freedom per job, making it well-suited to sweeps with few power caps; however, we found that this constant-elasticity model was a poor fit to our observed data, as shown in Fig. 9. Empirically, we find that all three options are highly correlated (see Fig. 10), but we select the proposed ratio definition for two reasons. First, it has a direct interpretation as ratio of power reduction to throughput loss, and the physical interpretation of the other two metrics is less clear. Second, it provides a clear segmentation between power inelastic (PFI > 1) and elastic (PFI ≤ 1), with a dynamic range (≈ [1, 2.5]). The area definition also segments (at zero), but we found that it has a very narrow range (≈ [−0.1, 0.1]).
T = k Pα fit Observed
0.2
log(T / Tmax )
0.0 −0.2 −0.4 −0.6 −0.8 −1.0 −1.0
−0.8
−0.6
−0.4
log(P / Pmax )
−0.2
0.0
Figure 9: Comparison of constant-elasticity fit with the observed power-performance curve for gpt-oss-20b fine-tuning on 8×H200s.
14
0.08 0.06
PFIarea
0.04 0.02 0.00 −0.02
PFICE
−0.1 −0.2 −0.3 −0.4
1.5
2.0
2.5
PFIratio (adopted)
−0.025 0.000
0.025
PFIarea
0.050
0.075
−0.4
−0.3
PFICE
−0.2
−0.1
Figure 10: Comparison of the empirical distribution of the three alternative PFI definitions on the training set cohort.
B
Detailed experimental setup
B.1
Infrastructure
We ran all experiments on Google Kubernetes Engine (GKE) using a3-ultragpu-8g nodes with 8× NVIDIA H200 GPUs (141 GiB HBM3e, 700 W TDP), fourth-generation NVLink, and a CX-7 RoCE fabric exposing eight GPUDirect-RDMA NICs per node over an MTU-8896 multi-VPC topology. Multi-node runs (16- and 32-GPU) are scheduled through Kueue with Dynamic Workload Scheduler (DWS) Flex-start or Spot capacity. We provisioned clusters in multiple GCP regions based on availability of Flex-start and Spot allocations.
B.2
Software stack
Workers ran a Docker image (rayproject/ray:2.53.0-py312-cu128) with Ray 2.53.0, PyTorch 2.10 (CUDA 12.8), Transformers ≥ 5.2, Accelerate ≥ 1.10, PEFT ≥ 0.17, FlashAttention 2.8.1 (cu12/torch2.10 prebuilt wheel), bitsandbytes ≥ 0.43.3, and torchtitan 0.2.2. Distributed training uses Ray Train’s TorchTrainer with NCCL backend; FSDP (full-shard, auto-wrap by transformer-decoder layer class) is used for dense jobs and for all LoRA SFT jobs. MoE pretraining is dispatched to a torchtitan subprocess that drives FSDP2 with native Expert Parallelism. We use a flat parallelism mesh: data_parallel_shard_degree equals world_size (FSDP shards across every rank), with expert parallelism nested inside the FSDP group at degree EP. GPU telemetry is harvested by an nvcr.io/nvidia/k8s/dcgm-exporter:3.3.7-3.5.0-ubuntu 22.04 sidecar deployed alongside every worker pod and accessed from the Ray container. Prior to collecting data, we verified that logging telemetry using DCGM does not materially affect job performance. 15
B.3
Models and datasets
Pretraining and SFT runs use FineWeb-Edu and UltraChat-200K, respectively, pre-tokenized into Ray Data shards in region-local object storage. gpt-oss-20b is dequantized from MXFP4-packed expert weights to BF16. We use the AdamW optimizer with default parameters. B.4
Power control mechanism
Per-GPU power caps are set using nvidia-smi -i <local_rank> -pl <watts>, called at the start of every run (after torch.cuda.set_device but before any CUDA context allocation). The cap is swept in 100 W increments from 200 W to 700 W on single-node (8-GPU) groups; the 16- and 32-GPU groups sweep 200 W, 500 W, 600 W, and 700 W only. We verify that each power cap was set correctly by reading DCGM_FI_DEV_POWER_MGMT_LIMIT and DCGM_FI_DEV_ENFORCED_POWER_LIMIT from DCGM during the steady-state phase of each run. B.5
Performance data collection
GPU metrics are scraped from the DCGM sidecar at 5 s intervals by a single local_rank=0 worker per node. System metrics (CPU, RSS, disk and network bytes) are sampled in-process via psutil, also every 5 s. Training metrics are logged on every step by a custom TelemetryCallback and include global step, loss, learning rate, observed tokens-per-second, achieved TFLOPS, and Model FLOPs Utilization (MFU). RDMA counters from the kernel’s InfiniBand interface are appended on RoCE-equipped nodes when present. B.6
Measurement protocol
Each configuration is run for 30 minutes. On the HuggingFace Trainer path (dense and LoRA jobs), evaluation is triggered at 10 minutes and a single distributed-checkpoint write at 20 minutes via an all-reduce-synchronized timer callback. The torchtitan EP path triggers eval and checkpoint after fixed step counts chosen so that both events occur at similar times to the dense/LoRA jobs (10 and 20 minutes, respectively). To isolate steady state from transients, we partition traces into three phases: warmup (frame-buffer usage below 10 GiB), startup (weights loaded but gpu_util = 0), and training (sustained gpu_util ≥ 50%). All aggregates are reported over the training phase only. Per-step throughput, MFU, and per-GPU power are aggregated as the mean of the top 10% of samples within the training phase to exclude checkpoints and evals. Runs at different power caps or hyperparameter values are executed sequentially on a single cluster after a 60 second delay, after verifying that GPU memory has been released. B.7
Compute resources
The 131 headline runs and 24 held-out validation runs that produce the figures and tables in this paper consume approximately 1,000 H200 GPU-hours, broken down by cluster size in Table 4. Table 4: Compute for paper-included runs. Wall-clock figures are the actual DCGM-bracketed durations from the master telemetry frame, not the 30 min training cap.
GPUs / cluster
Runs
Mean wall-clock (min)
GPU-hours
8 16 32
120 19 16
34.9 36.8 38.1
558 186 325
Total
155
35.6
∼1,069
The full project consumed additional H200 compute due to preliminary sweeps, hyperparameter tuning, and reruns of jobs that failed mid-training. Across all six profiling regions, 410 distinct H200 submissions produced DCGM telemetry and together consumed approximately 1,725 GPU-hours (Table 5). The 155 reported runs are roughly one third of those submissions and account for 62% of H200 GPU-hours. 16
Table 5: Compute resources used for this project. Cluster shape
DCGM-active runs
GPU-hours
H200, 8 GPUs H200, 16 GPUs H200, 32 GPUs
324 53 33
Total H200 (DCGM-active) Reported in paper
410 155
∼969 ∼272 ∼485
H100, 8 GPUs (reported in Appendix D.2)
34
∼1,725 ∼1,069 ∼149
Table 6: Power Flexibility Index for each of the headline sweep groups on H200. Each group runs a single job (fixed model, task, cluster size, and training configuration) at multiple power caps; # runs is the number of power-caps measured. 95% CI is estimated from 10,000 Monte Carlo samples for each run, with power and throughput measurements resampled from the empirical noise distribution. Job Dense FT H200x16 (Llama-3.1-70B) Dense FT H200x8 (Llama-3.1-70B, no compile) Dense FT H200x8 (Llama-3.1-70B, compile) Dense FT H200x16 (Qwen3-32B) Dense PT H200x8 (Llama-3.1-70B, compile) Dense PT H200x32 (Qwen3-32B, seq4096/bs8) Dense PT H200x32 (Qwen3-32B, seq4096/bs4) Dense PT H200x8 (Llama-3.1-70B, no compile) Dense PT H200x32 (Qwen3-32B, seq8192/bs8) Dense PT H200x8 (Qwen3-32B, no compile) Dense PT H200x32 (Qwen3-32B, seq8192/bs4) Dense PT H200x16 (Llama-3.1-70B) Dense PT H200x16 (Qwen3-32B) Dense PT H200x8 (Qwen3-32B, compile) MoE FT H200x8 (gpt-oss-20b, seq2048/bs8) MoE FT H200x8 (Qwen3-30B-A3B) MoE FT H200x16 (Qwen3-30B-A3B) MoE FT H200x8 (gpt-oss-20b, seq4096/bs8) MoE PT H200x8 (gpt-oss-20b, EP1, no compile) MoE PT H200x8 (gpt-oss-20b, EP1, compile) MoE PT H200x8 (Qwen3-30B-A3B, EP1) MoE PT H200x8 (gpt-oss-20b, EP8, compile) MoE PT H200x8 (gpt-oss-20b, EP4, compile) MoE PT H200x8 (Qwen3-30B-A3B, EP8) MoE PT H200x8 (gpt-oss-20b, EP8, no compile)
C
# runs
PFI
95% CI
3 6 6 4 6 4 4 6 4 6 4 4 4 6 6 6 4 6 6 6 6 6 6 6 6
1.10 1.22 1.27 1.52 1.17 1.23 1.25 1.26 1.27 1.30 1.30 1.34 1.47 1.51 1.79 2.00 2.20 2.67 1.24 1.35 1.36 1.43 1.45 1.52 1.57
[1.096, 1.102] [1.218, 1.225] [1.264, 1.269] [1.519, 1.527] [1.171, 1.177] [1.226, 1.230] [1.248, 1.251] [1.254, 1.260] [1.265, 1.270] [1.295, 1.301] [1.300, 1.305] [1.335, 1.342] [1.470, 1.478] [1.494, 1.518] [1.750, 1.833] [1.976, 2.028] [2.146, 2.244] [2.615, 2.729] [1.209, 1.268] [1.335, 1.358] [1.349, 1.370] [1.407, 1.454] [1.435, 1.476] [1.456, 1.597] [1.557, 1.576]
Data tables
Tables 6 and 7 provide the raw PFI data for each sweep group in the headline and validation sets, respectively, including a 95% confidence interval derived using Monte Carlo resampling from the empirical noise distributions derived from intra-run time-series variation (median standard error of the mean 0.04% for power and 0.05% for throughput). In addition, Figs. 11, 12, and 13 includes the raw and normalized power-performance curves for each job in Tables 6 and 7. 17
Dense
MoE
Dense FT H200x16 (Llama-3.1-70B)
Pretraining
Fine-tuning
Held out
Dense FT H200x8 (Llama-3.1-70B, no compile)
PFI = 1.10
PFI = 1.22
20
0.8 0.6 0.4
0
60
T /Tmax
40
1.0
Tput (k tok/s)
60
T /Tmax
Tput (k tok/s)
1.0
40 20
10
20
0.25
Dense FT H200x8 (Llama-3.1-70B, compile)
0.4
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.27
Dense FT H200x16 (Qwen3-32B)
0.25
20
0.8 0.6 0.4
0
60 40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
Dense PT H200x8 (Llama-3.1-70B, compile)
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.17
Dense PT H200x32 (Qwen3-32B, seq4096/bs8)
0.25
20
1.00
PFI = 1.23
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
1.00
PFI = 1.52
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
Tput (k tok/s)
60
0.50
P/Pmax
1.0
40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
Dense PT H200x32 (Qwen3-32B, seq4096/bs4)
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.25
Dense PT H200x8 (Llama-3.1-70B, no compile)
0.25
20
1.00
PFI = 1.26
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
0.6
0
Cluster power (kW)
40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
Dense PT H200x32 (Qwen3-32B, seq8192/bs8)
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.27
Dense PT H200x8 (Qwen3-32B, no compile)
0.25
20
1.00
PFI = 1.30
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
0.8
40 20
0.8 0.6 0.4
0 10
20
Cluster power (kW)
0.25
0.50
0.75
1.00
10
20
Cluster power (kW)
P/Pmax
0.25
0.50
0.75
1.00
P/Pmax
Figure 11: Power-throughput curves (raw and normalized) for each headline and validation run on H200.
18
Dense
MoE
Dense PT H200x32 (Qwen3-32B, seq8192/bs4)
Pretraining
Fine-tuning
PFI = 1.30
Held out
Dense PT H200x16 (Llama-3.1-70B)
PFI = 1.34
20
0.8 0.6 0.4
0
60
T /Tmax
40
1.0
Tput (k tok/s)
60
T /Tmax
Tput (k tok/s)
1.0
40 20
10
20
0.25
Dense PT H200x16 (Qwen3-32B)
0.4
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.47
Dense PT H200x8 (Qwen3-32B, compile)
0.25
20
0.8 0.6 0.4
0
60 40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
MoE FT H200x8 (gpt-oss-20b, seq2048/bs8)
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.79
MoE FT H200x8 (Qwen3-30B-A3B)
0.25
20
1.00
PFI = 2.00
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
1.00
PFI = 1.51
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
Tput (k tok/s)
60
0.50
P/Pmax
1.0
40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
MoE FT H200x16 (Qwen3-30B-A3B)
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 2.20
MoE FT H200x8 (gpt-oss-20b, seq4096/bs8)
0.25
20
1.00
PFI = 2.67
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
0.6
0
Cluster power (kW)
40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
MoE PT H200x8 (gpt-oss-20b, EP1, no compile)
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.24
MoE PT H200x8 (gpt-oss-20b, EP1, compile)
0.25
20
1.00
PFI = 1.35
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
0.8
40 20
0.8 0.6 0.4
0 10
20
Cluster power (kW)
0.25
0.50
0.75
1.00
10
20
Cluster power (kW)
P/Pmax
0.25
0.50
0.75
1.00
P/Pmax
Figure 12: Power-throughput curves (raw and normalized) for each headline and validation run on H200
(continued).
19
Dense
MoE
MoE PT H200x8 (Qwen3-30B-A3B, EP1)
Pretraining
Fine-tuning
Held out
MoE PT H200x8 (gpt-oss-20b, EP8, compile)
PFI = 1.36
PFI = 1.43
20
0.8 0.6 0.4
0
60
T /Tmax
40
1.0
Tput (k tok/s)
60
T /Tmax
Tput (k tok/s)
1.0
40 20
10
20
0.25
MoE PT H200x8 (gpt-oss-20b, EP4, compile)
0.4
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.45
MoE PT H200x8 (Qwen3-30B-A3B, EP8)
0.25
20
1.00
PFI = 1.52
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
Tput (k tok/s)
60
0.50
P/Pmax
1.0
40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
MoE PT H200x8 (gpt-oss-20b, EP8, no compile)
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.57
Dense FT H200x8 (MistralSmall-24B-Base-2501) [held out]
0.25
20
1.00
PFI = 1.31
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
0.6
0
Cluster power (kW)
40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
Dense PT H200x8 (MistralSmall-24B-Base-2501) [held out]
0.50
0.75
1.00
10
20
P/Pmax
Cluster power (kW)
PFI = 1.18
MoE FT H200x8 (gemma-4-26B-A4B) [held out]
0.25
20
1.00
PFI = 2.13
0.8 0.6 0.4
0
60
T /Tmax
40
0.75
1.0
Tput (k tok/s)
T /Tmax
60
0.50
P/Pmax
1.0
Tput (k tok/s)
0.8
40 20
0.8 0.6 0.4
0 10
20
0.25
Cluster power (kW)
0.50
0.75
1.00
10
MoE PT H200x8 (DeepSeek-V2-Lite) [held out]
20
Cluster power (kW)
P/Pmax
0.25
0.50
0.75
1.00
P/Pmax
PFI = 1.67
60
T /Tmax
Tput (k tok/s)
1.0
40 20
0.8 0.6 0.4
0 10
20
Cluster power (kW)
0.25
0.50
0.75
1.00
P/Pmax
Figure 13: Power-throughput curves (raw and normalized) for each headline and validation run on H200
(continued).
20
Table 7: Power Flexibility Index for the 4 held-out validation sweep groups on H200. Job
# runs
PFI
6 6 6 6
1.31 1.18 2.13 1.67
Dense FT H200x8 (Mistral-Small-24B-Base-2501) Dense PT H200x8 (Mistral-Small-24B-Base-2501) MoE FT H200x8 (gemma-4-26B-A4B) MoE PT H200x8 (DeepSeek-V2-Lite)
Table 8: PFI for six matched sweep groups on H100 and H200. All groups use 8 GPUs. Rank ordering is nearly preserved across generations (Spearman ρ = 0.943). Arch/Task
Model
EP
Runs (H100)
PFI (H200)
PFI (H100)
Dense PT Dense FT Dense FT MoE PT MoE PT MoE FT
Mistral-Small-24B-Base-2501 Llama-3.1-70B Mistral-Small-24B-Base-2501 gpt-oss-20b DeepSeek-V2-Lite Qwen3-30B-A3B
1 1 1 8 8 1
6 4∗ 6 6 6 6
1.18 1.22 1.31 1.43 1.67 2.00
1.26 1.19 1.33 1.38 1.51 1.93
∗ This group contains four runs rather than six: the 500 W and 200 W runs failed with transient errors and could not be
re-run within the available allocation.
D
Additional results
D.1
All DCGM metrics
For completeness, Fig. 14 extends Figs. 5 to all measured DCGM metrics. D.2
H100 validation
Our main characterization is confined to H200 GPUs. To test whether the qualitative patterns we report depend on that specific platform, we ran a small set of matched sweeps on H100 GPUs: six sweep groups (34 runs total), with at least one group in each of the four architecture×task cells (dense PT, dense FT, MoE PT, MoE FT). Table 8 reports these results along with those of the matched H200 sweeps. The rank ordering of PFI is largely preserved between H100 and H200 (Spearman ρ = 0.943), and we observe that the DCGM-to-PFI model fit on H200 data achieves R2 = 0.56 on the H100 data. The largest absolute difference in PFI between H100 and H200 is 0.16 and the median difference is 3.5%, which is small compared to the between-job PFI range of 1.10 to 2.67 on H200. We observe no systematic direction to the difference in PFI between platforms. The sample size presented here is not sufficient to claim generalizable results on H100 workloads. D.3
Sensitivity of PFI to the Power-Cap Sweep Grid
PFI averages the ratio of fractional power reduction to fractional throughput reduction over the sampled power caps, so its value depends on which caps are sampled. Because our 8-GPU groups sweep six caps (200–700 W) while the 16- and 32-GPU groups sweep only four (200, 500, 600, 700 W), comparing PFI between these groups confounds scale with the sweep grid. Table 9 quantifies the size of this effect by providing PFI for every 8-GPU group recomputed using only the four caps present in the coarse grid used for 16- and 32-GPU groups (excluding the 300 W and 400 W measurements). Restricting to the coarse grid increases PFI for all but one group (median increase 0.154, or 10.7%) but does not materially change the ordering of groups (Spearman ρ = 0.940). We draw two conclusions. First, absolute PFI values are comparable only between groups measured on the same grid, and we accordingly avoid quantitative comparisons of PFI magnitude between the 8-GPU groups and the 16-/32-GPU groups. Second, the rank ordering that our allocation policy relies on is robust to this choice of grid, so PFI remains usable as an ordinal signal across groups measured at different resolutions. While the PFI predictor presented in §5 is trained on the full dataset, our 21
Table 9: PFI for 8-GPU sweep groups computed on the full six-cap grid (200–700 W) and on the coarse four-cap grid (200, 500, 600, 700 W) used for the 16- and 32-GPU groups. Sweep group
PFI (6 caps)
PFI (4 caps)
∆
∆ (%)
Dense FT H200x8 (Llama-3.1-70B, no compile) Dense FT H200x8 (Llama-3.1-70B, compile)
1.22 1.27
1.37 1.44
+0.145 +0.177
+11.9 +14.0
Dense PT H200x8 (Llama-3.1-70B, compile) Dense PT H200x8 (Llama-3.1-70B, no compile) Dense PT H200x8 (Qwen3-32B, no compile) Dense PT H200x8 (Qwen3-32B, compile)
1.17 1.26 1.30 1.51
1.33 1.43 1.42 1.67
+0.155 +0.177 +0.121 +0.163
+13.2 +14.1 +9.3 +10.8
MoE FT H200x8 (gpt-oss-20b, seq2048/bs8) MoE FT H200x8 (Qwen3-30B-A3B) MoE FT H200x8 (gpt-oss-20b, seq4096/bs8)
1.79 2.00 2.67
1.98 2.22 3.32
+0.189 +0.221 +0.646
+10.5 +11.1 +24.2
MoE PT H200x8 (gpt-oss-20b, EP1, no compile) MoE PT H200x8 (gpt-oss-20b, EP1, compile) MoE PT H200x8 (Qwen3-30B-A3B, EP1) MoE PT H200x8 (gpt-oss-20b, EP8, compile) MoE PT H200x8 (gpt-oss-20b, EP4, compile) MoE PT H200x8 (Qwen3-30B-A3B, EP8) MoE PT H200x8 (gpt-oss-20b, EP8, no compile)
1.24 1.35 1.36 1.43 1.45 1.52 1.57
1.16 1.40 1.40 1.57 1.65 1.67 1.66
−0.082 +0.055 +0.037 +0.144 +0.194 +0.153 +0.092
−6.6 +4.1 +2.7 +10.1 +13.3 +10.1 +5.8
end-to-end results on power allocation in §6 demonstrate that this predictor is still useful in practice (likely because PFI remains correlated across different grids). D.4
Statistical tests for group difference
We test whether PFI differs across task/architecture regimes using the n = 25 headline sweep groups. Because dense pretraining and dense LoRA SFT are not separable in our data, we pool them and compare three groups: dense, MoE pretraining, and MoE fine-tuning. We use the Kruskal–Wallis H test followed by Dunn’s post-hoc test over all three pairs with Holm–Bonferroni correction (m = 3). The Kruskal–Wallis test rejects equality of the three groups (H = 12.37, df = 2, p = 0.0021). For Dunn’s test, only the dense vs. MoE fine-tuning difference survives p-value correction. MoE pretraining is not separable from either dense or MoE fine-tuning at this sample size. These results indicate that our sample size can statistically distinguish a single high-PFI cluster (MoE fine-tuning) from a continuum of lower-PFI jobs. Tables 10 and 11 provide group medians and pairwise comparisons, respectively. Table 10: PFI by task/architecture regime across the 25 headline sweep groups. Group
n
Median PFI
Mean PFI
Dense (PT + LoRA SFT) MoE pretraining MoE fine-tuning
14 7 4
1.27 1.43 2.10
1.30 1.42 2.17
Table 11: Dunn’s post-hoc test, all three pairs, Holm–Bonferroni corrected (m = 3). ∗ denotes pHolm < 0.05. Comparison Dense vs. MoE fine-tuning MoE pretraining vs. MoE fine-tuning Dense vs. MoE pretraining
n (14, 4) (7, 4) (14, 7)
22
z
praw
pHolm
Sig.
−3.44 −1.90 −1.64
0.0006 0.0568 0.1020
0.0017 0.1137 0.1137
∗
Table 12: Spearman rank correlation between the DRAM copy product and PFI as the k highestleverage sweep groups are removed, over the 25 headline groups. Leverage is computed from the predictor only. Groups are removed in leverage order: Qwen3-30B-A3B (MoE, fine-tuning, 8 GPU), gpt-oss-20b (MoE, fine-tuning, 8 GPU), Qwen3-30B-A3B (MoE, fine-tuning, 16 GPU), Qwen3-32B (dense, pretraining, 8 GPU), Llama-3.1-70B (dense, pretraining, 16 GPU). k dropped
n
ρ
0 1 2 3 4 5
25 24 23 22 21 20
0.535 0.477 0.406 0.321 0.438 0.463
Table 13: Spearman rank correlation between the DRAM copy product and PFI within architecture subgroups, over the 25 headline sweep groups. Subgroup Dense only MoE only All except MoE fine-tuning All groups
D.5
ρ
n
−0.178 0.509 0.347 0.535
14 11 21 25
Sensitivity analysis of fit PFI model
Table 12 provides a sensitivity analysis for the DRAM copy product-based predictor, dropping high-leverage sweep groups and providing the Spearman correlation between DRAM copy product and PFI at each step. Table 13 splits the correlation by architecture. We find no correlation within the dense subset, a positive but not significant correlation within the MoE subset (ρ = 0.509, n = 11, p = 0.11), and a significant correlation over all groups (ρ = 0.535, n = 25, p = 0.006). As a result, we find that contrast between architecture-task clusters drives correlation between PFI and GPU telemetry, rather than variation within clusters.
23
1.5
2.5
2.5
2.0
2.0
2.0
1.5
1.5
ρ=-0.27 (n=25)
2.0
2.0
2.0
2.0
1.5
PFI
2.5
PFI
2.5
PFI
1.5 ρ=-0.13 (n=25)
40 60 Mean mem-copy util (%)
0
ρ=-0.13 (n=25)
100 NVLink TX (GB/s)
0
2.5
2.0
2.0
2.0
2.0
1.5
1.5
ρ=-0.30 (n=25)
0.0
PFI
2.5
PFI
2.5
PFI
2.5
1.5
0
ρ=0.26 (n=25)
5 PCIe RX (GB/s)
50
ρ=0.79** (n=25)
100 VRAM used (GiB)
1600 1800 SM clock (MHz) 2.5
2.0
2.0
2.0
2.0
1.0
1.5 ρ=-0.37 (n=25)
60 GPU temp (°C)
1.5
1.5 ρ=0.54* (n=25)
70
PFI
2.5
PFI
2.5
PFI
2.5
1.5
ρ=0.00 (n=25)
10 20 30 DRAM × copy product
1.0
600 800 1000 Power per SM (W/%)
2.5
2.0
2.0
2.0
2.0
1.5
ρ=0.63** (n=25)
ρ=-0.67** (n=25)
0.4 0.6 DRAM/SM ratio
0.25 0.50 0.75 Tensor/SM ratio 2.5
2.0
2.0
PFI
2.5
1.5
PFI
2.5
PFI
2.5
1.5
ρ=-0.61** (n=25)
1 2 3 Arith. intensity (tensor/DRAM)
2.5
1.5
100 NVLink RX (GB/s)
1.5
ρ=-0.17 (n=25)
2.5 5.0 PCIe TX (GB/s)
600 Mean power (W)
1.5
ρ=0.51* (n=25)
90 Mean GPU util (%)
PFI
PFI
500
2.5
ρ=-0.03 (n=25)
PFI
ρ=-0.46* (n=25)
0.2 0.4 Mean DRAM active (%)
2.5
80
PFI
ρ=0.54* (n=25)
0.25 0.50 Mean tensor active (%)
1.5
PFI
1.5
ρ=-0.64** (n=25)
0.6 0.8 Mean SM active (%)
PFI
PFI
PFI
PFI
2.0
2.5
PFI
Dense MoE Validation
2.5
1.5 ρ=0.55* (n=25)
0.4 0.6 Copy/util fraction
ρ=-0.06 (n=25)
0 200 400 NVLink per SM (GB/s/%)
1.5 ρ=-0.26 (n=25)
0 500 NVLink per DRAM (GB/s/%)
ρ=0.79** (n=25)
0.6 0.7 Clock fraction of boost
Figure 14: PFI vs. all DCGM GPU metrics and composite features. Dashed lines show ordinary least-squares
fit; Spearman ρ and significance are annotated. ∗ q < 0.05, ∗∗ q < 0.01, where q is the Benjamini–Hochberg FDR-corrected p-value across all metrics in this figure.
24