arXiv:2605.16255v1 [cs.DC] 15 May 2026
Designing Datacenter Power Delivery Hierarchies for the AI Era Grant Wilkins∗
Fiodar Kazhamiaka
Alok Gautam Kumbhare
Stanford University Stanford, CA, USA
Microsoft Azure Research Redmond, WA, USA
Microsoft Azure Research Redmond, WA, USA
Chaojie Zhang
Ricardo Bianchini
Microsoft Azure Research Redmond, WA, USA
Microsoft Azure Research Redmond, WA, USA
Abstract Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major challenge for datacenter power delivery designers. As power densities increase, a datacenter designed for a different target density may strand power, i.e., may be unable to use all the power that its delivery hierarchy has provisioned. Designs must remain efficient over long datacenter lifetimes and multiple hardware generations. Power utilization is particularly important as grid power capacity is a scarce resource in the AI era. Designing an efficient power delivery hierarchy for the long run is difficult because rack placement feasibility, workload impact, and cost depend jointly on electrical topology, deployment granularity, placement policy, power oversubscription, and workload mix. Moreover, each of these factors evolve over time, have inter-dependencies across multiple resource dimensions, and generally do not lend themselves to closed-form analysis. To address this challenge, we develop a framework for evaluating datacenter power delivery designs using throughput, power, and cost metrics over realistic arrival, oversubscription, and decommissioning sequences. The framework combines projection models for GPU, compute, and storage deployments with operational factors grounded in production data from Azure. Our results show that multi-resource stranding materially changes deployable capacity, effective capital expenditure, and delivered performance, and quantify how rising density from rack- and pod-scale AI systems shapes these outcomes. For AI datacenter design, the relevant planning objective is not installed megawatts, but deployable capacity over time.
1
Figure 1. P99 of rack power density since 2020 for datacenter deployments, showing distinct accelerator generations and a widening gap between GPU and non-GPU power density. Density is normalized to the maximum P99 value observed in each quarter at Azure.
project rack- and pod-scale systems approaching 1 MW in a few years [34, 38, 55]. Figure 1 shows the same shift in production deployments at Azure, where accelerator rack power is rising much faster than general compute and storage racks. These denser systems also tighten infrastructure coupling, as power, cooling, and local interconnect increasingly scale together [12, 36, 37]. Datacenter halls are built with a fixed power delivery hierarchy that is chosen at construction time and persists for 15–25 years across multiple hardware generations [16, 31]. Power delivery design is thus a long-term commitment. A hierarchy chosen for today’s hardware must remain efficient for substantially denser systems that arrive years later. Worse, rack deployment feasibility is hierarchical, not siteor even hall-wide. A rack or pod must satisfy capacity and redundancy constraints at every level of the power delivery path [7, 54, 59]. As deployment quanta grow to consume a meaningful fraction of UPS, busbar, PDU, and cooling capacity, aggregate provisioned power becomes a misleading proxy for what the hall can still admit. A hall can retain substantial available power and still reject the next deployment because the remaining capacity is fragmented across domains. We refer to this unusable capacity as stranded.
Introduction
Motivation. Demand for generative AI is driving a rapid datacenter buildout. These new AI datacenters are substantially different than their cloud datacenter predecessors with respect to their power delivery hierarchies. In the web-services era, rack power density was often below 20 kW [5]. Today, AI accelerator racks exceed 150 kW, and public roadmaps ∗ Work partly completed while an intern at Microsoft Azure Research.
1
30
Block (8+2) Low Rack Power Growth
Distributed (4N/3) Low Rack Power Growth
20
401T 132T
10
19T 5T
0 12.0
12.5
13.0
13.5
14.0
14.5
Effective Cost of DC ($/W)
15.0
multi-resource placement constraints and combines infrastructure cost models with comparative throughput models for emerging accelerator generations. Using our framework, we evaluate many variants of two common classes of power delivery designs: distributed redundant and block redundant. Our results show that designs with similar provisioned capacity can diverge significantly in deployable capacity, effective cost, and delivered throughput. Figure 2 shows how important metrics such as throughput per watt and cost per watt can differ across a variety of power delivery designs, rack power projections, and LLM inference workloads: throughput per watt varies by more than 20x while cost per watt varies by more than 20%. The impact of these ranges is measured in billions of dollars. In summary, we make the following contributions:
Model Param Count
Power-Normalized Throughput (tok/s/W)
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
0.6T
Figure 2. Each marker represents a combination of datacenter design, workload, and rack density projections. Colors represent different LLM mixture of experts models being served across the fleet. Each combination is analyzed with our framework, and is compared on throughput of LLM inference per watt versus effective fleet cost. Highlighted points represent diverse workloads, labels describe the points’ design and power projection.
• We identify deployable power capacity over time, rather than installed MW or first-order $/W, as the key metric for evaluating AI datacenter power delivery designs. • We develop an evaluation framework that quantifies this metric under hierarchical multi-resource placement constraints with arrivals, oversubscription, and decommissioning. • We show that redundancy topology changes how capacity becomes stranded as deployment quanta grow: distributed-redundant designs degrade through fragmented headroom, and block-redundant through coarse capacity granularity. • We quantify when larger GPU pods improve inference efficiency enough to justify their higher deployment granularity under realistic power-delivery constraints.
Standard planning metrics do not capture this effect. Provisioned MW and CapEx per MW price installed electrical capacity, not the load a hierarchy can continue to admit after years of arrivals and partial filling [5]. A hall can therefore look efficient at commissioning and still perform poorly later, if future hardware cannot use the residual capacity structure left by earlier placements. This becomes harder to ignore as deployment quanta grow significantly and fast. The problem is also dynamic: hall efficiency depends on the sequence of hardware arrivals and retirements, power oversubscription (harvesting), and prior placements across electrical domains [7]. These interactions form a multi-resource, multiyear packing problem over power, liquid cooling, air cooling, and space, and does not admit a closed-form analysis. Workload throughput metrics (such as inference tokens per second), which ultimately capture the purpose of datacenters, add another dimension to the design problem. Throughput depends on the deployed hardware, and power delivery designs that do not support high power density may be cheaper to provision on a per-MW basis. This introduces a design trade-off in the tokens/sec/W metric as shown in Figure 2: the higher throughput enabled by hosting workloads on a high-power, tightly interconnected pod must be balanced against the increased infrastructure cost required to support such deployments. The size, architecture, and parallelism configurations of the models being hosted will influence this trade-off. Our work. In light of these gaps and challenges, we propose a framework for power delivery design and evaluation that accounts for all relevant factors to produce the most efficient design over the datacenter’s lifetime. We model the designs by the capacity they can actually deploy over time, not just the capacity they install at commissioning. Our framework simulates multi-year fleet evolution under hierarchical
2
Background
Modern datacenters are built to satisfy two requirements at once: high availability under equipment or utility failures, and sufficient distribution capacity to host rack deployments. Those requirements are related but not identical. A design may provision substantial aggregate power and still be unable to admit a new deployment due to reserve, row, or lineup constraints. This paper studies how that gap emerges over time as deployments are placed into a fixed hierarchy. 2.1
Power-Delivery Hierarchy and Feasibility
Datacenters are organized as one or more data halls served by a power-delivery hierarchy with some components shared across halls, (e.g., generators or batteries). Designers provision redundancy throughout this hierarchy so that failures, maintenance, or upstream outages do not interrupt service. We do not model outage probability directly. Instead, we study how standard Tier III/IV-style availability choices shape effective deployable capacity [54]. A typical power path, shown in Figure 3, includes stages such as Grid → Substation → Backup Generation/Storage → UPS → Switchboard → Row Distribution → Rack PSU → 2
ATS
ATS
ATS
ATS
UPS
UPS
UPS
UPS
Switchboard
Switchboard
Switchboard
Switchboard
Switchboard
Switchboard
Switchboard
Row
Row
Row
Row
Row
Switchboard
Row
UPS
UPS
UPS
UPS
Switchboard
Switchboard
Switchboard
Switchboard
Row
Row
Row
Row
Source to Row wiring determines redundancy
Row
Row
(a) Distributed Redundant (4𝑁 /3)
Busbars / PDUs Row
Row
Row
Row
Row
Row
Rack
Figure 3. Example of major components in a datacenter power-delivery hierarchy, from grid and generator/battery down to the rack level. Generic diagram not showing redundancy.
UPS
UPS
Switchboard
Switchboard
Row
Row
Row
Row
Row
Row
UPS UPS UPS UPS Figure 4. Schematic wiring differences between (a) distributed and (b) blockSwitchboard redundant designs. Switchboard In distributed deSwitchboard Switchboard signs, failover is managed by keeping reserve load on each UPS. In block designs, failover is managed by transferring Row Row Row Row Row Row load to the reserve UPS.
loss of any one line-up [59]. This makes reserve flexible, but it reduces total usable capacity within each line-up. In block redundancy (Figure 4(b)), some line-ups are used only for fail-overs. This avoids sharing reserve across active domains, and increases the total capacity in each line-up, but makes capacity coarser and less flexible to use across blocks [5, 39]. The tradeoff is therefore structural: distributed designs tend to fragment residual headroom across active domains, while block designs tend to quantize usable capacity at coarser units.
Supporting High-Power Racks
Typically, each row of racks is supplied by two busbars connected to independent line-ups. At 400 V AC distribution, individual busbar capacity is typically limited to less than 1 MW, owing to safety, installation, and supply-chain constraints. Supporting higher-power deployments within a single row therefore requires installing additional busbars in parallel. Notably, this approach can fragment capacity and reduce placement flexibility: high-power accelerators are preferentially deployed in costly high-density (HD) rows provisioned with greater power capacity, while lower-power CPU and storage servers are placed in low-density (LD) rows. 2.3
UPS Switchboard
(b) Block Redundant (3+1)
Server PSU, though the exact ordering varies across facilities [5, 19, 28, 39, 44, 58, 59]. We use line-up to refer to a common upstream electrical branch: a set of rows or racks that share the same upstream power-delivery equipment. Deployment feasibility is hierarchical, as a deployment must satisfy capacity and redundancy constraints at every level of the path, not only at the site interconnect. Racks are typically deployed in groups referred to as clusters that are in close proximity, ideally within a single row and connected to the same row-level network switch. This deployment pattern emphasizes row power capacity as a key constraint, where power capacity can become stranded in row-level pockets. 2.2
UPS Switchboard
2.4
Capacity Stranding
Stranded capacity is provisioned capacity that remains unused because another constraint binds first [28]. In datacenters, stranding is multi-level: a facility may retain headroom at the site or line-up level and still be unable to admit another rack because some lower-level constraint binds first. This matters because modern AI deployments are coarse placement units. They arrive as racks or pods, often with joint networking and cooling requirements [34, 36, 55]. As a result, deployability depends not only on how much infrastructure is installed, but also on how that capacity is partitioned across the hierarchy.
Redundancy as a Capacity Constraint
Power delivery hierarchies commonly use either distributed redundancy or block redundancy [5, 39, 59] to provide high availability. The relevant distinction, as rack power increases, is that the design affects how much residual headroom remains usable for new placements. In distributed redundancy, reserve is spread across active line-ups. An 𝑥𝑁 /𝑦 design provides 𝑥 total line-ups but only 𝑦 line-ups of supported load, so each parent must retain some capacity for fail-overs. For example, in a 4𝑁 /3 system, as shown in Figure 4(a), each lineup reserves 25% of its capacity so the system can survive the
3
How Power Delivery Topologies Strand Capacity
Datacenter electrical designs are often compared by installed high-availability (HA) capacity and CapEx per provisioned megawatt. We show why those commissioning metrics can be misleading, isolate and describe topology-specific mechanisms that lead to deployment inefficiencies, and motivate the lifecycle evaluation framework that follows. 3
1.0 0.8 0.6 0.4 0.2 0.0
3+1 4N/3
CDF
CDF
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
0.00
0.05
0.10
1.0 0.8 0.6 0.4 0.2 0.0
that affect stranded power are different across distributed and block designs, and are discussed next.
3+1 4N/3
3.2
0.00
0.05
In a distributed 𝑥𝑁 /𝑦 design, reserve is shared across active line-ups. Each line-up can only use a fraction of its rating for HA load, and the remaining must be available for failover. This makes reserve flexible, but it also makes placement require satisfying many simultaneous headroom constraints. Consider a deployment 𝑟 connected to 𝑘𝑟 ≥ 2 parents. If one parent fails, the surviving parents must absorb the deployment’s failover load. The required headroom per surviving parent is
0.10
Line-Up Stranding Fraction
Line-Up Stranding Fraction
(a) Single Data Hall
(b) Fleet-Wide, Lifecycle
Figure 5. CDF of UPS stranding under (a) single-hall Monte Carlo analysis and (b) the final state of an 8-year fleet-scale lifecycle simulation. The local view suggests that 4𝑁 /3 and 3+1 are similar. The lifecycle simulation separates them: 3+1 develops higher tail stranding and requires additional halls to serve the same deployed demand.
3.1
Distributed Designs: Reserve Fragmentation
Δ(𝑃𝑟 , 𝑘𝑟 ) =
𝑃𝑟 . 𝑘𝑟 − 1
(1)
Placement is feasible only if enough parents simultaneously have at least this much local headroom. Aggregate Í slack is not sufficient. A hall can have 𝑑 ℎ𝑑 > 𝑃𝑟 and still reject the deployment because the slack is spread across parents that are each individually too full. For example, a 10𝑁 /8 hall with ten 2.5 MW UPS units provides 20 MW of HA capacity. At 18 MW deployed uniformly, each UPS has 200 kW of headroom. A 650 kW rack with 𝑘𝑟 = 4 needs about 217 kW of headroom on each surviving parent, so placement fails despite 2 MW of aggregate remaining capacity. This mechanism becomes more important as deployment quanta grow. Larger racks and pods increase Δ(𝑃𝑟 , 𝑘𝑟 ), while partially filled halls reduce the local headroom available at each parent. Load-balancing heuristics can delay this failure by keeping headroom even across parents, but they cannot remove it once an indivisible deployment is large relative to the remaining local capacity.
A Tale of Two Designs
Consider a 4𝑁 /3 distributed-redundant hall (as shown in Figure 4(a)) and a 3 + 1 block-redundant hall (as shown in Figure 4(b)). At 2.5 MW per UPS line-up, both designs provide 7.5 MW of high-availability IT capacity and have similar baseline cost under our component model ($10M/MW for 4𝑁 /3 and $10.3M/MW for 3+1); on these static metrics, 4𝑁 /3 appears slightly preferable. A similar conclusion is reached when conducting a single-hall Monte Carlo analysis of rack placement: deployment traces of rack clusters can be generated based on projected distributions of power, cooling, and space requirements and placed into the hall until it saturates. An example of this analysis is shown in Figure 5(a), where the two designs exhibit similar line-up stranding, with 4𝑁 /3 appearing slightly better. The comparison changes once the same designs are evaluated over an 8-year fleet lifecycle. The lifecycle setting adds the effects that commissioning metrics omit: halls partially fill, hardware generations arrive with different power densities, racks are harvested or decommissioned, and new deployments must fit into the residual capacity left by earlier placements. In Figure 5(b), 3+1 develops materially higher tail stranding and requires more halls to serve the same demand. Static metrics suggest only a small difference: 3+1 costs about 3% more per provisioned MW than 4𝑁 /3. Over the 8year fleet lifecycle, however, that small first-order gap nearly doubles to a 5.8% difference in CapEx, as higher stranding in 3+1 requires 23 more halls to be built. The reason is not aggregate unused power, but placement feasibility within a partially filled hierarchy. As deployments arrive over time, admission depends on where residual headroom remains across specific UPS domains, rows, and line-ups, as restricted by the topology’s failover requirements. The mechanisms
3.3
Block Designs: Line-Up Quantization
Block-redundant designs fail differently. Reserve capacity is separated from active load, and block designs avoid the cross-parent reserve-fragmentation condition above. The downside to block designs is that usable capacity is coarser. A 2 MW deployment must fit inside the remaining capacity of a single active line-up, whereas in a distributed-redundant design this power can be split across two or more line-ups. Specifically, for a block with usable capacity 𝐶 and deployment power 𝑃𝑟 , the block admits ⌊𝐶/𝑃𝑟 ⌋ deployments. The leftover capacity is 𝜂 (𝑃𝑟 ) = (𝐶 − ⌊𝐶/𝑃𝑟 ⌋𝑃𝑟 ) /𝐶.
(2)
The key point is divisibility. Just below a threshold 𝑃𝑟 = 𝐶/𝑞, 𝑞 deployments fit. Just above the threshold, only 𝑞 − 1 fit, and the remainder of provisioned line-up capacity is insufficient for the next same-sized deployment. 4
Site Stranding (%)
80 60 40 20 0
4 Datacenter Design Evaluation Framework
Block Distributed
0
1000
2000
3000
Rack Power Density (kW)
4000
We now define the lifecycle model used to evaluate deployable capacity under hierarchical multi-resource placement constraints. The model captures three IT classes—GPU, general compute, and storage—and abstracts the arrival, placement, harvesting, and retirement processes that determine how halls fill over time. Input models for arrivals, cost, and workload impact are described in Section 5.
5000
Figure 6. Single-hall, single-SKU stranding under increasing deployment power. Each experiment fills one hall with repeated deployments of the same SKU and reports the capacity left undeployable at saturation. Distributed redundancy (4𝑁 /3) strands capacity when too few parents have enough simultaneous failover headroom. Block redundancy (3+1) strands capacity at divisibility thresholds of the line-up or UPS-block capacity. The diagonal oscillations are due to rowspace constraints.
3.4
4.1
Rack Resource and Lifetime Model
Power and cooling. Deployments are modelled at rack granularity.1 Each rack 𝑟 is characterized at installation time 𝜏 by power demand 𝑃𝑟 (𝜏) and a cooling demand vector 𝐶𝑟 (𝜏) derived from 𝑃𝑟 (𝜏). Cooling demand is split into air cooling for storage, general compute, and GPU networking, and direct-to-chip liquid cooling for GPU accelerators. Cooling demand is calculated from rack power using fixed conversions: 165 CFM/kW for air cooling and 2 LPM per rack for direct-to-chip liquid cooling [37]. Tile space and networking ports are modeled similarly. Availability and connections. Racks are deployed at one of two availability tiers. High-availability (HA) racks are admitted only when placement preserves the reserve required to survive any single line-up failure. Low-availability (LA) racks may consume reserve capacity, as in Flex [59], but are not guaranteed power through failures or maintenance. GPU pods require multiple independent feeds and sufficient busbar capacity, so they are restricted to highdensity rows. Post-2030 projections of GPU pod and rack power consumption (Appendix B.1) exceed the row power limit for our designs. If a pod or rack exceeds the busbar power limit, cross-row cables can be used to draw power capacity from adjacent rows. Harvesting and decommissioning. Racks are often over-provisioned at deployment and later power oversubscribed or harvested to better match observed utilization [25]. Decommissioned racks release their allocated resources. Lifetimes are modeled as N (𝜇𝑟 , 𝜎𝑟 ) using hardware-class-specific parameters (Section 5.2).
Mechanisms in Isolation
Figure 6 exposes these mechanisms with a single-SKU sweep. For each point on the x-axis, we instantiate a hall of a fixed topology and repeatedly place identical deployments of that power until placement fails. We then measure the fraction of provisioned capacity that remains unused but cannot admit another deployment. The block design shows sharp jumps because small changes in deployment power can cross a divisibility threshold. The distributed design changes more smoothly because failures arise from simultaneous headroom constraints across parents rather than a single block boundary. This experiment is intentionally simpler than a real fleet. It removes mixed SKUs, arrival order, harvesting, retirement, and workload performance. That simplification is useful because it separates the two structural causes of stranding. In distributed redundancy, stranding rises when no sufficiently large set of parents has enough simultaneous failover headroom. In block redundancy, stranding spikes when deployment power crosses a divisibility threshold of the usable block capacity. The same symptom—undeployable power— therefore comes from different structural causes. Real fleets smooth these patterns, but cannot undo them at high power densities. CPU and storage racks can absorb some residual fragments, and heterogeneous GPU generations reduce the sharpness of any single threshold. At the same time, online arrivals, indivisible deployment quanta, harvesting, and decommissioning make the residual capacity of each hall history-dependent. A design therefore cannot be evaluated only by installed MW or by a single saturatedhall experiment; it must be evaluated by how much capacity remains deployable after a sequence of placements and removals. The rest of the paper builds this lifecycle evaluation.
4.2
Placement and Scheduling
A deployed rack is assigned to a specific row and tile subject to hierarchical resource constraints. If no feasible placement exists in any active hall, a new hall is constructed instantly. This simplification of the hall commissioning process isolates the effects of the design on the end-to-end metrics, which are subject to noise from demand prediction errors. Placement constraints. The power hierarchy is modeled as a tree from substation to row. A placement is feasible if
1 For brevity, we use racks as the basic deployment unit. GPU pods are
composed of multiple networked racks that must be placed together. 5
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
Distributed (10N/8)
Block (8+2)
0.00
0.05
0.10
0.15
0.0
0.1
Single-Site Stranding Single-Site Stranding Placement Method Min Variance Min Utilization
Power (MW)
CDF
1.0 0.8 0.6 0.4 0.2 0.0
0.2
Design Specification
Arrival Simulator
Lineup Constructor
CPU
UPS
UPS
UPS
UPS
GPU
Switchboard
Switchboard
Switchboard
Switchboard
Storage
Row
Row
Row
Row
Row
Row
Placement Simulator
Random Round Robin
1. 2. 3.
Place racks via heuristic Perform rack lifecycle events Track DC telemetry
Performance Model
Figure 7. Line-up-level stranding in Monte Carlo simulation of a 10𝑁 /8 and 8 + 2 hall populated with storage, compute, and GPU racks under four online placement policies. Variance minimization yields the lowest stranding.
CapEx Calculator
Design Schema
• Fleetwide DC Count • Stranded Power per DC • Perf per $ per Watt
adding a rack does not exceed effective capacity at any ancestor node, where effective capacity is the residual capacity available after enforcing redundancy constraints. Under distributed 𝑥𝑁 /𝑦 redundancy, effective HA capacity at node ℓ is (𝑦/𝑥) · 𝐶 ℓ , while low-availability (LA) racks may use the full capacity 𝐶 ℓ . Cooling, space, and networking resources impose additional row-level constraints. Appendix C.1 gives the full ancestor-path feasibility condition and demand vector notation. Deployment quanta is modeled as the minimum number of same-SKU racks that must be placed together in one row, consistent with contemporary rack placement mechanisms [7] Placement policies. Rack placement decisions can have a large effect on stranding. We evaluate four online heuristics: min waste which uses a best-fit heuristic based on existing placements, random, round robin across rows, and variance minimization. Variance minimization places each deployment with the goal of minimizing imbalance across UPS domains to reduce stranded capacity, especially in distributed-redundant power delivery systems. We observe the lowest stranding from variance minimization; results from a Monte Carlo placement policy comparison is shown in Figure 7; variance minimization is used as the default policy throughout the rest of this paper. 4.3
Demand Projection
Figure 8. Evaluation pipeline for comparing datacenter power-delivery designs over the deployment lifecycle. Demand projections and design specifications drive a placement simulator, cost model, and performance model, which together produce fleet-level deployability, stranding, cost, and throughput metrics. stranding metrics isolate the portion of 𝑈𝑡(𝑚) that cannot be used because another constraint binds first. Stranding is therefore multi-dimensional: a row may retain 500 kW of electrical headroom, but if liquid cooling capacity is exhausted, that power is unavailable to accelerator deployments. We report two cost metrics. Initial $/MW is the CapEx of one hall normalized by its nameplate IT capacity (Section 5.3). Effective $/MW measures the infrastructure built per MW of deployed IT: Í Í Effective $/MW = 𝑛𝑖=1 𝐾𝑖 / 𝑛𝑖=1 𝑃ˆ𝑖 where 𝐾𝑖 is the CapEx of hall 𝑖 and 𝑃ˆ𝑖 is the final deployed IT MW hosted in hall 𝑖 at the end of the simulation horizon. The gap between these metrics captures how much provisioned infrastructure fails to translate into deployable load.
Stranding Metrics
4.4
A hall may retain unused power, cooling, or space and still be unable to admit another rack because some other constraint binds first. Stranded capacity is provisioned capacity that remains unused because it cannot be converted into deployed load under the current placement constraints. For each resource dimension 𝑚 ∈ {power, air cooling, liquid cooling, space}, let
Rack Placement Simulation
The framework uses two placement simulators: single-hall for mechanism isolation and fleet-scale for multi-year deployment dynamics. Single data hall. Single-hall simulation isolates architectural effects without fleet-level confounders. In each Monte Carlo trial, we instantiate one hall, place racks until 100 consecutive arrivals fail, apply harvesting, and resume placement until another 100 consecutive failures occur. Repeating this process over independently sampled arrival traces yields distributions of stranding and bottleneck behavior.
(𝑚) 𝑈𝑡(𝑚) = 𝐶 prov − 𝐿𝑡(𝑚) (𝑚) denote unused provisioned capacity at time 𝑡, where 𝐶 prov is total provisioned capacity and 𝐿𝑡(𝑚) is deployed load. Our
6
Data Hall Count
20
h Hig
CPU
m diu Me Low
GPU
10 0 0.0
Storage
0.2
0.4
0.6
0.8
Normalized Unused Power
Fleet Historical Data
(1) Projected Growth to Region
1.0
(1) Timeline (2) SKU Lifetimes (3) Region (4) Storage : Compute
(3) Schedule Metadata
(2) SKU Power Density Growth CPU
Our Simulation Arrival Simulator (1) Arrival rates (2) Deployment sizes (3) SKU distributions
Power (MW)
Figure 9. Validation of our simulator against historical rack placements in Azure over 6 years to a subset of both new and mature data halls, and comparing the simulated unusedpower distribution to the observed one. Unused power is normalized by the maximum observed value. We report unused rather than stranded power because some halls are not yet saturated.
GPU
Storage
Years
Figure 10. Deployment-trace generation pipeline. Stage (1) specifies class-level arrival envelopes, stage (2) assigns perSKU rack power, and stage (3) adds lifecycle metadata such as availability tier and retirement time.
This mode is useful for identifying capacity harmonics and resource-ratio mismatches and iterating on the design prior to running longer fleet simulations. Datacenter fleet. Fleet simulation places racks across multiple halls over a multi-year horizon. This exposes interactions absent in single-hall analysis, including shared arrival budgets across active facilities, cross-hall imbalance, and compounding stranding under heterogeneous arrival sequences. The pipeline has three stages, shown in Figure 8. First, an arrival simulator (§5.2) generates monthly rack sequences from long-range demand projections. Second, a datacenter constructor instantiates halls from a power-delivery specification. Third, a placement engine assigns racks to rows across active halls, opens new halls when existing capacity saturates, harvests power from racks after some time, and decommissions racks at end-of-life. Cost (§5.3) and performance (§5.4) models aggregate infrastructure CapEx and workload throughput over the simulation horizon. The fleet simulation captures interactions between hardware refresh, hall aging, and power-delivery design that do not appear in single-hall analysis. In particular, step changes in GPU rack power can interact with fixed line-up sizes and amplify stranding differences that appear small at single-hall scale. Validation with ground truth. We validate the simulator against observed power utilization distributions across data halls over 18 Azure datacenters. Historical deployment traces for each datacenter are used to populate simulated halls that match real designs. Figure 9 compares simulated and observed unused power for racks deployed from 2020 to 2026. We report unused rather than stranded power because some halls are not yet saturated. In some cases, ±1 halls are constructed in simulation than in reality, attributed to variance in stranding and harvesting at these sites. The simulator has a close distributional match to historical observations, with median unused power within 6% of observed.
5
Model Instantiation and Experimental Inputs
This section describes the configuration of the evaluation framework with grounded assumptions for demand, hardware evolution, infrastructure cost, and workload throughput. These inputs define the arrival traces, rack and pod characteristics, and comparative performance and cost models used to evaluate how candidate power-delivery designs behave under lifecycle deployment. 5.1
Arrival Envelopes and Deployment Trace Generation
We model lifecycle demand as a time-ordered sequence of rack deployments. These sequences are generated from arrival envelopes: monthly capacity targets for each hardware class. The envelopes capture aggregate effects of demand growth, procurement limits, and fleet planning assumptions without committing to a specific forecast. We consider three hardware classes from 2025 to 2035: accelerators, general compute, and storage. Each class is specified by an initial deployment level, a growth trajectory, and a capacity cap. These determine annual power targets, which are distributed into monthly budgets using seasonality weights stylized after historical Azure procurement cycles, following prior rack-procurement studies [7, 29]. Within each monthly budget, discrete rack arrivals are sampled. Each arrival carries a demand vector 𝑑𝑟 (𝜏) and relevant lifecycle metadata. The resulting traces preserve long-range growth trends while capturing the variance and heterogeneity in rack arrivals. 7
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
Compute
6
Density
Density
6 4 2
0 0.00 0.25 0.50 0.75 1.00
Normalized Rack Power
(a) General Compute
hardware classes, these projections draw on public vendor disclosures and industry reports [13, 26, 32, 34, 38, 49, 52, 55]. Harvesting and lifetimes Harvesting and retirement determine when occupied capacity returns to the fleet. One year after deployment, one can optionally harvest up to 15% of provisioned power and cooling from storage and general-compute racks and up to 10% from accelerator racks, approximating median historical rates observed at Azure. Rack lifetimes are drawn from N (7, 1) years for storage and general compute and N (5, 0.5) years for accelerator rack and pod deployments.
Storage
4 2 0
0.25 0.50 0.75 1.00
Normalized Rack Power (b) Storage
Figure 11. Normalized rack-power distributions for Azure general-compute and storage deployments since 2023. These results are clustered into empirical distributions of representative SKU groups for future trace generation. 5.2
5.3
The simulator determines how many halls of a given topology must be built to serve a deployment stream. To compare those halls on CapEx rather than nameplate capacity alone, we use a topology-based, per-component cost model. This preserves hall-specific effects while evaluating all designs under the same cost assumptions. Datacenter construction costs are proprietary, vendor- and site-dependent, hence our goal is not to predict exact project cost, but rather develop a comparative model. Published Tier III/IV surveys place non-IT infrastructure cost at roughly $7–12 M/MW [1, 5, 14, 16, 31, 50]. $10 M/MW is a reference baseline, consistent with Schneider Electric’s per-rack scaling model [30, 45] and recent industry surveys [14, 31]. This baseline includes electrical, mechanical, building-shell, and soft construction costs, includes a high liquid-cooling share for accelerators [9], and excludes IT equipment. Per-component costing Where published equipment pricing is available, we use documented ranges. The remaining electrical budget is allocated across generators, transformers, switchgear, and transfer switches. Table 6 summarizes all per-component costs. Because different power-delivery topologies imply different counts and ratings of UPS modules, switchboards, busway runs, and transfer equipment, the same nameplate IT capacity can map to different hall CapEx. These hall-level costs feed directly into the initial and effective $/MW metrics defined in Section 4.3.
Rack Resource Projections and Lifecycle Parameters
Arrival envelopes determine how much capacity enters the fleet, yet it is necessary to specify how each arriving deployment is assigned rack power, cooling demand, harvesting behavior, and lifetime. General compute and storage SKUs. General compute and storage racks exhibit substantial intra-class power variation, as shown in Figure 11. To preserve that heterogeneity, historical Azure rack-power distributions are clustered into representative SKU groups. Each cluster 𝑗 is defined by a scaling factor 𝛼 𝑗 ∈ (0, 1] and deployment probability 𝑝 𝑗 . Given a projected maximum rack power 𝑃max (𝜏, 𝑠) for hardware class 𝑠 at time 𝜏, we generate SKU powers as 𝑃sku,𝑗 (𝜏, 𝑠) = 𝛼 𝑗 𝑃max (𝜏, 𝑠),
Component-Based Infrastructure Cost Model
(3)
and sample arriving racks according to 𝑝 𝑗 . This preserves both class-level power growth and the within-class variation observed in past deployments. Appendix B.3 gives the nonaccelerator power projections. Accelerator rack case studies Accelerator rack power is more tightly coupled to hardware generation and system form factor, so it is modeled explicitly by year and scenario. 2 For each year and growth scenario, a rack TDP anchored to public hardware disclosures is assigned and extended using the per-package TDP growth model in Appendix B. Our near-term case studies use dense rack-scale systems. For later years, we include a denser rack case to capture the possibility that future accelerator deployments arrive at higher effective rack power [26, 56]. We also distinguish rack-scale accelerator arrivals from pod-scale arrivals. Rackscale arrivals are single rack-local deployment units. Podscale arrivals represent multiple accelerator racks deployed together on a shared low-latency pod fabric, with aggregate demand equal to the sum across constituent racks. Across all
5.4
Workload Throughput Model
Infrastructure differences matter only insofar as deployed hardware is performant. Therefore the infrastructure simulation is complemented with a workload model that maps deployed accelerator capacity to comparative inference throughput. We focus on LLM inference, which we expect to account for a substantial share of utilization in datacenters not specialized for training. The model is comparative rather than fully operational: it estimates how accelerator generation, rack density, local interconnect scale, and inter-domain communication translate into delivered tokens per second under representative inference workloads.
2We capture broader industry trends: rising accelerator package power,
denser rack integration, larger local high-bandwidth accelerator domains, and increasing sensitivity to inter-domain communication. 8
1000
6000
Low Medium High
2000 0
20 25 20 27 20 29 20 31 20 33
20 25 20 27 20 29 20 31 20 33
Year
Year
(a) GPU rack-scale systems
(b) GPU pod-scale deployments
60
Power (kW)
Power (kW)
Scenario
4000
0
40 20
30
Scenario
Low Medium High
20 10
0
Year
20 25 20 27 20 29 20 31 20 33
20 25 20 27 20 29 20 31 20 33
0
Year
(c) Compute
(d) Storage
Figure 12. Projected power-density trajectories for GPU racks, GPU pods, CPU compute racks, and storage racks.
Here C𝜙 (𝑚) and M 𝜙 (𝑚) are the per-token compute and 𝜙 memory costs implied by model 𝑚, and 𝑇comm (𝑚, 𝐷) captures tensor-parallel (TP) and expert-parallel (EP) communication on deployment 𝐷. A fixed serving batch size of 𝐵 = 256 is used to represent a high-throughput operating point. Prefill and decode differ in both memory traffic and communication structure: prefill amortizes weight traffic across the prompt batch, whereas decode depends on active weights and growing KV-cache reads over the generated sequence. Appendix A gives the full formulation. Communication model. Communication depends on how much EP traffic remains within a local high-bandwidth accelerator domain and how much spills across domains onto the cluster fabric. Larger local domains keep more traffic on the fast on-package or rack-local interconnect. Smaller domains, or larger models, incur more inter-domain communication. Appendix A.2 and Appendix A provide the corresponding capacity and model details. The throughput model provides the performance side of the comparison in Figure 2. It lets us ask how much local accelerator scale improves inference throughput per watt, while the placement simulator determines whether the resulting rack or pod can actually be deployed.
6
2000
Power (kW)
Power (kW)
Inference throughput is modeled as a mixture-of-experts (MoE) architecture3 [43, 51] running on accelerator deployment 𝐷. Hardware inputs such as local accelerator-domain size, HBM capacity, HBM bandwidth, and compute throughput are drawn from public hardware disclosures and from extrapolations used in prior TCO analyses [46]. The two standard inference phases are considered separately: prefill, which processes the input prompt in parallel and is typically compute-bound [15, 60], and decode, which generates tokens autoregressively and is often limited by HBM bandwidth and KV-cache movement [27, 48, 60]. Phase bottlenecks. For model 𝑚, deployment 𝐷, and phase 𝜙 ∈ {pre, dec}, throughput is limited by the slowest of three resources: compute, HBM bandwidth, and communication: ! HBM 𝐵𝐷 1 𝐹𝐷 𝜙 , , 𝜙 TPS (𝑚, 𝐷) = min 𝜙 . (4) C (𝑚) M 𝜙 (𝑚) 𝑇comm (𝑚, 𝐷)
Table 1. Evaluation setup. We compare redundancy topologies under three GPU power-density trajectories and pod compositions, holding demand, infrastructure granularity, and operational policy fixed. Parameter
Setting
Comparison axes (varied) Redundancy topology 4𝑁 /3 vs. 3+1 (7.5 MW HA) vs. 10𝑁 /8 vs. 8+2 (20 MW HA) GPU TDP trajectory Low / Medium / High (Fig. 12) GPU deployment unit Single rack; pods of 3–7 racks Infrastructure (fixed) Buildout horizon 2026–2034 Cumulative IT demand 10 GW: 6.0 GPU, 2.8 compute, 1.2 storage Electrical granularity 2.5 MW UPS; 625 kW LD row; 2.5 MW HD row HD and LD row count 3:2 ratio of LD to HD rows; (Appendix C.2).
Evaluation of Datacenter Designs 3. Can operational levers reduce stranding enough to change design rankings? 4. When do larger GPU pods deliver enough throughput gain to justify their larger placement quanta?
We evaluate how commissioning metrics evolve for different designs as hardware TDP grows and halls partially fill with deployments. The evaluation seeks to answer four questions: 1. How do designs with similar installed HA capacity and base $/W compare over an 8-year horizon in terms of cost, power capacity, and workload throughput? 2. Does fleet-scale stranding follow the topology-specific mechanisms from Section 3.4?
6.1
Experimental Setup
Table 1 summarizes the evaluation setup. The central question is whether designs with similar installed HA capacity remain similar once evaluated over a multi-year deployment lifecycle. We therefore match designs by nameplate capacity—4𝑁 /3 versus 3+1 at 7.5 MW, and 10𝑁 /8 versus
3 Many current frontier models use MoE architectures.
9
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
Low TDP 30 25 20 15 10 5 0 8 9 0 1 2 3 202 202 203 203 203 203 Date
3+1
Medium TDP
High TDP
8 9 0 1 2 3 202 202 203 203 203 203
8 9 0 1 2 3 202 202 203 203 203 203
Date Design
4N3
8+2
Date
10N8
Figure 13. Tail (P90) site stranding over time for blockredundant (3+1, 8+2) and distributed (4𝑁 /3, 10𝑁 /8) designs under Low, Medium, and High GPU TDP projections. Lines show the median across pod compositions (3–7 racks); bands span the min–max range. Designs that appear similar under static capacity metrics separate once evaluated over the deployment lifecycle. 16
Cost Source
Base Reserve Stranding
14
Cost ($/W)
6.2
Tail Site Stranding (%)
8+2 at 20 MW—and subject them to the same 10 GW demand stream.4 GPU deployments follow three power-density trajectories (Fig. 12); compute and storage use the Medium trajectory throughout. We report three classes of metrics at increasing levels of fidelity. Commissioning metrics (installed HA MW, base$/W) capture what a design looks like at construction. Lifecycle metrics (P90 site stranding, effective $/W, halls built) capture what it delivers after years of arrivals, harvesting, and retirements. Workload metrics (inference TPS/W) capture whether deployed hardware translates into useful throughput. For the workload study, we evaluate MoE inference spanning 0.6 T to 401 T parameters (as shown in Table 2), from models that fit within a single rack-local accelerator domain to models where expert-parallel communication benefits from pod-local placement. We also sweep GPU share from 40% to 80% of total fleet power. The qualitative design ranking is stable across this range; higher GPU shares widen the gap because fewer lowdensity racks remain to absorb residual capacity fragments.
12
How Static Metrics Change with Rack TDP Growth
10 1 0
We first ask whether static commissioning metrics remain predictive as accelerator power grows over the facility lifetime. If they do, then designs with similar nameplate MW and similar base $/W should remain close when evaluated over multi-year deployment traces. Lifecycle deployability separates designs. Figure 13 shows tail site stranding over time under low, medium, and high GPU power density trajectories from Figure 12. The key result is not just that stranding rises with TDP, but that designs with similar static HA capacity separate materially once future deployments must be placed into partially filled halls. Under the High trajectory, the 3+1 design exceeds 20% tail stranding by 2033, while 4𝑁 /3 remains below 10%. Increasing the number of line-ups helps both families, but even as 8+2 improves over 3+1, it still strands more capacity than 4𝑁 /3 under the more aggressive projections. The reversal arises because future deployments must fit into the residual headroom left by earlier ones. What matters is not aggregate remaining slack, but whether that headroom remains in the right UPS domains, rows, and line-ups after years of arrivals, harvesting, and retirements. Rising GPU TDP makes this distinction visible because larger deployment quanta consume a larger fraction of each electrical component in the distribution network. The deployability gap changes cost metrics. This separation is not only a utilization effect. As stranding rises, additional halls must be built to serve the same IT load, which appears directly as higher effective cost. Figure 14 decomposes effective cost above each design’s base $/W into
TDP Scenario 3+1
4N/3
Design
8+2
10N/8
Low Medium High
Figure 14. Incremental effective cost above each design’s base $/W. Bars decompose this excess into reserve cost and stranding-induced cost. Error bars show standard deviation across pod compositions. The main moving term is the cost of stranded capacity, not the nominal cost of reserve. reserve cost and stranding-induced cost. All designs begin with similar base costs, and reserve varies only modestly with redundancy architecture. The main source of variation is instead stranded capacity: infrastructure that was provisioned but cannot be converted into deployed IT load. As GPU power rises, designs with worse lifecycle deployability convert more provisioned capacity into unusable fragments, which raises effective cost above the base. This effect is largest in 3+1, smaller in 4𝑁 /3, and remains muted in 10𝑁 /8 across all three TDP projections. The implication is that designs that begin with similar static costs do not remain similar once evaluated by the deployable capacity they preserve over the fleet lifecycle. This is why designs separate along the cost axis in Figure 2. The difference is not only what the hall costs to build, but how much of that built capacity remains deployable after years of arrivals, oversubscription, and retirements. 6.3
Topology Mechanisms of Lifecycle Stranding
We have shown that static installed-capacity metrics misrank designs as hardware evolves. This subsection asks whether
4 Not indicative of any vendor, hyperscaler, or specific projection.
10
Design
3+1 4N3
40 20
Both Levers 20
Rack Power Density (kW)
Operational Lever
Deployment Quanta Harvesting
Figure 15. P90 tail stranding versus effective per-domain deployment power for 3+1 and 4𝑁 /3 across all GPU TDP scenarios and pod compositions. Dashed vertical lines mark 2.5 MW UPS-block quantization thresholds, around which 3+1 exhibits pronounced stranding increases.
3+1
Deployment Quanta Harvesting
Both Levers
0 500 1000 1500 2000 2500 3000 3500 4000 4500
10
0
Total Cost (%) 10N/8
Both Levers
20 Deployment Quanta Harvesting
10
0
10
0
Total Cost (%) 8+2
Both Levers 20
10
0
Total Cost (%)
20
Total Cost (%)
Figure 16. Change in total cost relative to the baseline fleet under the best setting from each operational lever family. Negative values indicate cost savings. Tuning reduces some costs, but does not change design outcomes.
the resulting lifecycle stranding follows the topology-specific mechanisms identified in Section 3.4. Figure 15 plots P90 tail stranding against effective perdomain deployment power across GPU TDP scenarios and pod compositions. The pattern matches the structural distinction from Section 3.4. In 3+1, high-stranding points cluster near the dashed 𝐶/𝑞 thresholds of a 2.5 MW UPS block. Just above such a threshold, one fewer deployment fits, and the residual capacity becomes provisioned but undeployable, consistent with the quantization effect in Figure 6. The 4𝑁 /3 design degrades differently. Stranding still rises with deployment power, but not around discrete thresholds. Instead, it increases as larger deployments make it harder to satisfy the multiple-parent redundancy constraint. Figure 15 also demonstrates that at lower TDPs, block redundancy is an effective design. At commissioning and at lower rack densities, block and distributed designs can look comparable, and block designs retain practical advantages for their simple failure modes. The gap appears when highpower AI deployments make placement granularity a firstorder constraint. This also explains the variation across pod compositions. Changing pod size changes the effective deployment quantum seen by the hierarchy. Some quanta align poorly with upstream electrical ratings and produce abrupt jumps in undeployable capacity. Others mainly tighten placement feasibility and produce a more continuous loss of deployability. The fleet-scale behavior therefore follows the same topologydependent mechanisms seen in our analysis. 6.4
4N/3
Deployment Quanta Harvesting
Operational Lever
Tail Site Stranding (%)
60
quanta reduce some admission failures because they make low-density racks easier to pack. This effect is more pronounced in single-site packing, but at fleet scale it translates into only about 1% fewer data halls because large accelerator racks remain difficult to place. Harvesting returns occupied power and cooling to the fleet, opening some additional placement opportunities, but the gains are modest and do not change the design ranking. More generally, operational tuning can reduce imbalance, return some capacity, or make individual arrivals easier to place, but it does not change the underlying placement structure imposed by hierarchical high-availability constraints. Service-model relaxations such as low-availability tiering [59] can reclaim reserve capacity, but they alter the comparison: in distributed designs, reserve is shared across active parents, whereas in block designs, low-availability load would sit on primary line-ups without locally provisioned failover. We therefore treat such relaxations as orthogonal to the structural question studied here. Even with these levers, once the hierarchy’s constraints bind, the remaining slack is structurally hard to use. Policy can improve utilization of residual capacity, but it does not materially change which future racks and pods the hierarchy can admit.
Operational Levers for Reducing Stranding
6.5
A natural question is whether operational levers that apply comparably across these hierarchies can recover enough lifecycle loss to change design performance. If so, adjusting deployment quanta or harvesting should materially reduce cost and narrow the gap across designs. Figure 16 reports the largest cost reduction achieved by each lever family relative to the baseline of 10 racks per deployment quantum and no harvesting. Smaller deployment
When do GPU pod networking gains survive deployability constraints?
Larger pods improve serving efficiency by keeping more EP communication within a local high-bandwidth domain, but they also arrive as coarser indivisible placement quanta [36, 38]. The question is whether that serving gain survives the lifecycle deployability penalty imposed by the hierarchy. Figure 17 makes that tradeoff explicit for two designs serving an MoE-132T workload (scaled factors of DeepSeek-R1’s 11
Throughput per Watt (Tok/s/W)
0.24 0.22 0.20 0.18 0.16
10N/8
8+2
6 racks 7 racks 4 racks 5 racks 3 racks 2 racks
4/5 racks 3 racks 2 racks
1 rack
12.8
Net Pod Benefit (%)
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
13.0
1 rack
13.2
Effective Cost of DC ($/W)
6 racks 7 racks
10N/8
40 20 0 20
19
3 racks
13.4
8+2
51
132
401
Model Size (T) Racks per Pod 4 racks
5 racks
19
51
132
Model Size (T) 6 racks
401 7 racks
Figure 18. Pod payoff across model sizes for 10𝑁 /8 and 8+2. Larger pods help only once their communication benefit exceeds their deployability cost. The later crossover in 8+2 reflects its larger deployability penalty.
Figure 17. Effective fleet cost versus power-normalized throughput under High GPU TDP growth for MoE-132T. Larger pods increase throughput but also raise effective cost by increasing deployment granularity. 10𝑁 /8 preserves more of the throughput gain at fleet scale because it admits larger quanta with less deployability loss.
how much of the serving gain remains after years of arrivals into a constrained hierarchy. Pod efficiency therefore helps only when two conditions hold: (1) the workload is communication-limited enough to benefit from a larger local domain, and (2) the hierarchy is fine-grained enough to admit the deployment quantum without stranding that gain. These conditions are intuitive, and our framework quantifies an estimate for cross-over points for the benefit of pods across model size that vary across datacenter designs. Appetite for the high end of future accelerator power densities may not depend solely on theoretical improvement in hardware performance, but also on how efficiently a datacenter can host them.
model parameters [20]). Moving to larger pods shifts both designs upward and to the right. The upward movement is the serving-side effect: more traffic remains within the local domain, so throughput per watt increases. The rightward movement is the infrastructure-side effect: larger quanta are harder to place into partially filled halls, so effective fleet cost rises. In 10𝑁 /8, larger pods preserve more of their serving benefit at fleet scale. In 8+2, the same performance gain is offset more heavily by the cost of admitting a coarser quantum. The difference is not the serving model, but the hierarchy’s ability to absorb larger placements without giving up as much deployable capacity. To expose the crossover directly, define pod payoff relative to a single-rack baseline:
7
Related Work
Datacenter design and rack placement. Prior work studies redundancy schemes, capacity provisioning, and rack placement, but usually treats the electrical hierarchy as fixed. Industry guidance and systems texts describe distributed (𝑥𝑁 /𝑦) and block (𝑁 +𝑦) redundancy and compare them mainly on first-order cost, complexity, and maintainability [2, 3, 19, 22, 23, 53]. A separate line of work formulates rack placement and expansion as online or stochastic packing, improving utilization through better placement policies, migration, or lifetime-aware allocation [4, 7, 11, 18, 29, 33, 35]. These systems optimize where to place load, not whether the hierarchy can admit it. Our work complements this literature by exposing how, as deployment quanta consume a meaningful fraction of rows, line-ups, or UPS blocks, power delivery design can be the dominant source of inefficiency. Datacenter power management and oversubscription. A large body of work improves utilization within an existing hierarchy through power capping, DVFS, storage, and oversubscription [17, 24, 25, 28, 42, 57, 58]. Other systems reclaim reserve or underused capacity within a fixed hierarchy by relaxing service guarantees or exploiting utilization headroom [6, 21, 41, 59]. Our analysis is orthogonal, as we characterize how hierarchical structure shapes deployability under a fixed service model.
1 + ΔTPS/W − 1, 1 + ΔCost where ΔTPS/W is the fractional throughput-per-watt gain from better networking and ΔCost is the fractional increase in effective fleet cost induced by the larger placement quantum. Positive payoff means the serving gain exceeds the deployability cost. Negative payoff means the communication benefit is outweighed by lifecycle deployment cost. Figure 18 shows that the crossover depends on both workload and hierarchy. For smaller models, most communication is already contained within a rack-scale domain and pods have little to offer for serving throughput, but still incur a placement penalty, so payoff remains near zero or negative. As model size grows, more EP traffic spills across domains, and payoff becomes positive. The crossover is also topology-dependent. In 10𝑁 /8, payoff becomes positive earlier and rises more uniformly with pod size because the hierarchy admits larger quanta with less deployability loss. In 8+2, the same pod sizes cross divisibility and placement thresholds sooner, leaving less net benefit after lifecycle deployment. What changes across designs is Pod Payoff =
12
AI infrastructure and power delivery. Recent work shows that AI training and inference change datacenter requirements through higher rack power, different temporal behavior, and tighter coupling between compute, cooling, and interconnect [8, 10, 39, 47, 48]. This has prompted new discussions of AI-oriented facility design, including denser cooling integration and alternatives such as 800V DC distribution [12, 36, 46]. The closest prior work [46] evaluates AI datacenters through aggregate hardware lifecycle cost; our works provides a complementary perspective on AI datacenter infrastructure.
8
European Conference on Computer Systems (Online Event, United Kingdom) (EuroSys ’21). Association for Computing Machinery, New York, NY, USA, 556–573. doi:10.1145/3447786.3456259 [7] Saumil Baxi, Kayla Cummings, Alexandre Jacquillat, Sean Lo, Rob McDonald, Konstantina Mellou, Ishai Menache, and Marco Molinaro. 2025. Online Rack Placement in Large-Scale Data Centers: Online Sampling Optimization and Deployment. arXiv:2501.12725 [math.OC] https://arxiv.org/abs/2501.12725 [8] Ricardo Bianchini, Christian Belady, and Anand Sivasubramaniam. 2024. Datacenter power and energy management: past, present, and future. IEEE Micro (2024). [9] Robert Bunger and Wendy Torell. 2019. Capital Cost Analysis of Immersive Liquid-Cooled vs. Air-Cooled Large Data Centres. Technical Report White Paper 282. Schneider Electric. Detailed CapEx comparison of 2MW datacenter configurations. Provides itemized infrastructure costs including generators, UPS, switchgear, and cooling subsystems.. [10] Jae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng, Nikhil Bansal, and Mosharaf Chowdhury. 2024. Reducing Energy Bloat in Large Model Training. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (Austin, TX, USA) (SOSP ’24). Association for Computing Machinery, New York, NY, USA, 144–159. doi:10.1145/3694715.3695970 [11] Maxime C. Cohen, Philipp Keller, Vahab Mirrokni, and Morteza Zadimoghadddam. 2017. Overcommitment in Cloud Services Bin packing with Chance Constraints. In Proceedings of the 2017 ACM SIGMETRICS / International Conference on Measurement and Modeling of Computer Systems (Urbana-Champaign, Illinois, USA) (SIGMETRICS ’17 Abstracts). Association for Computing Machinery, New York, NY, USA, 7. doi:10.1145/3078505.3078530 [12] Data Center Frontier. 2025. OCP Summit 2025 Highlights: Advancing Data Center Densification and Security. https://www.datacenterfrontier.com/design/article/55324586/ocpsummit-2025-highlights-advancing-data-center-densification-andsecurity Industry shift toward 800V DC power distribution for megawatt rack scales. [13] Datacenters.com. 2025. Next-Gen Processors: Redefining Data Center Performance in 2025. https://www.datacenters.com/news/next-genprocessors-how-they-re-redefining-data-center-performance Highperformance processors pushing rack densities beyond 80 kW require liquid cooling. [14] Dgtl Infra. 2024. How Much Does it Cost to Build a Data Center? https://dgtlinfra.com/how-much-does-it-cost-to-build-a-datacenter/ Component-level cost breakdowns for Tier III/IV facilities. [15] Kuntai Du, Bowen Wang, Chen Zhang, Yiming Cheng, Qing Lan, Hejian Sang, Yihua Cheng, Jiayi Yao, Xiaoxuan Liu, Yifan Qiao, Ion Stoica, and Junchen Jiang. 2025. PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles (Lotte Hotel World, Seoul, Republic of Korea) (SOSP ’25). Association for Computing Machinery, New York, NY, USA, 399–414. doi:10.1145/3731569.3764834 [16] Lisa Duignan. 2024. Data centre cost index 2024. https://www. turnerandtownsend.com/insights/data-centre-cost-index-2024/ Global construction cost benchmarks for data centers. [17] Daniel Ellsworth, Tapasya Patki, Swann Perarnau, Sangmin Seo, Abdelhalim Amer, Judicael Zounmevo, Rinku Gupta, Kazutomo Yoshii, Henry Hoffman, Allen Malony, Martin Schulz, and Pete Beckman. 2016. Systemwide Power Management with Argo. In 2016 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 1118–1121. doi:10.1109/IPDPSW.2016.81 [18] Marius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall, Po-Yen Chou, Richard Cornew, Adela Kabiljo, Ranjith Kumar S, Maroo Lieuw, Justin Meza, Scott Michelson, Thomas Rohloff, Hayley Russell, Jeff Qin, and Chunqiang Tang. 2023. Global Capacity
Conclusion
AI accelerator growth changes power-delivery design from a commissioning problem into a lifecycle deployability problem. A hall can retain substantial provisioned power and still be unable to admit future racks or pods because residual capacity is fragmented across the hierarchy. Our results show that this gap is large enough to change design rankings: topologies with similar installed capacity and base cost diverge as deployment quanta grow, with stranded capacity becoming the dominant source of effective cost. The design objective for AI datacenters is therefore not installed megawatts, but deployable capacity over time. Power hierarchies should be evaluated by how much useful accelerator capacity they continue to admit across hardware generations, placement histories, and workload requirements. As AI facilities scale, this gap becomes a first-order systems design constraint, not an accounting nuance.
References [1] AccuTech Communications. 2024. Best Data Center Build Out Cost: Top 5 Key Factors. https://accutechcom.com/data-center-build-outcost/ States Tier III construction typically $7–$9M per MW (illustrative benchmark).. [2] ASCO Power Technologies. 2019. Power Redundancy Schemes for Data Centers. Technical Report PS-WP-REDUNDANCY-DATA. Uptime Institute. https://www.se.com/sg/en/download/document/PS-WPREDUNDANCY-DATA/ [3] Victor Avelar, Patrick Donovan, Wendy Torell, and Maria A. Torres Arango. 2025. How 6 AI Attributes Change Data Center Design. Technical Report White Paper 110, v3. Schneider Electric. https://www.se.com/us/en/download/document/SPD_WP110_EN/ [4] Hugo Barbalho, Patricia Kovaleski, Beibin Li, Luke Marshall, Marco Molinaro, Abhisek Pan, Eli Cortez, Matheus Leao, Harsh Patwari, Zuzu Tang, Larissa Rozales Gonçalves, David Dion, Thomas Moscibroda, and Ishai Menache. 2023. Virtual Machine Allocation with Lifetime Predictions. In Proceedings of Machine Learning and Systems, D. Song, M. Carbin, and T. Chen (Eds.), Vol. 5. Curan, 232–253. https://proceedings.mlsys.org/paper_files/paper/2023/file/ 48eb2c79643df5cb7d125945238bd7d0-Paper-mlsys2023.pdf [5] Luiz Andr’e Barroso, Urs H"olzle, and Parthasarathy Ranganathan. 2019. The Datacenter as a Computer: Designing Warehouse-Scale Machines (3 ed.). Springer, Cham. XVIII, 189 pages. doi:10.1007/978-3031-01761-2 [6] Noman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin, Sree Kodak, and Rohit Jnagal. 2021. Take it to the limit: peak prediction-driven resource overcommitment in datacenters. In Proceedings of the Sixteenth 13
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
Management With Flux. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). USENIX Association, Boston, MA, 589–606. https://www.usenix.org/conference/osdi23/ presentation/eriksen [19] Xiaobo Fan, Wolf-Dietrich Weber, and Luiz Andre Barroso. 2007. Power provisioning for a warehouse-sized computer. SIGARCH Comput. Archit. News 35, 2 (June 2007), 13–23. doi:10.1145/1273440.1250665 [20] Daya Guo et al. 2025. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 8081 (2025), 633–638. doi:10.1038/s41586-025-09422-z [21] Nishant Gupta, Iyswarya Narayanan, Shivam Handa, Sayak Chakraborti, Pankit Thapar, Baohua Shan, Ariel Rao, Yuanlai Liu, Pengyuan Wang, Yuqing Wu, Qingyi Gao, Chris Chao-Chun Cheng, Sihan You, Louis Huang, Jingyuan Fan, Kenny Yu, Kevin Lin, Tengfei Mu, Parth Malani, Haiying Wang, Trey Lu, and Peter Zhang. 2024. Dynamic Idle Resource Leasing To Safely Oversubscribe Capacity At Meta. In Proceedings of the 2024 ACM Symposium on Cloud Computing (Redmond, WA, USA) (SoCC ’24). Association for Computing Machinery, New York, NY, USA, 792–810. doi:10.1145/3698038.3698537 [22] James Hamilton. 2009. Internet-scale service infrastructure efficiency. SIGARCH Comput. Archit. News 37, 3 (June 2009), 232. doi:10.1145/ 1555815.1555756 [23] Pearl Hu. 2019. Electrical Distribution Equipment in Data Center Environments (White Paper 61, Rev. 2). Technical Report. Schneider Electric. https://www.se.com/us/en/download/document/SPD_VAVR8W4MEX_EN/ Equipment-level per-kW ranges for MV/LV switchgear, transformers, PDUs, panels.. [24] Vasileios Kontorinis, Liuyi Eric Zhang, Baris Aksanli, Jack Sampson, Houman Homayoun, Eddie Pettis, Dean M. Tullsen, and Tajana Simunic Rosing. 2012. Managing distributed UPS energy for effective power capping in data centers. In 2012 39th Annual International Symposium on Computer Architecture (ISCA). 488–499. doi:10.1109/ISCA.2012.6237042 [25] Alok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde, Felipe Frujeri, Nithish Mahalingam, Pulkit A Misra, Seyyed Ahmad Javadi, Bianca Schroeder, Marcus Fontoura, et al. 2021. {Prediction-Based} power oversubscription in cloud platforms. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). 473–487. [26] Ming-Chi Kuo. 2025. NVIDIA AI Server Power Roadmap: Kyber’s Next-Generation Strategy from GPU/Rack-Level to Data-Center Scale. https://medium.com/@mingchikuo/nvidia-ai-server-powerroadmap-kybers-next-generation-strategy-from-gpu-rack-level-todata-center-e380b459e183 Industry analysis of NVIDIA’s reference design scope extending to entire data center. [27] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (Koblenz, Germany) (SOSP ’23). Association for Computing Machinery, New York, NY, USA, 611–626. doi:10.1145/3600006.3613165 [28] Yang Li, Charles R. Lefurgy, Karthick Rajamani, Malcolm S. AllenWare, Guillermo J. Silva, Daniel D. Heimsoth, Saugata Ghose, and Onur Mutlu. 2019. A Scalable Priority-Aware Approach to Managing Data Center Server Power. In 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA). 701–714. doi:10. 1109/HPCA.2019.00067 [29] Rui Peng Liu, Konstantina Mellou, Evelyn Xiao-Yue Gong, Beibin Li, Thomas Coffee, Jeevan Pathuri, David Simchi-Levi, and Ishai Menache. 2025. Efficient Cloud Server Deployment Under Demand Uncertainty. Manufacturing & Service Operations Management 27, 2 (2025), 425–440. arXiv:https://doi.org/10.1287/msom.2023.0372 doi:10.1287/msom.2023. 0372
[30] Kevin McCarthy and Victor Avelar. 2016. Comparing UPS System Design Configurations. Technical Report White Paper 75. Schneider Electric – Data Center Science Center. https://download.schneiderelectric.com/files?p_Doc_Ref=SPD_SADE-5TPL8X_EN Revision 4. [31] John McWilliams, Ethan Tribble, Adrian Conforti, and Jason DOrlando. 2025. Data Center Development Cost Guide 2025. https://cushwake. cld.bz/Data-Center-Development-Cost-Guide-2025 Market-level totals and cost drivers across U.S. regions.. [32] Chris Mellor. 2025. Power Consumption and Datacenters. https://blocksandfiles.com/2025/07/14/power-consumption-anddata-centers/ Dell’Oro analysis: AI workloads require 60-120 kW/rack for accelerated servers. [33] Konstantina Mellou, Marco Molinaro, and Rudy Zhou. 2024. The Power of Migrations in Dynamic Bin Packing. Proc. ACM Meas. Anal. Comput. Syst. 8, 3, Article 45 (Dec. 2024), 28 pages. doi:10.1145/3700435 [34] Timothy Prickett Morgan. 2025. Nvidia Draws GPU System Roadmap Out To 2028. https://www.nextplatform.com/2025/03/19/ nvidia-draws-gpu-system-roadmap-out-to-2028/ Rubin Ultra VR300 NVL576 consuming over 600 kilowatts, 21× performance of GB200. [35] Christopher Muir, Luke Marshall, and Alejandro Toriello. 2024. Temporal Bin Packing with Half-Capacity Jobs. INFORMS Journal on Optimization 6, 1 (2024), 46–62. arXiv:https://doi.org/10.1287/ijoo.2023.0002 doi:10.1287/ijoo.2023.0002 [36] NVIDIA Corporation. 2025. Building the 800 VDC Ecosystem for Efficient, Scalable AI Factories. https://developer.nvidia.com/blog/ building-the-800-vdc-ecosystem-for-efficient-scalable-ai-factories Technical blog detailing 800V DC power distribution for megawatt rack scales. [37] Open Compute Project. 2023. OAI System Liquid Cooling Guidelines. White Paper. Open Compute Project. https://www.opencompute.org/documents/oai-system-liquidcooling-guidelines-in-ocp-template-mar-3-2023-update-pdf [38] Dylan Patel, Daniel Nishball, Kimbo Chen, Wega Chu, Ivan Chiam, and Cheang Kang Wen. 2025. Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack. https://newsletter.semianalysis.com/ p/another-giant-leap-the-rubin-cpx-specialized-accelerator-rack. [39] Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Brijesh Warrier, Nithish Mahalingam, and Ricardo Bianchini. 2024. Characterizing Power Management Opportunities for LLMs in the Cloud. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (La Jolla, CA, USA) (ASPLOS ’24). Association for Computing Machinery, New York, NY, USA, 207–222. doi:10.1145/3620666.3651329 [40] Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Inference Using Phase Splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). 118–132. doi:10.1109/ISCA59077.2024.00019 [41] Leonardo Piga, Iyswarya Narayanan, Aditya Sundarrajan, Matt Skach, Qingyuan Deng, Biswadip Maity, Manoj Chakkaravarthy, Alison Huang, Abhishek Dhanotia, and Parth Malani. 2024. Expanding datacenter capacity with dvfs boosting: A safe and scalable deployment experience. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1. 150–165. [42] Ramya Raghavendra, Parthasarathy Ranganathan, Vanish Talwar, Zhikui Wang, and Xiaoyun Zhu. 2008. No "power" struggles: coordinated multi-level power management for the data center. In Proceedings of the 13th International Conference on Architectural Support for Programming Languages and Operating Systems (Seattle, WA, USA) (ASPLOS XIII). Association for Computing Machinery, New York, NY, USA, 48–59. doi:10.1145/1346281.1346289
14
[56] Jarred Walton. 2025. Nvidia Shows Off Rubin Ultra with 600,000-Watt Kyber Racks and Infrastructure, Coming in 2027. https://www.tomshardware.com/pc-components/gpus/nvidiashows-off-rubin-ultra-with-600-000-watt-kyber-racks-andinfrastructure-coming-in-2027 Kyber rack architecture targeting 600kW per rack with Rubin Ultra GPUs. [57] Di Wang, Chuangang Ren, Anand Sivasubramaniam, Bhuvan Urgaonkar, and Hosam Fathy. 2012. Energy storage in datacenters: what, where, and how much? SIGMETRICS Perform. Eval. Rev. 40, 1 (June 2012), 187–198. doi:10.1145/2318857.2254780 [58] Qiang Wu, Qingyuan Deng, Lakshmi Ganesh, Chang-Hong Hsu, Yun Jin, Sanjeev Kumar, Bin Li, Justin Meza, and Yee Jiun Song. 2016. Dynamo: facebook’s data center-wide power management system. In Proceedings of the 43rd International Symposium on Computer Architecture (Seoul, Republic of Korea) (ISCA ’16). IEEE Press, 469–480. doi:10.1109/ISCA.2016.48 [59] Chaojie Zhang, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang, Pulkit A. Misra, Rod Assis, Kyle Woolcock, Nithish Mahalingam, Brijesh Warrier, David Gauthier, Lalu Kunnath, Steve Solomon, Osvaldo Morales, Marcus Fontoura, and Ricardo Bianchini. 2021. Flex: High-Availability Datacenters With Zero Reserved Power. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). 319–332. doi:10.1109/ISCA52012.2021.00033 [60] Hengrui Zhang, Pratyush Patel, August Ning, and David Wentzlaff. 2025. SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference. arXiv:2510.08544 [cs.AR] https://arxiv.org/abs/ 2510.08544 [61] Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the 18th USENIX Conference on Operating Systems Design and Implementation (Santa Clara, CA, USA) (OSDI’24). USENIX Association, USA, Article 11, 18 pages.
[43] Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He. 2022. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale. https: //proceedings.mlr.press/v162/rajbhandari22a.html [44] Varun Sakalkar, Vasileios Kontorinis, David Landhuis, Shaohong Li, Darren De Ronde, Thomas Blooming, Anand Ramesh, James Kennedy, Christopher Malone, Jimmy Clidaras, and Parthasarathy Ranganathan. 2020. Data Center Power Oversubscription with a Medium Voltage Power Plane and Priority-Aware Capping. In Proceedings of the TwentyFifth International Conference on Architectural Support for Programming Languages and Operating Systems. New York, NY, USA, 497–511. https: //dl.acm.org/doi/abs/10.1145/3373376.3378533 [45] Max Smolaks. 2023. Data center costs set to rise and rise. https://journal.uptimeinstitute.com/data-center-costs-set-torise-and-rise/ Analysis of supply chain impacts on infrastructure costs. [46] Jovan Stojkovic, Chaojie Zhang, Inigo Goiri, and Ricardo Bianchini. 2025. Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework. arXiv:2509.26534 [cs.AI] https://arxiv.org/abs/2509.26534 [47] Jovan Stojkovic, Chaojie Zhang, Inigo Goiri, Esha Choukse, Haoran Qiu, Rodrigo Fonseca, Josep Torrellas, and Ricardo Bianchini. 2025. TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms. Association for Computing Machinery, New York, NY, USA, 1266–1281. https://doi.org/10.1145/3676641.3716025 [48] Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2025. DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). 1348– 1362. doi:10.1109/HPCA61900.2025.00102 [49] The Register. 2023. Intel and AMD Just Created a Headache for Legacy Datacenters. https://www.theregister.com/2023/01/19/intel_ amd_uptime_cooling/ AMD Epyc 4 at 400W and Intel Xeon Scalable at 350W TDP. [50] Thunder Said Energy. 2024. Economic costs of data-centers? https: //thundersaidenergy.com/downloads/data-centers-the-economics/ Cost breakdown analysis including mechanical systems. [51] Jesmin Jahan Tithi, Hanjiang Wu, Avishaii Abuhatzera, and Fabrizio Petrini. 2025. Scaling Intelligence: Designing Data Centers for NextGen Language Models. arXiv:2506.15006 [cs.AR] https://arxiv.org/ abs/2506.15006 Nvidia Announces Reference De[52] Tom’s Hardware. 2025. sign for Colossal Gigawatt-scale Omniverse DSX Data Centers. https://www.tomshardware.com/tech-industry/artificialintelligence/nvidia-announces-reference-design-for-gargantuangigawatt-scale-omniverse-dsx-data-centers-single-data-centerrequires-a-nuclear-reactors-worth-of-power-generation NVIDIA’s blueprint for gigawatt-class AI data centers with 1 megawatt server racks. [53] Wendy Torell. 2016. Cost, Speed, and Reliability Tradeoffs between N+1 UPS Configurations. Technical Report White Paper 234. Schneider Electric – Data Center Science Center. https: //www.apc.com/us/en/support/resources-tools/white-papers/costspeed-and-reliability-tradeoffs-between-n1-ups-configurations.jsp Revision 2. [54] W Pitt Turner IV, JH PE, PE Seader, and KJ Brill. 2006. Tier classification define site infrastructure performance. Uptime Institute 17 (2006). [55] Jarred Walton. 2025. Nvidia Announces Rubin GPUs in 2026, Rubin Ultra in 2027, Feynman Also Added to Roadmap. https://www.tomshardware.com/pc-components/gpus/nvidiaannounces-rubin-gpus-in-2026-rubin-ultra-in-2027-feynam-after Rubin NVL144 specifications: 3.6 EFLOPS FP4, 288GB HBM4, 13 TB/s bandwidth.
15
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
A
Performance Model Details
𝐷 be the number of accelerator packages in one Let 𝑁 pkg
local domain of deployment 𝐷, and let HBM𝐷 be HBM pkg capacity per package. We reserve a fraction 1 − 𝛼 of HBM for KV residency and runtime overhead, with 𝛼 = 0.7. The number of local NVLink domains required to host model 𝑚 on deployment 𝐷 is & ' 𝑊total (𝑚) 𝑁 dom (𝑚, 𝐷) = . (12) 𝐷 HBM𝐷 𝛼 𝑁 pkg pkg
We use a first-order comparative model of MoE inference throughput. For a given deployment, throughput is limited by the slowest of compute, HBM bandwidth, and communication. The model is used only to compare hardware and locality configurations in the fleet study, not to predict absolute serving latency. A.1
Per-Phase Throughput
For model 𝑚, deployment 𝐷, and phase 𝜙 ∈ {pre, dec}, we write ! HBM 𝐵𝐷 1 𝐹𝐷 𝜙 , , 𝜙 TPS (𝑚, 𝐷) = min 𝜙 , (5) C (𝑚) M 𝜙 (𝑚) 𝑇comm (𝑚, 𝐷)
We then approximate the fraction of expert-parallel traffic that leaves the local NVLink domain as 𝑓IB (𝑚, 𝐷) =
HBM is agwhere 𝐹𝐷 is deployment compute throughput, 𝐵𝐷 𝜙 𝜙 gregate HBM bandwidth, C and M are the per-token 𝜙 compute and memory costs of phase 𝜙, and 𝑇comm is the communication time. We assume 1 FMA = 2 FLOPs. Unless otherwise noted, we use FP8 weights (𝑏 𝑤 = 1 byte), FP4 activations and KV cache (𝑏 act = 𝑏 kv = 0.5 bytes), and serving batch size 𝐵 = 256. All models use 𝐾 = 2 routed experts per token and FF = 4𝑤. We use the same bottleneck form for prefill and decode; the phases differ only in weight traffic, KV traffic, and sequencelength dependence: C pre (𝑚) = 𝐿 4𝐾𝑤FF + 4𝑤 2 + 2𝑤𝑆𝑝 , (6) dec 2 C (𝑚, 𝑡) = 𝐿 4𝐾𝑤FF + 4𝑤 + 2𝑤𝑡 , (7)
𝑊total (𝑚) + 2𝐿𝑤𝑏 kv, 𝐵𝑆𝑝 𝑊active (𝑚) M dec (𝑚, 𝑡) ≈ + 2𝐿𝑤 (𝑡 + 1)𝑏 kv, 𝐵 2(𝑇𝐷 − 1) 𝜙 NTP (𝑚, 𝐷) = 𝐿 · 𝑤𝑏 act, 𝑇𝐷 M pre (𝑚) ≈
(13) otherwise.
𝜙 NTP (𝑚, 𝐷) 𝜙 𝑇TP (𝑚, 𝐷) = , NVL 𝐵𝐷
(8) 𝜙 𝑇EP (𝑚, 𝐷) = max
(9)
! 𝜙 𝜙 (1 − 𝑓IB (𝑚, 𝐷))NEP (𝑚) 𝑓IB (𝑚, 𝐷)NEP (𝑚) , , IB NVL 𝐵𝐷 𝐵𝐷 (15)
𝜙
𝜙
(14)
𝜙
𝑇comm (𝑚, 𝐷) = 𝑇TP (𝑚, 𝐷) + 𝑇EP (𝑚, 𝐷),
(10)
(16)
NVL and 𝐵 IB are local NVLink and remote InfiniBand where 𝐵𝐷 𝐷 bandwidth. The max term reflects concurrent local and remote transfers during the EP sublayer.
(11)
Here 𝐿 is the number of transformer layers, 𝑤 is hidden width, 𝑆𝑝 is prompt length, and 𝑡 is the effective context length during decode. 𝑊total counts all expert parameters, while 𝑊active counts the shared attention weights plus the routed experts touched by one token. Within one NVLink domain, shared attention uses tensor parallelism across 𝑇𝐷 packages, while MoE FFNs use expert parallelism. TP traffic stays on NVLink. EP traffic is split between local NVLink and remote InfiniBand according to the locality model below. Prefill amortizes total weight traffic across the prompt batch, whereas decode depends on active weights and growing KV-cache reads. A.2
𝑁 dom (𝑚, 𝐷) = 1,
This captures the first-order fact that once a model spans multiple domains, only a fraction 1/𝑁 dom of expert traffic can remain local. We use package counts throughout. HBM capacity, package TDP, and NVLink-domain membership are package-level quantities, so all capacity and communication terms are evaluated at the package/domain level. Given 𝑓IB (𝑚, 𝐷), we model TP and EP communication time as
𝜙
NEP (𝑚) = 2𝐿𝐾𝑤𝑏 act .
0, 1 , 1 − 𝑁 dom (𝑚, 𝐷)
A.3
Request-Level Throughput
For prompt length 𝑆𝑝 and output length 𝑆 out , we aggregate prefill, decode, and disaggregated KV transfer into a requestlevel throughput: 𝐵𝑆 out
TPS(𝑚, 𝐷) =
𝑆𝑝∑︁ +𝑆 out
.
𝐵𝑆𝑝 1 + + 𝑇KV TPSpre (𝑚, 𝐷) 𝑡 =𝑆 +1 TPSdec (𝑚, 𝑡, 𝐷) 𝑝
(17) Here TPSpre (𝑚, 𝐷) is evaluated at prompt length 𝑆𝑝 , and
Communication Locality Model 𝑇KV =
Communication depends on how much expert-parallel traffic remains within one local high-bandwidth domain. We estimate this from model fit in HBM.
2𝐿𝑤 𝑆𝑝 𝑏 kv 𝐵 transfer
captures KV transfer for disaggregated serving [40, 61]. 16
(18)
Table 2. Model configurations for the inference throughput study. 𝐿: transformer layers; 𝑤: hidden dimension; 𝐸: total experts; 𝐾: routed experts per token; 𝑆: evaluation context length. All MoE models use FF = 4𝑤. Model MoE-0.6T MoE-5T MoE-19T MoE-51T MoE-132T MoE-401T
𝐿
𝑤
𝐸
𝐾
𝑆
48 96 120 120 120 144
6 144 8 192 12 288 14 336 16 384 18 432
64 96 128 256 512 1 024
2 2 2 2 2 2
1 024 1 024 1 024 1 024 1 024 1 024
We vary package TDP across Low, Medium, and High scenarios: anchor 𝑃 pkg (𝜏, 𝑠) = 𝑃pkg (𝑠) (1 + 𝑔𝑠 )𝜏 −𝜏anchor ,
where 𝑠 ∈ {Low, Med, High} and 𝑔𝑠 ∈ {5%, 12.5%, 20%}. We hold performance to central projections so that scenario variation reflects power-density uncertainty rather than simultaneous changes in compute capability. For publicly disclosed near-term hardware, we use vendor-reported package-level anchors. For later years, we extrapolate FP4 FLOP/s, HBM bandwidth, and HBM capacity at constant annual rates of 30%, 15%, and 25%, respectively. Oberon is anchored at B200 in 2025 and Vera Rubin in 2026. Our later pod-scale study case is anchored at Rubin Ultra in 2027, held fixed through 2028, and extrapolated beginning in 2029. Deployment architecture specifies how packages are integrated into one deployment unit: package count, local NVLink-domain size, aggregate NVLink bandwidth, aggregate scale-out bandwidth, and non-package overhead power. These quantities change only at architectural transitions. Given package projections and deployment architecture parameters, we derive the rack-level quantities consumed by the throughput and placement models:
A.4 Model Inputs and Limitations The model consumes the following model-specific inputs: (𝐿, 𝑤,𝑊total,𝑊active ). For the model suite used in the paper, all models use 𝐾 = 2 routed experts per token and FF = 4𝑤. This model has three deliberate limitations. 1. It is a first-order comparative model, not a topologyaccurate runtime simulator. 2. Communication is modeled with bandwidth-time approximations rather than collective-specific kernels. 3. We do not model fine-grained overlap among TP communication, EP communication, and compute. A.5
Workload Model Suite
Table 2 lists the model configurations used in the throughput study. The MoE suite spans three orders of magnitude in total parameters, from a 0.6 T model whose experts fit within a single rack-local NVLink domain to a 401 T model that requires expert-parallel communication across multiple domains. All MoE models use top-𝐾=2 routing and FF=4𝑤. Expert counts grow with model size, which increases the fraction of traffic that spills onto inter-domain links for a given deployment architecture.
B
𝐹𝐷 (𝜏) = 𝑁 pkg · 𝐹 pkg (𝜏),
(20)
HBM HBM 𝐵𝐷 (𝜏) = 𝑁 pkg · 𝐵 pkg (𝜏),
(21)
𝐻𝐷usable (𝜏) = 𝛼 · 𝑁 pkg · HBMpkg (𝜏),
(22)
𝑃rack (𝜏, 𝑠) = 𝑁 pkg · 𝑃 pkg (𝜏, 𝑠) + 𝑃ovhd .
(23)
NVL and 𝐵 IB are taken directly from Table 3. We use Here 𝐵𝐷 𝐷 package counts throughout, since HBM capacity, package TDP, and NVLink-domain membership are package-level quantities.
B.2
Pod and Non-GPU Assumptions
A deployment pod is a co-procured, co-placed multi-rack unit. We distinguish between rack-scale deployment units and pod-scale deployment units. In the baseline model, each rack retains its own local NVLink domain, so pods change placement quantum but not rack-local communication structure: pod rack 𝑁 NVL = 𝑁 NVL . (24) Pod power is the sum of the constituent racks: ∑︁ 𝑃pod (𝜏, 𝑠) = 𝑃𝑟 (𝜏, 𝑠), (25)
Hardware and Cost Projections
We instantiate the fleet study with comparative projections for GPU package power, GPU package capability, deploymentunit architecture, non-GPU rack power, and facility infrastructure cost. Package-level trends determine compute, HBM bandwidth, HBM capacity, and TDP growth. Deployment architecture determines rack power, local-domain size, and communication bandwidth. These assumptions are used for sensitivity analysis, not product forecasting. B.1
(19)
𝑟 ∈ R pod
where R pod is the set of rack types in the pod. For a homogeneous pod, this reduces to 𝑁 racks 𝑃rack . General-compute racks are anchored at 20 kW in 2025 and grow at {3%, 5%, 8%} annually, reaching {26, 38, 52} kW by 2034. Storage racks are anchored at 15 kW in 2025 and grow at {2%, 4%, 6%} annually, reaching {18, 22, 26} kW by 2034. These trajectories define the non-GPU rack power inputs used by the SKU generation procedure. Unless otherwise
GPU Deployment Projections
A GPU package is the atomic unit of the projection model. A package may contain one or more compute dies together with its attached HBM. Package-level quantities determine TDP, compute throughput, HBM bandwidth, and HBM capacity. 17
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
Table 3. Deployment architecture parameters. NVLink values are aggregate unidirectional bandwidth per local NVLink domain; scale-out values are aggregate per deployment unit.
Architecture
Available
𝑁 pkg (pkgs)
DGX-H200 Blackwell–Oberon Vera Rubin NVL72 Kyber / Rubin Ultra
2024 2025 2026+ 2027+
8 72 72 144
Dies /pkg
NVL domain (pkgs)
NVL 𝐵𝐷 (TB/s)
IB 𝐵𝐷 (TB/s)
𝑃ovhd (kW)
1 1 2 4
8 72 72 144
3.6 64.8 259.2 750.0
0.4 7.2 14.4 57.6†
3 25 30 35
† Later pod-scale study assumption derived from public Rubin Ultra system disclosures under our unidirectional aggregate convention; detailed NIC topology
remains unsettled in public materials.
Table 4. Per-package performance projections used to instantiate the throughput model. Oberon is anchored at B200 (2025) and Vera Rubin (2026); the later pod-scale study case is anchored at Rubin Ultra (2027). Post-anchor extrapolation begins in 2029. Oberon Year
𝐹 (PF)
𝐵 HBM (TB/s)
2025 2026 2027 2028 2029 2030 2031 2032 2033 2034
10.0 50.0 50.0 50.0 65.0 84.5 109.9 142.8 185.6 241.3
8.0 22.0 22.0 22.0 25.3 29.1 33.5 38.5 44.2 50.9
Kyber / Rubin Ultra HBM (GB)
𝐹 (PF)
𝐵 HBM (TB/s)
HBM (GB)
192 288 288 288 360 450 563 703 879 1,099
100.0 100.0 130.0 169.0 219.7 285.6 371.3 482.7
32.0 32.0 36.8 42.3 48.7 56.0 64.4 74.0
1,024 1,024 1,280 1,600 2,000 2,500 3,125 3,906
Table 5. Derived rack power (kW) across growth scenarios. Anchor values follow announced or study-anchor specifications; later values are extrapolated using Eq. 23. Oberon (𝑁 pkg = 72)
Kyber / Rubin Ultra (𝑁 pkg = 144)
Year
Low
Med
High
Low
Med
High
2025 2026 2027 2028
157 160 166 173
180 178 197 218
203 196 226 262
— — 515 515
— — 600 600
— — 685 685
2029 2030 2031 2032 2033 2034
180 188 197 205 214 224
243 271 303 339 379 425
341 434 545 677 836 1,025
539 564 591 619 648 679
671 750 839 940 1,053 1,180
815 971 1,158 1,382 1,652 1,975
noted, compute and storage use the Medium scenario and GPU racks vary across scenarios. B.3
not intended to predict operator-specific build cost. Values include equipment and installation and are intended as representative installed costs rather than hyperscaler-specific internal costs.
Facility Cost Assumptions
The cost model is used only to compare designs under a common infrastructure baseline and to instantiate the initial and effective $/MW metrics in the main evaluation. It is 18
Table 6. Facility infrastructure cost assumptions per MW of IT capacity. Component
Cost/MW
UPS systems Battery systems Backup generators MV transformers MV switchgear LV switchboards Automatic transfer switches Static transfer switches Row distribution (PDUs/busway) Busbar overhead Cooling systems Facility shell, site & engineering Fit-out & other
$1,000,000 $275,000 $750,000 $120,000 $60,000 $150,000 $70,000 $250,000 $100,000 $6,000 $3,000,000 $1,800,000 $2,800,000
C
Datacenter Reference Designs, Placement, and Parameters
C.1
Placement Feasibility
reserves failover headroom within each active line-up. For a high-availability deployment, 𝐶 ℓeff =
This subsection formalizes the hierarchical placement constraint used by the simulator.
d𝑟 = (𝑃𝑟 , CFM𝑟 , LPM𝑟 , 𝑛𝑟 ) , where 𝑃𝑟 is power demand (kW), CFM𝑟 is air-cooling demand, LPM𝑟 is liquid-cooling demand, and 𝑛𝑟 is tile count. GPU pods may consume all four resources. General-compute and storage racks have LPM𝑟 = 0.
Non-power resources. Cooling and space constraints are enforced at row and ancestor nodes exactly as in Eq. 26. Networking is provisioned at build time and is not modeled as a binding online placement constraint. Space is limited by the fixed tile count per row.
Ancestor-path feasibility. We model the power-delivery hierarchy as a rooted tree whose internal nodes are distribution components with per-resource capacities. For a candidate row location ℓ, let path(ℓ) = {ℓ0, ℓ1, . . . , ℓℎ }
C.2
denote the ancestor path from the row (ℓ0 ) to the substation (ℓℎ ). Placement of deployment 𝑟 at location ℓ is feasible if and only if ∀ℓ𝑘 ∈ path(ℓ), ∀𝑚,
(27)
at each line-up node ℓ. Low-availability deployments may use the full rated capacity 𝐶 ℓ and therefore consume reserve capacity. For block redundancy (𝑁 + 𝑘), 𝑁 primary line-ups carry IT load and 𝑘 standby line-ups are reserved for failover. Each primary line-up may therefore be loaded to its rated capacity, so 𝐶 ℓeff = 𝐶 ℓ for all placed deployments. The standby lineups do not constrain placement at the row level, but they do contribute to total hall cost.
Demand vector. Each deployment unit 𝑟 has resource demand vector
𝐿ℓ(𝑚) + 𝑑𝑟(𝑚) ≤ 𝐶 ℓ(𝑚) ,eff 𝑘
𝑦 𝐶ℓ , 𝑥
Counting Rows
Reference designs are defined by a set of UPS line-ups, a partition into power domains, and balanced row-to-line-up wiring. Balance means that all distinct connection patterns allowed by a design appear equally often.
(26)
𝑘
source dimension 𝑚, 𝑑𝑟(𝑚) is the corresponding component of d𝑟 , and 𝐶 ℓ(𝑚) is the effective capacity at that node after 𝑘 ,eff accounting for redundancy constraints.
Block-redundant designs. In a block-redundant design, all rows in one power domain connect to the same set of active line-ups. If a hall has 𝑁 line-ups partitioned into 𝑘 power domains, then the row count in each class must be a multiple of 𝑁𝑘 .
Effective capacity under redundancy. Effective electrical capacity depends on redundancy topology and availability tier. For distributed redundancy (𝑥𝑁 /𝑦), a system with 𝑥 total line-ups and 𝑦 line-ups of usable high-availability capacity
Distributed-redundant designs. In a distributed-redundant design, balance requires all admissible connection combinations within a power domain to be represented equally. We use two row classes. Low-density rows connect to two up stream line-ups, so their count must be a multiple of 𝑁2/𝑘 .
where 𝐿ℓ(𝑚) is the current aggregate load at node ℓ𝑘 in re𝑘
19
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare, Chaojie Zhang, and Ricardo Bianchini
integer multiples of 𝑁2/𝑘 and 𝑁4/𝑘 that make the ratio of high-density to low-density rows as close as possible to the block-design reference while keeping total row counts comparable. We bias the reference designs toward more low-density rows because low-density rows absorb residual line-up capacity that cannot be consumed by high-power deployments once high-density rows saturate. Without enough low-density rows, capacity strands at the line-up level even when aggregate hall power remains available.
High-density rows connect to four upstream line-ups, so their count must be a multiple of 𝑁4/𝑘 . Row classes. Each row contains 24 rack positions. Lowdensity rows use two upstream feeds. High-density rows use four. This two-class model is a stylized approximation used to compare designs under common assumptions; it does not exclude other feed counts or row constructions. For block-redundant designs, we use a base hall with 6𝑁 low-density rows and 4𝑁 high-density rows. For distributedredundant designs, row counts must respect the balancedcombination rules above. We therefore choose the smallest
20