Conceptio › Archive › arXiv CS
arXiv CSopen access

Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

arXiv:2609.24639v1 [cs.DC] 21 Sep 2026

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference Mingyuan Yan

Haiyu Wang

Linxuan Biao

[email protected] Department of Electrical and Computer Engineering New York University Brooklyn, New York, USA

[email protected] Department of Electrical and Computer Engineering New York University Brooklyn, New York, USA

[email protected] Department of Electrical and Computer Engineering New York University Brooklyn, New York, USA

H. Jonathan Chao

Sai Qian Zhang

Wenqi Cui∗

[email protected] Department of Electrical and Computer Engineering New York University Brooklyn, New York, USA

[email protected] Department of Electrical and Computer Engineering New York University Brooklyn, New York, USA

[email protected] Department of Electrical and Computer Engineering New York University Brooklyn, New York, USA

Abstract Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption. Prefill–decode (PD) disaggregation has emerged as a prevalent architecture for large-scale inference serving. However, determining the appropriate numbers of prefill and decode instances is challenging because serving capacity depends jointly on workload characteristics, hardware constraints, queueing, and KV-cache reservations. Existing approaches largely rely on profiling and simulation, providing limited analytical insight into how provisioning decisions shape the tradeoff between serving capacity and power consumption. This paper develops an analytical framework for power-aware provisioning of PD-disaggregated AI inference. Given an inference workload and hardware, the framework models the serving capacity and average power consumption of a provisioned deployment. The serving-capacity model is derived from the joint distribution of input–output lengths and hardware compute and memory limits. In particular, it explicitly captures the coupling between prefill and decode induced by KV-cache reservations, as well as the impact of request queueing. On this basis, the power model determines perinstance power consumption as a function of normalized serving throughput. Together, the models determine the serving capacity– power Pareto front among candidate provisioned deployments, enabling the service provider to choose a provisioned deployment as the workload or available power changes.

1

Introduction

Electricity demand from datacenters is growing rapidly and is expected to add hundreds of terawatt-hours to annual electricity demand, with AI accounting for a substantial share of this growth [20]. Expanding electricity generation and grid capacity, however, typically takes much longer than constructing a datacenter [9, 21]. As a result, datacenter projects increasingly face delayed or denied grid connections [40, 48], as well as conditional interconnection agreements that limit when and how much power a datacenter may ∗ Corresponding author.

consume from the grid [10]. These constraints make two properties of computing loads increasingly valuable: efficiency in delivering a required service while consuming as little power as possible, and flexibility in reducing power demand while maintaining as much serving capacity as possible. Within AI workloads, inference is becoming a major and persistent source of electricity demand [42]. Both request volume and request size are increasing: large language model (LLM) based services are becoming routine [6, 11], agentic workflows turn one user task into multiple model calls [35, 56], and reasoning models use more tokens before producing an answer [17]. These trends make power-aware inference provisioning increasingly important, both to reduce the power required for a given serving capacity and to understand how much serving capacity can be maintained as power is reduced. An AI inference request is typically processed in two phases: prefill and decode. During prefill, the model processes the entire input sequence, constructs the key-value (KV) cache reused during decode, and generates the first output token. During decode, the model reuses that cache and generates the remaining output one token at a time [57]. Prefill is highly parallel and generally computebound, whereas decode repeatedly reads model weights and KVcache data and is generally limited by memory bandwidth. This difference motivates prefill–decode (PD) disaggregation [41, 61], which places the two phases in separate pools of model instances that can be provisioned independently. This architecture has consequently been widely used in a range of large-scale inference systems [11, 22, 32, 45, 49, 62]. Choosing how many prefill instances and how many decode instances to run is a central provisioning decision for PD-disaggregated AI inference. For a given workload and hardware platform, this choice determines the system’s power consumption and its serving capacity (i.e., the maximum request rate it can sustain). Overprovisioning provides serving capacity beyond the required level but incurs additional resource and power costs, whereas underprovisioning fails to meet the required serving capacity. The two pools must also be balanced: insufficient serving capacity in either phase creates a bottleneck and can leave resources in the other phase

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

underutilized. These tradeoffs motivate a joint characterization of serving capacity and power across provisioned deployments. Existing PD provisioning methods commonly evaluate candidate deployments using simulation [15, 19, 41, 61] or empirical performance measurements [31]. However, hardware profiling is timeconsuming, while simulation-based comparisons must be repeated when the request-length distribution changes. Analytical models reduce this dependence on empirical evaluation [4, 19, 37, 47, 61], but existing approaches do not capture the coupling between prefill and decode in a PD-disaggregated deployment. Consequently, they cannot characterize deployment-level serving capacity, and load-dependent power consumption is generally not modeled. This motivates the central question of this paper: Can we develop analytical models for PD-disaggregated AI inference that characterize how provisioning decisions shape the tradeoff between serving capacity and power? We therefore develop analytical models of serving capacity and average deployment power that apply across different workload distributions and hardware platforms. The key challenge in characterizing serving capacity is that prefill activity changes the cache memory available for decode. Each request reserves KV-cache space on its assigned decode instance before prefill begins. As a result, the decode-side KV-cache pool is shared by active decode requests and requests waiting for or undergoing prefill. This coupling limits the number of active decode requests, thereby affecting decode capacity and the characterization of serving capacity. To this end, we develop a serving-capacity model from the joint distribution of input and output lengths and the hardware compute and memory limits. The model accounts for the KV-cache space reserved while requests wait for or undergo prefill, which reduces the memory available to active decode requests. We model prefill and decode capacities separately and combine their aggregate capacities to obtain the end-to-end serving capacity. We then develop a power model that relates each instance’s power consumption to its normalized serving throughput. Together, the models determine the serving capacity–power Pareto front among candidate provisioned deployments. The front shows the minimum-power provisioned deployment for a required serving capacity and how much serving capacity can be maintained under a reduced power budget. We evaluate the models across different workloads, including workloads derived from the Mooncake [45] and Azure [41] production traces. The main contributions of the paper are summarized below: • Analytical serving-capacity model. We develop an analytical model of the serving capacity of PD-disaggregated inference from the joint distribution of request input and output lengths, as well as hardware compute and memory limits. The model explicitly captures the coupling between prefill and decode induced by KV-cache reservations, as well as the impact of request queueing. • Load-dependent power model. We develop separate power models for prefill and decode instances as functions of normalized serving throughput. The model captures how the provisioned numbers of instances and the serving load jointly determine the average power consumption of the deployment. • Serving capacity–power Pareto front. Combining the servingcapacity and power models yields a discrete Pareto front over

candidate provisioned deployments. This characterization quantifies the tradeoff between power consumption and serving capacity, providing a principled basis for grid-friendly inference control and provisioning decisions. • Provisioning for efficiency and flexibility. We use the Pareto front to identify minimum-power provisioned deployments for a required serving capacity and to quantify how much serving capacity can be maintained as available power is reduced. We evaluate these provisioning decisions across different workloads and candidate deployments against measurements.

1.1

Related Work

This work is closely related to topics on provisioning for PD-disaggregated inference, analytical models, and power-aware control. PD Disaggregation and Serving-Capacity Models. Under PD disaggregation [41, 61], the serving capacity of AI inference depends on the numbers of prefill and decode instances and the request workload. Existing provisioning methods commonly evaluate candidate deployments through simulation [15, 19, 41, 61] or throughput measurements [31]. However, hardware profiling [19, 61] can be time-consuming, while simulation-based comparisons must be repeated when the workload distribution changes. To reduce reliance on purely empirical evaluation, several works introduce analytical models for inference performance. DistServe [61] and BestServe [19] develop analytical models for batch execution times, but still rely on simulation to determine the request rates that meet latency requirements. A queueing-based method [31] derives the supported prefill request rate by applying an analytical queueing model to measured prefill performance, whereas it obtains decode throughput directly from measurements. There also exist other analytical models [4, 37, 47] for architectures that are not PD-disaggregated, but they cannot be directly applied to capture the coupling between the prefill and decode pools. In addition, these models alone are not sufficient to characterize the tradeoff between serving capacity and power. To address these gaps, this work develops analytical models for both PD serving capacity and power, while accounting for prefill–decode coupling induced by KV-cache reservations and the impact of request queueing. Power Modeling. Analytical power models for PD-disaggregated AI inference remain limited. Splitwise [41] determines the numbers of prefill and decode instances under a power budget, but models deployment power as proportional to the number of provisioned instances, without accounting for the effect of workload on power consumption. Several studies model the relationship between power consumption and batch size [8, 34, 51]. Other approaches estimate average GPU power using kernel-level information [30] or publicly available model and GPU specifications [14]. However, none of these models captures how the numbers of prefill and decode instances, together with the workload served by each pool, jointly determine the average power consumption of a PD-disaggregated deployment. A separate line of work models energy per request or token [13, 38, 43, 50, 52]; however, such energy models do not directly characterize the aggregate power demand of an inference deployment. In contrast, this paper models prefill and decode power separately as functions of normalized serving throughput, thereby

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

capturing how workload-dependent utilization of the two pools translates into deployment-level power consumption. Power-Aware Control of Inference Workloads. Runtime power control typically starts from an already provisioned inference deployment and adjusts its operating configuration in response to workload or power conditions. Some methods keep the number of instances fixed and control GPU frequency [36, 58], batch size [18], or quantization [12]. For PD disaggregation, runtime systems adjust the numbers of prefill and decode instances separately, to reduce energy consumption [5], meet power budgets [23, 33], or respond to feedback from the running fleet [22, 54]. Other approaches regulate the aggregate power demand through coordination with external energy resources, such as battery energy storage systems, uninterruptible power supplies (UPSs), and supercapacitors [26, 27, 55]. In contrast, this paper focuses on the provisioning configuration itself, using the numbers of prefill and decode instances as the control variables. This choice requires an analytical characterization of how each candidate provisioning configuration determines both serving capacity and power consumption, which enables systematic comparison among candidate deployments.

2 Problem Formulation 2.1 Provisioning and Workload An inference request contains an input sequence of ℓin tokens and produces an output sequence of ℓout tokens. AI inference proceeds in two phases. The prefill phase processes the complete input, constructs the KV cache, and produces the first output token. The decode phase then produces each remaining token in a separate forward pass. Under PD disaggregation, the two phases run on separate pools of model instances. Here, an instance is one running replica of the model on its assigned hardware. The service provider provisions 𝑛𝑃 prefill instances and 𝑛𝐷 decode instances. Together, these two numbers define a provisioned deployment (𝑛𝑃 , 𝑛𝐷 ). We label deployments as 𝑛𝑃 p𝑛𝐷 d; for example, 3p2d denotes a deployment with three prefill instances and two decode instances. The numbers of prefill and decode instances required depend on the inference workload. We denote the workload by W := (L, 𝜆). It consists of the joint distribution L of the input and output lengths (ℓin, ℓout ), and the mean request arrival rate 𝜆. The input length determines the computational work in prefill, while both the input and output lengths determine the work in decode. We assume that request arrivals are Poisson and that the workload remains stationary within each provisioning period. When the arrival rate or the length distribution changes, the provisioned deployment is chosen again for the new workload. The provisioned deployment needs to satisfy the service-quality requirement, which is defined by service level objectives (SLOs) and is commonly a bound on the latency its users experience. The SLOs mainly bound time to first token (TTFT), the time from the arrival of a request to its first output token, and time per output token (TPOT), the time between consecutive output tokens. A provisioned deployment meets its SLOs when both TTFT and TPOT stay below their bounds. Together, the workload W and the provisioned deployment (𝑛𝑃 , 𝑛𝐷 ) determine the inference system’s serving capacity, its

power consumption, and the latency its users experience. Provisioning chooses (𝑛𝑃 , 𝑛𝐷 ) for a given workload so that the fleet meets its SLOs while consuming as little power as possible.

2.2

Power-Aware Provisioning

The service provider seeks the provisioned deployment that provides the required serving capacity while consuming as little power as possible. For a provisioned deployment (𝑛𝑃 , 𝑛𝐷 ) under a workload W, let P (𝑛𝑃 , 𝑛𝐷 ; W) denote the average power of that provisioned deployment over a long serving window. Its serving capacity 𝜇 (𝑛𝑃 , 𝑛𝐷 ; W) is the maximum request rate it can sustain [1]. To provide the required headroom, the provisioned deployment must have a serving capacity of at least 𝜆min . The resulting provisioning problem is min P (𝑛𝑃 , 𝑛𝐷 ; W) 𝑛𝑃 ,𝑛𝐷 ∈Z+ (1) s.t. 𝜇 (𝑛𝑃 , 𝑛𝐷 ; W) ≥ 𝜆min . We model average power rather than instantaneous power, since short-term power fluctuations of individual instances tend to average out across a large provisioned deployment. Peak-power behavior is outside the scope of this paper. A provisioned deployment that satisfies this constraint is feasible. The constraint compares the serving capacity of a provisioned deployment with the required serving capacity of the workload and its SLOs. A change in the workload W requires solving (1) again. The required serving capacity 𝜆min is chosen based on the current request arrival rate and the headroom it requires to handle fluctuations in arrivals and satisfy the SLOs. We therefore define 𝜆min = 𝜌𝜆¯ , where 0 < 𝜌¯ < 1 is the maximum serving utilization chosen by the service provider. A lower 𝜌¯ reserves more headroom for tighter SLOs and variations in arrivals. Prior work uses values of 0.85 and 0.90 [7, 8, 37]. The parameter 𝜌¯ can also serve as a control parameter for power-responsive operation. When grid power is constrained, the service provider can temporarily reduce the re¯ thereby lowering the required served headroom by increasing 𝜌, serving capacity 𝜆min and enabling operation with fewer instances and lower power consumption, provided that the resulting service degradation remains within short-term tolerance.

2.3

Serving Capacity–Power Pareto Front

Although the provisioning problem (1) is an integer optimization problem, its decision variables consist only of the numbers of prefill and decode instances, (𝑛𝑃 , 𝑛𝐷 ). This low-dimensional structure makes it unnecessary to develop a dedicated integer optimization algorithm. Instead, the analytical models developed in Sections 3 and 4 can be used to obtain a discrete Pareto front between power and serving capacity, where the provisioning problem can then be reduced to a search over the resulting discrete Pareto front. Specifically, each candidate (𝑛𝑃 , 𝑛𝐷 ) corresponds to a point in the serving capacity–power plane, as illustrated in Figure 1. A deployment is Pareto-optimal if no other deployment provides at least the same serving capacity with lower power consumption. The collection of such deployments forms the discrete Pareto front, which characterizes the minimum power required to achieve different levels of serving capacity. Among the Pareto-optimal deployments satisfying the serving-capacity requirement 𝜆min , the one with the

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

6000

Power (W)

5000 4000 3000 2000 1000

3p5d

2p6d 3p2d

2p4d

1p6d

4p3d

3p4d

2p5d

1p7d

4p4d

5p3d 6p2d 5p2d

3p3d

4p2d

2p3d

1p5d

2p2d

1p4d

4p1d to 7p1d

2p1d

1p3d

3p1d

1p1d

4

6

8

Serving capacity (req/s)

10

12

Figure 1: Modeled serving capacity and power of candidate deployments under a fixed-length workload, with the corresponding Pareto front. Points are labeled by deployment. lowest power consumption is selected as the optimal provisioned deployment for solving (1). Thus, despite its integer formulation, the two-dimensional provisioning problem admits a simple geometric solution. However, the Pareto front may change with the operating conditions, particularly with the request workload distribution. For a fixed workload distribution, an empirical Pareto front can be constructed from measurements. Such empirical profiling, however, is time-consuming, and the number of candidate deployments increases rapidly with the number of provisioned instances. Repeating this process for every possible workload distribution is therefore impractical. To address this challenge, we develop analytical models for the power model P (𝑛𝑃 , 𝑛𝐷 ; W) and the serving-capacity model 𝜇 (𝑛𝑃 , 𝑛𝐷 ; W). Given an online workload distribution W, these models directly evaluate the power and serving capacity of each candidate deployment, allowing the corresponding Pareto front to be constructed without repeated empirical profiling. Figure 1 shows the serving capacity–power plane the models determine for one workload. In the next sections, we present these two models and derive their analytical forms.

Analytical Models

This section develops analytical models that relate a workload and a provisioned deployment to its serving capacity and average power. We first model the serving capacities of the prefill and decode pools, then characterize the KV-cache constraint that determines the operating batch, and finally derive the deployment-level serving capacity and power. Section 4 provides the detailed derivations and calibration procedures.

3.1

Prefill

Decode

resource demand peak performance utilization processing time

𝐹 (ops) 𝜋 (ops/s) MFU 𝑡𝑃

𝑄 (bytes) 𝛽 (bytes/s) MBU 𝑡𝐷

modeling unit # requests / unit # tokens / request

request 1 ℓin

batch iteration 𝐵 1

Provisioned deployments Pareto front

1p2d

2

3

Table 1: Serving-capacity model characterization.

Serving Capacity Model

Each request runs first on a prefill instance and then on a decode instance, with its KV-cache transferred between them. Modern KVtransfer backends [28, 41, 45] perform this transfer asynchronously and overlap much of it with prefill computation. We therefore assume that KV-cache transfer does not limit serving capacity and focus on the prefill and decode pools. Let 𝜇𝑃 and 𝜇𝐷 denote the serving capacities of each prefill and decode instance, respectively.

A deployment with 𝑛𝑃 prefill instances and 𝑛𝐷 decode instances therefore provides aggregate capacities of 𝑛𝑃 𝜇𝑃 and 𝑛𝐷 𝜇𝐷 . Since every request passes through both pools, the end-to-end serving capacity is limited by the bottleneck pool:  𝜇 = min 𝑛𝑃 𝜇𝑃 , 𝑛𝐷 𝜇𝐷 . (2) Prefill and decode are governed by different hardware bottlenecks. Prefill processes all input tokens in a single forward pass and is typically compute-bound, whereas decode generates one token per active request per iteration and repeatedly accesses model weights and KV-cache data, making it typically memory-bandwidth-bound. We therefore model prefill capacity from compute throughput and decode capacity from memory bandwidth. We use processing time as the intermediate variable for deriving serving capacity. For prefill, the processing time is the service time 𝑡𝑃 of a single request, whereas for decode, it is the duration 𝑡𝐷 of one batch iteration, during which each active request generates one output token. Despite this difference in interpretation, both phases follow the same roofline principle: processing time =

resource demand . peak performance × utilization

Serving capacity then follows by converting this processing time from the phase-specific modeling unit in Table 1 to requests per second. For compute-bound prefill, the resource demand is measured by the number of floating-point operations, denoted by 𝐹 (ops). The peak compute throughput, denoted by 𝜋 (ops/s), is a hardware parameter determined by the accelerator architecture and numerical precision. We represent the achieved fraction of peak compute throughput by an effective model FLOPs utilization (MFU), which aggregates the execution efficiency of linear-layer and attention computation. The resulting prefill service time is 𝑡𝑃 = 𝐹 /(𝜋 MFU). For bandwidth-bound decode, the resource demand is measured by the amount of data transferred from memory, denoted by 𝑄 (bytes). The peak memory bandwidth, denoted by 𝛽 (bytes/s), is likewise a hardware parameter. We represent the achieved memory bandwidth by an effective memory bandwidth utilization (MBU), which aggregates the effects of model-weight and KV-cache accesses. The resulting decode-iteration time is 𝑡𝐷 = 𝑄/(𝛽 MBU). Table 1 summarizes characterizations for the two models. Although prefill and decode share the same resource-based timing principle, they require different conversions from processing time to serving capacity, as derived next. 3.1.1 Serving Capacity of a Prefill Instance. Prefill is computebound [2, 44], so we quantify its processing demand in FLOPs.

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

Because matrix multiplications account for most of prefill computation, the demand 𝐹𝑟 for request 𝑟 with ℓin,𝑟 input tokens is 2 𝐹𝑟 = 2𝑁 ℓin,𝑟 + 𝑐 𝑎 𝐿𝑑 ℓin,𝑟 , | {z } | {z } linear layers

(3)

attention

where 𝑁 is the number of model parameters, 𝐿 is the number of transformer layers, and 𝑑 is the attention width. The linear term approximates the dense linear-layer cost as one multiplication and one addition per parameter per input token. The quadratic term comes from causal attention, and 𝑐 𝑎 depends on how the attention kernel applies the causal mask. Under the roofline model [53], the prefill instance runs at an achieved compute rate of 𝜋 MFU. Dividing this FLOP demand by the achieved compute rate 𝜋 MFU gives the prefill service time 𝑡𝑃,𝑟 : 𝐹𝑟 𝜋 MFU 2𝑁 𝑐 𝑎 𝐿𝑑 2 = ℓ . ℓin,𝑟 + 𝜋 MFU 𝜋 MFU in,𝑟 | {z } | {z }

𝑡𝑃,𝑟 =

=:𝑎𝑃

(4)

=:𝑏 𝑃

A prefill instance’s serving capacity is the reciprocal of its mean service time over requests: 1 = E𝑟 [𝑡𝑃 ] 𝜇𝑃

(5)

= 𝑎𝑃 E𝑟 [ℓin ] + 𝑏 𝑃 E𝑟 [ℓin2 ]. where E𝑟 [·] denotes the expectation over requests drawn from the length distribution L. Prefill capacity therefore depends on the first two moments of the input-length distribution. 3.1.2 Serving Capacity of a Decode Instance. At typical serving batch sizes, decode is memory-bound [44, 59]. Modern LLM serving systems commonly use continuous batching, in which completed requests leave the decode batch and requests that complete prefill are admitted to it between decode iterations. Consequently, the number of active requests varies over time, so decode capacity depends jointly on batching dynamics, request queueing, and perrequest processing. We show the resulting decode-capacity model below and elaborate its derivation in Section 4.1. Because decode is memory-bandwidth-bound, we model its processing demand in terms of memory traffic. We define the operating batch 𝐵 as the mean number of active requests across decode iterations. In each iteration, the decode instance reads the model weights once, requiring 2𝑁 bytes in bf16, and accesses the KV cache of every active request, with 𝜅 bytes transferred per context token. The mean memory traffic 𝑄¯ per decode iteration is therefore

h  i E𝑟 (ℓout − 1) ℓin + ℓout 2   ℓ¯ctx = . (7) E𝑟 ℓout − 1 The mean decode-iteration time 𝑡¯𝐷 is the mean memory traffic 𝑄¯ divided by the achieved memory bandwidth 𝛽 MBU: 𝑄¯ 𝑡¯𝐷 = 𝛽 MBU 𝜅 ℓ¯ctx 2𝑁 (8) + 𝐵. = 𝛽 MBU 𝛽 MBU | {z } | {z } =:𝑎𝐷

We next convert iteration time into request-level serving capacity. Each decode iteration generates one token for each active request and therefore produces 𝐵 output tokens on average. Since prefill generates the first output token, each request requires E𝑟 [ℓout − 1] decode tokens on average. The mean decode time per request is therefore 1 E𝑟 [ℓout − 1] 𝑡¯𝐷 . (9) = 𝜇𝐷 𝐵 The remaining unknown is the operating batch 𝐵, which is constrained by the available KV-cache memory. The next subsection outlines the model used to characterize 𝐵. 3.1.3 Operating Batch of a Decode Instance. One key challenge is that the operating batch 𝐵 is not a predetermined parameter. Modern serving systems commonly employ continuous batching, where the instantaneous batch size is determined dynamically to utilize the KV-cache capacity of a decode instance while reserving sufficient space for requests undergoing prefill to be admitted to decode upon completion [45]. For a request with input length ℓin , we denote this reservation by ℓin + 𝑅 token slots, where 𝑅 is the additional space reserved for subsequently generated output tokens. Prefill reservations reduce the KV-cache space available to active decode requests (Figure 2). The operating batch is therefore determined by both the KV-cache capacity and the reservations held during prefill. Let 𝐶 tok denote the total KV-cache capacity of a decode instance. Averaged over iterations, this capacity is divided among three components: (i) reservations held by requests waiting for or undergoing prefill, (ii) KV-cache space occupied by requests actively decoding, and (iii) a small amount of unused capacity when the next reservation cannot fit. Their mean occupancies satisfy 𝐶 tok ≈ 𝜇𝐷 (𝑡¯𝑊 E𝑟 [ℓin + 𝑅] + E𝑟 [𝑡𝑃 (ℓin + 𝑅)]) | {z } prefill occupancy

 + 𝐵 ℓ¯ctx + 𝑅 | {z } decode occupancy

𝑄¯ =

2𝑁 |{z} model weights

+ 𝜅𝐵 ℓ¯ctx , |{z}

(6)

KV cache

where ℓ¯ctx denotes the mean context length of an active request. The mean ℓ¯ctx differs from the mean context length of arriving requests: requests with longer outputs remain active for more decode iterations and are therefore sampled more often. Averaging over active request–iteration pairs gives

=:𝑏 𝐷

+

E𝑟 [(ℓin + 𝑅) 2 ] . 2 E𝑟 [ℓin + 𝑅] | {z }

(10)

unused capacity

where 𝑡¯𝑊 is the mean prefill waiting time, estimated using Kingman’s approximation as described in Section 4.2. Setting the prefill occupancy in (10) to zero gives the full-pool batch 𝐵 max , the operating batch when the KV-cache pool is fully available to active decode requests:  𝐶 tok − E𝑟 [(ℓin + 𝑅) 2 ]/ 2 E𝑟 [ℓin + 𝑅] 𝐵 max = . (11) ℓ¯ctx + 𝑅

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

Compute

Prefill instance Queue (tW) Request arrival

rn

⋯

r5

r4

Decode instance Continuous batching

Prefill (tP) r3

r2

r1

ra

rb

rc

rd

re

⋯

rf

Request completion

rz

r1: ℓin, ℓout

ra

r2: ℓin, ℓout

⋮ rn: ℓin, ℓout

Reserve

Memory r1

rb

Occupy

KV-cache memory of the decode instance (Ctok) Prefill reservations Active decode requests r2 rn ra rb rc ⋯ ⋯

∑r∈Pₖ (ℓin,r + R)

⋮ rz Unused

rz

∑r∈Dₖ (ℓctx,r,k + R)

Uk

Figure 2: Request processing and decode-side KV-cache allocation under PD disaggregation. Requests waiting for or undergoing prefill reserve space in the same KV-cache pool used by active decode requests.

3.2

Prefill Power (W)

With prefill reservations, the operating batch is smaller than 𝐵 max . Substituting the expressions for 𝑡¯𝐷 from (8) and 𝜇𝐷 from (9) into the memory balance (10) together with the waiting-time approximation yields a cubic equation in 𝐵. We show that this equation has a unique solution within the feasible operating regime, with details given in Theorem 4.3.

𝜆𝑃,𝑖 𝜆˜𝑃,𝑖 = , 𝜇𝑃

𝜆𝐷,𝑖 𝜆˜𝐷,𝑖 = max , 𝜇𝐷

400

Fixed, at capacity Fixed, below capacity Variable, at capacity

200 0.0

Power of a Provisioned Deployment

We next model the average power consumed by a provisioned deployment over a long serving window. The power consumption of an inference instance is closely related to the utilization of its underlying hardware resources. We find that a simple and effective characterization is the serving throughput of an instance relative to its serving capacity. Intuitively, an instance operating close to its serving capacity keeps the bottleneck hardware resource highly utilized, whereas an instance serving only a small fraction of its serving capacity leaves more of that resource idle. We therefore model power from the ratio between the instance serving throughput and its capacity. Let 𝜆𝑃,𝑖 denote the serving throughput of prefill instance 𝑖, and 𝜆𝐷,𝑖 that of decode instance 𝑖. We define their normalized serving throughputs as (12)

where 𝜇𝐷max denotes the full-pool decode capacity, defined as the serving capacity when the KV-cache pool is fully available to active decode requests, i.e., in the absence of prefill reservations. Equivalently, 𝜇𝐷max = 𝜇𝐷 (𝐵 max ), the decode capacity of (9) evaluated at the full-pool batch of (11). Requests in prefill reserve part of the KV-cache pool, so the operating batch of (10) is smaller than 𝐵 max . A decode instance may therefore consume less power than at saturation even when the provisioned deployment operates at its serving capacity. A value of 𝜆˜ = 0 corresponds to an idle instance, while 𝜆˜ = 1 corresponds to its reference capacity. Across both prefill and decode and all calibration workloads in Appendix B, per-instance power increases approximately linearly with normalized serving throughput at low to moderate load and then saturates at high load, as shown in Figure 3. We therefore

Decode

600

0.5

1.0

0.0

Normalized throughput λ ̃

0.5

1.0

Figure 3: Per-GPU power versus normalized serving throughput for prefill and decode. Points show measurements, and curves show the fitted capped-ramp models. ˜ using a capped linear function, model the per-instance power 𝑝 (𝜆) with separate parameters for the prefill and decode instances:  𝑝𝑠 (𝜆˜𝑠 ) = min 𝑝 0,𝑠 + 𝛾𝑠 𝜆˜𝑠 , 𝑝 sat,𝑠 , 𝑠 ∈ {𝑃, 𝐷 }, (13) where 𝑠 ∈ {𝑃, 𝐷 } indexes the prefill and decode instances, 𝑝 0,𝑠 is the static power floor, 𝛾𝑠 is the slope before saturation, and 𝑝 sat,𝑠 is the saturated power. The deployment power P is the sum of the power consumed by its prefill and decode instances: 𝑛𝑃 𝑛𝐷 ∑︁ ∑︁ P= 𝑝 𝑃 (𝜆˜𝑃,𝑖 ) + 𝑝 𝐷 (𝜆˜𝐷,𝑖 ). (14) 𝑖=1

𝑖=1

The sum in (14) holds for arbitrary per-instance serving throughputs. We assume that the load is balanced, so that the request rate 𝜆 is divided evenly among the instances of each pool, 𝜆𝑃,𝑖 = 𝜆/𝑛𝑃 and 𝜆𝐷,𝑖 = 𝜆/𝑛𝐷 , and the deployment power reduces to P =   𝑛𝑃 𝑝 𝑃 𝑛𝑃𝜆𝜇𝑃 + 𝑛𝐷 𝑝 𝐷 𝑛 𝜇𝜆max . When comparing provisioned de𝐷 𝐷 ployments, we evaluate this power at the serving capacity, 𝜆 = 𝜇, so that it depends only on (𝑛𝑃 , 𝑛𝐷 ) and the workload W.

4

Serving Capacity Model for Decode Instances

This section derives the serving-capacity model for decode instances introduced in Section 3. Throughout the analysis, E𝑟 , E𝑘 , and E𝑡 denote averages over admitted requests, decode iterations,

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

and continuous time, respectively. The analysis applies to arrival processes and workloads for which these averages converge to the corresponding expectations as the observation window grows.

4.1

mean iteration time 𝑡¯𝐷 gives 1 𝜇𝐷 |{z}

Decode Capacity Model

time per request

Unlike prefill, decode processes a changing set of active requests across iterations. This subsection provides a detailed derivation of the results presented in Section 3.1.2. Let D𝑘 denote the set of requests active in iteration 𝑘, and let 𝐵𝑘 = |D𝑘 | be the instantaneous batch size. We therefore define the operating batch as 𝐵 := E𝑘 [𝐵𝑘 ]. Each decode iteration reads the model weights once, requiring 2𝑁 bytes in bf16, and accesses the KV cache of every active request, transferring 𝜅 bytes per context token. For request 𝑟 ∈ D𝑘 , let ℓctx,𝑟,𝑘 denote its context length in iteration 𝑘, consisting of its input tokens and the output tokens generated before that iteration. The memory traffic in iteration 𝑘 is therefore ∑︁ 𝑄𝑘 = 2𝑁 +𝜅 ℓctx,𝑟,𝑘 . (15) |{z} 𝑟 ∈ D𝑘 model weights | {z } KV cache

Under the roofline model, dividing 𝑄𝑘 by the achieved memory 𝑄𝑘 bandwidth gives the decode-iteration time 𝑡𝐷,𝑘 = 𝛽 MBU . Averaging 𝑡𝐷,𝑘 over decode iterations gives the mean iteration time 𝑡¯𝐷 := E𝑘 [𝑡𝐷,𝑘 ] =

2𝑁 𝜅𝐵 ℓ¯ctx + . 𝛽 MBU 𝛽 MBU

(16)

By definition, ℓ¯ctx is the average over active request–iteration pairs, so the mean aggregate context length is 𝐵 ℓ¯ctx . Importantly, ℓ¯ctx is an active-request average rather than an average over arriving requests. Requests with longer outputs remain active for more decode iterations and therefore contribute more often to the aggregate context length. The following lemma expresses ℓ¯ctx in terms of the workload distribution. Lemma 4.1 (Mean active context length). Suppose each request has an input–output length pair (ℓin, ℓout ) that follows the workload distribution L, independently across requests. Then the mean context length over active request–iteration pairs satisfies Í  E𝑘 𝑟 ∈ D𝑘 ℓctx,𝑟,𝑘 ¯ ℓctx := E𝑘 [𝐵𝑘 ] 1 (17) E𝑟 [ℓout ] 2 Var(ℓout ) + Cov(ℓin, ℓout ) = E𝑟 [ℓin ] + + . 2 E𝑟 [ℓout − 1] | {z } | {z } arrival mean

=

residence correction

The variance term arises because requests with longer outputs appear in more active request–iteration pairs. The covariance term raises ℓ¯ctx when longer outputs tend to occur with longer inputs and lowers it when they tend to occur with shorter inputs [16, 47]. The detailed proof is given in Appendix A.1. The mean iteration time above describes the cost of one decode iteration. We next convert it into request-level serving capacity. Each decode iteration generates one token for every active request, or 𝐵 decode tokens on average. Since each request requires E𝑟 [ℓout − 1] decode tokens on average, the decode instance executes E𝑟 [ℓout − 1]/𝐵 iterations per admitted request on average. Multiplying by the

E𝑟 [ℓout − 1] · 𝐵 | {z }

𝑡¯𝐷 |{z}

,

iterations per request time per iteration

which is the decode-capacity relation in (9). The detailed derivation in Appendix A.2 accounts explicitly for the time-varying instantaneous batch size 𝐵𝑘 . Because this relation depends on decode execution through the mean iteration time 𝑡¯𝐷 , it also applies when the calibrated executiontime overheads in Appendix B.2 are included. The remaining quantity is the operating batch 𝐵, which we determine from the KV-cache memory balance next.

4.2

Operating Batch of a Decode Instance

This subsection provides a detailed derivation of the results presented in Section 3.1.3. We now derive the KV-cache memory balance in (10), which determines the operating batch 𝐵. After accounting for model parameters and runtime workspace, a decode instance has 𝐶 tok token slots available for its KV-cache pool, each with space for the key and value states of one token. In each decode iteration 𝑘, these slots are divided into three components. Let P𝑘 denote the admitted requests whose prefill has not completed; each request 𝑟 ∈ P𝑘 holds a reservation of ℓin,𝑟 + 𝑅 token slots. Each request 𝑟 ∈ D𝑘 active in decode occupies ℓctx,𝑟,𝑘 slots for its current context and reserves 𝑅 further slots. The remaining slots, denoted 𝑈𝑘 , are unused. At the start of iteration 𝑘 the three components fill Í Í the pool exactly, 𝐶 tok = 𝑟 ∈ P𝑘 (ℓin,𝑟 + 𝑅) + 𝑟 ∈ D𝑘 (ℓctx,𝑟,𝑘 + 𝑅) + 𝑈𝑘 . Averaging the number of slots in each component over decode iterations gives the mean KV-cache memory balance:  ∑︁     𝐶 tok = E𝑘  (ℓin,𝑟 + 𝑅)  𝑟 ∈ P𝑘    | {z } prefill occupancy

(18)  ∑︁    (ℓctx,𝑟,𝑘 + 𝑅)  + E𝑘 [𝑈𝑘 ] . + E𝑘  | {z }  𝑟 ∈ D𝑘   | {z } unused capacity decode occupancy

We next derive the three terms in (18): the mean prefill occupancy, the mean decode occupancy, and the mean unused capacity. Prefill occupancy. A request is admitted once KV-cache space has been reserved on its assigned decode instance. Its prefill waiting time 𝑡𝑊 is the time from admission to the start of prefill. Before prefill completes, each admitted request holds a KV-cache reservation on its assigned decode instance. Its contribution to prefill occupancy therefore depends on both the reservation size and its prefill holding time, the time spent waiting for and undergoing prefill. At the decode serving limit, each instance admits requests at its serving capacity. The following lemma relates the continuous-time mean occupancy to these request-level quantities. Lemma 4.2 (Mean prefill occupancy). Let P (𝑡) denote the admitted requests whose prefill has not completed at time 𝑡. Each request

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

holds ℓin +𝑅 token slots while waiting for and undergoing prefill. Then the continuous-time mean occupancy of these reservations is h ∑︁ i E𝑡 (ℓin,𝑟 + 𝑅) = 𝜇𝐷 E𝑟 [(𝑡𝑊 + 𝑡𝑃 )(ℓin + 𝑅)] . (19) 𝑟 ∈ P (𝑡 )

If, in addition, prefill serves requests in first-come, first-served (FCFS) order, request lengths are independent across arrivals, and 𝑡𝑃 = 𝑎𝑃 ℓin + 𝑏 𝑃 ℓin2 , then  E𝑟 [(𝑡𝑊 + 𝑡𝑃 )(ℓin + 𝑅)] = 𝑡¯𝑊 E𝑟 [ℓin ] + 𝑅  + 𝑎𝑃 E𝑟 [ℓin2 ] + 𝑅 E𝑟 [ℓin ] (20)  + 𝑏 𝑃 E𝑟 [ℓin3 ] + 𝑅 E𝑟 [ℓin2 ] , where 𝑡¯𝑊 = E𝑟 [𝑡𝑊 ] is the mean prefill waiting time. The occupancy relation (19) is Little’s law applied to KV-cache reservations, with reservation size as the weight [16, 25]. The factorization in (20) holds because, under FCFS and independent request lengths, a request’s waiting time is determined by the work already in the queue and is therefore independent of its own input length. The detailed proof is given in Appendix A.3. We estimate the mean prefill waiting time 𝑡¯𝑊 using Kingman’s approximation [24]:

Unused capacity. Requests reserve discrete amounts of KV-cache space, so the pool cannot always be filled exactly, and some space remains unused when the next reservation does not fit. We approximate the unused space by assuming that the reservation that does not fit in the remaining space is sampled in proportion to its size and that the KV-cache limit falls uniformly within it. The mean unused space is then E𝑟 [(ℓin + 𝑅) 2 ] . 2 E𝑟 [ℓin + 𝑅] Appendix A.4 gives the derivation. E𝑘 [𝑈𝑘 ] ≈

(24)

Operating batch. Combining the three occupancy terms gives the memory balance in (10). Let 𝑔(𝐵) denote the difference between the modeled KV-cache occupancy and the available pool capacity for a batch size 𝐵: 𝑔(𝐵) = 𝜇𝐷 (𝐵) (𝑡¯𝑊 (𝐵) E𝑟 [ℓin + 𝑅] + E𝑟 [𝑡𝑃 (ℓin + 𝑅)]) + 𝐵( ℓ¯ctx + 𝑅) + E𝑘 [𝑈𝑘 ] − 𝐶 tok . The operating batch satisfies 𝑔(𝐵) = 0, where the modeled occupancy matches the available capacity. The following theorem establishes its existence and uniqueness within the stable operating range.

(21)

Theorem 4.3 (Existence and uniqeness of the operating batch). If 𝐶 tok > E𝑘 [𝑈𝑘 ], then there is a unique operating batch 𝐵 > 0 satisfying 𝑔(𝐵) = 0 and 𝜌 𝑃 (𝐵) < 1.

with 𝜌 𝑃 = 𝑛𝐷 𝜇𝐷 /(𝑛𝑃 𝜇𝑃 ) the serving utilization of a prefill instance, and 𝐶𝑉𝑎2 and 𝐶𝑉𝑠2 = Var(𝑡𝑃 )/E𝑟 [𝑡𝑃 ] 2 the variation in its arrival intervals and prefill service times. The approximation allows general distributions of service times and interarrival intervals under i.i.d. assumptions. Here, prefill service time depends on input length, so we compute 𝐶𝑉𝑠2 from the workload’s input-length distribution. For Poisson arrivals assigned to prefill instances in round-robin order, 𝐶𝑉𝑎2 = 1/𝑛𝑃 . Lemma 4.2 gives a continuous-time average, whereas the KVcache memory balance in (18) averages occupancy over decode iterations. We approximate the latter by the former:

The theorem guarantees a unique operating batch within the stable prefill regime; the proof is given in Appendix A.5. The condition 𝐶 tok > E𝑘 [𝑈𝑘 ] is readily satisfied in the deployments considered here: the KV-cache pool accommodates multiple concurrent requests and is substantially larger than the mean unused space. The solution lies below both the full-pool batch 𝐵 max and 𝐵 𝜌 , the upper boundary of the range 𝜌 𝑃 (𝐵) < 1. We compute 𝐵 by solving 𝑔(𝐵) = 0 numerically within these bounds. Substituting the resulting batch into (9) gives the decode capacity, which determines the deployment serving capacity through (2).

𝑡¯𝑊 ≈

𝐶𝑉𝑎2 + 𝐶𝑉𝑠2 𝜌𝑃 · , 2 𝜇𝑃 (1 − 𝜌 𝑃 )

 ∑︁   ∑︁      E𝑘  (ℓin,𝑟 + 𝑅)  ≈ E𝑡  (ℓin,𝑟 + 𝑅)  . (22) 𝑟 ∈ P𝑘  𝑟 ∈ P (𝑡 )      This approximation is expected to be accurate when the decodeiteration duration is only weakly correlated with the prefill occupancy. The evaluation in Section 5 uses the complete model with this approximation. Substituting the waiting-time estimate in (21) into Lemma 4.2 and applying the sampling approximation in (22) gives the prefill-occupancy term in the memory balance. Decode occupancy. Each active request occupies KV-cache space for its current context and reserves an additional 𝑅 token slots. Averaging both parts over decode iterations gives the mean decode occupancy  ∑︁     E𝑘  (ℓctx,𝑟,𝑘 + 𝑅)  = 𝐵 ℓ¯ctx + 𝑅 , (23) 𝑟 ∈ D𝑘    since, by the definitions of 𝐵 and ℓ¯ctx , the mean total context length is 𝐵 ℓ¯ctx and the 𝐵𝑘 reservations of 𝑅 slots average to 𝐵𝑅.

5

Experiments

We first evaluate the accuracy of the analytical serving-capacity and power models against measurements across PD-disaggregated deployments, using one fixed-length workload and inference request samples from the Mooncake [45] and Azure [41] production traces. For serving capacity, we test generalization from fixed-length calibration workloads to variable-length workloads derived from production traces; for power, we test cross-workload and crossdeployment accuracy. We then compare the modeled and measured serving capacity–power Pareto fronts and evaluate whether they select the same minimum-power feasible deployment for a given serving-capacity requirement. Finally, we demonstrate two provisioning scenarios: reducing provisioned resources after the request rate falls, and temporarily reducing power under an acceptable serving-capacity threshold to respond to grid-side load-reduction signals.

5.1

Experimental Setup

All experiments serve Qwen3-32B [46] on H200 GPUs [39], with one GPU per instance. We use SGLang 0.5.9 [60], an open-source

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

# requests E𝑟 [ℓin ] E𝑟 [ℓout ]

Fixed-length

Mooncake

Azure

— 1024–8192 256

1000 8307 345

7200 1158 211

Table 3: Calibrated model parameters Symbol

Prefill

Serving capacity Hardware utilization Attention coefficient Per-iteration overhead Per-request overhead

MFU, MBU 𝑐𝑎 𝑡 iter 𝑡 req

0.67 2.17

Power ramp (W) Static power Slope Saturated power

𝑝0 𝛾 𝑝 sat

133 566 692

Decode

5000

Fixed-length Mooncake Azure

4000

6p2d 5p2d 4p2d 3p2d

3000

(1p2d)

1000

5p3d

(5p3d)

(5p3d)

5p3d

5p3d

5p2d (6p2d) (5p2d) (4p2d)

(5p2d)

(4p2d) (3p2d)

4p2d 3p2d (3p2d)

3p2d 3p1d (3p1d) (2p1d) 2p1d

(5p1d) 2p1d

3p1d (2p1d)

1p1d (1p1d)

1p1d

1

1p1d

(1p1d)

2

5

(6p2d) (5p2d) (4p2d)

5p2d

4p2d

(3p2d) (6p1d) (3p1d) (2p1d)

3p1d 2p1d

2000

(5p3d)

(1p1d)

10

Serving capacity (req/s)

20

Model Measured 40

0.77 1.00 ms 0.062 ms

448 458 678

LLM serving framework. For KV-cache transfers between prefill and decode instances, we use Mooncake [45] as the backend. Within the prefill and decode pools, requests are assigned to instances in round-robin order, so successive requests are distributed across the instances in turn. We conduct experiments with three types of workloads. (i) Fixedlength workloads, in which all requests have the same input and output lengths, using five discrete input lengths between 1024 and 8192 tokens and a fixed output length of 256 tokens. (ii) Request samples from the Mooncake conversation trace [45], collected from Moonshot AI’s Kimi chat service and containing long-context requests. (iii) Request samples from the Azure conversation trace [41], collected from a production LLM inference service on Microsoft Azure. Both traces record the input and output token counts of individual requests. Table 2 summarizes the requests used in our experiments. Section 5.2 evaluates one fixed-length workload and both trace-derived workloads across deployments with 2 to 8 instances. Each evaluated deployment operates at its serving capacity. Table 3 lists the calibrated model parameters used in the evaluation. For all serving-capacity and power comparisons, we exclude startup and shutdown transients from the measurement windows. Reported total power includes only GPUs assigned to instances currently in service. Appendix B provides the model and hardware constants and describes the calibration, measurement, and tracesampling procedures.

5.2

Power (W)

Table 2: Request counts and input and output lengths for the experimental workloads. The fixed-length column lists the lengths used across the fixed-length workloads; the trace columns report sample means.

Model Validation

Serving-capacity accuracy. We evaluate whether the analytical serving-capacity model, calibrated only on fixed-length workloads, remains accurate for variable-length workloads derived from production traces without further calibration. We first evaluate the fixed-length workload with 4096 input and 256 output tokens, which

Figure 4: Modeled and measured serving capacity–power Pareto fronts for the three workloads.

is included in the calibration workloads. Figure 4 compares the modeled and measured serving capacities for this workload. The mean absolute percentage error across all evaluated deployments is 1.2%. We then evaluate the Mooncake and Azure trace-derived workloads, neither of which is used to calibrate the serving-capacity model. We keep the calibrated model parameters fixed and update only the workload length moments for each workload. The mean absolute percentage errors are 3.0% for Mooncake and 1.6% for Azure. Power accuracy. We evaluate whether the same prefill and decode power models remain accurate across the evaluated workloads and deployments. Each model takes normalized serving throughput as input, and we keep its calibrated parameters fixed throughout the evaluation. The comparisons are visualized in Figure 4. For the fixed-length workload, the mean absolute percentage error in total GPU power across all evaluated deployments is 2.3%. The corresponding errors are 2.7% for Mooncake and 2.6% for Azure. These results show that the calibrated power models closely match measured deployment power across the evaluated workloads. Pareto fronts and deployment selection. Figure 4 shows that the modeled serving capacity–power Pareto fronts closely follow the measured fronts. Along the measured fronts, higher serving capacity comes with higher total GPU power. At comparable power levels, serving capacity is highest for Azure, followed by the fixedlength workload and Mooncake. This ordering is consistent with the mean input and output lengths in Table 2, since longer requests require more processing per request. We then evaluate whether the modeled and measured Pareto fronts lead to the same deployment choice for a given servingcapacity requirement. For each evaluated requirement, we select the minimum-power feasible deployment separately from the modeled and measured fronts. The selected deployments agree for 90% of the evaluated serving-capacity requirements for the fixed-length workload, 85% for Mooncake, and 86% for Azure. All disagreements

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

We use two scenarios to examine how provisioning changes can improve efficiency or provide temporary power flexibility. In the first, a decrease in request rate creates serving-capacity headroom that allows provisioned resources to be reduced while the deployment remains stable. In the second, the request rate remains unchanged and provisioned capacity is temporarily reduced below that rate, reducing power at the cost of transient service degradation. Both scenarios use the fixed-length workload with 4096 input and 256 output tokens and begin with three prefill instances and two decode instances (3p2d). To capture the transient effects of reconfiguration, we report TTFT and TPOT throughout each run. TTFT is measured from request arrival to the first output token and includes prefill-queue waiting time, prefill execution, and KV-cache transfer to the decode instance. TPOT is measured over individual intervals between consecutive output tokens rather than first averaging within each request, so that its p95 captures transient stalls. 5.3.1 Efficiency After a Request-Rate Drop. We examine whether a decrease in request rate creates sufficient capacity headroom to remove a decode instance. The request rate is initially set to 60% of the serving capacity of 3p2d and is then reduced to 3.1 req/s (Figure 5, left). Before any change in provisioning, the lower request rate alone reduces total power by approximately 270 W, while TTFT and TPOT remain nearly unchanged. This reduction reflects the load dependence of per-instance power: the same provisioned deployment consumes less power at a lower request rate. At 3.1 req/s, 3p1d still provides sufficient serving capacity, so one decode instance can be removed without overloading the resulting deployment. This reconfiguration reduces total power by an additional 560 W. The resulting 3p1d deployment operates at 60% of its serving capacity. Median TTFT remains approximately 0.55 s, while median TPOT increases from 23 to 28 ms as the remaining decode instance operates at higher load. Thus, after the requestrate reduction, capacity headroom can be converted into additional power savings without inducing queue growth; the main latency effect is a modest increase in decode-side TPOT. 5.3.2 Power Flexibility at a Fixed Request Rate. We examine how much power can be saved by temporarily reducing provisioned capacity below a fixed request rate, and how long the reduction can be sustained. We hold the request rate at 5.3 req/s, or 80% of the serving capacity of 3p2d. We then remove one decode instance for 10 min before restoring it (Figure 5, right). Removing the decode instance reduces total power by approximately 620 W but lowers serving capacity to 5.2 req/s, slightly below the offered request rate. With serving capacity below the request rate, requests accumulate in the queue and median TTFT increases from 0.7 to 8.5 s during the removal interval. Meanwhile, the remaining decode instance operates near its serving limit with a larger batch, increasing median TPOT from 27 to 39 ms. The much larger TTFT increase arises because requests must wait in the growing queue once the request rate exceeds serving capacity.

Request rate (req/s)

Power-Aware Provisioning

Power (kW)

5.3

Throughput

TPOT (ms) TTFT (s)

occur near a feasibility boundary, where the required serving capacity is within 4% of a deployment’s measured serving capacity.

Serving capacity

5

5

0 3

0 3

2 1 2

p50

2 20

p95

0

0

20

25

0

Arrival rate

0

10

20

Time (min)

30

0

0

10

20

Time (min)

30

Figure 5: Serving throughput, total GPU power, TTFT, and TPOT during decode-instance reconfiguration. Left: a decode instance is removed after the request rate decreases. Right: a decode instance is removed and restored at a fixed request rate.

Restoring the second decode instance raises serving capacity above the request rate and allows the accumulated queue to drain within approximately 1.3 min. TTFT, TPOT, and total power then return to their pre-removal levels. Thus, when removing resources would reduce serving capacity below the request rate, the deployment can still provide temporary power flexibility, but the reduction can be sustained only as long as the resulting queueing delay remains tolerable. Appendix D extends this behavior to a sequential scale-down experiment.

6

Conclusion

This paper develops an analytical framework for power-aware provisioning of PD-disaggregated AI inference. The framework characterizes how the workload distribution, hardware limits, and numbers of prefill and decode instances jointly determine serving capacity and average deployment power. In particular, the servingcapacity model captures the coupling between the two serving phases induced by KV-cache reservations, as well as the impact of request queueing, while the power model relates per-instance power consumption to normalized serving throughput. Combining these models yields a discrete serving capacity–power Pareto front over candidate provisioned deployments, providing a direct way to identify minimum-power deployments for a required serving capacity and to quantify the serving capacity available under different power limits. Experiments across fixed-length workloads and workloads derived from production traces show that the analytical models closely reproduce measured serving capacity and power while reusing the same calibrated parameters across workloads. The resulting modeled Pareto fronts also lead to provisioning decisions that largely

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

agree with those obtained from measurements. These results show that analytical characterization of the serving capacity–power relationship can reduce reliance on repeated profiling and simulation while providing a principled basis for efficient and grid-responsive operation of large-scale AI inference systems.

References [1] Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav S. Gulavani, Ramachandran Ramjee, and Alexey Tumanov. 2024. Vidur: A Large-Scale Simulation Framework for LLM Inference. In Proceedings of Machine Learning and Systems (MLSys). [2] Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI) (Santa Clara, CA, USA). USENIX Association, 117–134. https://www.usenix.org/ conference/osdi24/presentation/agrawal [3] Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. GQA: Training Generalized MultiQuery Transformer Models from Multi-Head Checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) (Singapore). Association for Computational Linguistics, 4895–4901. doi:10.18653/v1/2023.emnlp-main.298 [4] Ruicheng Ao, Gan Luo, David Simchi-Levi, and Xinshang Wang. 2025. Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints. arXiv:2504.11320 [cs.LG] [5] Omar Basit, Yunzhao Liu, Z. Jonny Kong, and Y. Charlie Hu. 2026. DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS. arXiv:2602.18755 [cs.DC] [6] Aaron Chatterji, Thomas Cunningham, David J. Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025. How People Use ChatGPT. Working Paper 34255. National Bureau of Economic Research, Cambridge, MA. doi:10. 3386/w34255 [7] Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, and Xue Liu. 2026. FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compressand-Route as Implementation Mechanism. arXiv:2603.16514 [cs.DC] [8] Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, and Xue Liu. 2026. inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference. arXiv:2603.16054 [cs.DC] [9] Xin Chen, Xiaoyang Wang, Ana Colacelli, Matt Lee, and Le Xie. 2025. Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects. arXiv:2509.07218 [eess.SY] [10] Commission for Regulation of Utilities. 2025. Large Energy Users Connection Policy. Decision paper CRU2025236. https://www.cru.ie/publications/28573/ [11] DeepSeek-AI. 2025. DeepSeek-V3/R1 Inference System Overview. Open Source Week, day 6. Retrieved August 30, 2026 from https://github.com/deepseekai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_ thing_deepseekV3R1_inference_system_overview.md [12] Bojun Du, Xiaoyi Fan, Ershun Du, Long Chen, Jianpei Han, Qingchun Hou, Ning Zhang, and Chongqing Kang. 2026. From Tokens to Energy Flexibility: Quantization-Enabled Demand Response for Data Centers with LLM Inference Workloads. arXiv:2606.18851 [eess.SY] [13] Mauricio Fadel Argerich, Jonathan Fürst, and Marta Patiño-Martínez. 2026. Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures. arXiv:2604.09048 [cs.DC] [14] Mauricio Fadel Argerich, Jonathan Fürst, and Marta Patiño-Martínez. 2026. WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs. arXiv:2607.02391 [15] Yicheng Feng, Xin Tan, Yangtao Deng, Yimin Jiang, and Hong Xu. 2026. Frontier: Towards Comprehensive and Accurate LLM Inference Simulation. arXiv:2605.21312 [cs.DC] [16] Robert G. Gallager. 1996. Discrete Stochastic Processes. Kluwer Academic Publishers, Boston, MA. doi:10.1007/978-1-4615-2329-1 [17] Daya Guo et al. 2025. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 8081 (2025), 633–638. doi:10.1038/s41586025-09422-z [18] Can Hankendi, Rana Shahout, Minlan Yu, and Ayse K. Coskun. 2026. PALS: PowerAware LLM Serving for Mixture-of-Experts Models. arXiv:2605.21427 [cs.AI] [19] Xiannan Hu, Tianyou Zeng, Xiaoming Yuan, Liwei Song, Guangyuan Zhang, and Bangzheng He. 2025. BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures. arXiv:2506.05871 [cs.LG] [20] International Energy Agency. 2026. Key Questions on Energy and AI. Technical Report. International Energy Agency, Paris, France. https://www.iea.org/reports/ key-questions-on-energy-and-ai

[21] Vincent Jacamon, Julie Dallard, and Thomas Spencer. 2025. Overcoming energy constraints is key to delivering on Europe’s data centre goals. IEA commentary. Retrieved August 30, 2026 from https://www.iea.org/commentaries/overcomingenergy-constraints-is-key-to-delivering-on-europe-s-data-centre-goals [22] Hongyi Jia, Jinghui Zhang, Lu Fang, Stephen Chen, Yan Cui, Ye (Charlotte) Qi, and Zijing Liu. 2025. Disaggregated Inference at Scale with PyTorch and vLLM. PyTorch Blog. Retrieved August 30, 2026 from https://pytorch.org/blog/ disaggregated-inference-at-scale-with-pytorch-vllm/ [23] Yiwei Jiang, Sangeeta Chowdhary, Nathaniel Morris, Rutwik Jain, Srilatha Manne, and Sam Bayliss. 2026. Power Aware Dynamic Reallocation for Inference. arXiv:2601.12241 [cs.DC] [24] J. F. C. Kingman. 1961. The Single Server Queue in Heavy Traffic. Mathematical Proceedings of the Cambridge Philosophical Society 57, 4 (1961), 902–904. doi:10. 1017/S0305004100036094 [25] Leonard Kleinrock. 1975. Queueing Systems, Volume I: Theory. Wiley-Interscience, New York. [26] Min-Seung Ko, Jae Woong Shim, and Hao Zhu. 2025. Mitigation of Datacenter Demand Ramping and Fluctuation using Hybrid ESS and Supercapacitor. arXiv:2512.08076 [eess.SY] [27] Can Emre Koksal, Richard A. Barry, and Artun Sel. 2026. A Theory of Probabilistic Power Provisioning for Data Centers with Distributed Energy Storage. arXiv:2608.12993 [cs.IT] [28] Ruiqi Lai, Hongrui Liu, Chengzhi Lu, Zonghao Liu, Siyu Cao, Siyang Shao, Yixin Zhang, Luo Mai, and Dmitrii Ustiugov. 2025. TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity. arXiv:2512.03416 [cs.DC] [29] Marco Laumanns, Lothar Thiele, Kalyanmoy Deb, and Eckart Zitzler. 2002. Combining Convergence and Diversity in Evolutionary Multiobjective Optimization. Evolutionary Computation 10, 3 (2002), 263–282. doi:10.1162/106365602760234108 [30] Kyungmi Lee, Zhiye Song, Eun Kyung Lee, Xin Zhang, Tamar Eilam, and Anantha P. Chandrakasan. 2026. EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads. In IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). 435–447. doi:10.1109/ISPASS69572. 2026.00049 [31] Luchang Li, Dongfang Li, Bozhao Gong, and Yu Zhang. 2026. SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference. arXiv:2603.04716 [cs.DC] [32] Rongzhi Li, Ruogu Du, Zefang Chu, Sida Zhao, Chunlei Han, Zuocheng Shi, Yiwen Shao, Huanle Han, Long Huang, Zherui Liu, and Shufan Liu. 2025. Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference. arXiv:2508.19559 [cs.DC] [33] Yueying Li, Jiayang Chen, Yuanfan Chen, Leo Han, Haoran Qiu, Esha Choukse, Rodrigo Fonseca, and Udit Gupta. 2026. PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response. arXiv:2608.21719 [34] Zhirui Liang, Jae-Won Chung, Mosharaf Chowdhury, Jiasi Chen, and Vladimir Dvorkin. 2026. GPU-to-Grid: Voltage Regulation via GPU Utilization Control. arXiv:2602.05116 [eess.SY] [35] Banruo Liu, Haoran Qiu, Íñigo Goiri, Rodrigo Fonseca, Ricardo Bianchini, and Esha Choukse. 2026. Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale. arXiv:2608.00101 [cs.AI] [36] Qunyou Liu, Darong Huang, Marina Zapater, and David Atienza. 2025. GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving. arXiv:2508.16449 [cs.PF] [37] Chengyi Nie, Nian Si, and Zijie Zhou. 2026. A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints. arXiv:2605.04595 [cs.LG] [38] Chenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao, Tongyang Wang, Xi Wang, and Yong Chen. 2026. TokenPowerBench: Benchmarking the Power Consumption of LLM Inference. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 32582–32590. [39] NVIDIA Corporation. 2023. NVIDIA H200 Tensor Core GPU. Product page. Retrieved September 18, 2026 from https://www.nvidia.com/en-us/data-center/ h200/ [40] Office of the Governor of Texas. 2026. Directive to the Public Utility Commission of Texas and ERCOT on data center interconnection. https://gov.texas.gov/uploads/files/press/Thomas_Gleeson_Pablo_Vegas_ Data_Centers_Directive_Letter_to_PUCT_ERCOT_August_2026_.pdf [41] Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Inference Using Phase Splitting. In Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA) (Buenos Aires, Argentina). IEEE, 118–132. doi:10.1109/ISCA59077.2024.00019 [42] David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean. 2022. The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink. Computer 55, 7 (2022), 18–28. doi:10.1109/MC.2022.3148714

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

[43] Hiari Pizzini Cavagna, Andrea Proia, Giacomo Madella, Giovanni B. Esposito, Francesco Antici, Daniele Cesarini, Zeynep Kiziltan, and Andrea Bartolini. 2026. SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference. In Proceedings of the 17th ACM/SPEC International Conference on Performance Engineering (ICPE). 83–95. doi:10.1145/3777884.3797011 [44] Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023. Efficiently Scaling Transformer Inference. In Proceedings of Machine Learning and Systems (MLSys). [45] Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot. In 23rd USENIX Conference on File and Storage Technologies (FAST) (Santa Clara, CA, USA). USENIX Association, 155–170. https://www.usenix.org/ conference/fast25/presentation/qin [46] Qwen Team. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] [47] Chendong Song, Meixuan Wang, Hang Zhou, Hong Liang, Yuan Lyu, Zixi Chen, Yuwei Fan, and Zijie Zhou. 2026. Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads. arXiv:2601.21351 [cs.LG] [48] State of New York. 2026. Executive Order No. 62: Establishing a Temporary Moratorium on Data Centers in New York While the State Develops Higher Standards for Data Center Development and Benefits Blueprint to Support Localities. https://www.governor.ny.gov/executive-order/no-62-establishingtemporary-moratorium-data-centers-new-york-while-state-develops [49] Boyu Tan, Jiarui Guo, Zongwei Lv, Haobo Sun, Tong Yang, Kan Liu, Xinfei Shi, Zetao Hu, Yaxin Yu, Chi Zhang, Jianning Zhang, Xi Yang, Wei Zhang, Bo Cai, Silu Zhou, Xiyu Wang, Na He, Yinghao Yu, Wending Bao, Guiyang Huang, Yuxing Yuan, Juncheng Yin, Nan Wang, Lin Yang, Zechao Zhang, Lu Chen, Guoding Li, Tao Lan, and Lin Qu. 2026. RTP-LLM: High-Performance Alibaba LLM Inference Engine. arXiv:2605.29639 [50] Tina Vartziotis, Rodopi Kosteli, Elli Danae Vartziotis, George Dasoulas, Michael Keckeisen, Konstantinos Skianis, Sotirios Kotsopoulos, and Francesca Dominici. 2026. From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs. arXiv:2607.26571 [51] vLLM Semantic Router Project. 2026. The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency. arXiv:2603.17280 [cs.DC] [52] Grant Wilkins, Srinivasan Keshav, and Richard Mortier. 2024. Offline EnergyOptimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems. In 3rd Workshop on Sustainable Computer Systems (HotCarbon). [53] Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: An Insightful Visual Performance Model for Multicore Architectures. Commun. ACM 52, 4 (2009), 65–76. [54] Yu Wu, Tongxuan Liu, Yuting Zeng, Siyu Wu, Jun Xiong, Xianzhe Dong, Hailong Yang, Ke Zhang, and Jing Li. 2025. Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture. arXiv:2505.11916 [cs.DC] [55] Yiheng Xie, Wenqi Cui, and Adam Wierman. 2026. Data Center Voltage RideThrough: Emerging Challenges and Opportunities. In Proceedings of the 17th ACM International Conference on Future and Sustainable Energy Systems. 379–385. [56] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations (ICLR) (Kigali, Rwanda). OpenReview.net, 33 pages. https://openreview.net/forum?id=WE_ vluYUL-X [57] Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and ByungGon Chun. 2022. Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI). USENIX Association, 521–538. https://www.usenix.org/ conference/osdi22/presentation/yu [58] Jiahuan Yu, Aryan Taneja, Junfeng Lin, and Minjia Zhang. 2026. VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing. In ISC High Performance 2026 Research Paper Proceedings (41st International Conference). IEEE, 1–19. doi:10.23919/ISC. 2026.11520495 [59] Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Zhe Zhou, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, Yan Yan, Beidi Chen, Guangyu Sun, and Kurt Keutzer. 2024. LLM Inference Unveiled: Survey and Roofline Model Insights. arXiv:2402.16363 [cs.CL] [60] Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. In Advances in Neural Information Processing Systems (NeurIPS) (Vancouver, BC, Canada), Vol. 37. 62557–62583. doi:10.52202/0790172000 [61] Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding

for Goodput-Optimized Large Language Model Serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI) (Santa Clara, CA, USA). USENIX Association, 193–210. https://www.usenix.org/conference/ osdi24/presentation/zhong-yinmin [62] Pengfei Zuo, Huimin Lin, Junbo Deng, Nan Zou, Xingkun Yang, Yingyu Diao, Weifeng Gao, Ke Xu, Zhangyu Chen, Shirui Lu, Zhao Qiu, Peiyang Li, Xianyu Chang, Zhengzhong Yu, Fangzheng Miao, Jia Zheng, Ying Li, Yuan Feng, Bei Wang, Zaijian Zong, Mosong Zhou, Wenli Zhou, Houjiang Chen, Xingyu Liao, Yipeng Li, Wenxiao Zhang, Ping Zhu, Yinggang Wang, Chuanjie Xiao, Depeng Liang, Dong Cao, Juncheng Liu, Yongqiang Yang, Xiaolong Bai, Yi Li, Huaguo Xie, Huatao Wu, Zhibin Yu, Lv Chen, Hu Liu, Yujun Ding, Haipei Zhu, Jing Xia, Yi Xiong, Zhou Yu, and Heng Liao. 2025. Serving Large Language Models on Huawei CloudMatrix384. arXiv:2506.12708

A Additional Derivations and Proofs A.1 Proof of Lemma 4.1 Proof. Let 𝐴𝐾 denote the number of requests admitted to a given decode instance during the first 𝐾 decode iterations. Reindexing the context contributions from iterations to requests gives

ℓ¯ctx :=

E𝑘

Í

𝑟 ∈ D𝑘 ℓctx,𝑟,𝑘

E𝑘 [𝐵𝑘 ] Í𝐾 Í 𝑘=1

= lim 𝐾→∞

1 = lim

𝑟 ∈ D𝑘 ℓctx,𝑟,𝑘

Í𝐾

𝑘=1 𝐵𝑘 Í𝐴𝐾 Íℓout,𝑟 −1 𝑟 =1

𝑗=1

(ℓin,𝑟 + 𝑗)

Í𝐴𝐾

𝑟 =1 (ℓout,𝑟 − 1)

𝐾→∞

2 E𝑟 =



hÍ

ℓout −1 𝑗=1 (ℓin + 𝑗)

i

E𝑟 [ℓout − 1]  i ℓout 3 E𝑟 (ℓout − 1) ℓin + 2 = E𝑟 [ℓout − 1]     E𝑟 [ℓout − 1] E𝑟 ℓin + ℓout + Cov ℓout − 1, ℓin + ℓout 2 2 = E𝑟 [ℓout − 1] 1 E𝑟 [ℓout ] 2 Var(ℓout ) + Cov(ℓin, ℓout ) 4 = E𝑟 [ℓin ] + + . 2 E𝑟 [ℓout − 1] | {z } | {z }

(25)

h

arrival mean

residence correction

Equation 1 counts the same context contributions request by request rather than iteration by iteration. A request contributes ℓout −1 terms, one per decode iteration, and its context length increases from ℓin + 1 to ℓin + ℓout − 1 over them. In 2 , boundary requests contribute a vanishing fraction of both sums as 𝐾 → ∞, so the request averages converge to expectations under the workload distribution [16]. Equation 3 sums this arithmetic sequence, yielding (ℓout − 1)(ℓin + ℓout /2). Finally, expanding the covariance in 4 gives (17), which completes the proof. □

A.2

Derivation of the Decode Serving Capacity

Proof. Let 𝐴𝐾 denote the number of requests admitted to a given decode instance during the first 𝐾 decode iterations. Each request requires ℓout − 1 decode tokens, while each decode iteration generates one token for every active request. Over a long serving

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

window, the mean decode time per admitted request is therefore

space–time occupancy over this interval from time to requests gives

 ∑︁    E𝑡  (ℓin,𝑟 + 𝑅)  𝑟 ∈ P (𝑡 )    ∫ 𝑇 ∑︁ 1 (ℓin,𝑟 + 𝑅) 𝑑𝑡 = lim 𝑇 →∞ 𝑇 0

𝐾 ∑︁

𝑡𝐷,𝑘 1 1 𝑘=1 = lim 𝐾→∞ 𝜇𝐷 𝐴𝐾

𝑟 ∈ P (𝑡 )

𝐾 ∑︁ ∑︁ © ª ℓctx,𝑟,𝑘 ® ­2𝑁 + 𝜅 𝑟 ∈ D𝑘 1 𝑘=1 « ¬ = lim 𝐾→∞ 𝐴𝐾 𝛽 MBU

2𝑁 𝐾 + 𝜅

2𝑁 𝛽 MBU

1 1 = lim (𝑡𝑊 ,𝑟 + 𝑡𝑃,𝑟 )(ℓin,𝑟 + 𝑅) 𝑇 →∞ 𝑇 𝑟 =1 # " 𝐴(𝑇 ) 1 ∑︁ 𝐴(𝑇 ) (𝑡𝑊 ,𝑟 + 𝑡𝑃,𝑟 )(ℓin,𝑟 + 𝑅) = lim 𝑇 →∞ 𝑇 𝐴(𝑇 ) 𝑟 =1

ℓctx,𝑟,𝑘

𝑘=1 𝑟 ∈ D𝑘

1 = lim 𝛽 MBU 𝐾→∞ 2 =

𝐾 ∑︁ ∑︁

𝐴(𝑇 ∑︁)

Í𝐾

𝑘=1 𝐵𝑘

lim

𝐴𝐾 !

𝐴𝐾

𝐾→∞

2 = 𝜇𝐷 E𝑟 [(𝑡𝑊 + 𝑡𝑃 )(ℓin + 𝑅)] . Í𝐾

lim 𝐾→∞

(27)

𝑘=1 𝐵𝑘

! (26)

𝐾

𝐴𝐾 ℓout,𝑟 ∑︁−1 ∑︁

(ℓin,𝑟 + 𝑗) 𝑟 =1 𝑗=1 𝜅 lim + 𝛽 MBU 𝐾→∞ 𝐴𝐾 " "ℓ −1 ## out ∑︁ 2𝑁 E𝑟 [ℓout − 1] 1 3 + 𝜅 E𝑟 = (ℓin + 𝑗) 𝛽 MBU 𝐵 𝑗=1

Equation 1 counts the same space–time occupancy request by request. Request 𝑟 holds ℓin,𝑟 + 𝑅 token slots for its prefill holding time 𝑡𝑊 ,𝑟 +𝑡𝑃,𝑟 , its waiting time plus its prefill service time. Requests whose reservations overlap either end of the observation window contribute only boundary terms, which vanish as 𝑇 → ∞. In 2 , the admission rate 𝐴(𝑇 )/𝑇 converges to the decode capacity 𝜇𝐷 under steady operation at the serving limit, while the request average converges to E𝑟 [(𝑡𝑊 + 𝑡𝑃 )(ℓin + 𝑅)]. This establishes (19). We next expand the expectation on its right-hand side:

  𝜅𝐵 ℓ¯ctx 2𝑁 4 E𝑟 [ℓout − 1] = + 𝐵 𝛽 MBU 𝛽 MBU E𝑟 [(𝑡𝑊 + 𝑡𝑃 )(ℓin + 𝑅)] 5 E𝑟 [ℓout − 1] 𝑡¯𝐷 . = 𝐵

= E𝑟 [𝑡𝑊 (ℓin + 𝑅)] + E𝑟 [𝑡𝑃 (ℓin + 𝑅)] 1 = 𝑡¯𝑊 E𝑟 [ℓin + 𝑅] + E𝑟 [𝑡𝑃 (ℓin + 𝑅)]   = 𝑡¯𝑊 E𝑟 [ℓin + 𝑅] + E𝑟 (𝑎𝑃 ℓin + 𝑏 𝑃 ℓin2 )(ℓin + 𝑅)  = 𝑡¯𝑊 E𝑟 [ℓin ] + 𝑅  + 𝑎𝑃 E𝑟 [ℓin2 ] + 𝑅 E𝑟 [ℓin ]  + 𝑏 𝑃 E𝑟 [ℓin3 ] + 𝑅 E𝑟 [ℓin2 ] .

(28)

Equation 1 is the mean decode time per admitted request over the observation window, in which boundary requests contribute a vanishing fraction as 𝐾 → ∞. Equations 2 and 3 relate iterationlevel and request-level counts of decode work. The total number Í of active request–iteration pairs, 𝑘 𝐵𝑘 , is also the total number of Í decode tokens generated. Hence 𝑘 𝐵𝑘 /𝐴𝐾 converges to E𝑟 [ℓout −1], Í while 𝑘 𝐵𝑘 /𝐾 converges to the operating batch 𝐵. The KV-cache contribution is reindexed by request as in (25). Substituting the definitions of ℓ¯ctx and 𝑡¯𝐷 in 4 and 5 then yields the capacity relation in (9). This relation also holds with calibrated overheads, since the token count is unchanged and the overheads are included in the mean iteration time 𝑡¯𝐷 . □

Under FCFS, an arriving request’s waiting time is determined by the work already in the queue. With request lengths independent across arrivals, it is therefore independent of that request’s own input length. Equation 1 uses this independence to factor E𝑟 [𝑡𝑊 (ℓin +𝑅)] as 𝑡¯𝑊 E𝑟 [ℓin +𝑅], with 𝑡¯𝑊 = E𝑟 [𝑡𝑊 ]. The resulting expression is (20), which completes the proof. □

A.3

A.4

Proof of Lemma 4.2

Proof. To compute the mean occupancy, consider continuous time 𝑡, and let 𝐴(𝑇 ) denote the number of requests admitted to a given decode instance during [0,𝑇 ]. Reindexing the accumulated

Derivation of the Unused-Capacity Approximation

Let 𝐴𝐾 denote the number of requests admitted to a given decode instance during the first 𝐾 decode iterations. Under the two steps

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

of the approximation, the mean unused space is 𝐴𝐾 ∫ ℓin,𝑟 +𝑅 ∑︁ 𝑢 𝑑𝑢 1 𝑟 =1 0 E𝑘 [𝑈𝑘 ] ≈ lim 𝐴𝐾 𝐾→∞ ∑︁ (ℓin,𝑟 + 𝑅)

prefill; MBU, 𝑡 iter , and 𝑡 req for decode; and the power-ramp parameters 𝑝 0 , 𝛾, and 𝑝 sat for each role. The fitted parameters are reused across provisioned deployments and workloads under the same model, hardware, and serving configuration. We first summarize the model and hardware constants, then describe the calibration method and report the calibration runs and fitted results. The parameter values used in the evaluation are listed in Table 3.

𝑟 =1 𝐴𝐾 ∑︁

(ℓin,𝑟 + 𝑅) 2 /2

𝑟 =1 = lim 𝐴𝐾 𝐾→∞ ∑︁

(29)

(ℓin,𝑟 + 𝑅)

𝑟 =1

2 E𝑟 [(ℓin + 𝑅) 2 ] . = 2 E𝑟 [ℓin + 𝑅] Equation 1 applies the two steps: the size weighting selects the reservation that does not fit, and the uniform position leaves on average half of it unused, so the mean unused space is the sizeweighted mean of 𝑠/2. Equation 2 divides numerator and denominator by 𝐴𝐾 and takes 𝐾 → ∞, which yields the ratio of workload expectations. The result is (24).

A.5

Proof of Theorem 4.3

Proof. On the stable range, increasing 𝐵 increases the decode capacity and therefore the prefill utilization. The waiting time in (21) increases with utilization, so both the prefill occupancy and the decode occupancy increase with 𝐵. The unused-capacity approximation in (24) is independent of 𝐵. Consequently, 𝑔(𝐵) is continuous and strictly increasing on this range. Explicitly,  𝑔′ (𝐵) = 𝜇𝐷′ (𝐵) 𝑡¯𝑊 E𝑟 [ℓin + 𝑅] + E𝑟 [𝑡𝑃 (ℓin + 𝑅)] ′ + 𝜇𝐷 𝑡¯𝑊 (𝐵) E𝑟 [ℓin + 𝑅] + ℓ¯ctx + 𝑅 > 0.

At 𝐵 = 0, both prefill and decode occupancies vanish, giving 𝑔(0) = E𝑘 [𝑈𝑘 ] − 𝐶 tok < 0. To find a positive upper value, we compare the full-pool batch with the queue-stability limit. The full-pool batch 𝐵 max of (11) is positive by assumption. Let 𝐵 𝜌 = sup{𝐵 > 0 : 𝜌 𝑃 (𝐵) < 1} denote the queue-stability limit, the upper end of the stable range. If 𝐵 max < 𝐵 𝜌 , the prefill queue remains stable at 𝐵 max . Substituting (11) into the residual gives 𝑔(𝐵 max ) = 𝜇𝐷 (𝐵 max ) (𝑡¯𝑊 (𝐵 max ) E𝑟 [ℓin + 𝑅] + E𝑟 [𝑡𝑃 (ℓin + 𝑅)]) > 0. If 𝐵 𝜌 ≤ 𝐵 max , then 𝜌 𝑃 (𝐵) → 1 as 𝐵 → (𝐵 𝜌 ) − . By (21), 𝑡¯𝑊 (𝐵) → +∞, while the definition of 𝜌 𝑃 gives 𝜇𝐷 (𝐵) → 𝑛𝑃 𝜇𝑃 /𝑛𝐷 > 0. Since E𝑟 [ℓin + 𝑅] > 0 and the remaining terms have finite limits, lim 𝐵→(𝐵 𝜌 ) −

𝑔(𝐵) =

lim 𝐵→(𝐵 𝜌 ) −

𝜇𝐷 (𝐵)𝑡¯𝑊 (𝐵) E𝑟 [ℓin + 𝑅] = +∞.

In either case, continuity and 𝑔(0) < 0 guarantee a root in 0, min(𝐵 max, 𝐵 𝜌 ) . Since 𝑔 is strictly increasing on the stable range, this root is unique. □

B

Model Calibration

The analytical models combine quantities obtained from the workload, model architecture, and hardware specifications with parameters estimated from measurements. We calibrate MFU and 𝑐 𝑎 for

B.1

Model and Hardware Constants

The H200 GPU [39] has a peak compute throughput of 𝜋 = 989 TFLOP/s and a memory bandwidth of 𝛽 = 4.8 TB/s. Qwen3-32B [46] has 𝑁 = 32.8B parameters and 𝐿 = 64 layers, with ℎ𝑞 = 64 query heads and ℎ kv = 8 KV heads of width 𝑑 head = 128 in groupedquery attention [3]. Two quantities follow from the architecture: the attention width 𝑑 = ℎ𝑞 𝑑 head = 8192, and the KV-cache bytes per context token 𝜅 = 4𝐿ℎ kv𝑑 head , since for each token the KV cache stores one key and one value per KV head in every layer, using two bytes per value in bf16.

B.2

Calibration Method

Prefill. Prefill calibration estimates MFU and 𝑐 𝑎 from the dependence of the prefill service time on input length. Because (4) contains both a linear and a quadratic term, the fit requires 𝑡𝑃 at several input lengths. For each input length, we saturate a single prefill instance while provisioning enough decode capacity to keep prefill as the bottleneck. The measured completion rate is then 𝜇𝑃 , and because each calibration workload uses a fixed input length, (5) reduces to 𝑡𝑃 = 1/𝜇𝑃 . We fit 𝑎𝑃 ℓin + 𝑏 𝑃 ℓin2 to the measured 1/𝜇𝑃 , with no intercept as required by (4). The fitted coefficients give MFU =

2𝑁 , 𝜋𝑎𝑃

𝑐𝑎 =

𝜋 MFU 𝑏 𝑃 . 𝐿𝑑

Decode. We calibrate the decode iteration time rather than decode capacity, because it also depends on the operating batch determined by the KV-cache memory balance of Section 3.1.3. At a fixed context length, the memory-traffic model (8) determines an iteration time that is linear in the batch, with intercept 𝑎𝐷 and slope 𝑏 𝐷 . Measurements approximately follow this dependence, but also indicate an iteration-wide overhead and a context-independent cost per active request. We account for these costs by adding 𝑡 iter to the intercept and 𝑡 req to the slope: 𝑎𝐷 =

2𝑁 + 𝑡 iter, 𝛽 MBU

𝑏𝐷 =

𝜅 ℓ¯ctx + 𝑡 req . 𝛽 MBU

(30)

These overheads may include fixed iteration costs, such as CPU– GPU dispatch and kernel launches, and context-independent perrequest costs, such as metadata processing. Both are empirical corrections and need not correspond to any single mechanism. Algorithm 1 uses these calibrated coefficients to compute the operating batch and decode capacity. To estimate these parameters, we first fit the dependence on batch size at each context-length setting, then fit the resulting slopes as a function of context length. Let 𝑗 = 1, . . . , 𝐽 index the context-length settings, and let ℓ¯ctx,𝑗 denote the context length of setting 𝑗. The calibration workloads use fixed input and output lengths, so Var(ℓout ) = Cov(ℓin, ℓout ) = 0; the residence correction in (17) therefore vanishes and ℓ¯ctx,𝑗 = ℓin,𝑗 + ℓout /2. At each setting

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

𝑗, we fit 𝑡¯𝐷,𝑗 = 𝑎𝐷,𝑗 + 𝑏 𝐷,𝑗 𝐵 to the measured mean iteration times, where 𝑎𝐷,𝑗 and 𝑏 𝐷,𝑗 are the fitted intercept and slope. Because the fitted intercepts vary only modestly across context lengths, we use their mean as the context-independent intercept assumed by the model. We then regress 𝑏 𝐷,𝑗 on ℓ¯ctx,𝑗 as 𝑏 𝐷,𝑗 = 𝑐 𝐷,1 ℓ¯ctx,𝑗 +𝑐 𝐷,0 , where 𝑐 𝐷,0 and 𝑐 𝐷,1 are the fitted intercept and slope. The three decode parameters are then 𝜅 , 𝛽𝑐 𝐷,1 𝑡 req = 𝑐 𝐷,0,

MBU =

𝐽

𝑡 iter =

2𝑁 1 ∑︁ 𝑎𝐷,𝑗 − . 𝐽 𝑗=1 𝛽 MBU

These values are substituted into (30) in all reported results. Power. We calibrate the power model by fitting the capped ramp of (13) to measured per-instance power at different values of normalized serving throughput. The measurements must cover both the region where power rises with the serving throughput and the region where it has saturated. Because the throughput at which power saturates is not known in advance, we fit 𝑝 0 , 𝛾, and 𝑝 sat jointly, using measurements from all calibration workloads to obtain one ramp for each role. The shared ramp per role assumes that the normalized serving throughput accounts for the workload dependence of average power.

B.3

Decode. At each of the five input lengths, every decode instance of every provisioned deployment in the serving-capacity sweeps provides one measurement: its logs record the number of active requests at fixed iteration intervals, so their mean over the measurement window is the batch 𝐵, and 𝐵 divided by the rate at which the instance generates tokens over the same window is its mean iteration time. Provisioned deployments of different shapes cover the range of batches: when decode limits a provisioned deployment, its instances run at the batch that the KV-cache pool allows, which is nearly the same for every such provisioned deployment at a fixed input length; when prefill limits it, its decode instances receive fewer requests and run at smaller batches. These instances may idle between iterations, and weighting each measurement by its generation rate reduces the influence of such idle-contaminated estimates. The three longer lengths draw on 24 provisioned deployments spanning every prefill and decode count within the 8-GPU budget, giving 60 decode instances each; the two shorter lengths draw on 9 provisioned deployments each, giving 22, so the first regression fits 60 or 22 measurements per length and the second fits the 5 slopes. The five slopes and intercepts give the MBU, 𝑡 req and 𝑡 iter of Table 3. The measured intercepts rise slightly from the shortest to the longest context length, which (30) does not represent; the mean absorbs this variation. The fitted MBU = 0.77 corresponds to 77% of peak memory bandwidth, a substantial fraction of the available bandwidth. The two overheads matter at different scales: 𝑡 iter adds only a few percent to the weight-read time, whereas 𝑡 req accounts for a substantial part of the per-request cost at short contexts but becomes less important as the context grows.

Calibration Runs and Results

We use fixed-length workloads to calibrate the serving-capacity model. Every request has 256 output tokens, while the input length takes one of five values: 1024, 1536, 2048, 4096, and 8192 tokens. Prefill calibration uses the three longer input lengths, and decode calibration uses all five. Power calibration additionally includes measurements from the request-rate sweep and from runs using workloads derived from the production traces. Prefill. At each of the three longer input lengths, we run 5 provisioned deployments, with 1 prefill instance and 1 to 5 decode instances, at an arrival rate above their serving capacity, 15 runs in all. We obtain the completion rate by dividing the decode tokengeneration rate over the measurement window by ℓout − 1, the number of tokens a request receives in decode, since prefill produces its first output token. The completion rate varies by under 2% across each ladder, so prefill is the bottleneck in every provisioned deployment, and its mean over the 5 provisioned deployments is 𝜇𝑃 at that input length. The fit to the three values of 1/𝜇𝑃 gives the MFU and 𝑐 𝑎 of Table 3. The fitted MFU = 0.67 corresponds to an effective compute rate of 67% of GPU peak. Ideal triangular causal attention evaluates ℓin (ℓin + 1)/2 query–key pairs, so its quadratic coefficient approaches 2 from above. The fitted 𝑐 𝑎 = 2.17 is modestly higher. The diagonal explains less than one percent of this gap over the calibration range; the remainder may reflect masked work in boundary tiles, softmax, and differences between the effective utilization of attention and linear layers. We therefore read 𝑐 𝑎 as an effective attention coefficient rather than an exact count of causal query–key pairs.

Power. Each measurement is one GPU in one run: its power is the mean over the measurement window, and its normalized serving throughput is the rate its instance served in that window divided by the serving capacity the model assigns to that role for that workload. The serving-capacity sweeps and runs using the trace-derived workloads provide measurements at serving capacity, with the nonbottleneck role of each provisioned deployment below it, while the request-rate sweep at ℓin = 4096, which runs 7 provisioned deployments at 5 fractions of their measured serving capacity, is the only data with both roles below serving capacity at the same time. We fit the three ramp parameters jointly by bounded nonlinear least squares. The intercept is bounded below by the power of a GPU that holds the model weights but serves no requests, 115 W on this hardware, because a ramp evaluated at zero rate describes such an instance rather than an idle machine. We use multiple starting points because some initializations place every measurement on the cap and leave 𝑝 0 and 𝛾 unconstrained; we retain the fit with the lowest error. The prefill ramp is fitted on 905 per-GPU measurements and saturates at a normalized serving throughput of 0.99, with a root-mean-square error of 19 W; the decode ramp is fitted on 781 and saturates at 0.50, with 49 W. The decode ramp therefore runs at its cap over the upper half of its rate range, which is why a serving-capacity error reaches the modeled power only in part. The measurements of the fixed-length workloads and the Azure trace-derived workload are placed against the two fitted ramps in Figure 3; the fitted ramps approximately describe the power measurements across the calibration workloads. In the case studies of

Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, and Wenqi Cui

Section 5.3, GPUs belonging to removed instances remain idle at approximately 120 W each; their power is excluded from the reported total.

of a disagreement is the relative distance from 𝜆min to the nearest measured serving capacity of any deployment, of which Section 5.2 reports the maximum over all disagreements.

B.4

Algorithm 1: ComputeServingCapacity Input: length distribution L; provisioned deployment (𝑛𝑃 , 𝑛𝐷 ); model, hardware, and serving constants 𝑁 , 𝐿, 𝑑, 𝜅, 𝜋, 𝛽, 𝐶 tok , and 𝑅; and calibrated MFU, 𝑐 𝑎 , MBU, 𝑡 iter , and 𝑡 req Return: prefill capacity 𝜇𝑃 , full-pool decode capacity 𝜇𝐷max , and deployment serving capacity 𝜇

Trace Request Samples

The trace-derived workloads are request samples drawn from the Mooncake and Azure conversation traces. Requests whose input length exceeds a threshold are excluded first, and the experimental requests are then resampled with replacement from the remaining ones. For Mooncake, the threshold is 38,000 input tokens, which leaves room for the output within the configured context window of 40,960 tokens; it excludes 647 of the 12,031 trace requests (5.4%), and 1,000 requests are resampled from the rest. For Azure, the threshold of 32,768 input tokens excludes none of the 19,366 trace requests, and 7,200 requests are resampled. For Azure, input and output lengths are resampled as pairs, which preserves their correlation; the Mooncake runs drew the two lengths independently, so the Mooncake samples have no input–output correlation.

1

2

/* Prefill capacity from the compute roofline Compute 𝑎𝑃 , 𝑏 𝑃 , and 𝜇𝑃 from (4) and (5)

*/

/* Decode coefficients from the memory roofline */ Compute ℓ¯ctx , 𝑎𝐷 , and 𝑏 𝐷 from (7) and (8), with the calibrated overheads of (30)

/* Operating batch from the memory balance */ Set 𝐶𝑉𝑎2 ← 1/𝑛𝑃 and 𝐶𝑉𝑠2 ← Var𝑟 (𝑡𝑃 )/E𝑟 [𝑡𝑃 ] 2 4 for each candidate 𝐵 selected by the bracketed root solver do 5 Compute 𝑡¯𝐷 and 𝜇𝐷 from (8) and (9) 6 Compute 𝜌 𝑃 ← 𝑛𝐷 𝜇𝐷 /(𝑛𝑃 𝜇𝑃 ) and 𝑡¯𝑊 from (21) 7 Evaluate the prefill occupancy, decode occupancy, and unused capacity of (10) 8 end 9 Set 𝐵 to the root at which the three occupancies sum to 𝐶 tok , and keep the corresponding 𝑡¯𝐷 and 𝜇𝐷 3

C

Algorithm for Computing Serving Capacity

Algorithm 1 summarizes the procedure for computing the serving capacity of a provisioned deployment under a given workload. Its inputs are the workload length distribution, the provisioned deployment, the model and hardware constants of Appendix B.1, and the calibrated parameters of Appendix B.2. The prefill capacity follows from the compute roofline alone, whereas the decode side has to be solved jointly: the operating batch 𝐵 determines the decode capacity and the prefill waiting time, and the reservation occupancy they imply in turn constrains 𝐵. The algorithm therefore solves the modeled memory balance of Section 4.2 for 𝐵, and the remaining quantities follow from it. The full-pool decode capacity 𝜇𝐷max is obtained by substituting the full-pool batch 𝐵 max of (11) into the decode-capacity relation (9); it provides the reference capacity by which the power model normalizes decode throughput in (12). In our implementation, the root of the memory balance is found with SciPy’s brentq, an implementation of Brent’s bracketed rootfinding method, on the bracket 0, min(𝐵 max, 𝐵 𝜌 ) of Section 4.2, with the upper endpoint chosen inside this interval where the residual 𝑔(𝐵) is finite and positive, at the default tolerance. The deployment-selection comparison of Section 5.2 applies the provisioning problem (1) twice for each required serving capacity 𝜆min : once with the modeled 𝜇 (𝑛𝑃 , 𝑛𝐷 ; W) and P (𝑛𝑃 , 𝑛𝐷 ; W), which gives the model’s choice, and once with their measured counterparts 𝜇 meas and P meas , which gives the measurement-based choice; each choice is the first deployment at or beyond 𝜆min along the corresponding front of Figure 4. The fronts of Figures 1 and 4 are 𝜖-Pareto fronts [29] with 𝜖 = 3% on both serving capacity and power, computed separately on the modeled and on the measured deployments; only deployments on a front are drawn, and the selection uses every evaluated deployment. The front of Figure 1 is computed for the fixed-length workload with 4096 input and 256 output tokens. The model’s choice is evaluated with its measured serving capacity and power. The requirement 𝜆min takes 400 values spaced logarithmically between the smallest and largest measured serving capacity of the workload; the agreement is the share of these values at which the two choices coincide, and the margin

/* Capacities of the provisioned deployment */ Compute 𝜇 ← min(𝑛𝑃 𝜇𝑃 , 𝑛𝐷 𝜇𝐷 ) by (2) max from (11) 11 Compute the full-pool decode capacity, with 𝐵 10

𝐵 max ← 12

𝐶 tok − E𝑘 [𝑈𝑘 ] , ℓ¯ctx + 𝑅

𝜇𝐷max ←

𝐵 max E𝑟 [ℓout − 1] (𝑎𝐷 + 𝑏 𝐷 𝐵 max )

return (𝜇𝑃 , 𝜇𝐷max, 𝜇)

D

Sequential Scale-Down Experiment

We extend the two provisioning scenarios of Section 5.3 with a sequential scale-down experiment at a fixed request rate. The experiment illustrates the boundary between efficient scale-down and overload: reducing provisioned capacity can lower power while the resulting deployment retains sufficient serving capacity, but once serving capacity falls below the request rate, further power reduction causes sustained queueing and rapidly increasing latency. The deployment follows 3p2d → 3p1d → 2p1d → 1p1d at 3.6 req/s, or 80% of the serving capacity of 2p1d (Figure 6). The first two resulting deployments remain on the Pareto front and retain sufficient serving capacity for this request rate. Removing a decode instance reduces total power by approximately 520 W and increases median TPOT from 24 to 29 ms, while median TTFT remains unchanged. Removing a prefill instance next reduces total power by approximately 180 W, less than the decode removal because the two remaining prefill instances now operate at 80% of their serving capacity. Median TTFT increases from 0.5 to 0.9 s, with individual-request TTFT reaching up to 6 s.

Analytical Power-Aware Provisioning for Prefill–Decode Disaggregated AI Inference

Power (kW)

Request rate (req/s)

Throughput

TPOT (ms) TTFT (s)

After the third removal, 1p1d has a serving capacity of 2.3 req/s, below the fixed request rate of 3.6 req/s. Serving throughput falls to the serving capacity, while requests accumulate in the queue because the request rate exceeds it. TTFT increases throughout the overloaded period, reaching a maximum of 320 s by the end of the run, approximately 14 min later. The 470 W power reduction therefore comes at the cost of a persistent throughput shortfall and a growing queue, so 1p1d cannot sustain the offered request rate. Thus, the final removal marks a qualitative change in operating regime: the earlier power savings preserve a stable deployment, whereas the final reduction is obtained only by leaving serving capacity below the request rate.

Serving capacity

Arrival rate

5 0 3 2 1 p50

250

p95

0 25 0

0

5

10

15

20

25

Time (min)

30

35

40

Figure 6: Three removals chained under one request rate (3p2d → 3p1d → 2p1d → 1p1d), the same rows as Figure 5.

Record · ID 1028651 · SHA-256 a12653cdacbff985
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.