1
ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters
arXiv:2609.15230v1 [cs.DC] 14 Sep 2026
Rui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang
Abstract—arge language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling.arge language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling.L Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to ServiceLevel-Objective (SLO) violations. In this paper, we study joint cooling–computing control for LLM inference: minimizing perjob GPU-plus-cooling energy while satisfying thermal safety and latency SLO constraints. We present ETCInfer, an energyefficient, thermal-aware scheduler that selects a pre-job Computer Room Air Conditioner (CRAC) setpoint and adapts perGPU frequency and micro-batch size during execution. ETCInfer builds compact physics-informed control models by calibrating GPU heat generation, chassis heat dissipation, CRAC power, and prefill/decode latency relations from telemetry. These models estimate hidden thermal states and time-to-throttle, enabling the scheduler to evaluate energy, temperature, and latency before applying an action. We formulate this joint setpoint–frequency– micro-batch control problem as a partially observable Markov decision process and design ETCAdapter, a learning-based controller that minimizes per-job energy under thermal safety and SLO constraints. We implement ETCInfer as a coordination layer over typical inference and cluster management stacks. Evaluation across real-trace simulation and validation experiments shows that ETCInfer reduces total job energy by up to 33.1%, thermal throttle exposure by up to 92.9%, and keeps SLO violation rates below 0.7% even at ambient temperatures up to 48◦ C. Index Terms—LLM Inference, Datacenter Scheduling, Energy Efficiency, Thermal Throttle, Parallel Processing, Distributed Inference, GPU Cluster
I. I NTRODUCTION Currently, large language models (LLMs) perform remarkably in natural language processing (NLP). Models like GPT from OpenAI and Gemini from Google have demonstrated impressive abilities across various tasks, including question answering, summarization, search, code generation, and induction [1], [2]. With rapidly advancing capabilities, LLMs are seeing surging user demand, but this growth also brings substantial computational costs for scaling inference on billionparameter models. To address these demands, companies are Rui Lu is with the Department of Computing, The Hong Kong Polytechnic University, Hung Hom 999077, Hong Kong. Emails: ([email protected]) Dan Wang is with the Division of Environment and Sustainability Academy of Interdisciplinary Studies, Hong Kong University of Science and Technology, Clear Water Bay, Hong Kong. Emails: ([email protected]) Rui Ge, Huanghuang Liang, and Xiaobo Zhou are with the State Key Laboratory of IOTSC, University of Macau, China. Rui Ge and Huanghuang Liang are also affiliated with the School of Computer Science, Wuhan University, China. Emails: (gerui, huanghuangliang, [email protected])
rapidly building new AI datacenters equipped with millions of high-performance GPUs such as Nvidia H100. For example, Microsoft plans to invest $80 billion in 2025 to expand its AI-enabled datacenters globally [3]. As a result, the energy consumption of AI datacenters has increased significantly, contributing substantially to their carbon footprint. Today, AI facilities are already a visible component of global electricity demand and are projected to reach several percent of worldwide consumption within the next decade. In AI datacenters, energy consumption is dominated by two closely coupled sources: computational hardware (GPUs) and the cooling systems required to dissipate the resulting heat. High-performance GPUs, such as the NVIDIA H100 SXM, can draw up to 700 W each, amounting to several megawatthours annually per GPU [4]. Since most consumed electrical energy ultimately becomes heat in the room, cooling infrastructure is essential for safe operation and hardware reliability, often consuming 30%–40% of total datacenter power [5], [6]. Raising the ambient temperature in the computer room is therefore an attractive knob to reduce cooling energy and carbon emissions. Recent data-center studies [7] indicate that increasing the supply or ambient temperature to around 41◦ C can cut cooling energy costs by up to 56% compared to conventional settings of 20–25◦ C. Correspondingly, several standards have been established to guide optimal cooling configurations [8]–[10], as summarized in Fig. 2. Singapore recommends 28–32◦ C for Level-4 datacenters [11], while the EU Code of Conduct allows higher operating ranges under appropriate reliability controls [9], [10]. Consequently, operating AI datacenters at 35–41◦ C has the potential to yield substantial long-term energy and carbon savings. However, operating LLM inference at warmer, energyefficient ambient setpoints introduces coupled thermal and performance challenges. High GPU throughput requires substantial power, generating massive heat. Cooling systems are typically designed for moderate ambient temperatures (around 20–30◦ C) with adequate inlet airflow. As the ambient temperature rises, the temperature gradient from chip to air narrows, reducing cooling effectiveness and accumulating heat. When junction temperature approaches the limit, the GPU triggers thermal throttling [12], [13], reducing clock frequency and voltage to protect itself, which directly lowers throughput until the temperature returns to a safe range. For heavy LLM inference jobs, repeated or extended throttling can severely impact time-to-first-token (TTFT), time-per-output-token (TPOT), and effective throughput, and accelerate device aging.
2
Fig. 1. The cooling systems in AI datacenters.
Existing work on AI datacenter energy mainly optimizes compute or cooling separately. On the compute side, DVFS and model compression lower GPU energy for LLM workloads [14]–[17]. On the cooling side, CRAC-oriented studies raise setpoints and refine cooling design to reduce cooling energy and emissions [7]–[11]. At the serving layer, thermalaware resource provisioning and schedulers such as TEAP and TAPAS [13], [18] emphasize host allocation, batching, and placement but do not jointly control facility-side setpoints with per-GPU thermal states. In ETCInfer, rack placement and airflow direction are treated as sources of heterogeneous inlet temperature and cooling delay, not as runtime control actions. These techniques provide useful building blocks, but their objectives remain separated across serving, device power, and facility cooling. As a result, current systems do not jointly optimize ambient setpoint, per-GPU heat dynamics, and job SLOs, leaving the design space where GPU frequency, microbatch size, and room setpoint interact, leading to three coupled challenges: Challenge 1 (Thermal–performance modeling). A thermal prediction model is necessary, which tightly links LLM latency to GPU temperature, especially near the throttle point. As the frequency drops to protect the device, job latency would extend. The key challenge is to predict time-to-throttle and performance from real-time sensor data. Challenge 2 (Joint computing–cooling energy modeling). A joint energy model on both GPUs and CRACs is necessary. Raising ambient setpoints reduces cooling power but increases execution time under throttling, canceling savings. The challenge is to couple dynamic GPU heat generation and dissipation so we can quantify GPU-plus-CRAC per-job energy under different control actions. Challenge 3 (SLO-aware thermal scheduling). An SLOaware scheduler is necessary that selects and operates at warm ambient setpoints. It also allocates jobs, chooses microbatch sizes and GPU frequency, and avoids throttling using incomplete thermal information, maintaining SLO guarantees. The objective is not to maximize the setpoint, but to find a safe operating point where cooling savings do not cause throttlinginduced latency violations. In this paper, we propose ETCInfer, an energy-efficient, thermal-aware cooling-joint scheduler for LLM inference in AI datacenters. ETCInfer serves as a coordination layer across LLM serving, GPU control, and room cooling, rather than replacing the serving engine, cluster manager, GPU driver, or CRAC controller. First, we develop a physics-informed
Fig. 2. Datacenter guidelines. Fig. 3. Breakdown of Green: candidate setpoints. datacenter energy.
modeling framework that couples GPU heat generation, airside heat dissipation, room-level cooling, and LLM workload latency to expose per-GPU thermal states, time-to-throttle, and SLO forecasts. Second, we formulate the joint control of ambient setpoint, GPU frequency, and micro-batch size as a partially observable decision problem and design ETCAdapter to optimize per-job energy under SLO and thermal-safety constraints. ETCInfer combines a joint cooling–computing formulation, physics-informed hidden-state estimation, and online integration with LLM serving controls. Third, we implement and evaluate it using trace-driven experiments in real-trace CFD simulation and validation experiments. Results show that ETCInfer reduces total job energy by up to 33.1%, shrinks thermal throttle exposure by up to 92.9%, and keeps SLO violation rates below 0.7% even at ambient temperatures up to 48◦ C. We also release the ETCInfer source code to support reproducibility. The contributions of this paper can be summarized as follows: • We formulate joint LLM scheduling and ambient temperature control as a partially observable decision problem whose actions include CRAC setpoint, GPU frequency, and microbatch size. • We develop physics-informed control models that couple heat generation, air-side heat dissipation, CRAC power, and workload latency to estimate hidden thermal states, time-tothrottle, and SLO risk. • We design ETCAdapter, an online controller that adapts frequency and micro-batch size from telemetry and learned latent dynamics while enforcing SLO and thermal-safety constraints. • We implement ETCInfer on typical LLM inference frameworks and integrate it into AI datacenter stacks. • Our experiments show ETCInfer significantly reduces energy and throttling while preserving strict SLOs. II. BACKGROUND A. Thermal Challenges in Modern GPUs. Heat generation in scaled semiconductors inside the GPU brings challenges. GPUs convert nearly all electrical power into heat as billions of switching elements toggle at high frequency. Heat flux rises with core frequency and voltage, both kept high to sustain throughput for compute-dense LLM workloads. Advanced semiconductor scaling further amplifies heat density, which increases thermal stress in future GPUs. Recent studies show that smaller transistors possess lower thermal mass and fewer heat-dissipation pathways, causing faster temperature accumulation and larger performance loss.
3
For example, 5nm GAAFETs trap more heat than 7nm FinFETs [19], while 7nm exhibit 12K self-heating, equivalent to 5nm structures reaching 17K. This 5K rise significantly affects performance, increasing gate delay by up to 39% at 5nm versus 25% at 7nm, intensifying leakage currents. Thermal Throttle of GPUs. When heat generation exceeds cooling capacity, junction temperature rises, eventually triggering thermal throttling. This reduces voltage and frequency, limiting GPU performance and throughput, especially for compute-intensive LLM inference. Prolonged throttling accelerates GPU aging, increasing risks like electromigration and joint fatigue. Studies show higher temperatures correlate with higher error rates and earlier failures [20]. Avoiding throttling improves throughput, energy efficiency, and hardware lifespan. Dynamic Voltage and Frequency Scaling (DVFS) on GPUs. DVFS is a common technique for managing GPU power and thermal behavior [21]–[23], following earlier GPU power/performance and GPGPU power-modeling work [24], [25]. It adjusts supply voltage and operating frequency to reduce heat generation at the source, lowering the likelihood of thermal throttling. On NVIDIA GPUs, operators can adjust core and memory clocks, voltages, and board power caps through NVML. By balancing power and performance, effective DVFS policies reduce unnecessary energy use and mitigate throttling, improving resource efficiency in AI datacenters. B. Heat Dissipation in Datacenters The heat dissipation for GPUs of a datacenter is shown in Fig. 1. Typically, servers host multiple identical GPUs (e.g., 6–8 cards) within a single server chassis. Several chassis are mounted in a rack, and racks are arranged in rows. Each chassis and GPU uses front-to-back fans that draw cold air from the cold aisle and exhaust hot air into the hot aisle. Most facilities implement hot-aisle/cold-aisle containment. Computer Room Air Conditioner (CRAC) units ingest hotaisle return air, remove heat, and deliver conditioned supply air to the cold aisles, closing the thermal loop in the datacenter. The fans on chassis and GPUs, typically are supplied by chassis vendors (e.g., Supermicro) and GPU vendors (e.g., NVIDIA) and are designed for ambient temperatures of 20–30 ◦ C. Each fan consumes only a few watts, which is negligible compared to GPU power, so their energy cost is often ignored, and they are assumed to operate at maximum RPM. However, their effectiveness decreases as ambient temperature rises [26], [27]. Cooling Guidelines for Datacenters. In datacenters, CRAC systems (approximately 35%) and servers/GPUs (approximately 55%) account for most energy usage [28], making ambient temperature a key control parameter. Lower setpoints increase onboard and chassis cooling capacity and reduce GPU thermal throttling, but excessive cooling wastes energy. Industry guidelines in Fig. 3 recommend fixed ambient ranges to balance reliability and efficiency: ASHRAE suggests 27 ◦ C, while the EU allows up to 35 ◦ C [8], [9]. Recent studies indicate that raising ambient temperatures toward 41 ◦ C can further save energy with minimal impact on hardware lifespan [7]. Efficient cooling reduces energy use and also carbon
emissions, enabling AI datacenters to support large-scale LLM workloads with lower environmental impact. C. Scheduling of LLM Inference LLM inference at datacenter scale requires a scheduler that places requests on heterogeneous GPUs, coordinates batching and model parallelism, and manages memory-intensive state such as KV caches under time-varying load [29]–[31], including prefill/decode and chunked-prefill scheduling systems [32]–[34]. The scheduler must balance user experience, cost, and infrastructure limits such as memory and bandwidth. However, most production schedulers are thermal-agnostic: they abstract GPUs as virtual compute units and ignore perdevice physical context such as inlet temperature, fan speed, and thermal headroom. At the application layer, platforms schedule virtual resources, such as containers and Pods in Kubernetes or Ray, while cluster managers that place physical GPUs, such as Kubernetes with the NVIDIA device plugin, SLURM, or LSF, usually lack feedback loops from CRACs and device telemetry for per-GPU thermal awareness. As a result, placement and scaling decisions are driven mainly by software metrics such as queue length, latency, and utilization rather than actual thermal conditions. Thermal Prediction–Driven Thermal-Aware Scheduling. While the energy/carbon-focused scheduler reduces power and emissions, it cannot prevent performance drops when GPUs approach thermal limits. If cooling is slow or inlet temperatures are high, die temperature rises, triggering thermal throttling, which reduces throughput and risks SLO violations. LLM inference amplifies this risk due to long, bursty decode phases that keep power high even at moderate utilization. Heterogeneous rack placement and airflow create uneven thermal conditions across replicas. Prolonged throttling accelerates wear, increases error rates, and threatens hardware lifespan and service reliability. These challenges motivate a thermal-aware scheduler that treats temperature and timeto-throttle as key signals, guiding placement, batching, and DVFS adjustments, while coordinating with facility controls to maintain throughput and stability under real heat constraints. III. S CHEDULER OVERVIEW A. Design Overview Modern LLM inference in AI datacenters follows a structured pipeline involving multiple schedulers and runtime engines, as shown in Fig. 4. This figure is a general multiservice deployment view: different Pods or virtual services may host different LLM services, while the concrete model choices used in evaluation are specified in Section VII-A. Before execution, the Service Provider defines each job class with detailed specifications, including TTFT and TPOT SLOs, input and output token limits, job categories such as code completion or dialog, and the target model family or service variant, e.g., Llama-3.1-Instruct. The Cluster Scheduler, typically implemented on Kubernetes or VMware, interprets these specifications, creates Pods, retrieves models from registries, and prepares the runtime environment, such as vLLM. During inference, it manages auto-scaling, load balancing,
4
Fig. 4. General multi-service LLM inference workflow in an AI datacenter. The figure illustrates the deployment architecture rather than the exact model set used in evaluation.
and placement. The Node Resource Manager assigns Pods to physical nodes, exposes hardware states, and enforces resource isolation. Incoming user queries arrive through the Inference Frontend, where the Requests Router groups and dispatches them to Pods based on job category, prompt length, and model choice. Cooling infrastructure maintains chassis temperature following site guidelines. In this paper, we design an energy-efficient, thermal-aware cooling-joint Scheduler, ETCInfer (Fig. 5), which raises ambient temperature to reduce cooling energy while preserving service quality and hardware lifetime. We build a thermalstate model linking GPU heat generation and dissipation to latency and energy. ETCInfer coordinates three controls: (1) Micro-batch control in serving Pods at the virtual machine layer, which shapes GPU load, heat, queuing, and SLOs because a larger micro-batch improves utilization and amortizes overhead but also increases per-step work, power, heat, and latency risk, (2) GPU-frequency control on physical machines via DVFS, which tunes power and throughput under thermal and SLO limits, and (3) CRAC setpoint control in the computer room, which adjusts ambient temperature within safety margins. These jointly predict and schedule thermal dynamics to meet service objectives. ETCInfer supplies coordinated control decisions to these layers, but it does not replace vLLM batching logic, Kubernetes/KServe placement and routing, GPU driver enforcement, or the facility CRAC controller. Assumptions. Our scheduling makes several practical assumptions: (1) Ambient setpoint changes take effect with delays, so chassis inlet temperature varies with airflow, rack position, and activity. (2) Chassis fans operate at maximum speed with negligible, constant power, and are excluded from optimization. (3) Each GPU follows a driver-level DVFS voltage–frequency curve, and we tune only frequency while voltage adjusts automatically. (4) Each inference job is routed to a Pod via the Requests Router and shares stable properties such as SLOs and input length. B. Scheduling Objectives Let Gj be the set of GPUs of the Pod assigned to job j. Let B represent the inference batch with size Nbatch from the Requests Router and I = {0, 1, . . . , Ij } index control and measurement instants. For GPU x ∈ Gj and time i ∈ I, sensors
x provide observation ōxi = (P̄ix , ūxi , T̄ix , T̄in,i ), including GPU x x power P̄i , GPU utilization ūi , and GPU core temperature T̄ix , x 1 . chassis inlet temperature T̄in,i Scheduling Space. 1) Ambient setpoint Tset ∈ [Tmin , Tmax ] is an initialization action before the job begins. It can only be configured, but cannot be adjusted in real-time. 2) PerGPU frequencies fix ∈ [fmin , fmax ] for all x ∈ Gj and i ∈ x I. 3) Each micro-batch allocated in GPU x has size Nmic and it must satisfy the SLO (SLOj ) forPthe job . The size of x the micro-batch satisfies that Nbatch = x∈Gj Nmic . A larger x Nmic assigns more token work to GPU x in one serving step. ETCInfer increases it when SLO slack and thermal headroom are sufficient, and decreases it when predicted temperature, throttling risk, or latency risk rises. The three controls operate on different time scales: setpoint is selected before execution because cooling reacts slowly, while frequency and microbatch size are adapted during execution because they directly change heat generation, throughput, queueing, and SLO risk. In summary, the actions at time i could be represented by
a0 = Tset ,
x ai = {fix , Nmic } for i > 0.
(1)
Models and constraints. Let Lj (·) denote the latency metrics, x e.g., TTFT and TPOT, and let Tthr be the throttle threshold. The scheduling problem for minimizing the overall job energy consumption E(j, Tset ) can be formulated as: min s.t.
E(j, Tset ) x Lj ({fix }, Nmic ) ≤ SLOj , x T̄ix ≤ Tthr ,
∀fix ∈ [fmin , fmax ], ∀ x ∈ Gj , ∀ i ∈ I. Here, T̄ix is a sensed safety variable at runtime rather than a x control variable. ETCInfer acts on Tset , fix , and Nmic to keep future sensed temperatures below the throttle threshold. Overall Job Energy Consumption. To compute the energy of LLM inference job j on its allocated GPUs and room cooling, let ∆ti = ti − ti−1 . The interval energy over period i ∈ I is X x ∆Ei (j, Tset ) = wi P̃CRAC,i−1 (Tset ) ∆ti + P̄i−1 ∆ti , (2) x∈Gj 1 In this paper, ¯ · denotes sensor readings, ˜· denotes physics-based estimates, and ˆ· denotes learned predictions. For example, T̄ix is sensed temperature, T̃ix is physics-estimated temperature, and T̂ix is model-predicted temperature.
5
quantities, and P̃CRAC,i (Tset ) is a room level quantity shared by all GPUs. Typically, õxi = {Q̇xG (i), Q̇xD (i), t̃xthr (i)}
(5)
where Q̇xG (i), Q̇xD (i) are the heat generation/dissipation rate of GPU x at time i. t̃xthr (i) is the time to trigger GPU throttle. Although those non-observable variants cannot be obtained directly, we could obtain them with physics-based estimation in the following section. Fig. 5. Workflow of ETCInfer Scheduler.
IV. P HYSICS - INFORMED MODELING where P̃CRAC,i−1 (Tset ) denotes the estimated electric power consumed by the room cooling system at the selected setpoint and IT heat load, and wi ∈ [0, 1] represents the fraction of shared CRAC power allocated to cooling, computed as wi = P tot x x∈Gj P̄i−1 /Pi−1 . For a dedicated testbed, set wi = 1. The total energy to complete job j at setpoint Tset is E(j, Tset ) =
Ij X
∆Ei (j, Tset ).
(3)
i=0
C. Scheduling Problem as POMDP We model the energy-efficient, thermal-aware cooling-joint LLM inference scheduling problem as a Partially Observable Markov Decision Process (POMDP), which handles control with hidden stochastic states and incomplete observations [35]. Scheduling is sequential: each interval-i action changes throughput and power, and affects future die temperature, time-to-throttle, SLO slack, and CRAC energy via thermal inertia. Thus, greedy interval control or fixed thresholds may save immediate energy yet trigger later throttling or SLO violations. With a fully observed thermal–workload state, the problem becomes an MDP whose transition depends on the current state, action, and stochastic workload. ETCInfer uses a POMDP because GPU and room sensors are noisy and delayed, and key variables are unobserved: GPU heat generation, heat dissipation, local inlet-air effects, and time-to-throttle. Actions, including ambient setpoint, GPU frequency, and micro-batch size, jointly affect heat, cooling, and latency. ETCInfer therefore combines model-based prediction with belief-state control under partial observability, delay, and model error. We use POMDPs to capture the partial observability. The contribution of ETCInfer lies in the ETCInfer state, action, constraint, and model design for coupled cooling and LLM serving. Compared with threshold control, deterministic optimization [36], MPC [37], robust MPC [38], [39], and filtered MPC [40], ETCInfer retains prediction but learns belief states and a world model from observations, adapting online to stochastic token lengths, telemetry delay, airflow heterogeneity, DVFS variation, and hidden thermal dynamics. The system state si at time i captures thermal and workload conditions as si =
h
i {ōxi , õxi }x∈Gj , ρ̄i , Tset , E1:i (j, Tset ), P̃CRAC,i (Tset )
(4)
where ōxi is the sensor observation of GPU x and includes sensed temperature used for safety checking, ρ̄i is the percentage of SLO left for this job, õxi contains GPU-related hidden
This section presents physics-based estimators that infer the non-observable variables needed for control. We integrate three layers, including GPU heat dynamics, facility cooling, and workload behavior based on physical rules. This unified view addresses what the model computes, how much heat the GPUs create, how much heat the room can dissipate, and how controls and scheduling trade performance against energy. Our model consists of five physics-informed components. 1) Heat Generation Model of GPU. We estimate board heat from electrical power using a calibrated, frequency-driven die-power model, mapping die power to board power in a lightweight hardware-aware manner. 2) Heat Dissipation Model of Cooling. We model the air-side path in a multiGPU chassis, linking fan capacity and airflow geometry to maximum removable heat and describing how shared cooling is distributed across GPUs. 3) CRAC Power Model. Roomlevel cooling power is expressed as a COP-based function of setpoint and IT heat with a clipped standby term. 4) Thermal Safety and Time to Throttle. Estimated temperature and removed heat to obtain the predicted time before throttling, ensuring safe operation. 5) Latency Estimation Model of SLOs. Prefill and decode latencies are combined into total latency and checked against SLOs, enabling evaluation of scheduling actions without executing the job. The power model follows frequency-driven GPU power modeling [21], [23]–[25], [41], the cooling and CRAC terms follow standard air-side heattransfer and COP-based cooling approximations [5], [6], [8], [9], [42]–[44], and the latency model follows roofline-style compute, memory, and communication accounting [29]–[34], [45]. ETCInfer calibrates their coefficients from telemetry and uses the calibrated models for online scheduling. A. Thermal Estimation Models To estimate GPU thermal states, we account for two processes: heat produced by the GPUs and heat removed by the cooling system. First, we build a heat generation model that estimates a GPU’s heat output for a given hardware and workload in § IV-A1. Then, we model heat dissipation through the onboard cooler to the ambient air. They allow us to estimate temperature over time. 1) Frequency-driven Heat Generation Model of GPU: In modern GPUs, electrical energy powers transistors and memory, with losses as Joule heat in the silicon, VRMs, and the board. A GPU card has hundreds of components, including VRAM, the GPU die, fans, VRMs, inductors, and capacitors.
6
Among these, the GPU die consumes most power, about 80– 85% in compute-heavy workloads, with VRAM at about 16%. Therefore, thermal output in watts is approximately equal to the instantaneous electrical power draw of the GPU board as: x Q̇xG (i) ∼ P̃ix = P̃die (i)/ϕdie ,
(6)
where ϕdie = 80% is the proportion of the overall power from a single GPU die. We use a physics-based model [21], [23]–[25], [41] that estimates dynamic power as a calibrated function of operating frequency. After per-SKU profiling with GPU telemetry, this compact model tracks the operating-pointdependent power used by ETCInfer’s scheduler. The physicsbased approach to estimate the GPU die power is: x P̃die (i) = β1x fix + β2x (fix )3 + β3x
(7)
with β1x , β2x obtained via calibration and β3x representing leakage power. Voltage effects are captured in the fitted β. 2) Heat Dissipation Model of Chassis Cooling: We model the air-side heat removal of a closed server chassis that contains multiple high-power GPUs, following standard controlvolume heat-transfer relations for datacenter cooling analysis [5], [44] and datacenter thermal-guideline/placement assumptions [8], [9], [42], [43]. The model connects fan specifications with the internal flow resistance of the chassis. The goal of this model is to estimate the heat that can be carried away by the airstream per unit time, i.e., the heat dissipation rate. Assumptions. 1) Fans operate at their maximum capability (fixed speed). 2) No interaction between fans, where Nfan fans provide Nfan times the single-fan capacity. 3) The chassis is a well-mixed control volume on the air side. 4) A bypass factor γ ∈ (0, 1] accounts for leakage and recirculation. Nfan fans at maximum capability. Let uface be the measured or rated face velocity at maximum speed and Afan the open flow area of the fan disc after hub and grill deductions. The total volumetric flow of Nfan fans is V̇N = Nfan V̇max ,
where V̇max = uface Afan .
(8)
The corresponding mass flow is x x ṁ = ρ(T̄in ) V̇N = ρ(T̄in ) Nfan V̇max .
(9)
Therefore, the air-side capacity rate of each GPU could be computed on average as: tot x C̃air (i) = γ ṁ cair (T̄in )
[W/K]
(10)
x where ρ(T̄in ) denotes air density as a function of inlet temperature and cair ≈ 1006 J kg−1 K−1 for dry air near room temperature. And for each share of GPU over this chassis cooling system, the capacity rate is
Q̇x (i) x tot . C̃air (i) = C̃air (i) P G y y Q̇G (i)
(11)
The airflow acts like a conveyor with capacity rate C̃air (t)[W/K]. Because heat must be generated before it can be removed, the heat dissipation rate is limited by both the generation and the air-side capacity, and is modeled as: x x Q̇xD (i) = min{Q̇xG (i), C̃air (i)[T̃ix − T̄in,i ]}
(12)
This bound enforces energy conservation and avoids unrealistically low outlet temperatures when cooling is saturated. 3) GPU Throttle Trigger Time Estimation: We adopt a lumped thermal model for GPU x, consistent with compact thermal modeling practice in datacenter thermal analysis [44]. The temperature from time i to the next time interval is estimated i + 1 as: x T̃i+1 = T̃ix + [Q̇xG (i) − Q̇xD (i)]/C x · ∆t,
(13)
x depends on the last-sampled temperature T̃ix with Here, T̃i+1 measurement interval ∆t = ti+1 − ti , subject to the heat capacity of this substance, i.e., the temperature increase of this substance given a certain amount of heat. Here, C x = mx cx is the lumped thermal capacitance of the GPU, with mx the effective thermal mass and cx the specific heat capacity. x Let Tthr denote the thermal throttle threshold. The estimated time to reach the throttle threshold from T̃ix is
t̃xthr (i) =
x Tthr − T̃ix · C x , if Q̇xG (i) − Q̇xD (i) ≥ 0. x Q̇G (i) − Q̇xD (i)
(14)
B. Power Model of CRAC We estimate the cooling device’s electric power when the air temperature setpoint Tset maintains room inlet temperatures within a target range. This model uses a COP-based approximation of datacenter cooling power with site-calibrated coefficients [6], [44]. The setpoint and server-plus-cooling assumptions follow datacenter cost models and thermal guidelines [8], [9], [42], [43]. Assumptions. 1) The IT equipment generates a sensible total amount of heat rate Q̇IT (i) that the CRAC removes. 2) Latent load is negligible in the data hall. 3) Efficiency changes with Tset at a constant slope in a short range around Tref . a) Reference Coefficient of Performance: COPref is the ratio of heat removed to electric power at Tref . A higher COPref means lower power for the same load. For example, a COP of 4 means the device removes 4 kW of heat by using 1 kW of electricity. COP rises when the temperature lift decreases, so increasing the evaporating or chilled-water temperature or lowering the condensing temperature typically improves COP. b) Temperature sensitivity: Let s denote the fractional change in cooling power per ◦ C. In practice, s ∈ [0.02, 0.05] and is calibrated from real data, as shown in Tab. I. c) Model of CRAC Power: The cooling system power at ti for setpoint Tset and heat removal Q̇IT (i) is: P̃CRAC,i (Tset ) =
Q̇IT (i) [ 1 − s (Tset − Tref ) ] . COPref
(15)
When raising the setpoint by ∆T > 0, the power reduces by approximately s × ∆T in percent. Lowering the setpoint increases power by the same rule. Eq.15 reflects the gain when the setpoint temperature rises. d) Choice of the setpoint Tset : Given an allowed inlet band x x Tmin ≤ T̄in ≤ Tmax , choose Tset so that T̄in ≤ Tmax . For planning, evaluate (15) at the selected Tset . For an upper bound on savings, we adopt Tset = Tmax .
7
TABLE I T YPICAL CRAC SAMPLES
System type DX CRAC (air cooled) Chiller + CRAH (air cooled) Chiller + CRAH (water cooled) High efficiency variable speed chiller
COPref 3.0 to 4.0 3.5 to 5.5 5.5 to 7.0 7.0 to 9.9
s per ◦ C 0.03 to 0.04 0.02 to 0.03 0.03 to 0.05 0.04 to 0.05
C. Latency Estimation Model for Multi-GPU LLM Inference To satisfy SLO constraints, estimating latency Lj (·) is necessary. We use a roofline-style estimator that separates compute-bound, memory-bound, and communication-bound terms, and then calibrates residual runtime overheads from measured traces [29], [30], [45], following common LLMserving latency, prefill/decode scheduling, and KV-cache considerations [31]–[34]. Generally, LLM inference consists of two stages: prefill and decoding, which recent serving systems schedule differently because of distinct batching, compute, and memory behavior [32]–[34]. 1) In the prefill stage, the model processes the entire input, computes layer activations, and builds the key–value (KV) cache for attention. The delay until the first output token appears is called TTFT. 2) In the decoding stage, the model generates tokens sequentially by reusing the KV cache from the prefill stage. Here, token latency is primarily determined by KV reads and increases with the effective context length. The time for each output token is called TPOT. In each stage, the time is the maximum of compute, memory, and any communication time due to parallelism. Given per-GPU frequencies fix ∈ [fmin , fmax ] for all x ∈ Gj and i ∈ I, the effective peak FLOPs NFx LOP S (f x ) x at the chosen precision can and peak VRAM bandwidth NBW be determined. The relationship between frequency and peak FLOPs is profiled as: x NFx LOP S ∝ Ncore · fix · 2,
(16)
x where Ncore is the number of GPU core of GPU x, and the factor 2 reflects fused multiply-add FLOP accounting in the profiled peak-FLOPs estimate. Then, let X X x Θtot = ηc NFx LOP S (f x ), Btot = ηm NBW , (17) x∈Gj
x∈Gj
where ηc , ηm ∈ (0, 1) capture kernel, utilization, and memory-efficiency losses. Let the model have NL layers, hidden size dmodel , attention width dattn (often dattn = dmodel ), and Nne non-embedding parameters. Each element uses selt bytes (e.g., 2 for BF16/FP16, 1 for INT8). At decode step k, the context length is: Nin,k = Nin + k − 1,
k = 1, 2, . . . , Nout .
(18)
a) Compute time: Given input prompt length Nin and output prompt length Nout , the standard FLOPs accounting for a forward pass is: Cpref (Nin ) = 2Nne Nin + 2NL dattn Nin (Nin + 1), N out X Cdec (Nout ) = 2Nne + 2NL dattn Nin,k k=1
(19) (20)
We keep Cdec in summation form because each generated token has a different context length, while Cpref can be expanded directly over the fixed input prompt. The corresponding compute-bound times are estimated by Cdec Cpref , t̃cmp,dec = . (21) t̃cmp,pref = Θtot Θtot b) Memory time (KV cache traffic): During the prefill, we write KV for every token and layer with Nbatch batch size: BytesKV,prefill = Nbatch · Nin · (2NL dmodel ) selt .
(22)
During the decode step k, we read the cached K and V of all Nin,k prior tokens and write the new token’s K and V: BytesKV,dec (k) = Nbatch (2NL dmodel ) selt (1 + Nin,k ). (23) Then, when generating Nout tokens steps, BytesKV,decode =
N out X
BytesKV,dec (k)
(24)
k=1
The corresponding memory-bound times are BytesKV,dec BytesKV,prefill , t̃mem,dec = . (25) t̃mem,prefill = Btot Btot c) Communication time (multi-GPU): Let tensor-parallel degree be pTP and pipeline-parallel stages be pPP with pTP pPP ≤ |Gj |. Tensor parallel. We abstract the intra-layer scheme that performs two all-reduces per layer as: Bytesact pTP − 1 t̃comm,TP ≈ 2NL · αTP log pTP + · , (26) βTP pTP where αTP and βTP are the interconnect latency and bandwidth, and the per-layer activation size is Bytesact = S Nbatch dmodel selt /pTP
(27)
where S = Nin for prefill and S = Nin,k for decode step k. x Pipeline parallel. The bubble factor for Nmic micro-batches: pPP − 1 , (28) ϕbubble = x Nmic + pPP − 1 which multiplies the dominant per-stage time. Thus, increasing x Nmic reduces the pipeline-bubble penalty, but it also increases the work assigned to GPU x in that step and can raise power, heat generation, and per-step latency. We collect all communication terms as t̃comm,prefill and t̃comm,dec . d) SLOs and total latency estimation: For each phase, we calculate the roofline maximum of compute and memory, and then add communication. With pipeline parallelism, multiply the dominant per-stage time by (1 + ϕbubble ). t̃prefill = max{t̃cmp,prefill , t̃mem,prefill } + t̃comm,prefill , t̃dec = max{t̃cmp,dec , t̃mem,dec } + t̃comm,dec .
(29)
The overall end-to-end latency could be estimated by: L̃tot = t̃prefill + t̃dec .
(30)
The SLO metrics of a job can be estimated by: ˜ ˜ TTFT ≈ δ0 + t̃prefill , TPOT = δ1 + t̃dec /Nout , (31) where δ0 and δ1 are calibrated offsets that represent queueing, dispatch, first KV allocation, and other runtime overheads.
8
V. ETCI NFER S OLUTION Our proposed solution, ETCInfer, comprises three core stages for energy optimization, detailed in Fig. 6. The solution contains three parts. First, we select a safe and promising temperature setpoint before each job. Second, we train an online adaptation model, ETCAdapter, that learns from logs and interactions with the simulator, informed by a learned world model. Third, we operate an online controller that adapts actions at every step using the belief over hidden states, while a safety layer enforces constraints. The solution uses standard belief-state and model-based control ideas, but adapts them to a joint control surface for LLM inference: ambient setpoint before execution, GPU frequency during execution, and micro-batch size at the serving layer. This design lets ETCInfer evaluate cooling energy, compute energy, SLO risk, and thermal safety in one control loop. The setpoint stage handles slow facility response, the online frequency and microbatch stage handles fast workload and thermal changes, and the safety layer prevents energy savings from violating latency or throttle constraints. A. Pre-job Temperature Setpoint Selection Before all jobs start, we first build a conservative starting policy π0 that already meets the SLO target and the thermal limit. This starting policy is independent of ETCInfer designs, which typically serve as the default scheduling strategy in the existing scheduler and ensure that the system can run within SLO constraints, regardless of whether ETCInfer is active or not. Next, we use it to construct a feasible set of setpoints and then select one that offers the potential for energy savings. Specifically, the starting policy π0 fixes device frequency and micro-batch size for each GPU x during a job. The starting x policy π0 records a safe and fixed frequency fsafe and balanced x micro-batch size Nsafe so that the latency satisfies SLOj and x the hardware temperature never exceed Tthr . For each candidate T ∈ [Tmin , Tmax ], iterate π0 with (13). Aggregate to the estimated job energy Ẽ(j, T |π0 ) =
I˜j (π0 )
X
wi P̃CRAC,i−1 (T ) +
i=0
X
x P̂i−1 ∆ti ,
(32)
x∈Gj
x and enforce the constraints Lj (π0 ) ≤ SLOj and T̂ix ≤ Tthr under the thermal dynamics. The feasible set is ST = {T ∈ [Tmin , Tmax ] : all constraints hold at risk level δ}. Thus, the candidate optimization space is filtered by the current job, hardware state, latency budget, and predicted thermal headroom before a setpoint is selected. We select ⋆ Tset = arg min Ẽ(j, T | π0 ). T ∈ST
(33)
After each job, the policy π0 is updated using the lowestenergy run, logging hardware specs, SLOj , and the setpoint x x Tset , and yielding updated fsafe and Nsafe . B. Training and Designs of ETCAdapter ETCAdapter is designed to learn an inner policy πin that adapts at run time to reduce the true energy E(j, Tset ) while
Fig. 6. ETCInfer three-stage energy optimization.
x keeping Lj (·) ≤ SLOj and T̄ix ≤ Tthr . A belief bi summarizes uncertainty over si and is updated by a filter
bi+1 = Upd(bi , ai , ōi+1 ), implemented as a Bayesian update. Learning process. We update πin via model-based RL in latent space. The world model gθ is trained on real and imagined data to capture thermal dynamics (13). Short-horizon action sequences are simulated to compute cumulative rewards with penalties on latency Lj and predicted temperature T̂ix . The policy is refined using advantage-weighted regression with x with a trust-region constraint, and a barrier ensures T̄ix < Tthr high probability. Model Module Designs Four neural modules implement the belief, the one-step predictor, the policy generator, and the critic. 1) Belief encoder ϕ: The belief encoder maps the history up to step i, denoted hi = {o1 : i , a1 : i−1 , Tset }, and outputs a compact latent state zi = ϕ(hi ) that summarizes the posterior belief over hidden system variables. This latent state conditions the one-step predictor gθ , the policy πin , and the critic V . The encoder is trained jointly with gθ to retain information relevant to temperature, power, and control. Training minimizes a reconstruction and prediction loss on the next observation, temperatures, and powers, with a temporal smoothness regularizer on the latent dynamics: h 1:X 1:X 2 − Ti+1 ∥2 Lenc = E ∥ôi+1 − oi+1 ∥22 + ∥T̂i+1 {z } | | {z } optional observation temperature i 1:X 1:X 2 + ∥P̂i+1 − Pi+1 ∥2 + α E ∥zi+1 − zi ∥22 . (34) | {z } power
When raw observations are unavailable, the first term is omitted, leaving only prediction and smoothness losses, encouraging zi to be a Markovian summary for control. 2) World model gθ as a one-step predictor: The world model gθ inputs the current latent zi , the control action ai , the setpoint Tset , and the step size ∆t. It predicts and outputs the 1:X next latent zi+1 , the temperatures T̂i+1 , and the GPU power 1:X P̂i+1 . When needed, it also produces the next observation ôi+1 . The cooling power PCRAC (Tset ) is computed by a separate analytic model and is not a prediction of gθ .
9
x . The policy is trained with advantageobserved power P̄i−1 weighted regression under a trust region, plus a small behaviorcloning term that pulls it toward high-quality short-plan actions evaluated in gθ : Lπ = − E vi log πin (ai | zi , oi , Tset ) + βKL KL πin ∥ πref + λbc E ∥ai − aplan ∥22 , (37) i
Algorithm 1: Online operation with energy- and latency-aware scoring ⋆ Input: Tset , policy πin , critic V , world model gθ , x safety residual S, thresholds {Tthr }, job SLOj . ⋆ Output: Realized job energy E(j, Tset ), logs for future refresh. ⋆ 1 Fix Tset = Tset and start the job; 2 for i = 0, 1, . . . until the job ends do 3 Measure oi , update bi , and compute zi = ϕ(bi ); 4 Generate action sequences guided by πin ; 5 Score each sequence with the return from r(bi , ai ; Tset ) and the constraint costs based on L̃tot and the predicted Tix ; 6 Select the first action from the best safe sequence and apply the safety residual to obtain ãi ; 7 Execute ãi , log (oi , ai ), and update the model and critic on a small batch; 8 if predicted latency threatens SLOj then 9 shrink the batch size limits and/or increase the frequency within safe bounds, or lower the switching penalty. x 10 if predicted or sensed temperature approaches Tthr then 11 reduce frequency and/or shrink micro-batch limits to lower heat generation before throttling.
where vi = exp(Ai /τ ) are critic-derived weights, πref is the previous policy for stable updates, and aplan are actions chosen i by the planner within the learned world model. 4) Critic V : The critic takes as input the latent state zi , the observable information oi , and the selected setpoint Tset . Its role is to provide a scalar estimate V (zi ) of the discounted return starting from step i. This value function supplies the advantage signal that is used to update the policy. Training of the critic follows a temporal difference scheme with value expansion on imagined rollouts produced by the world model gθ . For a rollout horizon H, we construct the target Gi =
and we optimize LV = E (V (zi ) − stop grad(Gi ))2 .
⋆ ) by (3) from the realized logs; Compute E(j, Tset ⋆ 13 return E(j, Tset ) and the new logs.
(39)
The critic employs short imagination inside gθ so that the value targets are consistent with the learned dynamics and with the energy accounting. In the policy update, we use the advantages Ai together with the weights vi obtained from the critic.
The model is trained with a simple objective that matches its predictions to measured values and keeps them consistent with the thermal physics. The learning objective is h 1:X 1:X 2 1:X 1:X 2 Lwm = E ∥T̂i+1 − Ti+1 ∥2 + ∥P̂i+1 − Pi+1 ∥2 i h i2 1:X 1:X + β ∥ôi+1 − oi+1 ∥22 + λphys E T̂i+1 − T̃i+1 . (35) 2
The first terms are standard prediction errors for temperature, power, and observations o. The physics term enforces agreement with the thermal model, which improves robustness at unseen operating points. To quantify epistemic uncertainty, we also fit an ensemble {gθ(k) }K k=1 on bootstrap splits. The variance across members guides risk-aware setpoint selection and provides signals for safety training. 3) Online Policy Generator πin : The policy generator inputs the current latent state zi , the observable information ōi , and the selected setpoint Tset , producing the control action ai . To align policy learning with the energy objective, we define the step reward as: X
γ h ri+h + γ H V (zi+H ), Ai = Gi − V (zi ). (38)
h=0
12
r = −(wi P̃CRAC,i−1 (Tset ) +
H−1 X
x P̃i−1 )∆t − λsw ∆a2i ,
(36)
C. Online end-to-end operation The goal of the online controller is to operate at the selected ⋆ and at every step adapt the action according to the setpoint Tset current belief, so that the energy is directly optimized and all modeled latency and temperature constraints remain satisfied. At time i, the system receives the observation oi and updates the belief bi together with the latent representation zi = ϕ(bi ). Based on this latent state, the controller generates several candidate action sequences. Each candidate is evaluated by its short return r(bi , ai ; Tset ) and by the constraint costs that are obtained from the predicted latency L̂b and the predicted temperatures Tix from (13). For the micro-batch component, x candidate actions with larger Nmic are preferred only when the predicted SLO slack and thermal headroom can absorb the x extra per-step load. Otherwise, the controller shrinks Nmic or raises frequency within safe bounds. The controller selects the action that is safe and has the best score, executes this action on the system, and records the transition. The logged samples are then used to refresh the fast learners online.
x∈Gj
where ∆a2i = ∥ai − ai−1 ∥22 represents the change in control actions, and λsw is the switching smoothness weight that controls how strongly the policy penalizes action changes. A larger λsw ensures smoother frequency switches and batching updates. Note that during inference, we always use the
VI. I MPLEMENTATION We implement ETCInfer through three components: ETCInfer Scheduler, ETCAdapter, and Ambient Setpoint Control. Together, they connect serving control, learning-based adaptation, and ambient-temperature actuation.
10
A. ETCInfer Scheduler Implementation ETCInfer uses common inference and cluster stacks: vLLM for serving, Kubernetes for cluster management, KServe as the inference interface, and NVIDIA DCGM with a Kubelet extension to expose hardware status and frequency control. ETCInfer acts as an external coordination controller: it reads SLO metadata and telemetry, writes safe control decisions, and leaves native execution to the serving engine and cluster manager. We released the ETCInfer source code.2 LLM service layer. We deploy vLLM for high-throughput multi-GPU inference. vLLM exposes a runtime control file that ETCInfer updates to select micro-batch size. We extend the KServe router to obtain SLO metadata from requests, enabling ETCInfer to set SLO targets while KServe preserves routing fairness. ETCInfer runs as a Kubernetes controller for Pod creation without modifying core scheduling. Hardware access and control. We integrate NVIDIA DCGM and a Kubelet plugin to collect per-GPU power, temperature, and utilization. A sidecar reports telemetry every control interval. GPU frequency is exposed through NVML functions. The plugin restricts changes to safe ranges and restores baseline settings after jobs finish or risks appear. Execution flow. For each job, ETCInfer selects placement and an ambient setpoint, then KServe starts vLLM. During execution, ETCInfer performs stepwise control: it ingests telemetry, updates the physics model, and adjusts micro-batch size and GPU frequency while enforcing SLO and thermal constraints. B. ETCAdapter Implementation ETCAdapter is implemented in Python using PyTorch and runs separately from ETCInfer, communicating through lightweight RPC to avoid interfering with vLLM. Its encoder, latent updater, world model, policy, and critic are separate torch.nn.Modules built from MLP blocks with linear layers, PReLU, and layer normalization. A 128-dimensional CPU latent state is maintained. Device and rack embeddings are attention-pooled. A 20k-transition replay buffer supports micro-batch training. We pretrain the encoder and world model on six hours of traces using Adam (10−3 , batch 256). Policy and critic use advantage-weighted regression (5 × 10−4 , batch 128). Online, ETCAdapter updates only final layers with one to two gradient steps every two seconds. C. Ambient Setpoint Control Implementation We combine CoolSIM CFD modeling with real cluster traces to evaluate airflow, cold supply, hot return, and recirculation under LLM loads. ETCInfer selects a continuous inlet setpoint, and CFD computes the resulting temperature field. The CFD solver is used for calibration and trace-driven ambient evaluation, not as a blocking component in the online vLLM request path. We also validate behavior on a small server by adjusting chassis inlet airflow with PT100-sensor feedback. 2 https://anonymous.4open.science/r/ETCInfer-2761/
TABLE II W ORKLOAD PROFILES AND SERVICE - LEVEL - OBJECTIVE THRESHOLDS
Inference Task Chat Dialog Trace Arrival Pattern SLO (TTFT) SLO (TPOT) SLO (E2E)
OASST1 Poisson ≤800 ms ≤50 ms ≤4 s
Code Assist
Summarization
PromptSet Bursty with spikes ≤1500 ms ≤60 ms ≤8 s
WebGPT Poisson ≤2000 ms ≤80 ms ≤12 s
VII. E VALUATION AND S IMULATION A. Evaluation Setup We evaluate ETCInfer using real-trace CFD simulation and validation experiments, covering energy, SLOs, and thermal safety across workloads and ambient conditions. The goal is to test whether ETCInfer improves efficiency under realistic facility and serving-stack constraints. Validation testbeds. We validate ETCInfer on a small-scale physical workstation with an Intel Core i9-13900K CPU, 128 GB RAM, four RTX3090 GPUs, and four RTX4090 GPUs (V 1, V 2), each with 24 GB VRAM. Models are quantized to fit memory. The tower chassis has front intake, rear exhaust, and 120 mm fans inside a 2 m × 2 m enclosure. ETCInfer selects continuous inlet setpoints from 18◦ C to 48◦ C. CFD simulator configuration. To capture airflow and recirculation, we model two 2 m × 2 m × 2.4 m rooms, R1 and R2. R1 has strong containment and uniform inlet temperatures, while R2 has partial containment, stronger recirculation, and higher inlet temperatures at the same setpoint. Each chassis includes 8 NVIDIA H100-class GPU heat sources and a 350 W Intel Xeon 8480 CPU. GPU heat follows measured board-power traces scaled to the H100 SXM power class [4]. We use open LLM prompt traces and a production cluster trace: OASST1 [46], PromptSet [47], WebGPT [48], and Alibaba2020 [49]. We derive arrivals, parameterize power/temperature dynamics, and replay setpoints from 18◦ C to 47◦ C. LLM inference models. We use Llama-3.1-Instruct [50] for Chat Dialog (CD) and search summarization (SUM), and Code-Llama-Instruct [51] for code assist (CA). These traces cover distinct token statistics, arrival patterns, and workloadspecific SLOs in Table II. Metrics. In our evaluation, we measure energy from three views: average job overall energy, average job computing energy, and average CRAC energy. All values are with respect to (wrt) the default vLLM setting under an 18◦ C ambient setpoint. For request r in workload class j, let ar be the arrival time at the inference frontend, τr,k be the emission time of the k-th output token, and Nrout be the number of generated tokens. We define the latency metrics as TTFTr = τr,1 − ar , τr,Nrout − τr,1 TPOTr = , Nrout − 1 Le2e r = τr,Nrout − ar .
Nrout > 1,
The workload class j specifies thresholds (θjTTFT , θjTPOT , θje2e ), as listed in Table II. Request r
vLLM-38 TAPAS
vLLM-48 DLLM
ETCInfer GLLM
90 80 70 60
CD
CA
SUM
LLM Inference Task
100
vLLM-28 DSO
vLLM-38 TAPAS
vLLM-48 DLLM
ETCInfer GLLM
90 80 70 60
CD
CA
SUM
LLM Inference Task
(a) R1
120 110 100 90 80 70 60
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
(b) R2
Overall Job Energy (%)
vLLM-28 DSO
Overall Job Energy (%)
100
Overall Job Energy (%)
Overall Job Energy (%)
11
120 110 100 90 80 70 60
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
(c) V1
(d) V2
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
120 110 100 90 80 70 60
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
(a) R1
120 110 100 90 80 70 60
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
(b) R2
Computing Job Energy (%)
vLLM-28 DSO
Computing Job Energy (%)
120 110 100 90 80 70 60
Computing Job Energy (%)
Computing Job Energy (%)
Fig. 7. Comparison of overall job energy consumption across four scenarios (R1, R2, V1, V2) and three tasks (CD, CA, SUM).
120 110 100 90 80 70 60
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
(c) V1
(d) V2
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task (a) R1
100 90 80 70 60 50 40
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task (b) R2
100 90 80 70 60 50 40
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
Shared CRAC Energy (%)
vLLM-28 DSO
Shared CRAC Energy (%)
100 90 80 70 60 50 40
Shared CRAC Energy (%)
Shared CRAC Energy (%)
Fig. 8. Comparison of computing energy consumption across four scenarios (R1, R2, V1, V2) and three tasks (CD, CA, SUM).
100 90 80 70 60 50 40
vLLM-28 DSO
vLLM-38 TAPAS
CD
vLLM-48 DLLM
CA
ETCInfer GLLM
SUM
LLM Inference Task
(c) V1
(d) V2
Fig. 9. Comparison of shared CRAC energy consumption across four scenarios (R1, R2, V1, V2) and three tasks (CD, CA, SUM).
satisfies the latency SLO only if TTFTr ≤ θjTTFT , ≤ θje2e . We measure serving TPOTr ≤ θjTPOT , and Le2e r performance through metric-specific SLO violation rates for TTFT, TPOT, and end-to-end latency. For a metric m, the violation rate is the fraction of requests whose measured value exceeds the corresponding threshold θjm . Thermal safety is reported as throttle exposure time. Baselines. We compare ETCInfer with recent LLM serving systems and with commonly used thermal and energy control strategies. Those baselines without specification would execute in the default temperature of 28◦ C. • vLLM Default is a throughput-first serving baseline that uses the original vLLM framework running under fixed ambient conditions {18◦ C, 28◦ C, 38◦ C, 48◦ C}, denoted as vllm18/28/38/48. It serves as the baseline with stable environment settings where cooling is independent from GPU scheduling. • DSO [52] is a GPU energy-efficiency optimizer that fuses static program information with runtime signals for DVFS decisions. In our comparison, it controls device-side frequency but does not coordinate with ambient setpoints or airflow models. • TAPAS [13] is a thermal- and power-aware scheduler. It predicts power and device temperature and allocates requests to a cooler rack to avoid thermal violations within SLOs. It focuses on device safety and power efficiency, but does not adjust ambient conditions or use airflow models.
• DLLM [53] is a cluster energy control approach using elastic reconfiguration and GPU frequency selection. The method adapts GPU assignment and DVFS to lower energy use while preserving service quality. Its control stays inside the cluster and does not incorporate facility-side thermal management. • GLLM [17] is an energy-aware pruning method for LLMs. We include it as a model-side efficiency baseline that lowers compute demand, but it does not perform runtime placement, GPU-frequency control, or ambient setpoint coordination. B. Evaluation Results 1) Improvement of Energy Efficiency: We compare ETCInfer with existing baselines across R1, R2, V1, and V2. Energy is normalized to vLLM-18 in each scenario, so Fig. 7–9 report relative job energy. In R1 and R2, cooling is a large share of total energy. Higher-ambient vLLM settings reduce CRAC power but increase compute energy because jobs run longer. Thus, total savings are limited, and long traces such as SUM can even consume more energy; the best vLLM saving is 18.2%. DSO, TAPAS, DLLM, and GLLM slightly reduce compute energy but leave cooling mostly unchanged, limiting gains to about 12.0%–20.1%. ETCInfer achieves the largest savings in R1 and R2 by coordinating device-side and facility-side actions. It lowers GPU power in non-critical phases, cutting compute energy by at least 12.5%, and raises ambient setpoints safely,
12
LLM Inference Task
LLM Inference Task
(a) R1
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (TTFT) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (TTFT) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (TTFT) (%)
SLO Violate Rate (TTFT) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
LLM Inference Task
(b) R2
LLM Inference Task
(c) V1
(d) V2
Fig. 10. Comparison of SLO violation rates for TTFT across four scenarios (R1, R2, V1, V2) and three tasks (CD, CA, SUM).
LLM Inference Task
LLM Inference Task
(a) R1
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (TPOT) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (TPOT) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (TPOT) (%)
SLO Violate Rate (TPOT) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
LLM Inference Task
(b) R2
LLM Inference Task
(c) V1
(d) V2
Fig. 11. Comparison of SLO violation rates for TPOT across four scenarios (R1, R2, V1, V2) and three tasks (CD, CA, SUM).
LLM Inference Task
LLM Inference Task
(a) R1
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 7.0 6.0 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (Total Time) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 7.0 6.0 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (Total Time) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 7.0 6.0 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
SLO Violate Rate (Total Time) (%)
SLO Violate Rate (Total Time) (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 7.0 6.0 5.0 4.0 3.0 2.0 1.0 0.0 CD CA SUM
LLM Inference Task
(b) R2
LLM Inference Task
(c) V1
(d) V2
Fig. 12. Comparison of SLO violation rates for overall end-to-end latency across four scenarios (R1, R2, V1, V2) and three tasks (CD, CA, SUM).
5.0 0.0
CD
CA
SUM
LLM Inference Task (a) R1
10.0 5.0 0.0
CD
CA
SUM
LLM Inference Task (b) R2
15.0 10.0 5.0 0.0
CD
CA
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 20.0
Throttle Exposure Time Rate (%)
10.0
15.0
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 20.0
Throttle Exposure Time Rate (%)
15.0
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 20.0
Throttle Exposure Time Rate (%)
Throttle Exposure Time Rate (%)
vLLM-18 vLLM-28 vLLM-38 vLLM-48 DSO ETCInfer GLLM TAPAS DLLM 20.0
SUM
LLM Inference Task (c) V1
15.0 10.0 5.0 0.0
CD
CA
SUM
LLM Inference Task (d) V2
Fig. 13. Comparison of throttle exposure time across four scenarios (R1, R2, V1, V2) and three tasks (CD, CA, SUM).
reducing cooling energy by 49.9%. Together, these effects save up to 33.1% total energy in R2. In V1 and V2, GPU power dominates, so ambient-only control is weaker and vLLM28/48 can exceed vLLM-18. ETCInfer still improves compute efficiency and preserves cooling savings. 2) Improvement of SLO violation rate: We evaluate TTFT, TPOT, and end-to-end SLO violations in Fig. 10–12; lower is better. vLLM-18 keeps violations low, with 0.3% total violations on CD and at most 0.5% on long SUM traces. Raising ambient without thermal awareness quickly worsens latency. In R1 and R2, vLLM-28/38 often reach 0.6%–3.8% violations, while vLLM-48 exceeds 7.9% in the worst R2 case. The effect is stronger in V1 and V2, where long requests already push GPUs near compute and thermal limits. DSO reduces power but slows decoding, causing about 3× the
vLLM-18 violation rate in R2 and V2. TAPAS, DLLM, and GLLM improve over vLLM-38/48, usually staying within 0.6%–2.0%, but do not recover vLLM-18 quality. ETCInfer keeps violations close to vLLM-18 in all scenarios: totals stay within 0.2% of vLLM-18 in R1/R2 and below 0.7% at the highest setpoint. 3) Improvement of Hardware Thermal Safety: We measure thermal safety by throttle exposure time in Fig. 13, i.e., the fraction of execution above a high-temperature threshold. Higher exposure indicates sustained thermal stress and greater aging or shutdown risk. Higher-ambient vLLM settings sharply increase exposure, especially in R2 and V2 and on the long SUM trace. DSO suppresses spikes but extends runtime, so heat still accumulates. TAPAS, DLLM, and GLLM reduce exposure but remain riskier than ETCInfer. In the worst
(a) Overall Energy
R2 V2 Testbed Scenario
ETCInfer Freq_only Static 10.0 8.0 6.0 4.0 2.0 0.0
Setpoint_only Rule-based
R2 V2 Testbed Scenario
(b) SLO Violation Rate (c) Throttle Expose Time
DS
NTF
RS
SS
80 60 40 20 0
R2 V2 Testbed Scenario
(a) Overall Energy
5.0
DS
NTF
RS
SS
4.0 3.0 2.0 1.0 0.0
R2 V2 Testbed Scenario
8.0
DS
NTF
RS
SS
R2 V2 Testbed Scenario
(a) Overall Energy
5.0 4.0 3.0 2.0 1.0 0.0
Safe Fixed-28C
Fixed-38C Max_Temp
R2 V2 Testbed Scenario
Safe Fixed-28C 6.0
Fixed-38C Max_Temp
4.5 3.0 1.5 0.0
R2 V2 Testbed Scenario
(b) SLO Violation Rate (c) Throttle Expose Time
TABLE IV L ATENCY- ESTIMATOR PREDICTION ERRORS ON WORKLOAD TRACES .
4.0
Workload
2.0
CD CA SUM
0.0
R2 V2 Testbed Scenario
(b) SLO Violation Rate (c) Throttle Expose Time
TABLE III VALIDATION COMPARISON WITH MODEL - BASED CONTROL ALTERNATIVES . Avg. energy saving
Avg. SLO violation
Avg. throttle-exposure reduction
18.6% 22.1% 25.4% 27.6% 31.8%
1.84% 1.26% 0.92% 0.74% 0.48%
54.3% 63.5% 72.8% 78.9% 89.7%
Greedy threshold Deterministic optimization MPC Robust MPC+filter ETCInfer
Fixed-38C Max_Temp
6.0
Fig. 16. Effects of SLO and thermal-safety knobs.
Control method
Safe Fixed-28C
Fig. 15. Effects of the ambient setpoint setup.
Throttle Exposure Time Rate (%)
100
SLO Violation Rate (%)
Average Energy (%)
Fig. 14. Effects of joint control.
100 80 60 40 20 0
Throttle Exposure Time Rate (%)
Setpoint_only Rule-based
SLO Violation Rate (%)
R2 V2 Testbed Scenario
ETCInfer Freq_only Static 5.0 4.0 3.0 2.0 1.0 0.0
Average Energy (%)
Setpoint_only Rule-based
Throttle Exposure Time Rate (%)
ETCInfer Freq_only Static 100 80 60 40 20 0
SLO Violation Rate (%)
Average Energy (%)
13
case, V2 with vLLM-48 on SUM spends 17.1% of execution above the threshold, while ETCInfer reduces this to 1.2%, a 92.9% drop. ETCInfer consistently achieves the lowest throttle exposure below 1.6% in R2 and V2. C. Ablation Study We analyze ETCInfer components to measure the effect of each design choice. All ablations use R2 and V2, averaged over three traces. 1) Effectiveness of Joint and Adaptive Control: We evaluate joint ambient-setpoint, workload-scheduling, and onlineadaptation control in Fig. 14. ETCInfer is compared with Setpoint only, Freq only, Rule-based, and Static. Full ETCInfer has the lowest total energy by reducing compute and CRAC costs. In R2, Setpoint only leaves energy 7.9% higher and increases SLO violations by 40.0%. In V2, Freq only is 6.9% worse and increases throttling. Rule-based performs between single-actuator baselines and Static, while Static still uses 3.7% more energy, showing the online adaptation benefits. 2) Impact of Ambient Setpoint Setup: We compare four pre-job setpoint policies in R2 and V2: Safe, fixed 28◦ C, fixed 38◦ C, and Max Temp. Fig. 15 reports total, compute, and CRAC energy, SLO violations, and throttling. Safe nearly minimizes energy while keeping violations and throttling low. 2 In (a), gray stripes denote computing energy (bottom) and CRAC energy (top); in (b), they denote TTFT and TPOT. This notation also applies to Figs. 15–17.
TTFT MAPE
TPOT MAPE
E2E MAPE
1.7% 2.2% 2.4%
2.1% 2.5% 2.8%
1.5% 1.9% 2.2%
Max Temp saves only 2.6% and 2.3% total energy in R2 and V2, but raises SLO violations by 1.2× and throttle exposure by over 65.7%. Fixed-28C improves reliability, reducing SLO violations by 12.0% and throttling by 16.7% in R2, but increases total and CRAC energy by 7.9% and 6.9%. Fixed38C saves only 1.1–1.3% energy while raising violations by 59.8% and throttling by 40.5–63.7%. Thus, fixed setpoints cannot balance energy and reliability. 3) Sensitivity to SLO and Thermal Safety Knobs: We compare default safety (DS), strict safety (SS), relaxed safety (RS), and no ttt feature (NTF) in Fig. 16. SS tightens SLO budgets by 10% and lowers the thermal limit by 3◦ C; RS expands SLOs by 20% and raises it by 3◦ C. DS balances energy, violations, and throttling. SS uses 4.9% more energy but reduces violations by 26.5%. RS saves 3.7% energy but increases violations and near-limit time. NTF worsens all metrics, confirming the importance of time-to-throttle features. 4) Comparison with Model-based Control Alternatives: Table III compares ETCInfer with Greedy threshold, Deterministic optimization, MPC, and Robust MPC+filter. All share telemetry, actions, SLOs, and R2/V2 traces; results are averaged across scenarios and traces. Model-based baselines improve over greedy thresholding, but ETCInfer gives the best energy–SLO–thermal balance. Deterministic optimization is brittle under state-estimation error. MPC improves stability but degrades when hidden thermal states or token latency drift from the calibrated model. Robust MPC+filter reduces violations but conservative. ETCInfer performs better as its belief encoder and learned world model adapt from observation/action history, while the safety layer rejects risky actions. 5) Sensitivity to Inference Model Family: We test model sensitivity on Chat Dialog in Table V, using locally quantized Qwen2.5-Instruct [54], DeepSeek-R1-Distill-Qwen [55], and Mistral-Instruct [56]. Workload traces, SLOs, thermal scenarios, telemetry, and scheduler code are unchanged; only model architecture and token latency differ. ETCInfer preserves the Llama-family trend: it reduces energy and throttling while keeping SLO violations below 0.7%. Thus, its gains depend on telemetry, token-latency estimates, and thermal response,
14
TABLE V S ENSITIVITY TO INFERENCE MODEL FAMILY ON C HAT D IALOG . Inference model family
Energy saving
SLO violation
Throttle reduction
Qwen2.5-Instruct DeepSeek-R1-Distill-Qwen Mistral-Instruct
30.4% 28.7% 26.9%
0.43% 0.51% 0.47%
88.6% 85.2% 82.4%
not one checkpoint. 6) Accuracy of Prediction Telemetry: We validate thermal and power predictors against measured traces. For R2 and V2, die-temperature RMSE stays below 2.2◦ C, inlet-temperature RMSE below 1.2◦ C, power RMSE below 4.1%, and timeto-throttle error within 5.5% for most jobs. We validate the calibrated prefill/decode latency estimator on held-out workload traces. Table IV reports mean absolute percentage error (MAPE) for TTFT, TPOT, and end-to-end latency. The largest MAPE is 2.8%, indicating that the estimator is accurate enough to reject actions that threaten SLO constraints. Consistently, Figs. 10–12 show that total SLO violations stay below 0.7% under high-ambient operation when the estimator guides scheduling decisions. VIII. R ELATED W ORKS LLM inference scheduling spans SLO-centric, energy/carbon-aware, and thermal-aware methods, which together motivate coordinated compute–thermal–cooling control. SLO-centric Scheduling. SLOs specify latency and availability targets tied to user experience and pricing. SLOaware schedulers manage admission, batching, and placement using in-flight batching, speculative decoding, and KV-cache optimization [31]. Examples include iteration-level scheduling in Orca, prefill/decode disaggregation in DistServe, and chunked-prefill scheduling in Sarathi-Serve [32]–[34]. Production systems follow the same goal: vLLM improves throughput with PagedAttention, and TensorRT-LLM improves utilization through kernel fusion and quantization [29], [30]. These systems sustain throughput and tail latency, but usually treat GPUs as thermally stable and attribute slowdowns to queueing or compute contention. ETCInfer instead uses temperature trends, thermal headroom, and time-to-throttle to guide batching and DVFS under thermal variability. Energy and Carbon-centric Scheduling. Another line of work optimizes energy or emissions under latency/deadline constraints [22], [57]. Examples include data-center rightsizing [58], geo-distributed scheduling with grid coordination [59], low-carbon shifting of delay-tolerant jobs [60], adaptive DVFS and power capping [61], carbon-aware load balancing [62], locality-aware data-operator scheduling [63], and HPC energy plugins [64]. These methods reduce energy or emissions while bounding SLO violations, but often assume cooling reacts independently or quickly to power changes. ETCInfer models coupled compute–cooling dynamics and jointly controls ambient setpoint, DVFS, and batching. Thermal-aware Scheduling. Thermal-aware scheduling prevents heat-induced performance loss. Prior work shows sustained DNN workloads can overheat GPUs, motivating heuris-
tic and RL schedulers that switch GPU/NPU execution by thermal state [65]. Automotive SoC studies address CPU–GPU thermal coupling through balanced assignment, co-scheduling, thermal-server abstractions, and utilization bounds [66], [67]. Thermal constraints also affect data centers: classic studies incorporate cooling cost and spatial thermal effects into placement and consolidation [42], [43]. Related work studies 3D-stacked LLM memory scheduling [68], coordinated cooling and compute management for AI datacenters [69], and thermal modeling [70] and energy-saving techniques for cloud datacenters [44]. TAPAS uses historical thermal signals for VM placement and routing [13]. TAWS integrates voltage/frequency behavior into batching and placement under higher ambient-temperature standards [71]. ETCInfer further couples ambient setpoints, GPU DVFS, and micro-batch sizing to reduce energy while preserving TTFT, TPOT, and thermal safety. IX. C ONCLUSION In this paper, we present ETCInfer, a thermal-aware scheduler that jointly manages ambient temperature, GPUs, and micro-batch sizing to minimize inference job energy under thermal and latency constraints. By modeling coupled compute–cooling dynamics and leveraging learning-based control, ETCInfer adaptively selects safe setpoints and schedules in response to workload and thermal dynamics. Extensive experiments across real-trace simulations show that ETCInfer significantly reduces total energy, thermal throttling, and SLO violations compared to SOTA baselines. Our results emphasize the need for coordinated control in future AI datacenters and highlight the potential of physics-informed learning in achieving energy-efficient, thermally safe LLM inference. R EFERENCES [1] Z. Jiang, H. Lin, Y. Zhong, et al., “Megascale: Scaling large language model training to more than 10,000 gpus,” in Proc. of the 21st USENIX Symposium on Networked Systems Design and Implementation, 2024. [2] P. Patel, E. Choukse, C. Zhang, et al., “Characterizing power management opportunities for llms in the cloud,” in Proc. of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, 2024. [3] B. Smith, “A golden opportunity for american ai.” Microsoft On the Issues, 2025. [4] NVIDIA Corporation, “NVIDIA H100 Tensor Core GPU.” https://www. nvidia.com/en-us/data-center/h100/, 2024. Official product specifications. H100 SXM maximum thermal design power is listed as 700 W. [5] L. Su, K. Dong, Q. Sun, et al., “Research progress on energy saving of data center cooling system,” Advances in New and Renewable Energy, vol. 7, no. 1, pp. 95–106, 2019. [6] R. Wang, Z. Cao, X. Zhou, et al., “Green data center cooling control via physics-guided safe reinforcement learning,” in Proc. of the 15th ACM International Conference on Future Energy Systems, 2024. [7] Y. Zhang, H. Li, and S. Wang, “The global energy impact of raising the space temperature for high-temperature data centers,” Cell Reports Physical Science, vol. 4, no. 10, 2023. [8] ASHRAE Technical Committee 9.9, Thermal Guidelines for Data Processing Environments. ASHRAE Datacom Series, Peachtree Corners, GA: American Society of Heating, Refrigerating and Air-Conditioning Engineers (ASHRAE), fifth edition, revised and expanded ed., 2021. [9] M. Acton et al., “2024 best practice guidelines for the eu code of conduct on data centre energy efficiency,” JRC Technical Note JRC136986, European Commission, Joint Research Centre, Ispra, Italy, 2024. European Code of Conduct for Energy Efficiency in Data Centres, 15th Edition.
15
[10] M. Acton, J. Booth, and D. Paci, “Best practice guidelines for the eu code of conduct on data centre energy efficiency,” JRC Technical Report EUR40267 EN EUR40267, Publications Office of the European Union, Joint Research Centre, European Commission, Mar. 2025. [11] Singapore Standards Council, “Singapore Standard SS 697:2023: Deployment and Operation of Data Centre IT Equipment under Tropical Climate.” Standard, June 2023. [12] S. Kim, K. Bin, et al., “zTT: Learning-based DVFS with zero thermal throttling for mobile devices,” in Proc. of ACM International Conference on Mobile Systems, Applications, and Services, 2021. [13] J. Stojkovic, C. Zhang, Í. Goiri, et al., “Tapas: Thermal-and poweraware scheduling for llm inference in cloud platforms,” in Proc. of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, 2025. [14] Z. Lai, K. T. Lam, C.-L. Wang, et al., “Latency-aware dynamic voltage and frequency scaling on many-core architectures for data-intensive applications,” in Proc. of 2013 International Conference on Cloud Computing and Big Data, pp. 78–83, IEEE, 2013. [15] G. Xiao, J. Lin, M. Seznec, et al., “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International conference on machine learning, pp. 38087–38099, PMLR, 2023. [16] R. Jin, J. Du, W. Huang, et al., “A comprehensive evaluation of quantization strategies for large language models,” in Findings of the Association for Computational Linguistics, pp. 12186–12215, 2024. [17] C. Tian, X. Qin, and L. Li, “Greenllm: Towards efficient large language model via energy-aware pruning,” in Proc. of the 32nd International Symposium on Quality of Service, pp. 1–2, IEEE, 2024. [18] D. Zhang, H. Xia, X. Wang, et al., “Thermal elasticity-aware host resource provision for carbon efficiency on virtualized servers,” IEEE Transactions on Computers, 2025. [19] V. A. Chhabria and S. S. Sapatnekar, “Impact of self-heating on performance and reliability in finfet and gaafet designs,” in International Symposium on Quality Electronic Design, pp. 235–240, IEEE, 2019. [20] G. Ostrouchov, D. Maxwell, R. A. Ashraf, C. Engelmann, M. Shankar, and J. H. Rogers, “Gpu lifetimes on titan supercomputer: Survival analysis and reliability,” in International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–14, IEEE, 2020. [21] J. Guerreiro, A. Ilic, N. Roma, et al., “Modeling and decoupling the gpu power consumption for cross-domain dvfs,” IEEE Transactions on Parallel and Distributed Systems, vol. 30, no. 11, pp. 2494–2506, 2019. [22] Q. Wang, X. Mei, H. Liu, Y.-W. Leung, Z. Li, and X. Chu, “Energyaware non-preemptive task scheduling with deadline constraint in dvfsenabled heterogeneous clusters,” IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 12, pp. 4083–4099, 2022. [23] S. M. Nabavinejad, S. Reda, and M. Ebrahimi, “Coordinated batching and dvfs for dnn inference on gpu accelerators,” IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 10, pp. 2496–2508, 2022. [24] S. Hong and H. Kim, “An integrated gpu power and performance model,” in Proceedings of the 37th Annual International Symposium on Computer Architecture, pp. 280–289, 2010. [25] J. Leng, T. Hetherington, A. ElTantawy, S. Gilani, N. S. Kim, T. M. Aamodt, and V. J. Reddi, “GPUWattch: Enabling energy optimizations in GPGPUs,” in Proceedings of the 40th Annual International Symposium on Computer Architecture, pp. 487–498, 2013. [26] M. Zapater, O. Tuncer, et al., “Leakage-aware cooling management for improving server energy efficiency,” IEEE Transactions on Parallel and Distributed Systems, vol. 26, no. 10, pp. 2764–2777, 2015. [27] J. Yao et al., “Adaptive power management through thermal aware workload balancing in internet datacenters,” IEEE Transactions on Parallel and Distributed Systems, vol. 26, no. 9, pp. 2400–2409, 2015. [28] I. Riu, D. Smiley, S. Bessasparis, et al., “Load growth is here to stay, but are data centers: Strategically managing the challenges and opportunities of load growth,” white paper, Energy & Environmental Economics, 2024. [29] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with PagedAttention,” in Proc. of the ACM SIGOPS 29th Symposium on Operating Systems Principles, pp. 611–626, 2023. [30] NVIDIA Corporation, “TensorRT-LLM.” https://github.com/NVIDIA/ TensorRT-LLM, 2025. [31] J. Jiang, Y. Chen, Z. Zhang, B. He, P. Luo, M. Lu, Y. Chen, H. Zhang, J. Du, D. Huang, and Y. Lu, “Efficient kv cache spillover management on memory-constrained gpu for llm inference,” IEEE Transactions on Parallel and Distributed Systems, vol. 37, no. 1, pp. 90–105, 2025. [32] G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for Transformer-Based generative models,” in
16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), (Carlsbad, CA), pp. 521–538, USENIX Association, July 2022. [33] Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “DistServe: Disaggregating prefill and decoding for goodputoptimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24), (Santa Clara, CA), pp. 193–210, USENIX Association, July 2024. [34] A. Agrawal, N. Kedia, A. Panwar, J. Mohan, N. Kwatra, B. Gulavani, A. Tumanov, and R. Ramjee, “Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24), (Santa Clara, CA), pp. 117–134, USENIX Association, July 2024. [35] L. P. Kaelbling, M. L. Littman, and A. R. Cassandra, “Planning and acting in partially observable stochastic domains,” Artificial Intelligence, vol. 101, no. 1–2, pp. 99–134, 1998. [36] S. Boyd and L. Vandenberghe, Convex Optimization. Cambridge University Press, 2004. [37] J. B. Rawlings, D. Q. Mayne, and M. M. Diehl, Model Predictive Control: Theory, Computation, and Design. Nob Hill Publishing, 2nd ed., 2017. [38] A. Bemporad and M. Morari, “Robust model predictive control: A survey,” in Robustness in Identification and Control (A. Garulli, A. Tesi, and A. Vicino, eds.), vol. 245 of Lecture Notes in Control and Information Sciences, pp. 207–226, Springer, 1999. [39] D. Q. Mayne, M. M. Seron, and S. V. Raković, “Robust model predictive control of constrained linear systems with bounded disturbances,” Automatica, vol. 41, no. 2, pp. 219–224, 2005. [40] R. E. Kalman, “A new approach to linear filtering and prediction problems,” Journal of Basic Engineering, vol. 82, no. 1, pp. 35–45, 1960. [41] V. Kandiah, S. Peverelle, M. Khairy, et al., “Accelwattch: A power modeling framework for modern gpus,” in Proc. of the IEEE/ACM International Symposium on Microarchitecture, 2021. [42] J. Moore, J. Chase, P. Ranganathan, and R. Sharma, “Making scheduling “cool”: Temperature-Aware workload placement in data centers,” in 2005 USENIX Annual Technical Conference (USENIX ATC 05), (Anaheim, CA), USENIX Association, Apr. 2005. [43] E. Pakbaznia and M. Pedram, “Minimizing data center cooling and server power costs,” in Proceedings of the 2009 ACM/IEEE International Symposium on Low Power Electronics and Design, pp. 145–150, 2009. [44] J. Lin et al., “Thermal modeling and thermal-aware energy saving methods for cloud data centers: A review,” IEEE Transactions on Sustainable Computing, vol. 9, no. 3, pp. 571–590, 2023. [45] S. Williams, A. Waterman, and D. Patterson, “Roofline: An insightful visual performance model for multicore architectures,” Communications of the ACM, vol. 52, no. 4, pp. 65–76, 2009. [46] A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z. R. Tam, K. Stevens, A. Barhoum, D. Nguyen, O. Stanley, R. Nagyfi, et al., “OpenAssistant Conversations: Democratizing large language model alignment,” in Advances in Neural Information Processing Systems, 2023. [47] K. Pister, D. J. Paul, P. Brophy, and I. Joshi, “PromptSet: A programmer’s prompting dataset,” 2024. [48] R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al., “WebGPT: Browser-assisted question-answering with human feedback,” 2021. [49] Alibaba Group, “Alibaba cluster trace 2020: Gpu and heterogeneous cluster telemetry.” https://github.com/alibaba/clusterdata, 2020. [50] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al., “The Llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. [51] B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, et al., “Code Llama: Open foundation models for code,” arXiv preprint arXiv:2308.12950, 2023. [52] Q. Wang, L. Li, W. Luo, et al., “Dso: A gpu energy efficiency optimizer by fusing dynamic and static information,” in Proc. of 32nd International Symposium on Quality of Service, pp. 1–6, IEEE/ACM, 2024. [53] J. Stojkovic, C. Z, Í. Goiri, et al., “Dynamollm: Designing llm inference clusters for performance and energy efficiency,” in IEEE International Symposium on High Performance Computer Architecture, 2025. [54] Qwen Team, “Qwen2.5 technical report,” arXiv preprint arXiv:2412.15115, 2024. [55] DeepSeek-AI, “DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning,” arXiv preprint arXiv:2501.12948, 2025.
16
[56] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al., “Mistral 7b,” arXiv preprint arXiv:2310.06825, 2023. [57] D. Wang, B. Liu, R. Lu, Z. Zhang, and S. Zhu, “StoreLLM: Energy efficient large language model inference with permanently pre-stored attention matrices,” in Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems, pp. 398–406, 2025. [58] S. Albers and J. Quedenfeld, “Algorithms for right-sizing heterogeneous data centers,” in Proc. of the 33rd ACM Symposium on Parallelism in Algorithms and Architectures, pp. 48–58, 2021. [59] H. H, Y. W, et al., “Coordinating workload scheduling of geo-distributed data centers and electricity generation of smart grid,” IEEE Transactions on Services Computing, vol. 13, no. 6, pp. 1007–1020, 2017. [60] P. Wiesner, I. Behnke, D. Scheinert, K. Gontarska, and L. Thamsen, “Let’s wait awhile: How temporal workload shifting can reduce carbon emissions in the cloud,” in Proc. of the 22nd International Middleware Conference, pp. 260–272, 2021. [61] Y. Sun, Z. Ding, et al., “Learning-enabled adaptive power capping scheme for cloud data centers,” IEEE Transactions on Smart Grid, vol. 16, no. 6, pp. 4755–4767, 2025. [62] W.-T. Lin, G. Chen, and H. Li, “Carbon-aware load balance control of data centers with renewable generations,” IEEE Transactions on Cloud Computing, vol. 11, no. 2, pp. 1111–1121, 2022. [63] L. Cheng, Y. Wang, Q. Liu, et al., “Network-aware locality scheduling for distributed data operators in data centers,” IEEE Transactions on Parallel and Distributed Systems, vol. 32, no. 6, pp. 1494–1510, 2021. [64] A. Aaen S, M. Albano, and S. Xavier, “Automatic energy-efficient job scheduling in hpc: A novel slurm plugin approach,” in Proc. of the SC’23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, pp. 1831–1838, 2023. [65] T. Tan and G. Cao, “Thermal-aware scheduling for deep learning on mobile devices with npu,” IEEE Transactions on Mobile Computing, vol. 23, no. 12, pp. 10706–10719, 2024. [66] Y. Lee, K. G. Shin, and H. S. Chwa, “Thermal-aware scheduling for integrated cpus–gpu platforms,” ACM Transactions on Embedded Computing Systems, vol. 18, no. 5, pp. 1–25, 2019. [67] Y. Lee, “Thermal-aware design and management of embedded real-time systems,” in 2021 Design, Automation & Test in Europe Conference & Exhibition, pp. 1252–1255, 2021. [68] S. He, P. Yan, et al., “Tasa: Thermal-aware 3d-stacked architecture design with bandwidth sharing for llm inference,” arXiv preprint arXiv:2508.07252, 2025. [69] N. B. Abera and Y. Chen, “Coordinated cooling and compute management for ai datacenters,” arXiv preprint arXiv:2601.08113, 2026. [70] H. Zhang, S. Zhu, R. Lu, and D. Wang, “HotGPU: A thermal profile dataset for immersion-cooling ai datacenters,” in Proceedings of the 5th Workshop on Sustainable Computer Systems, 2026. [71] R. Lu and D. Wang, “A thermal-aware workload scheduler for highperformance llm inference in cooling-regulated datacenters,” ACM SIGENERGY Energy Informatics, vol. 5, no. 2, pp. 98–104, 2025.