arXiv:2609.14775v1 [cs.DC] 13 Sep 2026
CATS: A Carbon-Aware Task Simulator for Reducing AI Data Center Emissions Dayuan Chen
Ziliang Zong
Computer Science Department Texas State University San Marcos, TX, USA [email protected]
Computer Science Department Texas State University San Marcos, TX, USA [email protected]
Abstract—The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper. First, we compile a global alignment dataset unifying 140 operational and planned cloud regions across 8 major providers with fiveminute carbon-intensity traces for 145 grid regions from 2022 to 2024, revealing that 50% of current sites lie in medium-to-high carbon-intensity grids, indicating a siting-carbon mismatch and unrealized carbon reduction potential. Second, we develop CATS (Carbon-Aware Task Simulator), a flexible trace-driven framework that profiles six AI inference tasks across multiple GPU types, synthesizes realistic diurnal curve, geographical and task mixes, and SLA constraints, and evaluates spatial and temporal schedulers against two baselines while reporting comprehensive metrics including carbon emissions, energy consumption, runtime, queue delay, and hardware utilization. Third, we quantify achievable CO2 savings under realistic constraints: in a 24-hour trace with 600,000 tasks at fleet utilization of 0.37, spatial shifting reduces CO2 by 38.4% versus speed-first baseline, while temporal shifting yields 16% savings with bounded SLA violations at 3.27%. These results advocate locating future data centers in low carbon-intensity grids and demonstrate that carbon-aware scheduling on today’s fleets can achieve substantial operational emissions reduction. Index Terms—Carbon Awareness, Sustainability, cloud, workload shifting
I. I NTRODUCTION Cloud computing has enabled AI applications to scale to trillion-query services within a few years, driving rapid adoption of GPU- and TPU-accelerated data centers (DCs). In the United States alone, DC electricity use reached 176 TWh (4.4% of national load) in 2023 and is expected to double or triple by 2028 as AI services multiply [1]–[3], with similar trajectories in Europe and Asia [2]–[4]. As GPUs and TPUs draw an order of magnitude more power than traditional CPUpowered servers, operational emissions from AI inference have become a pressing concern that cannot be ignored [5]–[7]. While a single AI inference requires far less computing and energy than model training, the volume of inference requests is massive. OpenAI revealed in early 2024 that about 100 billion words were generated every day. With an inference cost of $30 per million tokens, their annual inference cost exceeds
the estimated training cost of $200 million1 . This indicates that AI inference accounts for a major share of data center workload, energy use, and carbon emissions [8]. Operational emissions of an inference task are the product of its energy consumption and the carbon-intensity (CI) of the grid at the location and time where computation occurs. Two approaches exist to reduce these emissions: decreasing energy consumption or lowering carbon-intensity [9], [10]. Optimizing energy consumption includes improving the energy efficiency of all IT hardware and supporting infrastructure. Reducing CI can be achieved by strategically choosing locations in cleaner grid regions or through shifting tasks spatially and temporally to windows with lower grid emissions. This work focuses on CI reduction approaches for AI inference tasks. While low-CI data center siting (i.e. placing DCs in regions with low average grid CI) has been studied in academic literature [11], the methodology remains largely theoretical. Industry DC providers often emphasize renewable energy procurement, but rarely disclose the actual CI of the grids supplying their DCs [12]–[17]. Thus, a comprehensive, globalscale study on how existing data centers align with clean grids is essential but remains absent. Workload shifting has also been presented, with several research proposing carbon-aware spatial and temporal scheduling across DCs [18]–[21]. However, current approaches have several limitations. First, candidate data centers are often restricted to a single provider or limited to grids where CI data is available. Hardware capacity constraints are often overlooked, resulting in simplistic or absent modeling of queuing delay. Furthermore, many works do not consider realistic workload dynamics and arrival patterns. As a result, a comprehensive, flexible simulator for carbon-aware workload shifting that can model realistic hardware constraints, queuing behavior, and arrival patterns remains to be developed. To address these gaps, we analyze the relationship between global grid carbon-intensity and existing as well as proposed data center locations, revealing how current deployments align with low-CI grids, and offering guidance for future carbonaware site selection. In addition, we present CATS, a Carbon1 Inference cost for output only for gpt-4-0125-preview model, assuming 75 tokens per 100 words.
Aware Task Simulator with plug-in shifting schedulers that supports any cloud providers, incorporates global grid carbonintensity data, and enables flexible modeling of GPU types and hardware capacity constraints while respecting service level agreements (SLAs). CATS considers queuing delay, task mixes, and arrival patterns, and provides comprehensive reporting of key operational and sustainability metrics (e.g. carbon emissions, energy consumption, runtime, queuing delay, and GPU utilization) to enable detailed evaluation of carbon-aware task scheduling policies under realistic operational conditions. Our paper makes three major contributions: 1) We compile a multi-cloud siting-carbon alignment dataset by joining 140 cloud regions with five-minute marginal carbon-intensity traces for 145 grid regions from 2022 to 2024. We show that 50% of existing sites reside in medium to high carbon-intensity grids, indicating a siting-carbon mismatch and highlighting opportunities for emission reduction. 2) We develop the Carbon-Aware Task Simulator (CATS), which supports various AI inference tasks, multi-GPU profiling (runtime/energy), diurnal and geographicallyskewed arrivals with task mix and SLAs. CATS is a discrete-event simulator that replays traces against two baselines (Speed-First and Carbon-First) and two carbonaware policies (Spatial Shifting and Temporal Shifting). 3) We demonstrate the effectiveness of CATS by quantifying the CO2 savings and performance trade-offs under realistic constraints. In a 24-hour, 600,000-task trace at fleet utilization near 0.37, spatial shifting reduces CO2 by 38.4% relative to Speed-First, and temporal shifting reduces CO2 by 16.1% relative to Speed-First. The reset of the paper is organized as follows. Section II reviews related work on carbon-aware data center siting and design, carbon-intensity signals, sustainable AI, and task scheduling. Section III details our data sources and preprocessing pipeline and presents the siting-carbon alignment analysis across the U.S., EU (incl. UK), and Australia. Section IV details the CATS architecture and design. Section V provides experiments and evaluation on realistic traces across four schedulers. Section V concludes our study. II. R ELATED W ORKS A. Carbon-Aware Data Center siting and Design Decisions of where to build data centers played a crucial role in determining their carbon footprints. A growing body of work studied where to build and how to design data centers to lower operational emissions. Ayyildiz proposed a multi-criteria decision framework that explicitly scored regions by renewable availability and grid factors to guide the siting of data centers [11]; Similarly, Wang developed a comprehensive approach to optimize data center carbon emissions through siting and operational configuration [22]; Al-Ayyoub modeled growth decisions under energy and cost constraints [23]. Other studies proposed integrated optimization of data center configuration
and capacity shaping to minimize carbon output. Acun et al. integrated workload, grid signals and hardware choices to plan and operate data centers that maximized renewable energy usage and minimized carbon impact [24]; Lin showed that dynamically adapting data center power usage and regional load shifting could raise renewable utilization [25]; McMullen quantified and compared environmental improvements on site renewables for data centers [26]. In industry, cloud providers committed to decarbonization [12]–[16]. These pledges reinforced carbon-aware data center siting and design, but once a facility was built, it inherited the temporal variability of its grid and is tied to the local carbon profile, and these data center could not by themselves address the temporal variability of power generation or the needs of legacy data centers in carbon-intensive regions. These gaps motivated complementary operational solutions such as carbon-aware workload shifting to continuously optimize carbon efficiency after deployment, which is the focus of our study. B. Carbon-Intensity Data and Marginal Emissions Rates Operational decisions required accurate data about the electricity’s carbon-intensity over time and location. WattTime’s “Marginal Operating Emission Rate (MOER) Methodology” [27] formalized marginal signals for accurate real-time, location-aware accounting of IT workloads’ carbon impact. MOER represented the incremental CO2 emissions caused by an additional unit of power demand, contrasting with Average Operating Emission Rates (AOER), which simply divided total emissions by total generation across all resources, regardless of which plants are actually affected by a change in demand. Unlike AOER, which served as a static attributional measure, MOER is a consequential indicator–it reflects the real-world causal impact of actions such as shifting demand or integrating renewables, because only the marginal plants adjust their output. MOER thus enabled accurate identification of optimal time periods when shifting workloads can directly lead to emission reductions, facilitating effective scheduling of flexible workloads to align consumption with periods of lower marginal emissions. Comparative studies showed that using AOER could mislead online scheduling, while MOER better captured real abatement potential [28]. Empirical work demonstrated the preference of MOER for real-time control [19], [29]–[32]. Our study adopts MOER for all scheduling and accounting. C. Sustainable Cloud Computing for AI Recent studies showed that cloud computing accounted for 2.5% to 3.7% of global CO2 emissions, which was greater than the global emissions created by commercial airline flights (roughly 2.4%) [33], [34]. The scale and growth of AI further pushed data center demand, making sustainability a real concern. Patterson et al. quantified training footprints and showed that careful choice of model, data center, and hardware had the potential to reduce carbon impact by 1000x. Analyses on real LLM deployments revealed that serving/inference could
dominate energy consumption than model training, shifting the optimization from training alone to end-to-end or serving specific operations [10]. Everman showed that hardware and model selection and scheduling materially shifted serving-time emissions without impacting performance [35]. D. Carbon-Aware Workload Scheduling A mainstream research direction in sustainable computing was carbon-aware scheduling, which dynamically adjusted where and when workloads are executed to take advantage of cleaner energy. Early pioneering studies demonstrated the potential of temporal shifting (delaying or advancing flexible task to greener times) and spatial shifting (routing tasks to greener geographic locations) approaches. For example, [36] showed that intelligently routing traffic across a set of distributed data centers could substantially increase the use of renewable energy. Qureshi et al. demonstrated the economic and energy upside of geographic shifting [37]. Chien analyzed the carbon impact in AI inference, and showed the potential of geographical workload shifting [38]. On the temporal side, researchers developed schedulers that shifted deferrable batch tasks to off-peak or high-renewable periods. For example, Google’s carbon-intelligent computing platform shifted daily non-urgent computing to times when each data center’s local grid was cleaner [18]. [25] explored adjusting data center capacity and load distribution in response to grid conditions, and found that regional load shifting combined with throttling could help integrate more renewable energy. Academic efforts generalized carbon-aware scheduling for various platforms: from cloud batch schedulers, web service, to inference serving [30], [39]–[41]. E. Cloud Simulation Tool General purpose cloud simulation toolkits such as CloudSim [42] provided a comprehensive framework to model and simulate cloud computing environments. While these simulators were flexible and widely used, they were not tailored to the carbon-aware scheduling or to the GPU-intensive characteristics of modern AI inference workloads. In contrast to prior simulators, we introduce CATS, a trace-driven carbon-aware task simulator for cloud-scale AI inference. Our goal is to provide a reproducible framework for carbon-aware AI and data center operations that supports carbon-intensity signals, heterogeneous GPUs, multiple AI workloads, and pluggable scheduling policies. First, we quantify how today’s multi-provider cloud footprint aligns with low-carbon grids by fusing public siting metadata with carbon-intensity traces and reporting siting distributions across carbon bands. Second, using CATS, we estimate the nearterm decarbonization headroom without moving facilities by replaying realistic AI inference traces and comparing two baselines against carbon-aware spatial and temporal schedulers under capacity and delay constraints. Together, these results clarify where siting suffices, where it does not, and how much additional CO2 reduction is achievable through operational workload shifting.
III. DATA C ENTER A LIGNMENT A NALYSIS In this section, we conduct a comprehensive global analysis using datasets including 140 operational and planned cloud regions across 8 major providers (e.g. Amazon AWS, Microsoft Azure, and Google Cloud etc.) with carbon-intensity dataset spanning 145 grid regions across the global from 2022 to 2024. A. Data Sources and Preprocessing We collect cloud data center metadata from TeleGeography’s Cloud Infrastructure Map [43], a widely used and publicly accessible resource that keeps tracks of global cloud infrastructure. For each site the dataset records the cloud service provider (CSP), the metro area (city, country), the number of availability zones (AZs), and an operational vs. planned flag. The dataset includes 347 data centers consist of 8 major CSPs representing 2025. For comparison, we also collected the 2022 snapshot from the same source, which contains 293 data centers from 7 CSPs. This three-year gap allows us to identify emerging data centers and analyze whether recent expansion aligns with low-carbon grids. Figure 1 summarizes provider counts and net changes.
Fig. 1. Cloud infrastructure by provider, 2022 vs 2025. Bars show 2022 site counts (blue) and net change to 2025 (orange = increase; hatched = decrease). Labels give 2025 totals with the net change in parentheses. Based on TeleGeography’s Cloud Infrastructure Map, the combined number of DCs of the eight CSPs expanded from 293 in 2022 to 347 in 2025.
For carbon-intensity signal we use Marginal Operating Emission Rates (MOER) from WattTime [27], a recognized source employing empirical modeling approaches. MOER measures the extra carbon emissions generated caused by the additional electricity consumed in a grid (in lbsCO2 /MWh), which is the real-time emissions rate that changes every five minutes. The model comprehensively map grid load changes to various factor that directly affect local grid emissions, such as fuel mix, imports and exports, and renewable output, making it more representative for operational workload updates. This contrasts with the average operating emission rates, which simply relate total emissions with total grid generation and therefore unable to adapt scenarios that partially shifting the demand. Using MOER lets us better evaluate places and windows where executing flexible workloads directly reduces emissions. We obtain five-minute MOER traces for 145 grid regions spanning the United States (114), Europe including the UK (25), and Australia (6), covering January 2022 to April 2025. Throughout the paper, unless otherwise noted, “carbonintensity” (CI) refers to MOER.
Fig. 2. Grids 2022-2024 average carbon-intensity with cloud data centers in three areas. The maps are divided into grid regions, the color in each region represents grid carbon-intensity, greener grids have lower carbon-intensity and hence less carbon footprint per energy consumption. The markers scattered across maps show the location and number of data centers, marker with colored circle indicate new data centers built since 2022. The two highlighted grids include the proposed data centers for Stargate Project (Abilene, TX) and xAI (Memphis, TN).
For global data center locations, we employ Geoapify Location Platform’s geocoding API [44] to convert each data center metro area to latitude/longitude and map them to grid regions using WattTime’s GeoJSON boundaries. For five-minute CI data, we convert the timestamp from grids’ local time to UTC time for cross-region alignment. We address the only observed gap by forward-filling a missing window in WEM grid (Western Australia, from Nov. 7, 2024 22:00 to Nov. 8, 2024 00:30). The result dataset includes 122 cloud data centers in 2022 and 140 in 2025 with CI signals. B. Alignment Analysis Figure 2 visualizes available grids average carbon-intensity (lbsCO2 /MWh) from 2022 to 2024 and overlays the numbers of cloud data centers at each location in the United States, Europe, and Australia. Grids color from green to red shows their average CI. Greener grid regions indicate a cleaner grid and redder colored regions denote grids with more per unit carbon emissions. The markers scattered across maps show the number of data centers in this location, and for better visualization, data centers close to each other are combined and report in a single marker (distance threshold: 150 km for U.S. and Australia, 300 km for EU including UK). Markers circled by colored outline denote regions with at least one added data centers since 2022. Furthermore, we highlight two grids with red boundaries indicating two emerging data centers (the Abilene “Stargate” site in ERCOT NORTHCENTRAL in Texas and the Memphis xAI in TVA in Tennessee). In all three maps, both existing data centers before 2022 and new builds since 2022 are not concentrated in low carbon regions. In the United States (Fig. 2a showing 72 collected data centers), more data centers are on east and west coasts. There are 12 data centers located in the CAISO NORTH grid (approximately 748 lbsCO2 /MWh), representing about
one-sixth of the total and situated in a relatively greener region with lower carbon intensity. In contrast, 22 are in PJM DC and 9 in PJM SOUTHWEST OH - both regions with much higher carbon intensity, around 1250 lbsCO2 /MWh. This means running the same workload in PJM grid would emits roughly 67% more CO2 than in greener west coast grids if identical energy are used. For data centers emerge since 2022, there are 4 in PACW (near Washington, CI: 1274), 2 in PJM CHICAGO (CI: 1190), and 1 in SPP SIOUX (in Iowa, CI: 1122). All of them have an average of over 1100 CI, which is not carbon efficient. The two newly built or underconstruction large AI data centers - Stargate in Abilene, Texas (grid CI: 1,100) and another by xAI in Memphis, Tennessee (grid CI: 1,177) - are also not siting in low carbon intensity grids. In Europe including UK (Fig. 2b with 54 collected data centers), providers deploy across a wide CI spread rather than clustering in the cleanest grids. DE (Germany) and UK have most data centers (10 each, CI in DE: 1709, CI in UK: 929), followed by FR (France, 6, CI: 847). Of all data centers, there are 17 (31%) located in regions with less than 900 CI, 18 (33%) reside in over 1200 CI grids. There are 17 new builds since 2022 across these area, 9 (53%) below 900 CI, 1 between 900 and 1200 CI, and 7 located in regions with an average of more than 1200 CI (4 in ES, CI:812, 3 in FR:847, 2 in IT:850, 2 in UK:929, 1 in IE:1103, 1 in RS:1391, 2 in DE:1709, 1 in SE:1735, 1 in PL:1909). We can see in this area, providers tend to place new data centers in greener regions. In Australia (Fig. 2c with 14 collected data centers), all facilities reside in two relatively carbon-intensive mainland regions (10 in NEM NSW with 1607 lbsCO2 /MWh, 4 in NEM VIC with 1236 lbsCO2 /MWh)). There are 2 new builds still in existing grids. Overall, existing data centers and recent expansion patterns
TABLE I N UMBER OF DATA CENTERS IN EACH CI RANGE CI (lbsCO2 /MWh)
Azure
Oracle
AWS
Alibaba
Tencent
IBM
Huawei
US
EU+UK
AUS
Total
<900 900-1200 1200-1500 >1500
8 13 9 8
7 13 8 5
4 9 7 5
5 2 13 4
2 1 2 2
2 0 2 1
1 2 2 2
0 1 0 0
12 22 38 0
17 19 1 17
0 0 4 10
29 41 43 27
Column Total
38
33
25
24
7
5
7
1
72
54
14
140
do not correct the carbon-siting mismatch. Table I classifies 140 facilities into four CI ranges for eight CSPs and three geographies (US, EU+UK, AUS). Using CI < 900 lbsCO2 /MWh as “low-carbon”, only 29 out of 140 (21%) sites are in low-CI grids. 70 out of 140 (50%) are in ≥ 1,200 side (mediumhigh to very high), and the remaining 41 out of 140 (29%) fall in 900-1,200. By region, the U.S. tilts toward the 1,2001,500 CI interval (39 sites) and no data center in very high CI grid; Europe places a substantial share in ≥ 1,500 (17 sites); Australia also skewed to > 1,500 (10 sites). By provider, no major CSP places a majority of its footprint in the < 900 range; each maintains sizable presence in ≥ 1,200 grids, with mix differences across providers but a shared pattern of weak low-CI concentration. Two conclusions are derived from this global analysis. First, today’s multi-provider footprint is misaligned with clean grids: half of sites in these three areas sit in ≥ 1,200 CI grids, while only around one-fifth are in < 900 CI grids. Second, recent growth has not improved this alignment, while Europe show an aware of siting over 50% new data centers in low-CI grids. It is important for providers to take grid carbon-intensity into considerations when planning for future data centers. On the other hand, the results also motivate us to find solutions, under current data center siting, on carbon-aware scheduling across existing regions. In the following sections, we analyze the potentials of carbon reduction through workload shifting.
IV. CATS D ESIGN AND A RCHITECTURE In this section, we present the design and architecture of Carbon-Aware Task Simulator (CATS), an essential tool to support the evaluation of various carbon-aware scheduling algorithms under realistic operational conditions. CATS supports various AI inference task types, multi-GPU profiling, diurnal and geographically-skewed arrivals with task mix and per-task SLAs (delay limits). CATS consists of five components: (1) a trace generator that synthesizes AI inference task traces; (2) a virtual cloud infrastructure that simulates data centers and GPU pools; (3) a discrete-event simulation engine that directs tasks to target resources; (4) a carbon accounting module that measures pertask carbon emissions at both scheduling time and execution time, and (5) a scheduling policy module that implements four scheduling policies.
A. Modeling the AI Inference Trace To model AI inference traces, CATS takes as input the benchmark runtime and energy data for each task–GPU pair. Users specify a few parameters: simulation start time (UTC), total duration (hours), number of tasks, task and region mix weights (each summing to 1), per-task delay limits (“HH:MM:SS”), and a 24-hour arrival curve (a vector of hourly weights). These settings can be defined in a single YAML file. Using this configuration and the benchmark table, CATS generates a reproducible, time-ordered inference trace. Each trace entry includes an arrival timestamp, task type, origin region, delay limit, and per-GPU mean runtime and energy values for scheduling. Trace synthesis proceeds in four steps. First, the 24-hour diurnal vector is rotated to align with the simulation start hour, normalized over the simulation horizon, and used to allocate the total task count across hours. Second, within each hour the minute-level arrival counts are drawn from a Poisson distribution with mean equal to the hour’s target divided by 60, and then adjusted to match the hourly total. Third, within each minute the tasks’ arrival timestamps are sampled uniformly over the 60-second window, task types and origin regions are sampled independently from the normalized task mix and region mix lists, and each task is attached with its task-specific delay limit together with the per-GPU mean runtime and energy from the benchmark table. Fourth, all generated events are sorted by time, assign sequential event IDs, and written out as a CSV trace file. To ensure reproducibility, we apply a deterministic seeding scheme: for each hour index h the trace generator sets the random seed to the user-provided base seed plus h before drawing Poisson minute counts and sampling task and region assignments. Given the same YAML configuration and base seed, the generator produces identical traces. All timestamps in the trace are stored in UTC so the simulator can replay tasks in time order and align them with grid carbon-intensity signals. Because each trace records the per-GPU mean runtime and energy metrics, schedulers can later compare carbon impact and performance between devices without rerunning any profiling. B. Modeling the Cloud Infrastructure To model the cloud infrastructure, CATS requires the GPU capacity of each data center. This configuration is provided in the same YAML file used for trace generation configuration. CATS models the infrastructure with three levels: (i) multiple
data centers; (ii) within each data center, one or more GPU pools (one pool per GPU type with a fixed number of identical devices); and (iii) single-GPU execution, where each task occupies exactly one device for its runtime. Within each GPU pool, CATS maintains the total number of GPUs of that type, the number of devices currently running tasks, a normal queue to hold tasks when no free GPU available, a min-heap that tracks the predicted finish time of each device in the pool, a priority queue used by temporal scheduling policy to hold tasks that should run immediately when capacity becomes available, which can skip the normal queue, and a counter that tracks the total estimated runtime of tasks currently in the priority queue. For fast predictions of start and finish times, every pool initializes a min-heap of “next-free” times with one entry per device, set to the simulation start time. Assigning a task updates the entry who finishes earliest to the task’s predicted finish time. Figure 3 illustrates this earliest-finish assignment within a pool. The orange bars cross horizontal axis indicate current running tasks on each GPU, and the numbered white bars show the already scheduled tasks on each GPU, the numbers denote the order of tasks were scheduled when the scheduler prioritizes task latency. When a new task (task 8) arrives at this GPU pool, it will be placed on the device whose finish time is the smallest in the min-heap, so that it starts immediately after task 2 completes.
Fig. 3. Task scheduling in a GPU pool. Bars cross horizontal axis indicate scheduled task on each GPU, the orange regions indicate current running tasks on each GPU. A min-heap is used to place incoming tasks to the earliest available GPU, as showed by numbered bars. The new task 8 is scheduled to the left most GPU, which will be executed after task 2 is completed.
C. Scheduling Engine and Workflow We implement a customized, trace-driven discrete event simulator (DES) as the scheduling engine to orchestrate between the arrival trace, the cloud infrastructure model, the scheduling policy, and the carbon accounting module. The DES consumes the synthetic trace in time order, maintains a global event heap, and repeatedly pops the next event until no events remain. For each event, it queries the scheduler for placement decisions, updates pool capacities and queues while
enforcing constraints, logs per-task carbon and performance metrics, and outputs a trace where each task has its actual start time, end time, and assigned GPU pool. Given a fixed trace, configuration, and scheduling policy, the engine is fully deterministic.
Fig. 4. Carbon-Aware Task Simulator (CATS) Scheduling Engine workflow.
The synthetic trace is kept in a min-heap over time and each is labeled with an event type, initialized to ARRIVAL event. The event calendar contains three event types: ARRIVAL, DEFERRED, and COMPLETION, as illustrated in Figure 4. 1 , the simulator calls the scheduler with the task, On arrival ⃝ a snapshot of all data-center states, and current time. The scheduler selects a target region and GPU pool and returns either an immediate placement or a deferred placement within 2 . For an immediate placement, the the task’s delay limit ⃝ DES refreshes the target pool’s min-heap of “next-free” times 3 , and, if a device is available it starts the task immediately ⃝ updates task predicted finish time, and pushes a completion 4 ; otherwise, the event at the at the predicted finish time ⃝ 5 . For a deferred decision, task joins the pool’s normal queue ⃝ the simulator inserts a DEFERRED event for this task at the 6 . requested release time ⃝ 7 , the task should run When a DEFERRED events fires ⃝ as soon as possible to keep the carbon benefits. The DES again checks the target pool: if a GPU is idle it starts the task 8 ; otherwise, the task is placed in the pool’s immediately ⃝ 9 . Both normal and priority queues are first-inpriority queue ⃝ first-out (FIFO), but when capacity frees, the simulator looks at the head of each queue and applies an earliest-deadline-first rule so that deferred tasks with earlier deadlines can cut ahead of non-deferred tasks while never preempting a running task. 10 , the simulator frees one GPU in the On completion ⃝ corresponding pool, triggers the carbon accounting module to compute the task’s average carbon-intensity and emissions over its execution window, and updates performance statistics. It then checks the pool’s queues: if there are waiting tasks, it 11 , selects the next one using the same earliest-deadline rule ⃝ 12 , then scheduling a new and starts this task immediately ⃝
COMPLETION event. Throughout the run, the DES logs pertask events and scheduler decision, per-pool state changes for audits. D. Scheduling Policies Our goal is to minimize the operational CO2 by making scheduling decisions that exploit variation in carbon-intensity across space and time. Within CATS, scheduling policies are implemented as pluggable modules that share the same interface: on each task arrival the scheduler observes the current data center state, chooses where to run (spatial shifting across data centers) or when to start (temporal shifting at local data center), and GPU type selection when multiple node types are available, while respecting capacity and SLA constraints, and return either a immediate placement decision or a deferral placement decision. Figure 5 illustrates the general view of this workload scheduling. We implement Speed-First and Carbon-First as two baseline scheduling algorithms, and Spatial Shifting and Temporal Shifting as two carbon-aware workload shifting algorithms.
Spatial shifting policy. The spatial shifting policy extends the Carbon-First baseline by allowing routing across data centers. For each region-GPU candidate, it predicts a start/finish window using that pool’s queue, computes the associated emissions, and selects the lowest-emission candidate that respects the task’s delay limit. To avoid shifting for negligible gains, it applies minimum relative and absolute saving gates: by default it only shifts away from the origin if the best remote option reduces predicted emissions by at least ϵs,rel = 0.05 (5%) and by more than a small absolute threshold. If the best origin option has zero predicted emissions (e.g. during a zerocarbon window) or no SLA-safe remote candidate exists, the task stays local; if neighter local nor remote can meet the delay limit, the scheduler return the best origin option with SLA violation. Figure 6 shows persistent inter-regional carbonintensity gaps that create opportunities for spatial shifting.
Fig. 6. Carbon-intensity (CI) Trends Across Selected Regions. Three-year carbon-intensity (3-day moving average) for four U.S. regions. Persistent interregional gaps indicate substantial headroom for carbon-aware spatial shifting.
Fig. 5. Workload shifting illustration. A task is initially sent to region A servers, spatial shifting redirects the task to another region while temporal shifting defers the task for execution.
Speed-First baseline policy. In the speed-first baseline (baselinespeed ), tasks always stay at their origin data center. At each task arrival, the scheduler ranks GPU types by predicted finish time, computed as each pool’s next available time plus the task’s mean runtime on that GPU. It selects the pool with the earliest predicted finish, breaking ties by shorter runtime. If the chosen pool has an idle GPU, the task starts immediately; otherwise, it is queued, and the pool’s min-heap is updated accordingly. Carbon-First baseline policy. The carbon-oriented baseline (baselinecarbon ) also keeps tasks in their origin region. For each GPU type there, it predicts the task’s start and finish window based on the pool’s queue, then estimates emissions by multiplying the region’s carbon intensity over that window by the task’s mean energy. It selects the option with the lowest predicted emissions, breaking ties by earlier finish time. If no option meets the delay limit, it falls back to the lowestemission choice even if this causes an SLA violation. As before, tasks either start immediately on an available GPU or join the normal queue.
Temporal shifting policy. The temporal shifting policy keeps tasks in their origin region but searches within each task’s delay limit for a future release time that minimizes predicted emissions. For each local GPU type, it predicts earliest queue-aware start time, scans candidate start times up to the latest SLA-safe release, and evaluates emissions for each start/finish window. A candidate is admissible only if it (i) passes a rate cap α limiting how many GPU-seconds of work can be newly deferred out of the current five-minute window of the pool (d, g), (ii) satisfies a capacity ledger ρ that caps the deferred GPU-seconds landing in each future per-minute bucket over a 24-hour horizon, and (iii) respects a backlog cap β bounding the total deferred GPU-seconds across that horizon. These three guards are maintained independently for every region-GPU pair. The policy also enforces minimum relative and absolute saving gates (e.g., ϵt,rel = 0.05 and ϵt,abs = 0.5 (gCO2 )) to avoid deferring for trivial savings. If a candidate passes the savings gates and all three guards, the scheduler returns a “defer-until” decision. The DES will then places the task at the chosen release time, using the priority queue with earliest-deadline-first between queues to protect SLAs. If no admissible deferral is found, the policy falls back to the previous Carbon-First placement at arrival. Figure 7 illustrates the intraday carbon-intensity swings that enable temporal shifting.
energy on that GPU pool (converted to MWh) and the timeweighted average carbon-intensity in its execution region over task’s runtime. For a candidate (d, g) with predicted window [sj , sj + rj,g ], we estimate: Z sj +rj,g 1 Ej,g pred ) CId (t)dt (1) CO2 (j; d, g) = ( 3.6 × 109 rj,g sj Fig. 7. Intraday carbon-intensity over a 24-hour period for four regions. Pronounced within-day swings enable carbon savings via bounded temporal shifting at the original site.
E. Carbon Emission Accounting CATS computes operational emissions by aligning each task’s execution window with the carbon-intensity series of its execution region and multiplying by the task’s measured energy. Let D be the set of data centers. Each data center d ∈ D maps to a grid region with carbon-intensity CId (t) (lbsCO2 /MWh) sampled every five minutes. Data center d supports one or more GPU pools, each pool contains a fixed number of identical devices of a given GPU types. Tasks arrive as a fixed trace J . A task j ∈ J has arrival time aj , origin data center oj , a task type, and per-GPU benchmarks: mean runtime rj,g (seconds) and mean energy Ej,g (joules) for each GPU type g. Each task has a delay limit ∆j and is non-preemptive, occupying exactly one GPU for its entire runtime. The system parameters used in this subsection are summarized in Table II. TABLE II N OTATION AND PARAMETERS Notation
Description
System Configuration D Set of data centers CId (t) Carbon-intensity signal in data center d at time t Nd,g Number of GPUs of type g at data center d Workload Characteristics aj Arrival time of task j ∆j Delay limit of task j sj Predicted start time of task j dj Selected data center for task j gj Selected GPU pool for task j rj,g Mean service time (seconds) for task j on GPU type g Ej,g Mean energy consumption for task j on GPU type g
On arrival, the scheduler finds the optimal choice and returns the selected data center dj , GPU pool gj , a flag indicating whether the task can start immediately or should enqueue, and, for the temporal policy, whether the task should be deferred. The choice is determined by evaluating the task’s predicted start time sj at each candidate (d, g) (sj is candidatespecific, i.e., sj,(d,g) ), accounting for queuing delay (for nontemporal policies), or by searching (sj , g) pairs, the potential release times at each local GPU pool (temporal policy), and computing the predicted carbon emissions, the product of its
where CId is in lbsCO2 /MWh and Ej,g is in joules. The factor 3.6 × 109 converts joules to MWh so the results is in pounds of CO2 . In practice, the integral is evaluated by summing over the five-minute carbon-intensity signals that overlap the task’s runtime, this proportionally accounts for partial windows at both ends. When determining the start time sj , the scheduler tries to satisfy: (2) sj = max {aj , next free(d, g)} sj + rj,g ≤ aj + ∆j
(3)
If no candidate satisfies the delay limit boundary, the scheduler falls back to the best available option and records an SLA violation. For temporal shifting policy, at placement time, the scheduler estimates the start time using the earliest free GPU from the min-heap and, if priority queue is not empty, it adds an aggregate correction sj += prio q runtime / Nd,g . This adjustment averages the priority work across the pool and therefore is an approximation of how priority and normal queues will actually interleave. And this indeterminacy on which task start first may cross its deadline even though both looked safe at placement, which means the realized start time can exceed sj , we record this as prediction error. Our objective is to minimize total operational emissions: X min CO2 (j) (4) j∈J
Decision-time comparisons use COpred as above, while actual 2 accounting at completion replaces [sj , sj + rj,g ] with the realized start and finish time. This optimization is subject to three hard constraints enforced by the schedulers: (i) Capacity: at any time, the number of concurrent active tasks in pool (d, g) never exceeds its pool capacity; (ii) SLA/delay limit: for all tasks, sj +rj,g ≤ aj +∆j . (except when a policy explicitly allows SLA violation as a fallback); and (iii) Admissible regions: each task may be restricted to a subset Aj ⊆ D. In our experiments, Aj = oj for both baselines and the temporal scheduler, and Aj = D for the spatial scheduler. V. E XPERIMENTS AND E VALUATION In this section, we use CATS to evaluate four scheduling policies on realistic AI inference workloads. We first profile six types of AI inference tasks across four different GPUs, then synthesize a 24-hour trace with configurable task mix, region mix, and diurnal pattern, and finally replay the trace through the discrete-event simulator described in Section IV using two baselines and two carbon-aware shifting policies.
A. AI Inference Task Profiling We profile six widely used AI inference tasks types: text generation, text-to-speech, text-to-image, image captioning (image-to-text), image-to-image, and text-to-video. For each task family, we select multiple open-source models hosted on Hugging Face to span different compute footprints (e.g., three text-generation models from 2.7B to 13B parameters, lightweight and XL variants for diffusion models, and two textto-video models). Table III lists the models and their aliases. TABLE III AI I NFERENCE TASKS AND M ODELS USED
Task type
Model
Alias
Text Generation
meta-llama/Llama-2-13b-chat-hf mistralai/Mistral-7B-Instruct-v0.3 microsoft/phi-2
TG–Llama TG–Mistral TG–phi2
Text to Speech
suno/bark speechbrain/tts-tacotron2-ljspeech & speechbrain/tts-hifigan-ljspeech
Text to Image
stabilityai/stable-diffusion-xl-base-1.0 OFA-Sys/small-stable-diffusion-v0
T2I–sdxl T2I–small
Image to Text
Salesforce/blip2-flan-t5-xl nlpconnect/vit-gpt2-image-captioning
I2T–blip2 I2T–vigpt2
Image to Image
stabilityai/stable-diffusion-xl-base-1.0 OFA-Sys/small-stable-diffusion-v0
I2I–sdxl I2I–small
Text to Video
damo-vilab/text-to-video-ms-1.7b cerspense/zeroscope v2 576w
T2V–ms T2V–zeroscope
TTS–bark TTS–tacotron2
All models are deployed on four types of GPUs in the cloud: A100, A6000, H100, and GH200. Detailed host and device specifications are shown in Table IV. Each task is executed on a single GPU. TABLE IV H ARDWARE S PECIFICATION FOR B ENCHMARKS
Node
Host Specification
GPU Information
NVIDIA A100
AMD EPYC 7J13 30 vCPU & 220 GB RAM 512 GB SSD
1 x A100 40GB SXM4 Ampere
NVIDIA GH200
Neoverse-V 64 vCPU & 432 GB RAM 4 TB SSD
1 x GH200 96GB/480GB Hopper
NVIDIA A6000
AMD EPYC-Rome 14 vCPUs & 100 GB RAM 512 GB SSD
1 x A6000 48GB Ampere
NVIDIA H100
Intel(R) Xeon(R) Platinum 8480+ 26 vCPU & 225 GB RAM 1 TB SSD
1 x H100 80GB PCIe Hopper
Within each task family, we fix the inputs across models to make comparisons meaningful. For tasks requiring text input (text generation, text-to-speech, text-to-image, and text-
to-video), the input prompts average 29 tokens2 . For imagebased tasks (image-to-text and image-to-image), we use a single 224×224 color JPEG image (25.7 KB). The imageto-image task additionally includes a guiding text prompt. Outputs are stored in their native formats: text generation and image captioning results are logged as text in JSON; text-to-speech outputs are saved as WAV files (22.05 kHz for Tacotron2, 24 kHz for Bark); text-to-image and image-toimage outputs are saved as PNGs; and text-to-video results are exported as MP4 files at 16 fps. CATS records only runtime and device energy for analysis. For each model–GPU pair, we perform two warm-up iterations followed by ten measured runs, recording end-to-end runtime and on-device energy via NVML. We average the ten runs and compute mean power as energy divided by runtime. Results in Table V indicate that text-to-video workloads are the most compute- and energy-intensive, while image captioning, image-to-image, smaller text-generation models, and Tacotron2 are the fastest and least energy-demanding. Hopper devices (H100/GH200) generally outperform Ampere devices (A100/A6000) in speed. H100 consumes the least energy for most tasks, whereas A100 often achieves the highest power efficiency. In subsequent experiments, the per-pair mean runtime is used as the job “size” signal, and mean energy is used for emission accounting in combination with carbonintensity time series. B. Workload and Trace Configuration We generate a 24-hour trace with N = 600,000 tasks using the trace generator described in Section IV-A. Since detailed production mixes are rarely public, we adopt a plausible composition dominated by text generation: over 60% of tasks are text-generation, 15% are text-to-speech, and the remainder are distributed across image and video tasks, with larger models favored within each category. Table VI summarizes the per-model shares and per-task delay limits. Small tasks have tight deadlines, while larger tasks, such as text-to-video, are allowed to be deferred for up to three hours. We generate task arrivals uniformly across four U.S. grid regions: CAISO North, ERCOT Austin, PJM DC, and MISO Mason City, with each region contributing one-quarter of all arrivals. Under the spatial policy, tasks may later be routed to other regions. Task arrivals within each region follow a fixed diurnal pattern in local time. Prior studies [45]–[47] have shown that AI inference requests exhibited strong diurnal variations, with peak traffic more than twice that during offpeak hours. Our diurnal cycle, shown in Figure 8, reflects higher traffic during daytime with a mid-day peak and lower load after midnight. C. System Configuration and Policy Settings We deploy the trace on a four-region GPU fleet modeled as in Section IV-B. Each region hosts 65 GPUs with the same heterogeneous mix: 15×H100, 25×A100, 10×H200, 2 Token counts are collected using the OpenAI tokenizer for GPT-4o and GPT-4o mini at https://platform.openai.com/tokenizer.
TABLE V AI I NFERENCE TASK B ENCHMARK
Runtime (s)
Task Type TG–Llama TG–Mistral TG–phi2 TTS–bark TTS–tacotron2 T2I–sdxl T2I–small I2T–blip2 I2T–vigpt2 I2I–sdxl I2I–small T2V–ms T2V–zeroscope
Energy (J)
Power (W)
A100
A6000
GH200
H100
A100
A6000
GH200
H100
A100
A6000
GH200
H100
8.15 6.23 4.98 39.90 0.98 8.94 1.72 0.32 0.15 5.52 1.21 442.86 799.59
10.18 7.53 5.64 42.23 1.10 14.67 1.99 0.41 0.43 5.80 1.28 484.99 882.83
7.50 5.82 4.35 34.69 0.89 7.10 1.47 0.34 0.11 4.58 1.05 352.22 693.04
5.09 3.39 2.85 17.52 0.93 6.35 1.05 0.16 0.12 2.63 0.66 197.14 324.04
1,549.50 903.01 514.04 2,747.55 63.18 3,154.43 294.25 30.68 9.90 370.58 121.15 26,173.68 57,316.20
2,915.87 1,748.79 917.78 4,913.62 113.32 4,276.16 556.71 60.78 42.72 622.81 177.55 43,160.25 86,561.45
2,518.05 1,728.90 1,139.53 8,802.31 221.42 3,506.97 572.77 95.56 30.28 1,165.91 293.52 83,512.17 166,077.96
1,133.20 636.61 370.43 2,243.61 91.20 2,022.75 285.89 23.27 13.58 301.94 97.87 19,078.01 34,020.76
190.12 144.96 103.23 68.85 64.30 352.93 228.78 96.01 66.77 67.17 99.86 59.10 71.68
286.56 232.12 162.82 116.35 103.43 291.45 279.82 149.79 100.22 107.47 138.52 88.99 98.05
335.63 296.81 261.68 253.74 249.22 493.91 390.78 279.66 263.11 254.31 280.36 237.10 239.64
222.69 187.48 130.15 128.07 97.94 318.53 272.45 147.54 109.47 114.82 148.09 96.77 104.99
Note: All results in this table are benchmarked on actual GPUs and models.
TABLE VI TASK M IX AND D EFER L IMIT P Share ( = 1.0)
Delay limit (HH:MM:SS)
TG–Llama TG–Mistral TG–phi2
0.28 0.28 0.06
00:03:30 00:02:30 00:02:00
TTS–bark TTS–tacotron2
0.10 0.05
00:15:00 00:00:30
T2I–sdxl T2I–small
0.08 0.03
00:05:00 00:01:00
I2T–blip2 I2T–vigpt2
0.04 0.02
00:00:30 00:00:30
I2I–sdxl I2I–small
0.04 0.01
00:02:00 00:00:30
T2V–ms T2V–zeroscope
0.005 0.005
03:00:00 03:00:00
Task-Model Alias
(seconds), N the total number of tasks, and r̄ denote the mixweighted mean runtime per job, computed from the task mix and per-GPU benchmarks. For our configuration, r̄ = 13.70s. The fleet capacity C and trace-required load L are: C = G · H,
L = N · r̄
(5)
We define the fleet-average utilization as: L (6) C To make r̄ explicit and reusable, we first compute a per-GPUtype mix-weighted runtime: X r̄g = pt · rt,g (7) µavg =
t∈T
where pt is the task mix weight and rt,g is the measured mean runtime of task-model t on GPU type g. Let αg denotes the fraction of tasks that run on GPU type g (in our experiment αg is weighted by the GPU capacity), the fleet-level mean runtime is: X r̄ = αg · r̄g (8) g∈G
Fig. 8. Diurnal arrival curve. We simulate the pattern with more tasks arriving during the day, especially around noon, and low usage after midnight.
and 15×A6000, for a total of G = 260 GPUs. The simulation horizon is H = 24 hours (86,400 seconds) starting at 2024-0413 00:00:00 UTC. Carbon-intensity signals for the four regions come from five-minute signals over that day. We measure capacity and load using GPU-seconds. Let G be the total number of GPUs in the fleet, H the simulation horizon
The resulting fleet-average utilization is µavg ≈ 0.37. Because arrivals are diurnal, we also report a peak-hour utilization estimate. Using the 24-hour weight vector, the ratio of the peak hour to the average hour is Mpeak = 1.84 (normalized peak hour weight over average weight), the peak utilization is µpeak ≈ Mpeak · µavg = 0.67. Queuing theory has shown that system utilization must be kept well below 100% to avoid exponentially increasing wait times, particularly in systems with high variance in request arrivals and service times [48], [49]. We therefore configure our simulation with an average utilization of µavg = 0.37 and peak utilization of µpeak = 0.67. This conservative provisioning is necessary for the diurnal arrival pattern and the heterogeneous task mix, both of which contributing to
high coefficients of variation in arrivals and service times, as observed in production LLM workloads [45], [47]. For both spatial and temporal shifting policies, we set the relative saving gate ϵs,rel = ϵt,rel = 0.05, which requires a ≥ 5% predicted CO2 reduction to shift, and the absolute saving gate ϵs,abs = ϵt,abs = 0.02(gCO2 ). The absolute saving gate threshold is determined based on the Speed-First baseline scheduling results (avg= 0.15g/task, p95= 0.34g/task), calibrated to be roughly 10% of the average per-task savings under the Speed-First baseline. For temporal policy guards, we set rate cap α = 0.10, capacity ledger ρ = 0.10, and backlog cap β = 0.30. D. Results and Analysis
Fig. 9. Task distribution across GPU type on two baseline policies. Each row represent one GPU pool and columns indicate different regions, SpeedFirst policy prefer faster GPUs such as H100 and GH200, while Carbon-First policy prefer energy efficient GPUs like H100 and A100.
We evaluate four schedulers on the same trace at a fleetaverage utilization of ∼ 0.37 with a heterogeneous GPU mix: two baselines (Speed-First and Carbon-First) and two carbonaware policies (Spatial Shifting and Temporal Shifting). Table VII summarizes the system-level outcomes. In a heterogeneous GPU-type setting, the Carbon-First baseline lowers total CO2 by 14.7% (251.45 to 214.57 lbs) and energy by 14.9% (902.92 to 768.68 MWh) relative to the SpeedFirst baseline, but at a latency cost with per-task runtime increases from 8.84 to 53.50 seconds and queue wait from 0.56 to 44.92 second. The Spatial Shifting policy achieves the lowest carbon emissions with a 38.4% CO2 reduction relative to Speed-First (241.45 to 154.91 lbs) and a 27.8% reduction to Carbon-First (from 214.57 lbs). This comes with higher per-task latency (136.32 s runtime and 126.18 s queue wait). Total energy consumption is similar to the Speed-First baseline (-1%, 902.92 to 893.80 MWh), but 16.3% higher than CarbonFirst, due to routing tasks to cleaner regions that may run on less-efficient GPUs. The policy shifts around 59% of tasks across regions. The Temporal Shifting policy delivers a 16.1% CO2 reduction compare to Speed-First (251.45 to 210.83 lbs) and is roughly on par with Carbon-First on both CO2 (-1.7%, 214.57 to 210.83 lbs) and energy consumption (-0.1%, 768.68 to 767.94 MWh). Per-task latency is higher than under CarbonFirst (+24.3% in runtime and +28.9% in queue wait). With the configured guards and saving gates, the SLA violation rate is 3.27%, with a deferred fraction of 0.4%. Scheduler overheads are negligible, for the most time-consuming policy (Spatial Shifting), with 2.32 ms/task, the overhead is less than 0.002% of its mean runtime (136.32 seconds). When comparing the two baselines (Speed-First and Carbon-First), both execute tasks locally, the performance gap arises from GPU selection. Figure 9 visualizes the task distribution across GPU types under the two baselines. In Speed-First (left), H100s dominate (59.3% to 60.7% in every region), GH200s handle another large fraction (28.3% to 29.6%), and 10.8% to 12.4% of tasks run on A100s, while A6000s nearly unused. Carbon-First redistributes toward lower-energy devices: H100 and A100 collectively handle over 99% of the tasks. This device-mix shift explains the 14.9% energy drop and 14.7% CO2 reduction in Table VII. The
figure also shows that ERCOT Austin and PJM DC retain a small GH200/A6000 footprint while there is none in CAISO North and MISO Mason City. The reason is that CarbonFirst policy minimizes CI(t) × energy upon task arrival. When queues on H100/A100 are long and local carbon-intensity is rising, starting immediately on a less efficient GPU can beat waiting into a dirtier interval, as ERCOT Austin and PJM DC have more volatile intraday profiles (cf. Fig. 7). This task distribution pattern shows that on heterogeneous fleets, energyoriented device choice alone can deliver material carbon gains, mostly by shifting tasks from GH200 to A100 while keeping H100 saturated, and that carbon-intensity volatility leads to residual use of less-efficient GPUs. With the Spatial Shifting policy, tasks may be executed in a region different from their origin when the destination’s CI(t) × energy is lower. Figure 10 visualizes the resulting flows. The trace starts with an even origin split (25% per region). After routing, ERCOT Austin and CAISO North become net importers, executing 30.9% and 27.1% of all tasks, respectively, while PJM DC and MISO Mason City are net exporters, ending at 23.5% and 18.5%. This rebalancing is consistent with their carbon-intensity profiles on the study day, where ERCOT Austin and CAISO North have, on average, lower carbon-intensity than the other regions and ERCOT Austin shows several intraday troughs, so directing arrivals there yields the 38.4% total CO2 reduction reported earlier. Notably the flow pattern is many-to-many: each origin sends work to all four destinations and each destination receives work from all four origins, this breadth comes from the interleaving carbon-intensity over time. Figure 11 plots the number of active tasks (left axis) and each region’s carbon-intensity (right axis). The dashed baselines (Speed-First, Carbon-First) largely follow the diurnal arrival pattern. In contrast, the Spatial policy (solid green) tends to anti-correlates with carbon-intensity: it ramps up when the local grid is cleaner and backs off when it is dirtier. In CAISO North (panel a), carbon-intensity is lowest near 00-01 UTC; the Spatial policy immediately drives the region to its capacity limit (65 concurrent tasks). In ERCOT Austin (panel b), extended low-CI intervals appear around 04-06, 12-14, 17-
TABLE VII R ESULTS ACROSS FOUR SCHEDULING POLICIES Policies Baselinespeed Baselinecarbon Spatial Temporal
Total CO2 (lbs)
CO2 /task (gram)
Total Energy (MWh)
Runtime/task (second)
Queue Wait/task (second)
SLA Violation Rate (%)
Shifted Rate (%)
Sch. Time (ms/task)
251.45 214.57 154.91 210.83
0.19 0.16 0.12 0.16
902.92 768.68 893.80 767.94
8.84 53.50 136.32 66.51
0.56 44.92 126.18 57.92
3.27
59.88 0.4
0.06 0.58 2.32 0.88
Fig. 10. Spatial shift flow across regions. This Sankey diagram visualize task flow under spatial shifting policy, the left side is task original regions and right side is destination regions where task actually runs in.
our configuration, this yields an additional 1.7% CO2 reduction over the Carbon-First baseline, by deferring 0.4% of tasks (2,427 of 600,000). Unlike the other policies, temporal scheduling introduces SLA breaches. These SLA violations happen because deferred tasks in the priority queue and nondeferred tasks in the normal queue contend for the GPU pool. In our run, the SLA-violation rate is 3.27% (19,602 tasks): 97 violations are deferred tasks that slipped behind normal tasks, and 19,505 are non-deferred tasks that were further delayed when priority tasks overtook their place.
Fig. 12. Temporal shifting detailed on two regions. Each plot shows the number of deferred tasks (green bars: in, orange bars: out) and the carbonintensity (gray line) in each 5-minute window. Fig. 11. Spatial shift scheduling in a 24 hour window across four regions. Each plot shows the number of active tasks under three shifting policies (blue dashed: Speed-First, orange dashed: Carbon-First, green solid: Spatial Shifting) and the carbon-intensity (gray line with filled color). Two baseline policies follow task arrival pattern while spatial shifting policy schedule more tasks at low-CI window, capped by GPU capacity (65 GPUs per region).
20, and 22-23 UTC; during these windows, the Spatial curve reach close to the capacity limit, indicating continuous full utilization when ERCOT is the cleanest option. This behavior is consistent with the policy objective. From 01-04 UTC in PJM DC (panel d) and from 16-23 UTC for both CAISO North (panel a) and MISO Mason City (panel c), we observe nonzero Spatial activity even when their CI is relatively high. This reflects regional capacity limits: when the cleaner regions (e.g. ERCOT, sometimes CAISO) are saturated yet demand remains high, the scheduler must place some tasks in the next-beset regions to meet SLA constraints. Under temporal shifting, tasks may be deferred to a later start time within the same region when doing so lowers carbon emission while respecting the task’s delay limit. In
Figure 12 visualizes five-minute-bucket deferral activities for two regions with significant carbon-intensity variability. Bars above zero (Deferral-In) count tasks that begin in the bucket due to earlier deferrals; bars below zero (DeferralOut) count tasks deferred out of the bucket at arrival. The blue line shows the region’s carbon-intensity. When carbonintensity drops between consecutive five-minute intervals, we consistently see a surge of Deferral-Out in the higher-CI bucket followed by a surge of Deferral-In in the next lowerCI bucket. This pattern is visible in both ERCOT Austin and PJM DC and is precisely the behavior the temporal policy is designed to induce: shift a portion of arriving tasks across near-term buckets to catch cleaner windows without overwhelming future capacity. Temporal shifting is effective only under strict conditions. First, it needs frequent, sizable CI swings. When the grid is flat, deferring adds queuing risk with trivial carbon benefit. Second, it requires sufficient slack, short, interactive tasks with tight delay limits rarely benefit. Third, it is operationally complex because it depends on carefully tuned guard thresholds, and contention between the priority and normal queues can cause violations.
Decision cost scales with search scope and carbon accounting, but remains negligible. As shown in Table VI, SpeedFirst inspects only the origin’s GPU pools without carbon accounting (0.06 ms/task; 34.9 s total for 600k tasks). CarbonFirst stays local and evaluates predicted CO2 for each GPU type (0.58 ms/task; 346 s). Spatial expands the candidate set to all regions, computing CO2 per region-GPU pair (2.32 ms/task; 1,394 s), which corresponds roughly to a 4× increase over the local-only policies in our case. Temporal remains local but scans near-term windows up to each task’s delay limit (0.88 ms/task; 531 s). Even the slowest policy accounts for less than 0.002% of mean service time. We therefore treat scheduling time as trivial and exclude it from latency/energy/CO2 comparisons. To isolate routing effects from device heterogeneity, we repeat the study with four regions each hosting an identical pool of 65×A100 GPUs, keeping the workload, diurnal curve, delay limits, and scheduling thresholds unchanged. The results are listed in Table VIII. Under this single-GPU configuration, TABLE VIII H ETEROGENEOUS VS S INGLE GPU TEST RESULTS GPU pool/region
Policy
Carbon (lbs)
Energy (MWh)
25xA100, 15xA6000 15xH100, 10xGH200
Speed-First Carbon-First Spatial Temporal
251.45 214.57 154.91 210.83
902.92 768.68 893.80 767.94
65xA100
Speed-First Carbon-First Spatial Temporal
283.38 283.38 205.35 277.35
1,007.83
every task runs on the same device, total energy is identical across policies, and the two baseline policies become the same. The energy increase observed in heterogeneous fleet stems from GPU mix: when the scheduler minimizes CO2 , it sometimes choose cleaner grids on less energy-efficient GPUs. The carbon benefits persist: the Spatial policy lowers CO2 by 27.5% relative to the baseline, while the Temporal policy yields a ∼ 2% reduction. These results suggest that energy increases appear when device efficiency varies and can be constrained by an energy-aware strategy if desired. VI. C ONCLUSIONS AND F UTURE W ORK In this paper, we propose to reduce carbon emissions of AI data centers via carbon-aware siting and scheduling. Using TeleGeography locations aligned with high-resolution marginal carbon-intensity data, we visualize and quantify siting-grid misalignment across the U.S., EU (incl. UK), and Australia and observed that existing and recent built data centers are not preferentially placed in low-carbon-intensity regions. We then design the Carbon-Aware Task Simulator (CATS) and use it to quantitatively evaluate the impact of spatial and temporal shifting on carbon emission reduction. Our results show that spatial shifting yields significant carbon savings: with around 60% tasks shifted, it reduces
CO2 by 28% relative to carbon-first and 38% over speed-first. Temporal shifting reduces 1.7% versus carbon-first and 16.1% against speed-first, deferring around 0.4% of total tasks, but introduced 3.3% SLA violations due to task contentions. Both policies have minimal scheduling overhead (a few milliseconds per task). This work has several limitations that suggest clear directions for future research. First, it relies on historical carbonintensity data; we plan to incorporate short-term forecasts and evaluate robustness to forecast errors. Second, we omit financial costs; a multi-objective formulation that jointly optimize carbon, energy, latency, and financial cost is a natural extension. Third, we assume fixed GPU capacity; enabling elastic capacity in response to arrival patterns is another important next step. R EFERENCES [1] A. Shehabi, A. Hubbard, A. Newkirk, N. Lei, M. A. B. Siddik, B. Holecek, J. Koomey, E. Masanet, D. Sartor et al., “2024 united states data center energy usage report,” 2024. [2] I. Electricity, “Electricity 2024: Analysis and forecast to 2026,” International Energy Agency: Paris, France, 2024. [3] ——, “Electricity 2025: Analysis and forecast to 2027,” International Energy Agency: Paris, France, 2025. [4] ——, “Energy and ai,” International Energy Agency: Paris, France, 2025. [5] Q. Chen, J. Wang, and J. Lin, “Generative ai exacerbates the climate crisis,” Science, vol. 387, no. 6734, pp. 587–587, 2025. [6] P. M. Tagarro, “Forget the future, ai is causing harm now,” Science, vol. 388, no. 6747, pp. 595–595, 2025. [7] S. Luccioni, Y. Jernite, and E. Strubell, “Power hungry processing: Watts driving the cost of ai deployment?” in Proceedings of the 2024 ACM conference on fairness, accountability, and transparency, 2024, pp. 85– 99. [8] E. Erdil, “Optimally allocating compute between inference and training — epoch.ai,” https://epoch.ai/blog/optimally-allocating-computebetween-inference-and-training, [Accessed 31-10-2025]. [9] N. Jegham, M. Abdelatti, L. Elmoubarki, and A. Hendawi, “How hungry is ai? benchmarking energy, water, and carbon footprint of llm inference,” arXiv preprint arXiv:2505.09598, 2025. [10] D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon emissions and large neural network training,” arXiv preprint arXiv:2104.10350, 2021. [11] E. Ayyildiz, B. Yildirim, and N. Aydin, “Location selection methodology for data center with renewable energy integration,” Renewable Energy, p. 123270, 2025. [12] Meta, “2024 Sustainability Report - Meta Sustainability — sustainability.atmeta.com,” https://sustainability.atmeta.com/2024sustainability-report/, [Accessed 20-07-2025]. [13] Amazon, “2024 Amazon Sustainability Report — sustainability.aboutamazon.com,” https://sustainability.aboutamazon.com/2024report, [Accessed 20-07-2025]. [14] Apple, “Apple 2025 Environmental Progress Report,” https://www.apple.com/environment/, [Accessed 20-07-2025]. [15] Microsoft, “2025 Environmental Sustainability Report — Microsoft — microsoft.com,” https://www.microsoft.com/en-us/corporateresponsibility/sustainability/report/, [Accessed 20-07-2025]. [16] “2025 Environmental Report - Google Sustainability — sustainability.google,” https://www.sustainability.google/reports/google2025-environmental-report/, [Accessed 20-07-2025]. [17] I. Schneider and T. Mattia, “Carbon accounting in the cloud: a methodology for allocating emissions across data center users,” arXiv preprint arXiv:2406.09645, 2024. [18] A. Radovanović, R. Koningstein, I. Schneider, B. Chen, A. Duarte, B. Roy, D. Xiao, M. Haridasan, P. Hung, N. Care et al., “Carbonaware computing for datacenters,” IEEE Transactions on Power Systems, vol. 38, no. 2, pp. 1270–1280, 2022.
[19] E. Zhang, D. Wu, and J. Boman, “Carbon-aware workload shifting for mitigating environmental impact of generative ai models,” in 2024 IEEE International Conferences on Internet of Things (iThings) and IEEE Green Computing & Communications (GreenCom) and IEEE Cyber, Physical & Social Computing (CPSCom) and IEEE Smart Data (SmartData) and IEEE Congress on Cybermatics. IEEE, 2024, pp. 446–453. [20] A. Souza, S. Jasoria, B. Chakrabarty, A. Bridgwater, A. Lundberg, F. Skogh, A. Ali-Eldin, D. Irwin, and P. Shenoy, “Casper: Carbonaware scheduling and provisioning for distributed web services,” in Proceedings of the 14th International Green and Sustainable Computing Conference, 2023, pp. 67–73. [21] T. Sukprasert, A. Souza, N. Bashir, D. Irwin, and P. Shenoy, “On the limitations of carbon-aware temporal and spatial workload shifting in the cloud,” in Proceedings of the Nineteenth European Conference on Computer Systems, 2024, pp. 924–941. [22] F. Wang, C. Lv, and J. Xu, “Carbon awareness oriented data center location and configuration: An integrated optimization method,” Energy, vol. 278, p. 127744, 2023. [23] M. Al-Ayyoub, M. Wardat, Y. Jararweh, and A. A. Khreishah, “Optimizing expansion strategies for ultrascale cloud computing data centers,” Simulation Modelling Practice and Theory, vol. 58, pp. 15–29, 2015. [24] B. Acun, B. Lee, F. Kazhamiaka, K. Maeng, U. Gupta, M. Chakkaravarthy, D. Brooks, and C.-J. Wu, “Carbon explorer: A holistic framework for designing carbon aware datacenters,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, 2023, pp. 118–132. [25] L. Lin and A. A. Chien, “Adapting datacenter capacity for greener datacenters and grid,” in Proceedings of the 14th ACM International Conference on Future Energy Systems, 2023, pp. 200–213. [26] M. McMullen and A. P. Wemhoff, “Data center environmental burden reduction through on-site renewable power generation,” ASME Journal of Engineering for Sustainable Buildings and Cities, vol. 5, no. 2, p. 021001, 2024. [27] WattTime, Oct 2022, [Accessed 02-07-2025]. [Online]. Available: https://www.watttime.org/app/uploads/2022/10/WattTime-MOERmodeling-20221004.pdf [28] T. Sukprasert, N. Bashir, A. Souza, D. Irwin, and P. Shenoy, “On the implications of choosing average versus marginal carbon intensity signals on carbon-aware optimizations,” in Proceedings of the 15th ACM International Conference on Future and Sustainable Energy Systems, 2024, pp. 422–427. [29] A. Jagannadharao, N. Beckage, D. Nafus, and S. Chamberlin, “Timeshifting strategies for carbon-efficient long-running large language model training,” Innovations in Systems and Software Engineering, vol. 21, no. 2, pp. 517–531, 2025. [30] M. Chadha, T. Subramanian, E. Arima, M. Gerndt, M. Schulz, and O. Abboud, “Greencourier: Carbon-aware scheduling for serverless functions,” in Proceedings of the 9th International Workshop on Serverless Computing, 2023, pp. 18–23. [31] J. Dodge, T. Prewitt, R. Tachet des Combes, E. Odmark, R. Schwartz, E. Strubell, A. S. Luccioni, N. A. Smith, N. DeCario, and W. Buchanan, “Measuring the carbon intensity of ai in cloud instances,” in Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, 2022, pp. 1877–1894. [32] O. Corradi, “Marginal vs average: which one to use for real-time decisions,” URL https://www. electricitymaps. com/blog/marginal-vsaverage-real-time-decision-making, 2023. [33] J. Overton, “Issue brief— the growth in greenhouse gas emissions from commercial aviation (2019, revised 2022),” 2022. [34] H. Ferreboeuf, F. Berthoud, P. Bihouix, P. Fabre, D. Kaplan, L. Lefèvre et al., “Lean ict: Towards digital sobriety,” Report for the Think Tank The Shift Project, vol. 6, 2019. [35] B. Everman, T. Villwock, D. Chen, N. Soto, O. Zhang, and Z. Zong, “Evaluating the carbon impact of large language models at the inference stage,” in 2023 IEEE international performance, computing, and communications conference (IPCCC). IEEE, 2023, pp. 150–157. [36] Z. Liu, M. Lin, A. Wierman, S. H. Low, and L. L. Andrew, “Greening geographical load balancing,” ACM SIGMETRICS Performance Evaluation Review, vol. 39, no. 1, pp. 193–204, 2011. [37] A. Qureshi, R. Weber, H. Balakrishnan, J. Guttag, and B. Maggs, “Cutting the electric bill for internet-scale systems,” in Proceedings of the ACM SIGCOMM 2009 conference on Data communication, 2009, pp. 123–134.
[38] A. A. Chien, L. Lin, H. Nguyen, V. Rao, T. Sharma, and R. Wijayawardana, “Reducing the carbon impact of generative ai inference (today and in 2035),” in Proceedings of the 2nd workshop on sustainable computer systems, 2023, pp. 1–7. [39] N. Asadov, V. C. Coroamă, M. Franzil, S. Galantino, and M. Finkbeiner, “Carbon-aware spatio-temporal workload shifting in edge–cloud environments: A review and novel algorithm,” Sustainability, vol. 17, no. 14, p. 6433, 2025. [40] S. Qi, H. Moore, N. Hogade, D. Milojicic, C. Bash, and S. Pasricha, “Casa: A framework for slo-and carbon-aware autoscaling and scheduling in serverless cloud computing,” in 2024 IEEE 15th International Green and Sustainable Computing Conference (IGSC). IEEE, 2024, pp. 1–6. [41] Y. G. Kim, U. Gupta, A. McCrabb, Y. Son, V. Bertacco, D. Brooks, and C.-J. Wu, “Greenscale: Carbon-aware systems for edge computing,” arXiv preprint arXiv:2304.00404, 2023. [42] R. N. Calheiros, R. Ranjan, A. Beloglazov, C. A. De Rose, and R. Buyya, “Cloudsim: a toolkit for modeling and simulation of cloud computing environments and evaluation of resource provisioning algorithms,” Software: Practice and experience, vol. 41, no. 1, pp. 23–50, 2011. [43] TeleGeography. [Online]. Available: https://www.cloudinfrastructuremap.com/ [44] “Geoapify Location Platform: Maps, Geocoding, Routing, and APIs — geoapify.com,” https://www.geoapify.com/ , [Accessed 07-10-2025]. [45] Q. Weng, W. Xiao, Y. Yu, W. Wang, C. Wang, J. He, Y. Li, L. Zhang, W. Lin, and Y. Ding, “{MLaaS} in the wild: Workload analysis and scheduling in {Large-Scale} heterogeneous {GPU} clusters,” in 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22), 2022, pp. 945–960. [46] Y. Wang, Y. Chen, Z. Li, X. Kang, Y. Fang, Y. Zhou, Y. Zheng, Z. Tang, X. He, R. Guo et al., “Burstgpt: A real-world workload dataset to optimize llm serving systems,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, 2025, pp. 5831–5841. [47] Y. Xiang, X. Li, K. Qian, W. Yu, E. Zhai, and X. Jin, “Servegen: Workload characterization and generation of large language model serving in production,” arXiv preprint arXiv:2505.09999, 2025. [48] O. J. Boxma, J. Cohen, and N. Huffels, “Approximations of the mean waiting time in an m/g/s queueing system,” Operations research, vol. 27, no. 6, pp. 1115–1127, 1979. [49] M. Harchol-Balter, Performance modeling and design of computer systems: queueing theory in action. Cambridge University Press, 2013.
Dayuan Chen received his M.S. in Computer Science from University of Texas at Dallas in 2021. He is now a Ph.D. student at Texas State University. His research focuses on sustainable AI development, including efficient and carbon-aware Large Models fine-tuning and inference scheduling.
Ziliang Zong received his Ph.D. degree from Auburn University and is a Professor of the Computer Science Department at Texas State University. His research focuses on Energy-Efficient Computing and Systems, including Green Software Design, Green AI, Green Cloud, and Green Data Center. He served as a member of the Green Software Foundation, Associate Editor of the Sustainable Computing Journal, co-chairs and committee members of numerous conferences and workshops in high performance computing, green computing, cloud computing, and edge computing.