Energy-Aware Computing in the Year 2026 Roblex Nana Tchakoutea,∗,1 , Claude Tadonkia a Mines Paris - PSL University
arXiv:2605.24569v1 [cs.DC] 23 May 2026
Centre de Recherche en Informatique, 35 rue Saint-Honoré, Fontainebleau, 77300, France
ARTICLE INFO
ABSTRACT
Keywords: Energy-Aware Computing Cloud-Edge Continuum High-Performance Computing (HPC) Green AI Carbon Footprint Power Capping Job Scheduling Energy Efficiency
High-Performance Computing (HPC) has recently entered the Exascale era, and considerable efforts are underway to fully harness this power for large-scale applications, such as cutting-edge generative AI (training and exploitation). The corresponding energy consumption is very high, and forecasts are alarming, making this metric a critical systemic bottleneck. Addressing this issue presents a genuine challenge for the entire cloud-edge-HPC continuum at all scales, from low-power IoT microcontrollers to multi-megawatt data centers. Beyond financial costs, green computing is driven by considerations related to climate change and environmental concerns such as carbon footprint (𝐶𝑂2 𝑒), as well as constraints on energy production and supply, leading to a real need to regulate information and communication technology (ICT) activities. This article presents a comprehensive overview of energy-efficient computing, incorporating the most recent and significant contributions. Based on this exploration of the state of the art, we design and describe a holistic taxonomy of the aforementioned publications, structured around various perspectives, including hardware and software aspects, measurement instrumentation, software optimizations, dynamic task scheduling, voltage scaling, workload consolidation, federated learning, and cooling. Particular emphasis is placed on large-scale AI, which receives significant attention due to its considerable resource requirements. We conclude with an analysis of a forward-looking roadmap that considers the main perspectives of sustainable computing.
1. Introduction The trajectory of computer architecture has long been shaped by two foundational principles: Moore’s Law, which predicted the doubling of transistor density every two years, and Dennard Scaling, which ensured that total chip power remained roughly constant as transistors shrank. Together, these laws underpinned decades of increasing performance gains at manageable energy costs, which is no longer the case due to the fundamental power wall that now governs the design and operation of modern computing infrastructures. In addition to theoretical peak performance, an important focus is now on energy production/supply, heat dissipation, and environmental sustainability [1]. The Exascale era, currently led by landmark systems including El Capitan (≈30 MW), Frontier (≈25 MW), Aurora (≈39 MW), JUPITER Booster (≈16 MW) [2], has pushed facility power draws to unprecedented levels, making electricity and cooling costs the dominant factor in total cost of ownership (TCO). At this scale, even marginal improvements in energy efficiency translate directly into significant financial savings and a noticeable reduction in carbon emissions. The emergence of Generative Artificial Intelligence (GenAI) and Large Language Models (LLMs) has significantly exacerbated the (projected) energy crisis [3]. Training trillion-parameter models requires large-scale clusters that run continuously for several months. High-performance AI is currently powered by high-end accelerators with thermal design powers (TDPs) greater than 1000W per chip (e.g., NVIDIA Blackwell, AMD MI300X) [4, 5]. Thus, mitigating ∗ Corresponding author
∗∗ Principal corresponding author
ORCID (s): 0009-0009-8636-8466 (R.N. Tchakoute)
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
the energy requirement and the carbon footprint associated with large-scale AI has become a top priority [6]. Intelligent applications are increasingly migrating toward the "Edge" in order to reduce latency. This introduces a two-way energy context: embedded systems and the Internet of Things (IoT). Here, energy is a critical resource. Microcontrollers and embedded AI accelerators (e.g., Coral Dev Boards, Nvidia Jetson) must run inference tasks within milliwatt-scale power limits in order to cope with the limited energy budget, since they are battery-powered. Thus, modern computing ecosystems must be evaluated as a unified Cloud-Edge-HPC continuum. Historically, power efficiency was hardware-centric. However, achieving true sustainability requires a cross-stack approach that combines hardware optimizations with advanced system-level mechanisms such as energy-aware job scheduling, dynamic resource allocation, and intelligent power capping [7]. Evaluating Sustainable AI inherently requires balancing its predictive capabilities with the ecological consequences of its carbon footprint [8, 9]. Recent meta-synthesis studies have demonstrated a crucial need for sustainability-related investigations, and most qualitative surveys have focused on the suitability of all-inclusive approaches in which all energy-accountable hardware and software aspects are analyzed as a whole [10]. Concerns related to the energy consumption of IT activities can/should be considered from a broader perspective, including the three main phases: energy production, ITrelated activities, and disposal/recycling, as illustrated in Figure 1. Energy-efficient computing, in all its aspects and paradigms, is associated to the intermediate phase only, and this survey describes the landscape of related contributions.
Page 1 of 26
Energy-Aware Computing in 2026
Figure 1: The whole cycle of energy consumption and related consequences.
This paper aims to provide a structured and commented view of the landscape of the major scientific and technical contributions within the state-of-the-art of power-aware computing in the Computing continuum. We provide and articulate the following viewpoints: • Comprehensive Taxonomy: An up-to-date taxonomy of energy-related studies spanning edge microcontrollers, cloud computing infrastructures, AI accelerators, and Exascale supercomputers. • Cross-Stack Profiling: A comparative analysis of hardware and software energy profiling tools, distinguishing between in-band estimations and out-ofband measurements. • Code Optimization: A review of static and dynamic energy optimization strategies, with an emphasis on job scheduling, power capping, and workload consolidation. • Sustainable AI: An explicit focus on the carbon footprint of cutting-edge/large-scale AI through the review of frameworks for tracking and mitigation of energy-related metrics associated to LLM training and inference. • Cooling and conventional rules: An evaluation of liquid cooling technologies and a prospective analysis of conventional rules and technical challenges related to sustainable computing.
chose an objective filtering pipeline that covers the full Cloud-to-Edge continuum, integrating chronological bibliometric dynamics and high-impact industrial reports.
2.1. Popularity-Driven Filtering and Selection Our starting pool is a set of highly cited papers that explicitly focus on sustainable hpc-cloud-Edge computing, cutting-edge AI workloads, power capping, and hardware telemetry. Among these contributions, we selected a corpus of roughly 250 pivotal references, prioritizing the most recent ones (preferably post-2023). We applied the following two rules: • Journal/Conference Impact: We consider mainly Q1 journal papers (via SCImago/JCR) and CORE A/A* conferences (e.g., SC, ISCA, MICRO, ASPLOS, EuroSys), the ranking being considered at the publication date. • Empirical Validation Rule: Regarding works on theoretical energy models or optimization frameworks, we have considered those that provide empirical evaluation on real devices or validated simulators.
2.2. Integration of High-Impact Grey Literature In AI and large-scale HPC, the state-of-the-art is usually established by practitioners before being considered (in a more formal way ) in academic journals. Following this observation, we selected a set of non-academic literature:
2. Selection and Methodology
• Industrial Reports: Technical documentation, architectural whitepapers, and carbon footprint reports released by leading vendors (e.g., NVIDIA’s Blackwell specifications [4], Google’s TPU architectures [11], and Meta’s open-source LLM reports [3]).
To ensure a rigorous and reproducible overview of the state of the art, we adopt a multi-stage hybrid Systematic Literature Review (SLR) methodology. Given the rapid evolution of AI hardware, relying solely on data from classical databases might lead to outdated configurations. Thus, we
• Institutional Frameworks: Macro-environmental data and reports from governmental institutions, including data center forecasts of the International Energy Agency (IEA) [1] and objectives of the Paris Agreement [6].
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 2 of 26
Energy-Aware Computing in 2026
Figure 2: The landscape of energy-aware computing in the Cloud-Edge-HPC continuum.
2.3. Overview of the Major Selected Contributions We made a specific methodological choice regarding the selection of the papers for our corpus as previously described. We acknowledge that this choice may have resulted in the omission of genuine/valuable/valued contributions (less recent ones, books, blogs, conferences/journals with unmatched rank, open archives). The reader should be aware of this deliberate choice that we made. However, we think that it is highly likely that most of the aforementioned papers are cited in those selected for this survey.
Cloud Computing (46), Green Computing (41), Power Demand (27), Energy-Aware Scheduling (26), Data Centers (26), and Power Management (25). Figure 3 provides a visual summary of the keywords according to their magnitude.
2.4. Keyword-based Analysis of our corpus To identify the dominant research themes within the studied corpus, we analyzed the keywords provided by the authors. After processing and normalizing the extracted keywords, we obtained 2,133 terms with redundancies and 949 unique terms (45%). We also considered the case of semantic equivalence. This aggregation revealed that Energy Efficiency was the most prominent topic (99 occurrences), followed by Energy Consumption (77), Sustainability (67), and Carbon Footprint (61). Other recurrent themes included R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Figure 3: Panorama of the keywords from our corpus.
Page 3 of 26
Energy-Aware Computing in 2026 Topic
References
Profiling methodology Profiling tools Monitoring tools Hardware design
[12];[13];[14], [15], [16], [17]
Green software/OS & Code Power/perf predictive modeling
Job scheduling & VM placement
Others optimization approaches
Cooling Technologies Special case of AI
[15], [18], [19], [20], [21] [22], [23], [24] [25], [26], [27], [28], [29], [30], [31], [32], [33], [34], [35], [36], [37], [38], [39], [40], [41], [42], [43], [44], [45], [46], [47] [48], [49], [50], [51], [52], [53], [54], [55], [56], [57], [10], [58], [59], [60], [61], [62], [63], [64], [65], [66], [67], [68], [69], [70] [71], [24], [72], [73], [74], [75], [16], [76], [77], [78], [79], [80], [81], [82], [83], [34], [75], [74], [35], [36], [84], [85], [86], [87], [88], [89] [90], [91], [92], [78], [93], [94], [95], [96], [97], [98], [99], [100], [101], [102], [103], [104], [105], [106], [107], [108], [109], [110], [111], [112], [113], [114], [115], [116], [117], [118], [119], [120], [121], [122], [123], [124], [125], [126], [127], for AI ([128], [129], [130]) [131], [132], [133], [134], [135], [136], [137], [138], [139], [140], [141], [142], [143], [144], [145], [146], [147], [148], [149] [76], [40], [150], [134], [151], [152], [85], [107], [153], [154] [155] [156], [157], [15], [30], [50], [51], [31], [158], [159], [160], [9], [161], [162], [163], [164], [165], [166], [167], [168], [84], [169], [161], [170], [171], [172], [173], [87], [51], [42], [174], [46]
Table 1 Overview of Recent Energy-Aware Computing Studies
3. Main Energy-Related Metrics To optimize energy across the Cloud-Edge continuum, researchers rely on fine-grain metrics besides the classical electrical power (Watts) to encompass thermal dissipation, water consumption, and carbon emissions (over the whole life cycle). This extended focus implies adapting traditional HPC benchmarking so as to consider large-scale applications like those from Generative AI and large-scale infrastructures like Exascale machines.
3.1. Key Metrics for Large-Scale Infrastructures In extreme-scale computing, facility-level power consumption is the main critical ceiling. The November 2025 TOP500 list illustrates this unprecedented energy budget: El Capitan system achieves 1.809 ExaFlops drawing 29.68 Megawatts (MW) [2], while JUPITER Booster (the first European Exascale system) achieves 1.000 ExaFlops at a noteworthy 63 GFlops/Watt. Table 2 provides sample data related to processing performance and energy cost, which R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Table 2 Estimated Hourly Electricity Cost of Exascale Supercomputers (Based on Top500 Data - Nov. 2025) Machine EL CAPITAN FRONTIER AURORA JUPITER
Host Region USA (California) USA (Tennessee) USA (Illinois) Germany (Jülich)
Sustained Perf. 1.809 EFLOPS 1.353 EFLOPS 1.012 EFLOPS 1.000 EFLOPS
Power (MW) 29.68 MW 22.78 MW 38.69 MW 15.87 MW
$ / kwh 0.148 0.148 0.148 0.285
Cost / Hour $4,392 $3,371 $5,726 $4,522
illustrate the significant operational expenditure (OpEx) that justifies the need for energy efficiency [175, 176]. The total hourly cost is calculated using the average price per kWh within [2023;2026] [177]. Thermal Design Power (TDP), which now routinely exceeds 1000W for AI chips such as the NVIDIA B200 [4], governs the magnitude of cooling requirements. At the individual facility level, Power Usage Effectiveness (PUE) remains the standard [178], but is no longer sufficient from the isolation point of view. The shift toward extreme-density AI clusters requires a huge amount of chilled water, which promotes Water Usage Effectiveness (WUE) to the rank of a critical metric. In water-scarce regions, minimizing the volume of water evaporated per kilowatt-hour (L/kWh) is currently prioritized even above PUE [179]. Furthermore, optimizing modern data centers requires auditing their direct supply architecture; Power Supply Efficiency (PS-Loss) captures the resistance conversion of Uninterruptible Power Supplies (UPS) [142], where active load-balancing can reduce lifecycle expenditures independently of IT scheduling [143, 98].
3.2. AI-Related Energy-Carbon Data The traditional FLOPS metric is no longer the only focus when it comes to large-scale AI workloads. The complexity analysis of GenAI considers specific metrics such as Energy per Token (Joules/Token) for inference, and Training Throughput Efficiency (Tokens/Second/Watt) for training. Hybrid benchmarks such as Carbon-Aware Accuracy (Carburacy) evaluate hyperparameter sweeps simultaneously across cognitive effectiveness (F1/BLEU) and eco-sustainability curves, identifying models that achieve competitive accuracy at the lowest 𝐶𝑂2 cost [171]. Mostly focusing on raw power consumption obscures the environmental impact, which is quantified through the operational Carbon Footprint represented by the Carbon Intensity (𝑘𝑔 𝐶𝑂2 𝑒∕𝑘𝑊 ℎ). Table 3 provides the hourly footprints of the latest Exascale architectures based on regional grid intensities [180]. This dichotomy creates distinct contextual optimization scenarios: Cap-Ex Power Limits (maximizing FLOPS/W under strict utility caps), Renewable-Abundance (focusing purely on time-to-solution when powered by 100% renewables, this is the case of the LUMI system), and GridAware Migration (report or migration of cloud-based AI training for sustainability consideration).
Page 4 of 26
Energy-Aware Computing in 2026 Table 3 Estimated Hourly 𝐶𝑂2 Footprint of Leading Supercomputers Machine Power Grid Power (MW) 𝑘𝑔(𝐶𝑂2 𝑒) / kWh Total 𝑘𝑔(𝐶𝑂2 𝑒) / Hr EL CAPITAN (USA) Mixed Renewable 29.68 MW 0.220 ∼ 6, 529 FRONTIER (USA) Regional Average 22.78 MW 0.360 ∼ 8, 200 AURORA (USA) High Fossil 38.69 MW 0.320 ∼ 12, 380 JUPITER (DE) Mixed European 15.87 MW 0.380 ∼ 6, 030 LUMI (FI)* 100% Hydro/Nuclear 7.10 MW 0.050* ∼ 355 *LUMI operates on nearly 100% renewable energy; its lowest carbon footprint illustrates the impact of the carbon factor.
4. Architectures and Energy Trade-offs The end of Dennard scaling has led to the emergence of highly specialized architectures within the hardware landscape. Performance is no longer bounded by transistor density, but rather by power delivery and heat dissipation. An energy-aware computing investigation requires examining processors across the continuum: from battery-powered embedded systems to power-hungry (megawatt-scale) processors [181].
4.1. The Edge and IoT: Milliwatt-Scale Inference Within the edge ecosystem (remote sensors, drones, robotics), devices strictly target space, wattage, and performance (SWaP) envelopes. Ultra-low-power microcontrollers (MCUs) natively utilize ratio-based energy pre-assignment to distribute micro-joules proportionally to the priority level of the tasks. In energy-aware systems, schedulers dynamically adapt to intermittent RF/solar availability. By embedding ultra-lightweight adaptive checkpoints into NonVolatile RAM (NVRAM) a few milliseconds before voltage failure, Edge devices process Deep Neural Networks via transient energy recovery without having to restart [182]. In order to maximize inference-per-watt, edge platforms leverage Coarse-Grained Reconfigurable Architectures (CGRAs) and approximate computing. By modeling hardware error as an acceptable "additive noise", compilers schedule the execution of some blocks in a non-deterministic way on specialized co-processors, thus sacrificing numerical precision constraints in favor of energy reduction [33, 36]. As edge computing is subject to an increasing demand (e.g., Raspberry Pi 5 has a power peak of 15W), SingleBoard Computers tend to rely on Heterogeneous Multicore Processors (big.LITTLE architectures). In this context, non-stationary workloads require predictive learning-based orchestration [75] and decentralized power-management, thus allowing local components to autonomously regulate operating states [41, 43, 45]. A specific case of extreme-edge computing is that of Low Earth Orbit (LEO) satellites, which suffer from high operational thermal throttling. To bypass the limited solar energy budget, LEO architectures employ Layer-Wise CoInference Scheduling, splitting the neural network layers across satellite radio links and offloading tensor computations on Earth-bound servers [183, 129]. Furthermore, processing software updates on these decentralized sensors requires microservice containerization, yet traditional abstractions like Docker are well-known for being detrimental w.r.t power-constrained silicon. Compiling frameworks that directly target bare-metal devices (e.g., ONNX) R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
or that use low-memory WebAssembly (WASM) yields significantly efficient on-device inference [184]. Finally, as Edge AI pushes Multi-Processor System-on-Chips (MPSoCs) to their thermal limits, internal Network-on-Chip (NoC) wiring becomes a primary energy drain. State-ofthe-art NoC schedulers (e.g., IntelliNoC) deploy continuous Q-learning agents natively on silicon to apply fine-grained DVFS, thus autonomously shutting down hardware routing components to prevent localized overheating [96, 185, 26, 135].
4.2. Cloud Infrastructures and Network Overheads At the data center level, the energy battle pits traditional x86 architectures against ARM-based hyperscaler chips. To cope with the highly increasing power penalty of highclock frequencies, x86 architectures (e.g., Intel Xeon 6 Sierra Forest, AMD EPYC Turin) follow the path of heterogeneous design, deploying up to 288 Efficiency Cores (E-cores) per socket [186, 187]. Although highly efficient at the core level, these dense sockets carry 500W TDPs, thus testing the air cooling system. In contrast, driven by the RIKEN Fugaku supercomputer, cloud providers increasingly deploy custom ARM RISC silicon (AWS Graviton4, NVIDIA Grace). By eliminating legacy decode overheads and integrating LPDDR5X memory, ARM architectures maximize predictable performance-per-Watt [188]. Beyond the basic processors, the internal infrastructure of the Cloud represents a considerable energy vector. Current workflows tend to generate massive "Coflows." To reduce DCN power waste, server-centric networks and Software-Defined Networking (SDN) controllers dynamically map traffic dependencies, intentionally collapsing idle optical paths and packet-switching routing tables so as to minimize the networking footprint by up to 40% without detrimental packet losses [79].
4.3. AI Super-Accelerators and Exascale Horizons LLM training is driving a noteworthy shift in hardware design. While GPUs in data centers show an operational TDP of 300W, modern reticle-limit devices (e.g., NVIDIA Blackwell B200, AMD MI300X) draw from 1000W to 1200W [4]. Evaluating raw TDP hides a vulnerability induced by Transformers: Voltage Emergencies (𝑑𝑖∕𝑑𝑡). The highly synchronous, bursty execution of Attention layers currently creates unprecedented microsecond spikes, thus destabilizing power delivery networks and forcing a conservative static guard-band throttling that degrades the overall energy-to-solution margin. Domain-specific accelerators (e.g., Google TPUs) address this through fine-grained gating, selectively powering down inactive functional units cycle-by-cycle [189, 11]. Novel accelerator architecture such as GCNAX, which tailors the compute engine, buffer structure, and size based on specific Graph convolutional neural networks (GCNs) dataflow, was proposed [47]. Advances in interconnects (e.g., NVLink) optimize massive die-to-die tensor communication [29, 30, 190], while granular voltage scaling at the Page 5 of 26
Energy-Aware Computing in 2026 Table 4 Current Leading AI Accelerators (Training & Inference Engines - 2025/2026) Accelerator NVIDIA B300 (Ultra)
TDP (W) 1400 W
Peak Precision FP4, FP6, INT8
Memory 288GB HBM3e
AMD MI400X
1600 W
FP4, FP8, FP16
432GB HBM4
Google TPU v7
157 W
BF16, FP8, INT8
192GB HBM3e
Intel Gaudi 3
900 W
BF16, FP8
128GB HBM2e
Noteworthy Energy Features Dynamic tensor scaling; standardizes liquid cooling Advanced multi-die packaging isolates logic from I/O Systolic arrays minimize logic decode footprint Integrated RoCE networking eliminates NIC overhead
level of individual computational operators allows precise adaptation to workload characteristics [86, 191]. Looking toward 2026 and post-Exascale deployments, systems like Alice Recoque (AMD MI430X/SiPearl Rhea2 - France) and Discovery (ORNL - USA) rely on data-driven operational analytics to strictly enforce facility energy budgets [192, 193]. To power these cloud systems, accelerators have evolved towards "Rack-Scale" paradigms (e.g., the 120 kW NVIDIA GB300 NVL72). Hardware selection requires dealing with a compromise between energy efficiency, precision (FP4 natively drastically reduces energy waste), and software lock-in (ASICs provide superior FLOPS/W but restrict developers to specific compiler stacks).
4.4. Quantitative Overview of Hardware Trends
% of Surveyed Papers
A statistical analysis of recent literature in computer hardware (𝑁 = 43 selected papers) reveals a decisive shift away from traditional CPU. Approximately 42% of recent research works are dedicated to AI Super-Accelerators (e.g., GPUs, TPUs, Processing-in-Memory, and Wafer-Scale architectures), thus reflecting the industrial response to thermal concerns when processing large-scale workloads like those related to LLMs. The Edge and IoT continuum accounts for 33% of the literature, with a major emphasis on Sub-Watt NPUs and Battery-Free RF-harvesting chips. General-purpose cloud CPUs, Chiplets (UCIe), and network interconnects cover the remaining 25%. This distribution empirically confirms that the end of Dennard Scaling pushed toward prioritizing extreme-edge and extreme-accelerator heterogeneity over general-purpose silicon.
42 40
33 25
20 0 Accelerators
Edge & IoT CPUs & Networks
Figure 4: Distribution of recent hardware-related energy-aware studies across the computing continuum.
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
5. Energy and Carbon Profiling Tools Dynamic energy optimization strategies across the CloudEdge-HPC continuum require real-time, fine-grained telemetry. As hardware diversity expands, the abstraction gap between software workflow and hardware profile is widening. The profiling ecosystem can be categorized into three paradigms: silicon-level middleware, institutional facility monitoring, and cloud-level carbon estimators.
5.1. Silicon Sensors and Abstraction Middleware Modern processors incorporate native power sensors to manage power-related aspects. On bleeding-edge 3nm FinFET nodes, On-Die Power Sensors (ODPS) use machine learning calibration to basically limit sub-50𝜇s over-current excursions before passing abstractions to the OS [194]. To make these raw registers available to developers, middleware tools consider unified interfaces for machine-specific MSRs. Furthermore, execution frameworks that take parallelization into account (e.g., PARMA) model the hidden parallelism of dynamic workloads on the fly, thus adjusting core voltage domains to improve Energy-Delay-Product (EDP) by nearly 50% without invasive code instrumentation [73]. Hardware-Accelerator Telemetry: The NVIDIA Management Library (NVML) is a low-level API that provides instantaneous access to power-related values and programmatic power capping. However, recent profiling work reveals critical flaws: tools like nvidia-smi sample only ∼25% of actual runtime on heavily burdened A100/H100 GPUs, systematically underestimating total job energy due to clock frequency aliasing [13]. CPU Telemetry (RAPL): Intel’s RAPL and AMD’s APM stand as the reference of in-band profiling [195]. Beyond structural omissions, generic hardware limiters are "application-blind". When a mixed-tenancy node reaches an artificial power cap, RAPL indiscriminately throttles the fastest-running cores, thus aggressively stifling latencycritical high-priority applications as well as demanding software-controlled "Per-Application Power Shares" [196, 197]. Unified Middleware: LLNL’s Variorum [198] provides a vendor-agnostic C-API to dynamically cap power consumption across heterogeneous nodes. Similarly, the ALUMET energy-measurement framework [12, 199] provides an extensible, asynchronous Rust-based pipeline designed to minimize the intrusive overhead of high-frequency RAPL polling. Built upon LIKWID [20], the EE-HPC energy-related framework [23] provides hybrid HPC centers with job-specific monitoring, pairing phase detection with MPI/ OpenMP tuning. Data-Dependent Power Fluctuations: Macroscopic static modeling tends to be chronically inaccurate due to a strong dependency on static inputs. Benchmarking proves that feeding functionally dissimilar vectors (e.g., vision inputs vs. sparse tensors) diverts the GPU instruction branching, thus generating wild internal silicon footprints even when raw FLOPS remain identical [200, 201]. Similarly, comparative database profiling (MongoDB vs. PostgreSQL) Page 6 of 26
Energy-Aware Computing in 2026
5.2. Institutional-level Facility Monitoring To avoid software overloads, national power grid initiatives use out-of-band physical architectures. Platforms like D.A.V.I.D.E. integrate hardware sensors directly into compute backplanes for high-frequency telemetry without context-switching overheads [40]. Facility monitoring is mostly based on the Prometheus/Grafana stack and TimeSeries Databases (TSDBs). Deployment examples include CEEMS (Jean Zay supercomputer), which scrapes metrics at 10-second intervals so as to cross-reference jobs with the French carbon factors [22]; and Grid’5000 (Kwollect), which synchronizes out-of-band wattmeters with in-band RAPL data. High Definition Energy Efficiency Monitoring (HDEEM) uses FPGA probes and resistive PCIe shunts to capture transient micro-spikes that cannot be caught up by RAPL [14]. However, running 100,000 GPUs generates terabytes of TSDB metrics daily, thus creating a "Big Data" I/O bottleneck.
5.3. Virtualization Wall and Carbon Estimation Public cloud providers restrict access to hardware telemetry because of side-channel threats. Power-consumption alerts can be exploited to infer sensitive system behavior [32, 204]. The GateBleed attack exposed the fact that hardware power-gating recovery delays allow malicious intruders to reconstruct hidden Transformer routing paths and proprietary datasets, thus proving that telemetry breaks privacy [39]. R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
In such a constrained context, direct measurement is replaced by analytical estimation via Performance Monitoring Counters (PMCs) [89]. Lightweight Python libraries (e.g, CodeCarbon, Eco2AI) fall back to analytical estimation when hardware registers are unavailable. Recent methodologies utilize LLMs to predict power variations and drive HPC scheduling directly [78]. To handle operational uncertainty of the electrical grid, frameworks like U-DUCT introduce information-gap theory to sweep server uncertainty bounds without assuming precise power distributions [105]. For the sake of consistency, container-aware protocols like Phoeni6 ensure transparent and reproducible neural network benchmarking across different clusters [17, 205]. Going beyond passive measurement, Zeus [206] dynamically navigates the Pareto frontier between DNN training time and energy, thus actively adjusting GPU batch sizes at runtime.
5.4. Trends on Measurement & Profiling The evaluation of profiling methodologies across our corpus of 33 articles reveals the considerable impact of the "virtualization wall". While previous research studies heavily prioritized physical out-of-band monitoring, only 21% of post-2023 papers focus on hardware shunts or TSDB facility monitors. In addition, 45% of the literature now strictly evaluates Software-Level Carbon Estimators and Machine Learning predictive telemetry (e.g., LLM-based power prediction and CodeCarbon). In-band silicon sensors and abstraction middleware (e.g., RAPL, eBPF, ALUMET) belong to the remaining 34%. This proves a pivotal softwareengineering observation: because cloud hyperscalers prohibit hardware telemetry for security, researchers are increasingly pushed towards purely virtual and predictive carbon estimation. % of Surveyed Papers
illustrates that identical operational schemes can exhibit severe energy divergence based on the underlying query logic [202]. The Observer Effect: The middleware query relies on the execution of monitoring threads by the CPU. Querying MSRs above 100 Hz interrupts the application threads, pollutes the memory caches, and consumes some power, thus affecting the workload’s efficiency. To cope with this, state-of-the-art predictive techniques leverage NLP-adapted Deep Learning approaches. Frameworks like DeepPM autovectorize compiler-generated binaries into Transformer models in order to forecast power consumption, thus bypassing hardware execution [72]. Similarly, EffiCast utilizes sequence-based LSTM/Transformers to resolve dynamic L3-cache contention and project microsecond-latency V/f scaling without static inputs [82]. Standardized ML Telemetry & Edge AI: To deeply figure out measurement inconsistencies, the MLPerf Power methodology has been standardized [203]. This was provided through documenting over 1,800 reproducible profiles across 60 systems [160]. The CARAML suite similarly unifies assessments across different Exascale architectures [162, 174]. Scaling estimation to multi-node domains introduces some profiling bottlenecks; Exascale systems now deploy Real-Time AI-Powered Monitoring via neural classification models on peripheral nodes to extrapolate and forecast power deviations, thus preventing thermal faults autonomously [24, 21, 87].
60 45 40
34 21
20 0 Estimators
Silicon & eBPF Out-of-Band
Figure 5: Overview of energy measurement methodologies.
6. Software-Level Energy Optimization While high-performance chips define minimum energy consumption, a system’s carbon footprint is largely determined by its software. Thus, this carbon footprint can only be mitigated through advanced and comprehensive software optimization: from code compilation to operating system power management, including distributed coarse-grain scheduling.
Page 7 of 26
Energy-Aware Computing in 2026
6.1. Algorithmic and Code-Level Optimizations
conclude that there is a frequent failure to optimize transient memory thresholds, thus producing solutions that draw significantly higher power consumption than manual refactoring [49, 55]. UI/UX and Thread Governance: The Software Engineering community agrees on the significant power burden of executing heavy interfaces at the system-level; "PersonaBased UI/UX" avoids continuous screen-refresh operations in order to respect the power limits of the device [67, 108, 68]. On MPSoCs with heterogeneous edges, interfaces like HARP distribute asynchronous threads following a specific scheduling algorithm [210]. Energy optimization on large-scale computing clusters requires platform-specific co-tuning; JiT (Just in Time) compiled frameworks such as JAX have been shown to exhibit deep operational instability when forced to the lowest dynamic frequencies, which does not appear to be the case with TensorFlow [211].
Modern design mandates "energy clarity", thus promoting formal energy interfaces that provide maximum power consumption estimates before the integrated circuit is put into operation [207]. Advanced software engineering incorporates model-driven engineering (e.g., RADIANCE+) to prototype abstract algorithmic configurations and extract power consumption estimates before compilation [69, 108]. Data-Oriented ETL and DataFrames: Preprocessing represents a significant and continuous source of energy consumption. Evaluation of data management frameworks such as Pandas, Vaex, and Dask reveals that multithreaded memory access graphs (Dask) optimize parallelizable filtering, while lazy evaluation processors (Vaex) minimize active RAM consumption during transpositions. In distributed NoSQL ledgers (e.g., Cassandra), adjusting "Consistency Levels" fundamentally acts as a physical power multiplier, mitigating redundant network transmissions [208]. Bottomup profiling successfully parameterizes complex relational joins across distributed databases to predict computationally heavy query pathways [202]. Tiling and Precision Scaling: Source-code loop transformations such as tiling and fusion aim to reduce cache misses and redundant memory accesses. Energy-Aware tiling structurally forces auto-tuning compilers to prioritize power grids that constrain GPU register-leakage rather than raw execution latency [147]. Precision scaling (FP32 to INT8/FP4) decreases memory bandwidth requirements and ALU power activation. Frameworks and JVMs: Empirical studies show different DL frameworks exhibit measurable efficiency discrepancy [50, 51, 209]. At runtime, models compiled with TensorFlow frequently exhibit lower memory-related costs than those obtained with PyTorch, while JVM configurations aggressively shift static consumption from the host [53, 18]. Message Passing Power-Overhead: MPI processes consume dynamic power during wait times related to communication operations. The (COUNTDOWN) tool switches active CPUs into ultra-low ACPI states during these wait times, thus recovering up to 37% of power at the node level [52]. Furthermore, Green Linear Algebra (GreenLA) restructures the core semantic of the BLAS to minimize VRAM accesses, thus reducing power consumption by 13 [60]. Automatic kernel optimization through auto-tuning identifies optimal frequency and memory constraints before execution [83]. The lack of parallelization of the processor’s data loaders could lead to underloading of the energy-intensive active accelerators, thus increasing the number of joules per epoch. Energy Leaks in AI-Based Systems: The analysis of code repositories reveals pervasive power leaks (suboptimal API configurations). The lack of (efficient) parallelization of the processor’s data loaders could lead to underloaded power-hungry (1000MW) accelerators, thus increasing Joules-per-epoch. [64, 66]. Furthermore, evaluations of LLM-assisted coding (with GitHub Copilot, for instance)
At runtime, the operating system manages electrical physics in real-time using three main mechanisms: Dynamic Power Capping, Workload Consolidation, and CarbonAware Scheduling. DVFS and Q-Learning: Dynamic Voltage and Frequency Scaling (DVFS) exploits the available computational headroom. Advanced schedulers handle dynamic frequency management following the model of the NP-hard variablesize bin packing puzzle [94]. At runtime, MPI "ApplicationOblivious" trigger DVFS step-downs automatically [80]. In edge environments running on Linux, Deep Reinforcement Learning (DRL) adaptively orchestrates DVFS states across asynchronous workloads [48]. Q-Learning schedulers continuously build state-action matrices to track the arrival of tasks based on power availability, thus avoiding blockages related to greedy heuristics [116, 106]. Phase-Aware Core/Uncore Scaling: Modern architectures allocate substantial power to "uncore" components (caches, interconnects). Phase-aware approaches explicitly coordinate core/uncore adjustments [213, 57, 145].
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 8 of 26
6.2. Energy-Aware Compilation Traditional compilers (e.g., GCC, LLVM) target performance, but minimizing execution time does not linearly translate into energy minimization [147]. Energy-efficient compilers use instruction scheduling to minimize the Hamming distance between operands (their values), thereby reducing transistor switching. The MLIR framework optimizes tensor graphs based on the thermal profiles of the hardware, by inserting power preventive cut-off instructions. Datasets like Greenlight allow profiling the hardware footprint of TensorFlow API calls, thus noticing unpredictable energy spikes [21]. For a unify consoderation of extreme heterogeneity, SYCL C++ toolchains (e.g., SYnergy) incorporate explicit power limits at compile time, assigning finegrained core/memory scaling limits on a per-kernel basis [138, 51, 212].
6.3. Dynamic Power Management
Energy-Aware Computing in 2026
Job-Level Power Allocation (SLURM/PBS): Facility operators enforce strict power limits [214]. Adaptive strategies combine power capping and runtime profiling [131]. In congested environments, coordinated control mitigates transient peaks [215]. Middleware such as BAR/BARMAN intercepts OpenMP tasks to dynamically redistribute unused micro-power among coexisting applications [216]. Imposing strict limits on dense nodes causes cascades of flow limitation. Multi-CPU/GPU Reinforcement Power-Capping: Imposing strict limits on dense nodes causes a series of limitations. RL frameworks such as PowerCoord actively reallocate power budgets between host CPUs and PCIe accelerators based on phase-prediction [217]. Recent benchmark evaluations of the FRONTIER supercomputer demonstrate that Frequency Capping systematically outperforms raw power-capping on memory-bound applications [218, 148].
6.4. Cloud/Edge Workload Management Since static power leakage pushes hyperscalers to follow a dense stacking strategy for the management of containers, an accurate mathematical calculation of the multi-objective (performance and energy) mapping requires computationally expensive back-end heuristics such as hyperheuristics for Cloud Scheduling Problems (HHCSP) [100]. Virtualization and Learning-Based Orchestration: Energy-efficient cloud operation relies on dynamic workload consolidation [77, 219, 220]. Federated learning techniques enable coordinated scheduling across geo-distributed infrastructures [97, 221]. To address latency and energy overheads [158], DRL optimizes migration decisions [141, 123], with renewable energy-aware geo-distributed job assignment [99]. Microarchitectural innovations enable deeper low-power states with reduced wake-up latency [222, 223]. Orchestrators that use Flower Pollination Optimization have demonstrably reduced carbon emissions by up to 48.60% during Virtual Machine Placements (VMP) [121, 88]. Green Bin Packing & Quality-Aware Consolidation: The online bin packing approach treats volatile VMs through Green Bin Packing heuristics [224]. EcoCore and AgilePkgC intricately co-adjust CPU voltages under the awareness of tail-latency [223, 225]. Topology-Aware approaches seamlessly estimate CPU thermal limits and network switch hopdistances [125, 111, 126]. To solve the NP-hard bin packing problem, hypervisors use evolutionary algorithms that mathematically combine QoS constraints and power consumption limits [124, 226]. Dynamic Cloud/Edge Heuristics: Modern schedulers run complex heuristics such as modified artificial bee colony algorithms, binary search [227], differential evolution [119], and genetic algorithms [114, 228]. Container-based methodologies enforce hypervisor live-migrations to strictly match dynamic network limits [139]. The task scheduling community relies heavily on Particle Swarm Optimization (PSO)
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
[120, 122]. PSO-based multi-objective frameworks explicitly maintain swarm diversity through discrete/quantum variants [127]. By continuously adjusting particle speed, orchestrators quickly resolve VMP matrices (in real-time), performing hot migrations to optimize host storage while preventing SLA violations. Hybrid intelligent swarms (e.g., Cost-Energy Aware Spider Monkey Optimization) perform an exact routing of payloads to facilities experiencing maximum localized underutilization [149]. Accelerators Consolidation & Serverless Cold-Starts: Dynamic thread block migration redistributes workloads across GPUs [61]. Spatial partitioning (MIG slicing) enables multi-tenant execution, but at the cost of noticeable power and thermal interference across logically isolated workloads [229, 230]. Serverless computing, considering Function as a service (FaaS), introduces significant cold-start overheads [231]. Asymmetric multicore architectures assign latencycritical tasks to P-cores and keep-alive mechanisms to Ecores [232]. Antagonism Between SLAs and Sustainability: Flow network algorithms ensure SLA compliance alongside efficient energy management [90, 233]. Cuckoo Search algorithms maximize resource packing, thus preventing violations during DVFS throttling [163]. However, several analyses confirm a deep uncorrelation: almost all ML engineers fail to implement green metrics into production code due to a lack of standardized guidelines [63]. Carbon-Aware Scheduling: Optimizing for carbon requires considering spatio-temporal variability [91]. Temporal shifting postpones flexible workloads to low-carbon periods [101], while spatial routing directs computation where the availability of renewable energy is higher [140, 234]. Multi-agent RL techniques incorporate dynamic signals to optimize trade-offs [98, 156, 110]. Simulation platforms validate scalable abstractions [54], while edge devices dynamically offload workloads based on carbon network conditions [144]. Deliberately pausing/resuming scientific processing based on the network’s marginal carbon intensity (MCI) reduces the carbon footprint by more than 80% [235]. However, optimization exclusively based on MCI has flaws; indeed, missing the evaluation of absolute emission curves leads data centers to unexpected congestion [236]. Hypervisors such as Ecovisor and GreenFlow transmit physical carbon data directly via the API, ensuring that AI containers self-regulate [237, 164, 128]. Hypervisor Benchmarking and Cloud Service Brokerage (CSB): The design of the basic hypervisor (KVM vs. Xen) is inherently subject to power consumption fluctuations [238]. Empirical studies explicitly show the sharp decrease in yields associated with carbon-aware load shifting; frequent temporal shifting leads to tasks generating a significant carbon penalty from the use of intermediate network equipment [239]. For the purpose of standardization, CSBs deploy mathematical pricing frameworks to incentivize tenants to minimize computing peaks [240].
Page 9 of 26
Energy-Aware Computing in 2026
6.5. Time-to-Solution vs. Energy-to-Solution Modern scheduling techniques mainly seek efficiency. Optimizing for energy leads to instant execution of workloads in order to achieve maximum performance efficiency. Conversely, optimizing for carbon leads to delaying workloads following the availability of sustainable grids [146]. Race-to-Halt Paradigm: If a task is CPU-bound, the most energy-efficient strategy is to run at maximum clock frequency so as to finish it quickly and switch to a sleep mode. Stochastic queuing-network frameworks mathematically model cloud VM arrival rates, thereby consistently proving that Race-to-Halt minimizes footprints [81]. The Software Configurable Paradox: Reliable investigations show that optimizing the configuration matrix for the absolute fastest binary runtime frequently spikes maximum baseline voltage states, thus generating highly asymmetric electricity consumption overheads [242]. Time vs. Energy: The basic time/energy correlation seems to be broken during memory-bound workloads, where boosting the CPU frequency has no effect on time reduction, but increases power consumption. Linear Programming algorithms bound the Pareto front for heterogeneous tasks, empowering precise trade-offs [95, 92]. When it comes to managing time/energy trade-offs between independent workloads, energy-efficient schedulers route algorithmic queues to cores experiencing minimal physical thermal stress [118]. Distributed Scaling Imbalance (Straggler Effect): The workload imbalance causes the fastest processors to go into standby mode at the synchronization barriers (i.e. power wastage) [170]. Coordinated control strategies (selective frequency boost or controlled downscaling) align execution phases. Reliability-Performance Degradation Loop: Aggressive DVFS increases the vulnerability to soft errors. Stateof-the-art dispatchers seek a balance between execution time, power, and reliability [104, 112, 113]. Due to task dependencies and increased execution time, an energy-efficient scheduling of precedence-constrained applications within the context of SLAs on DVFS-enable clusters has been proposed by Hu et al. [115]. Failure-Aware VM Consolidation and Asynchronous checkpoints leverage exponential smoothing to predict impending hazards, thereby lowering overall operational energy by 34% without jeopardizing task recovery [71, 243, 244].
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Following what has been said, energy optimization requires explicitly weighing the cost of static leakage against the exponential penalty of dynamic frequency scaling.
6.6. Trends on Software Optimization An analysis of the selected papers within the literature on software optimization (𝑁 = 52 papers) reveals a sharp systemic shift toward distributed orchestration. Nearly 46% of the contributions focus on virtual machine consolidation and carbon-aware scheduling, mostly by leveraging metaheuristics like Particle Swarm Optimization (PSO) and Deep Reinforcement Learning (DRL). Dynamic Power Management (DVFS) and node-level Power Capping cover 31% of the corpus, while pure algorithmic and code-level optimizations (e.g., compiler loops, dataframes, "Energy Smells") account for 23%. This uneven distribution illustrates the fact that, while code efficiency remains fundamental, modern academic consensus favors mitigating static leaks from macro-infrastructures through intelligent cloud orchestrators. % of Surveyed Papers
Carbon-Aware DevOps: Carbon-aware computing extends to continuous integration (CI) & continuous delivery/deployment (CD) pipelines and GitOps frameworks [70]. By combining execution tracing and energy consumption monitoring, CI/CD systems directly associate variations in energy consumption with specific code changes, thus structurally preventing the deployment of suboptimal code in a production environment [61, 15, 65]. Industrial case studies on developers’ awareness of green software design and strategies to promote it within the company are a way to encourage green software design [241].
60 46 40
31 23
20 0 Consolidation
DVFS & Cap. Code & Compilers
Figure 6: Prevalence of software-level energy optimization strategies within recent literature.
7. The Energy Wall of Generative AI The impressive emergence of Generative Artificial Intelligence (GenAI) has significantly magnified the level of resources needed in data centers. Since 2012, the computing resources needed to train advanced AI systems is doubling approximately every 3.4 months [245]. As of early 2026, due to the transition toward trillion-parameter models together with the pervasive deployment of autonomous agentic workflows, the energy consumption of large-scale AI has moved from an operational challenge to a macro-environmental crisis. While researchers insist on foundational manifestos that chart the course towards "green AI", the rapid growth of needed capabilities requires an immediate architectural revolution that affects IT budgets, data storage and global community incentive structures to ensure the sustainability of the AI ecosystem [246]. To analyze this landscape, the AI energy footprint must be strictly divided into two categories using a formalized terminology: Red AI versus Green AI. The term "red AI" refers to any pipeline whose structure requires the use of very large computing resources for the sole purpose of improving performance accuracy by a tiny percentage, thus Page 10 of 26
Energy-Aware Computing in 2026
prioritizing performance over thermodynamic consequences [167]. Conversely, "green AI" systematically imposes efficiency requirements alongside accuracy criteria. To structure this quest for sustainability, the software engineering community has listed robust "green architectural tactics". However, empirical analysis of reference frameworks warns of extreme friction during commercial adaptation; Assessments of more than 160 active open-source ML ecosystems consistently reveal a significant gap between the existence of these ecological tactical protocols and the willingness of AI engineers to integrate them beyond basic performance metrics [247, 56].
7.1. The Global Carbon Cost of Models Training Training a dense state-of-the-art Large Language Model (LLM) requires processing tens of trillions of tokens across highly synchronized GPU clusters. A 100,000-GPU cluster consumes more than 100 to 150 Megawatts (MW). Training these dense models structurally requires fragmenting the model state across multiple devices using Pipeline Parallelism (PP) and Tensor Parallelism (TP) (e.g., via MegatronLM). Recent assessments have proven that arbitrary sizing of these parallel strategies strongly influences the overall efficiency of the cluster. Not optimized tensor parallelism often traps clusters in intensive NVLink synchronization phases, thus consuming a huge amount of inter-GPU static power. Future energy-aware compilers must integrate hyperparameter auto-tuners that calculate exact thermal and synchronization impedance before allocating LLM partition chunks, effectively slashing the hidden network energy cost associated with misaligned topologies [248]. A key operational strategy for green AI training is to mitigate peak fossil-fuel consumption in the electricity grid. Traditionally, schedulers deploy Suspend/Resume architectures (completely pausing the cluster and creating checkpoints when green energy drops). However, this severely penalizes the return on investment in equipment. Current state-of-the-art platforms bypass crude suspension entirely in favor of Carbon Elasticity. By adapting in real time and transparently the actual dimensions of distributed ML training (by adjusting batch size and number of active workers according to the carbon intensity of the network), elastic AI environments enable additional savings of 37% to 51% on the total carbon footprint without suffering crippling synchronization delays. [249]. Rather than training AI models with petabytes of highly redundant token distributions, significant power savings are increasingly achieved before VRAM allocation through data-centric AI. Modern evolution-based and data-centric training topologies use "elite sampling" engines to mathematically isolate only hypervariant feature data within a training corpus. By dynamically compressing to only 10% the footprint of the data used, data centers completely avoid redundant epoch execution, thus significantly reducing the continuous uptime of the servers by 98% without substantially penalizing the accuracy of multi-class AI inference [166, 168]. R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
While large language models is a dominant topic in public discourse, industrial hyperscalers devote the vast majority of their continuous inference energy to running Deep Learning techniques in Recommender Systems (RecSys). Systematic evaluation pipelines allow for rigorous tracing of the Pareto frontier between the accuracy of RecSys clicks and carbon emissions. The evaluation of dozens of distinct algorithmic architectures demonstrates that achieving the last 1% of recommendation accuracy frequently results in a penalty of roughly 40% in carbon emissions. By natively mapping these boundaries, engineers can mathematically select slightly smaller matrices that drastically reduce the company’s operational carbon emissions [250]. Finally, a major aggravating factor in GenAI’s carbon crisis is the low sharing of academic and corporate artifacts. When training large language models specifically designed for software engineering, the initial energy expenditure is noticeable. Large-scale literature reviews reveal that nearly 40% of research institutions and companies do not share their compiled artifacts or model weights, forcing researchers to rerun redundant and extremely computationally intensive training paradigms [58].
7.2. What About Inference ? While pre-training dominates headlines, hyperscaler telemetry from 2024–2026 indicates that Inference now accounts for the vast majority of the overall AI energy footprint. It is very important to characterize and understand the differences between training and inference power consumption patterns [159]. Given that Service Level Objectives (SLOs) [132] for public LLM endpoints require extremely low latency targets, commercial clusters often keep a large number of GPUs at high-voltage during idle states (because wake-up overhead could impact performance), just waiting for the next avalanche of prompts [251]. To resolve this ongoing energy drain, automated environment orchestrators (e.g., DynamoLLM) use massive search space formulations to continuously and asynchronously reconfigure dynamic LLM instances. By performing automatic and real-time adjustments of cluster dimension, clock frequency, and tensor model parallelism in response to fluctuations in inference queues, modern streaming platforms manage to save up to 50% of their total operational carbon footprint without ever violating underlying web latency service level objectives (SLOs). [164]. Furthermore, recent sophisticated energy profilers unequivocally demonstrate that the power consumption profile of inference cannot be considered as uniform. Processing an LLM typically follows two phases: the prefill phase and the decode phase. The prefill phase (which processes the user prompt) mainly involves compute-bound matrix-matrix multiplications, which therefore instantly puts the GPU cores at their highest voltage [84]. Conversely, the Decode phase (which generates tokens one-by-one) translates into strictly memory-bandwidth bound operations, leaving the
Page 11 of 26
Energy-Aware Computing in 2026
ALU unsolicited while wasting a huge amount of static highbandwidth memory (HBM) energy. New-generation inference schedulers proactively apply decoupled DVFS policies based on tokenization state, maximizing clock frequencies during the prefill phase and significantly reducing logic voltages during the decode phase to reduce inference carbon by more than 25% [172, 130]. To prevent hallucinations and access real-time corporate data, enterprises primarily deploy LLMs via RetrievalAugmented Generation (RAG) pipelines. From an energy perspective, a RAG pipeline creates a severe infrastructure strain: before the LLM generates a response, the system must execute an approximate nearest neighbor (ANN) search across a massive vector database, which involves a large pool of compute nodes that continuously execute memoryintensive tasks. Integrating these retrieved documents into the LLM significantly expands the context window. Maintaining the Key-Value (KV) Cache in the GPU’s VRAM prevents the model from recomputing tokens, but consumes massive dynamic power. A massive paradigm shift revealed in early 2026 directly links this caching physics to external grids through Carbon-Aware Prompt Caching. Because the network is heavily powered by fossil fuels, instead of blindly discarding cached requests via generic LRU algorithms, sophisticated broadcast platforms artificially extend the retention lifetime of massive request matrices directly in VRAM, thus ensuring that subsequent identical requests retrieve pre-computed tokens to physically avoid massive, carbon-emitting redundant recalculations [84, 172]. The most advanced AI paradigms rely on autonomous agents and multi-step reasoning models. Rather than directly answering a prompt, an agent breaks it into sub-tokens, perform searching (e.g., over the web), generates and runs intermediate codes, performs self-evaluation, and so on (iterates). This introduces a hidden Recursive Energy Multiplier. A single user interaction might trigger 50 background inference calls before a final answer is produced. Advanced 2026 architectures (such as the PEARL framework) explicitly perform Performance/Energy-Aware Routing for LLMs. By dynamically intercepting prompts and distributing them asymmetrically between massive monolithic models and lightweight edge models — evaluating semantic complexity in real-time along with active clusters power — orchestrators maintain optimal reasoning accuracy while efficiently reducing the recursive energy multiplier [230, 173]. Historically, the primary constraints in Deep Learning were execution time (latency) and the amount of GPU VRAM. To deploy increasingly large models, algorithmic compression techniques have been developed. Initially designed to minimize time-to-solution through the corresponding hardware speed-up, these techniques were subsequently redesigned to address energy-to-solution. The fundamental principle underlying this adaptation is the physics of silicon: reading data from HBM memory consumes much more joules than the actual ALU calculation.
However, each new generation of techniques introduces serious technical limitations and energy trade-offs. Pruning and Network Compression: Earliest compression techniques consisted of algorithmically identifying and removing redundant weight matrices. By removing certain parameters, the total number of multiplication-accumulation (MAC) operations decreases. However, early "unstructured" pruning randomly removed weights, thus creating a Sparsity Trap. The systolic arrays of modern GPUs are highly optimized for dense matrix operations. An unstructured sparse matrix forces the GPU to load zeros into memory and multiply by zero anyway, which does not allow for any energy savings on modern accelerators. Knowledge Distillation: Distillation trains a lightweight "Student" model to mimic a large "Teacher" model. Replacing, for example, a 1000W cloud inference query with a 5W Edge NPU query on a model with 8 billion parameters reduces operational energy consumption per token by nearly 200 times. However, distillation only shifts the energy load backwards, requiring a huge initial carbon investment to execute the Teacher’s forward passes, which often requires millions of Edge inferences to compensate. Neural Architecture Search (NAS): Hardware-aware Neural Architecture Search (HW-NAS) injects physical constraints directly into the reward function to discover ideal silicon topologies. However, the search process evaluates thousands of permutations, emitting carbon equivalent to that of five standard cars [155]. To overcome this problem, modern frameworks rely on training performance estimation (TPE), using early stop variants to abruptly interrupt training as soon as energy thresholds diverge from accuracy slopes, thus recovering up to 90% of the search phase energy [252]. Beyond simply rejecting parameters, NAS pipelines actively replace fundamental mathematical operations. Advanced architectures assign custom hardware-friendly inference operators (such as deep-shift or multiplication-free kernels) directly into the DNN structure layer by layer, thus outperforming comparable mixed-precision algorithms by prioritizing the simplification of operations [157]. Precision Scaling and Quantization: Instead of removing parameters, quantization reduces their bit-width. The energy cost of arithmetic (resp. memory) operations scales quadratically (resp. linearly) with bit-width. Recent hardware advances have introduced precision-tailored architectures (e.g., 5.3-bit Block Floating Point formats) that definitively conquer the historic sub-8-bit dynamic range limitations [27]. However, as the community attempts to push quantization to extreme limits, a structural limitation in the representation arises. Disruptive 2025 frameworks such as LUT-DLA break this limit by substituting scalar computation with vector-quantized Look-Up Tables (LUT), resulting in an impressive area efficiency gain of ≈ 146× [28]. Mixture of Experts (MoE) and Speculative Decoding: Recognizing the limits of extreme quantization has led to considering the alternative of dynamic sparsity at runtime. MoE uses a router to activate only a sparse subset of "Expert"
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 12 of 26
7.3. Evolution of Compression in Green AI
Energy-Aware Computing in 2026
7.4. Decentralization: Fog and Mobile Edge The ultimate optimization of the infrastructure is to offload heavy computing loads from centralized hyperscalers to intermediate Mobile Edge Computing (MEC) and Fog platforms. Handling the extreme mobility of nodes and decentralized energy ceilings requires replacing hierarchical managers with complex algorithmic coordination. The use of rigorous game theory approaches to offload computation to the mobile cloud allows heterogeneous devices with limited autonomy (battery-powered) to act as autonomous economic agents [253, 254, 255]. By mathematically modeling the significant penalty of transmission bandwidth delays compared to the relief provided by offloading localized computations [256], MEC architectures routinely calculate localized Nash equilibria [137]. This structurally prevents decentralized systems from saturating intermediate cloud proxies with computations, thus ensuring optimal and uniform battery survival margin across thousands of highly asynchronous edge appliances [257]. Small language models (SLMs), typically less than 10 billion parameters and optimized by quantization, allow efficient on-device inference on consumer platforms. Empirical lifecycle case studies validate this transition, demonstrating that migrating generative AI inference from centralized cloud infrastructures directly to edge platforms can generate energy savings of over 90%, while significantly reducing water consumption and carbon footprint through the elimination of network and cloud cooling costs [102, 103, 258]. Expanding Edge capabilities requires that local mobile clusters not only run inferences, but continuously customize models via parameter-efficient fine-tuning (PEFT) mechanisms, such as low-rank adapting (LoRA), physically bypassing full gradient weight updates to safely retrain models without inducing thermal throttling. Federated Learning (FL) prevents centralized data concentration at the expense of a detrimental "wooden barrel
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
effect" regarding carbon emissions. Bidirectional transmission of gigabytes of gradient updates over high-frequency RF networks typically consumes much more radio power than localized computations [17, 259, 165]. To address this problem in a targeted way, orchestrators perform continuous monitoring based on game theory (e.g., via MARL controllers) to adjust the inclusion of edge-client parameters. Frameworks like CLOVER replace random client inclusion with rigorous participant selection based on their carbon footprint. By mathematically evaluating the instantaneous carbon states of the local power grid alongside the computing speeds of the devices, the orchestrators selectively collect data that forces global convergence while artificially reducing FL carbon emissions by 25% [260, 261]. Historically, carbon-aware spatial shifting was strictly limited to the Cloud. Innovative research conducted in mid2025 establishes the effectiveness Mesoscale Spatial Shifting. Architectures such as the CarbonEdge framework dynamically track localized variations in "mesoscale" carbonintensity within interconnected geographic regions. By intelligently micro-migrating CDN workloads between geographically adjacent peripheral hubs with low transient carbon differences, modern orchestrators are able to achieve emission reductions of more than 78% without triggering latencies of more than 5 milliseconds [262].
7.5. Main Trends in Green AI By analyzing the rapidly expanding field of green AI (45 papers), the bibliometric trend reflects the commercial pivot from model creation to model deployment. While pretraining of massive reference models initially dominated the discourse, currently only 27% of our selected papers target optimizations of the training phase (e.g., Carbon Elasticity, Data-centric pruning, ...). Conversely, 42% of recent studies explicitly target inference optimizations, encompassing RAG KV-Cache management, Agentic routing, and algorithmic compression (MoE, Quantization). The remaining 31% addresses Decentralization, heavily exploring Federated Learning limits and Small Language Models (SLMs) on Edge devices. This confirms that the main challenge of current research (2025 and 2026) is to minimize the continuous and ubiquitous energy consumption of GenAI services. % of Surveyed Papers
sub-networks per token, thereby decoupling the total storage capacity from the active compute footprint. However, this creates a VRAM Leakage Paradox: Although MoE reduces the power consumption of the dynamic ALU, it requires the entire massive model to reside in the HBM, resulting in significant static power loss, thus greatly limiting its effectiveness on power-constrained deployments. Approximate Accelerators for Transformers: For spatial mapping engines such as Vision Transformers (ViTs), general approximate computing has shown a negative impact on deep-learning convergence. Advanced methodologies (e.g., the TransAxx framework) now securely insert approximate arithmetic multipliers directly into ViT execution paths. Using Monte Carlo Tree Search (MCTS) algorithms allow to efficiently search the space of possible configurations to perform Vision Transformer (ViT) with structurally defective mathematical chips, thereby enabling an noteworthy reduction in hardware power consumption while preserving visual quality [37].
42 40
31
27
20 0 Inference & RAG Edge & FL
AI Training
Figure 7: Quantitative overview of Green-AI literature.
Page 13 of 26
Energy-Aware Computing in 2026
8. Advanced Cooling Systems As time goes on, the abstractions of software scheduling and algorithmic compression eventually run up against the unavoidable laws of thermodynamics. All the electrical energy consumed by a computer system is eventually dissipated as heat. Historically, data center cooling was considered a secondary issue at the facility level. Today, the thermal density of Exascale and GenAI workloads has transformed cooling into the primary bottleneck, determining whether a processor can keep running at its peak FLOPs.
8.1. Air Cooling and Rack Density The thermal capacity of air is fundamentally insufficient for modern high-density architectures. Traditionally, the power consumed by Computer Room Air Conditioning (CRAC) units represents up to 40% of a facility’s total energy. To combat this via software, conventional "ThermalAware Schedulers" proactively balanced CPU loads spatially across different racks to eliminate physical hot spots, thus allowing CRAC compressors to safely step down their power levels. However, the shift from provisioning for a single server to that of large AI infrastructures has shattered these parameters at the rack-scale. A single NVIDIA GB300 (Blackwell Ultra) NVL72 rack integrates 72 GPUs and 36 CPUs, drawing more than 120 kW in a single physical cabinet. Trying to cool a 120 kW rack by pure ventilation would require localized winds of exceptional magnitude, which would create acoustic damage and consume more energy for the fans than for the processors. Consequently, the industry has unanimously recognized that for any rack exceeding roughly 35 kW, air cooling is no longer a viable solution.
8.2. Advanced Liquid Cooling Paradigms To manage the massive heat flows associated with AI processing, facilities have turned to liquid cooling, which, compared to air, has thermal capacities up to 4 times greater and thermal conductivities 20 to 100 times higher. Liquid Cold Plate and the Temperature Paradox: D2C (or Cold Plate) technology circulates a cold liquid directly over micro-convective blocks mounted on CPU, GPU, and memory chips. To maximize facility-level PUE, data centers increasingly supply "hot water" (18◦ 𝐶 to 25◦ 𝐶) rather than chilled water, thereby significantly reducing chiller plant power. However, this introduces a serious systemic paradox: the use of warmer inlet water reduces the energy consumption of cooling but inadvertently causes temperature spikes in the silicon chips. This physically exacerbates the leakage current of the transistors and suddenly triggers a localized DVFS limitation, thus introducing serious execution delays in synchronized MPI workloads that often wipe out the initial cooling energy savings [76]. Intelligent Cooling Controllers and Joint ThermalWorkload Orchestration: Traditionally, computer room air conditioning units (CRAC units) operated from static thresholds unrelated to server states. Nowadays, temperature is recognized as an active variable that has a direct impact R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
on hardware transistor leakage. Systematic assessment of temperature-aware power models demonstrate that failing to account for the increase in ambient temperature of the incoming air significantly distorts theoretical CPU power consumption predictions. This has led to the deployment of "Joint Data Center Cooling and Workload Management". By merging these logical flows using deep reinforcement learning (DRL), schedulers algorithmically assign high-intensity thermal threads to nodes that exhibit optimal cooling airflow, thereby reducing the overall cooling power of the installation independently of the hardware’s P-states [134, 151, 150]. Single-Phase/Two-Phase Immersion: The most extreme cooling method involves fully immersing the server’s motherboard in a dielectric (non-conductive) fluid. In twophase immersion, the fluid boils directly upon contact with the silicon chip, exploiting the latent heat of vaporization to achieve unprecedented thermal densities with near-zero pumping power, thus enabling PUEs close to 1.01. Critical Limitation (The PFAS Regulation Crisis): While two-phase immersion was considered the future of green data centers, it has collided with a recent regulatory hurdle. The dielectric fluids initially required for boiling (e.g., 3M’s Novec) are primarily perfluoroalkyl and polyfluoroalkyl substances (PFAS) — "forever chemicals" exhibiting very high environmental and biological toxicity. Due to severe restrictions and bans from the European Union and the US Environmental Protection Agency (EPA) on PFAS manufacturing, the viability of two-phase immersion has plummeted, forcing hyperscalers to revert to highly complex D2C water configurations. [152].
8.3. Water and Power Usage Effectiveness Optimizing for power usage efficiency (PUE) raises a major environmental paradox regarding water consumption. Modern data centers often achieve exceptional PUE not through refrigeration compressors, but rather through evaporative cooling towers. The development of very high-density AI clusters requires enormous quantities of chilled water, which constantly evaporates into the atmosphere to release the heat dissipated from the data center. Merely estimating the raw volume of Water Usage Effectiveness (WUE) does not take into account the geographic hydrology. To address this critical gap, modern supercomputing assessments (such as the ThirstyFLOPS framework [85]) introduce the Adjusted Water Impact (AWI) indicator. The AWI takes into account the high spatial and temporal volatility of local water stress, demonstrating that a supercomputer consuming 1 million liters of water in the rainy Pacific Northwest region is ecologically harmless, while the same consumption in an arid environment would cripple local public infrastructure. Thus, modern green planning advocates actively anticipating periods of inactivity to offset intensive processing only during droughts [85]. In waterstressed regions, municipalities are increasingly refusing building permits for liquid-cooled data centers, despite their low carbon footprint. Therefore, the use of closed-loop "dry coolers" is now mandatory, which inherently leads to higher Page 14 of 26
Energy-Aware Computing in 2026
consumption of electrical energy (thus degrading the PUE) but saves water.
8.4. Heat Recovery and Energy Reuse As ICT-related electricity needs approach 4% or 5% of global demand, governments are no longer viewing data centers as mere consumers, but rather as potential localized thermal power plants. The focus has shifted from minimizing cooling overhead to maximizing the Energy Reuse Factor (ERF). A high heat output (e.g., a 60◦ −80◦ 𝐶 liquid from the cooling of an Exascale system) is potentially very valuable. Thanks to heat exchangers, this wasted energy can be used for external infrastructures: District Heating: Supercomputers such as LUMI (Finland) integrate their liquid cooling loops directly into Kajaani’s district heating network, thus taking advantage of fossil fuel combustion to heat thousands of local residential homes during the winter. Distributed Edge Heating (e.g., Qarnot Computing): Innovative European architectures completely decentralize data centers by installing cloud-based computing nodes (viewed here, from a thermal perspective, as electric radiators) directly within residential buildings. This "computingbased heating" approach achieves near 100% energy efficiency, thus offsetting the cooling costs of the facilities and that of associated infrastructure.
8.5. Oversizing and Simulation-Driven Design In the exascale ecosystem, the evaluation of millions of active sensor parameters reveals a rigid limit of data centers: facilities tend to have more installed computing devices than the electrical grid can run simultaneously. The operation of these legally "oversized" systems requires the direct integration of neural networks into the infrastructure of the facilities. Frameworks such as the Adaptive Power Management System (DPS) explicitly map workloads in real time via ML analysis, securely transferring power credits from idle peripheral logic cores exclusively to critical processing queues in order to keep peak Exaflops below regulatory power thresholds [107, 175]. An embarrassing reality arises from the mismatch between the rapid evolution of equipment and the long life cycles of infrastructure. A physical data center is amortized over 15 to 20 years, while new generations of AI accelerators are produced every 12 to 18 months. Upgrading a 2015 air-cooled facility to accommodate a 120 kW liquidcooled Blackwell rack (planned for 2026) is structurally and financially prohibitive. Future next-generation AI workloads demand bespoke "AI factories." Designing these extreme thermal infrastructures precludes any empirical approach in favor of formal methods. Global hyperscalers are increasingly relying on holistic operational simulators, a major one being the Carbon Explorer framework. This topology integrates the interaction between local geography, PUE degradation, and daytime renewable energy to mathematically advise on structural deployment trade-offs between battery CapEx and direct utility energy consumption, thus R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
enabling early exploration of sustainability-focused design space even before concrete realizations [263].
9. Power-aware Computing Perspectives The shift from a paradigm focused on absolute performance to one that considers strict energy constraints has highlighted corresponding conceptual gaps in modern computing. On the cusp of the exascale era and with the expansion of generative AI, addressing the energy crisis requires moving beyond iterative software optimization alone. Based on critical limitations identified across the Cloud-Edge-HPC continuum, we have outlined five major and disruptive future directions encompassing hardware, software telemetry, supply chain auditing, and global legislation. These aspects will shape the future of computing through 2030.
9.1. Advanced Measurement: OS and On-chip As highlighted in Section 5, current high-frequency monitoring relies heavily on user-space polling or kernellevel interrupts. This induces the "observer effect": continuously polling hardware registers, such as MSRs, frequently triggers context switches, thus saturating CPU caches, interrupting AI workloads, and artificially amplifying nodelevel power consumption, simply to measure it. For largescale telemetry with no overhead, the community is rapidly turning to the following two fundamental paradigms. Kernel-Bypass via eBPF: Extended Berkeley Packet Filter (eBPF) technology (e.g., CNCF Kepler) allows researchers to securely integrate event-driven polling programs directly into the Linux kernel, without requiring custom kernel modules. Currently, frameworks using eBPF represent the core of the state-of-the-art for containerized cloud deployments. By intercepting OS-level events, eBPF architectures accurately map the energy footprint of highly ephemeral microservices and serverless FaaS workloads that cannot be detected by traditional MSR polling, thus eliminating blind spots in energy consumption measurements of modern virtualized grids [109]. In-Network Telemetry via DPUs: To eliminate the monitoring task at the host CPU, Data Processing Units (DPUs) and SmartNICs (e.g., NVIDIA BlueField, AMD Pensando) are absorbing the telemetry stack. DPUs aggregate, compress, and stream time-series data directly into facility management pipelines via a dedicated PCIe path, thereby ensuring that host CPUs are exclusively dedicated to computation.
9.2. In-Memory Computing and Hardware Co-Design For decades, von Neumann architectures have separated the arithmetic logic unit (ALU) from main memory. In GenAI inference paradigms, transferring terabytes of LLMrelated data via PCIe buses and memory interconnects consumes roughly two orders of magnitude more energy than that of the arithmetic operations. Next-generation hardware must clearly eliminate this "memory wall". Page 15 of 26
Energy-Aware Computing in 2026
Memory Disaggregation via Compute Express Link: Historically, administrators oversized local server RAM to prevent AI workloads from crashing (due to lack of memory), resulting in terabytes of unused but powered DRAM throughout the facility. Currently, the transition to the CXL interconnect standard enables memory disaggregation, grouping petabytes of memory into distributed chassis that are dynamically accessible by any processor via PCIe logic. Decoupling memory allocations from physical hosts safely eliminates current leakage due to "unused memory," thereby reducing the ongoing carbon footprint of RAM (integrated and operational) in a data center by more than 25% [169, 161]. Near-Data Processing (NDP): Processing-in-Memory (PIM) and in-storage processing (ISP) mitigate data movement by offloading the instructions to the cores integrated with the main memory level. PIM architectures (e.g., UPMEM DIMMs and HBM-PIMs) fundamentally reverse the conventional hardware paradigm by moving the ALU directly inside the DRAM arrays, thereby reducing parasitic capacitance energy from the memory buses [264]. However, a significant loss of efficiency occurs during the conversion of memory voltages. Representing the elite architectural limits, the systems now utilize Analog-Mixed-Signal (AMS) capacitive coupling directly inside standard 6T and 8T SRAM matrices. The close association of XNOR multipliers with shared in-memory analog-to-digital converters (ADCs) [265] enables remarkable advances in the computation of binarized neural networks, achieving macroscopic efficiencies of nearly 490 TOPS/W through the complete elimination of column-dense read paths [266]. Chiplets and Heterogeneous Integration: Instead of manufacturing massive monolithic chips prone to significant current leakage, silicon is fragmented into "chiplets" using the UCIe (Universal Chiplet Interconnect Express) standard. This allows manufacturers to combine high-performance logic components etched at 2 nm with I/O and analog components etched at 6 nm, thus offering more concentrated power and more stable consumption. Architectural carbon footprint modeling frameworks, such as the Architectural 𝐶𝑂2 Footprint Tool (ACT), ECO-CHIP, and 3D-Carbon, empirically prove that the migration to heterogeneous integration (HI) strategies reduces embodied carbon emissions by 30% compared to traditional monolithic systems [35, 267, 268]. Wafer-Scale Systems and Carbon Dependency: While the traditional industry is turning to UCIe chiplets, the extreme counter-paradigm involves monolithic wafer-scale integration (WSI). Architectures like the Cerebras WaferScale Engine (e.g., the CS-3) completely abandon chiplets, sculpting an entire supercomputer from a single continuous silicon wafer. Recent life cycle assessments of the total Carbon-Delay Product (tCDP) have empirically concluded that, despite the huge manufacturing challenges of waferscale processors, their complete elimination from off-chip diffusion dies generates a design space with uniquely optimized carbon performance [38]. R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Photonic computing: Replacing copper interconnects with photonic (light-based) data buses significantly reduces energy consumption related to static impedance, thus overcoming memory bandwidth limitations. Next-generation frameworks (such as the open-source tool EPiCarbon) demonstrate that photonics far surpasses traditional semiconductors structurally: manufacturing the hardware requires at least 4.1 times less manufacturing energy (embedded carbon) while offering significantly higher yields than equivalent-sized CMOS chips [34]. By considering moderate and aggressive photonic scaling, frameworks such as Albireo and ROCKET show that EDP and Energy can be reduced by at least 229× when compared to state-of-the-art electronic accelerators and GPUs on AI workloads [46, 269]. Neuromorphic Computing vs Analog Computing: For edge endpoints, pulsed neural networks mimic the biological brain by performing calculations only during discrete events. This enables the design of true "event-driven silicon", reducing standby power consumption to a few microwatts [25]. Furthermore, analog computing natively processes deep neural networks in the continuous load domain. Recent advances in mixed-signal and charge-based multiplyadd cores allow matrix algebra to be run at the fundamental thermal noise limit, thue leveraging large onboard capacitor arrays to routinely reduce computing power consumption by more than 37% compared to equivalent dense digital logic paths.
9.3. Operational Carbon vs Embodied Carbon Until now, the green software community has primarily focused on operational carbon (emissions mainly related to operational electricity consumption). However, with the rapid decarbonization of power grids, the operational footprint of a cluster ideally tends toward zero. Consequently, the predominant environmental impact of computing is shifting toward embodied carbon, which corresponds to the significant greenhouse gas emissions associated with semiconductor manufacturing, rare earth mining, server production, and transportation activities[270, 271]. Life cycle assessment (LCA) studies unequivocally state that the exponential proliferation of custom AI architectures necessitates a unified Pareto LCA frontier; hyperscale environments must quantitatively balance the operational carbon savings of an ASIC with the much larger carbon footprint required for semiconductor manufacturing. Replacing server arrays every two to three years structurally prevents data centers from achieving carbon neutrality, regardless of their power usage efficiency (PUE) [272, 273]. Lifecycle-Aware Software and System Sustainability: Software overload and unoptimized libraries inevitably lead to premature hardware obsolescence. We anticipate a paradigm shift toward software lifecycle engineering, where software is explicitly optimized to maintain backward compatibility, thereby extending the operational lifespan of a data center’s IT equipment from 3 to 7 years [59, 274]. To address this problem structurally, the software engineering community is mandatorily integrating sustainability across Page 16 of 26
Energy-Aware Computing in 2026
four key dimensions: environmental, economic, social, and technical. Integrating global sociological and economic constraints directly into developers’ workflows ensures that they prioritize compatibility with moderately innovative architectures by deliberately limiting algorithmic feature extensions that absolutely require next-generation hardware. Analytical Modeling and Rebound Effects: Establishing rigorous parameterized baselines (such as the FirstOrder analytical CArbon modeL (FOCAL) [74]) inherently isolates rebound effects (the Jevons Paradox). In-depth assessments clearly highlight that when cloud instances achieve significant efficiency gains, the absolute capacity of data centers tends to increase primarily to absorb the financial benefit associated with the energy margin. New analytical criteria make it possible to decouple this cycle by mathematically defining the strict limits necessary to transform "weakly" sustainable operational accelerators into closed-loop and strictly controlled life cycle assessments (LCAs), thus guaranteeing overall societal reduction margins in relation to the expansion of digital consumption habits [136, 42, 275, 276, 117].
9.4. Compliance with Standards and Regulations The development of green IT is driven more by regulatory incentives than technological ones. Sustainable IT will now move from a voluntary initiative, driven by corporate social responsibility (CSR), to a globally audited regulatory compliance framework. Macro-Policy and Granular Reporting: Under the European Energy Efficiency Directive (EED), data centers with computing power exceeding 500 kW are legally required to publish comprehensive sustainability indicators. In parallel, the Corporate Sustainability Reporting Directive (CSR) subjects IT companies to rigorous emissions audits related to scopes 1, 2, and 3 [277]. Current corporate sustainability reports fundamentally obscure critical architectural realities beneath broad overviews. In-depth analyses of computer hardware reveal a growing and unreported disparity: the absolute carbon footprint of local integrated circuits is increasing independently year after year. To put an end to this "greenwashing", future compliance initiatives must require extremely accurate reporting of hardware components [278, 279]. The EU AI Act and Ethical Prompts: Specific to the software field, the final clauses of the European AI Act require creators of general-purpose AI (GPA) to explicitly document the energy consumption and overall ecological footprint associated with training their fundamental models [280]. According to our vision of the 2030 Sustainable Development Roadmap, governance is undeniably shifting from training centers to the user interface. New paradigms of "ethical prompt engineering" (EPE) integrated into interactions impose auditable compliance controls specific to recovery architectures. EPE structurally limits unrestricted LLM requests by treating user prompts as versioned constraints specifically designed to block the execution of
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
resource-intensive or redundant models before they trigger the hardware logic layer [281, 62]. Digital Sustainability Labels: The criteria for acquiring computer equipment are crystallizing around ecolabels (e.g., EPEAT and Energy Star). In order to fill this regulatory gap at the software level, international consortia have published the "Green Software Measurement Model" (GSMM) [16]. This validated reference model standardizes how auditors empirically and dynamically extract resource efficiency gains from isolated software products. Furthermore, running models like the GSMM enables the imminent transition to software sustainability labels; European and international legislation now requires enterprise software suites to actively calculate and display, in real time, "green labels" similar to energy labels for appliances [278]. Economic Catalysts and Carbon Market Exchanges: To enforce global thermodynamic limits, providers are natively integrating energy-efficient algorithmic structures into their middleware. Emerging cloud brokers are deploying fully automated carbon trading data markets. Drawing inspiration from Nash auction systems, geodistributed Kubernetes clusters actively trade and bid for prioritized execution windows, exclusively during periods of high renewable energy production (e.g., local solar power peaks). By monetizing carbon allowances, the hypervisor infrastructure inherently prevents network capacity overflows and incentivizes edge users to dynamically suspend lower-priority tasks in exchange for direct financial dividends, thus realizing "green AI" through a rigid market capitalization [93, 153, 282]. Furthermore, large-scale operations conducted onboard the Fugaku supercomputer have deployed robust and operational "incentive-based energy efficiency programs" [133]. By directly returning to the teams of supercomputers, who voluntarily activated deeper P states in their scientific codes, a digital currency corresponding to their exact use, the system administrators have concretely validated that gamification on an institutional scale effectively redirects collective behaviors in order to optimize planetary impact. Microgrids and Wind Energy Deployments: The development of zero-carbon integrations requires the dismantling of traditional facilities. To address the drastic restrictions on distribution lines that are hindering the global energy transition to wind power, the macroeconomic frameworks of July 2025 advocate the dynamic deployment of high-performance data center equipment directly within local wind farms [283, 154]. Placing the AI training logic close to the power generation source acts as a highly elastic storage buffer. This proximity configuration systematically absorbs and modulates the raw transient power well before it is transmitted over high-voltage municipal power lines, thus minimizing extreme losses due to voltage conversion and significantly increasing the overall capacity factors of the power grid. Operators are deploying architectures such as PowerMorph as part of demand management contracts. By precisely manipulating the electrical footprint of QoSaware servers to the microsecond, the data center acts as an electrical regulation well, rapidly stabilizing the AC Page 17 of 26
Energy-Aware Computing in 2026
frequency of the regional electrical grid for significant financial gains. [284]. Within the framework of pan-European projects (Green.Dat.AI), strict structural "data spaces" harmonize semantic datasets well before the activation of deep learning architectures [285].
10. Conclusion Over the past decade, energy consumption has shifted from a secondary operating expense to the primary physical, financial, and environmental bottleneck in the computing ecosystem. As this comprehensive survey demonstrates, the challenge of energy-aware computing is ubiquitous across the entire cloud-edge-HPC continuum. From low-power microcontrollers performing milliwatt-scale inferences to exascale supercomputers of hundred-megawatt AI data centers, the laws of thermodynamics impose absolute limits on performance scaling. The rapid proliferation of generative artificial intelligence, LLM inference pipelines, and agentic AI workflows has accelerated the energy crisis. The traditional response, which relies solely on semiconductor manufacturing to produce low-power transistors, is no longer sufficient. Genuine sustainability requires a comprehensive and cross-cutting overhaul of the entire ecosystem. We have established that cutting-edge chips rely on extreme hardware heterogeneity, memory disaggregation (CXL), and drastic precision reduction (FP4/LUT) to maintain AI throughput. However, to measure this efficiency, the software community must overcome virtualization-related hurdles by deploying overhead-free eBPF polling tools, predictive tracing models for transformers, and reproducible carbon tracking that avoids the standard overhead of memory jump registers (MSRs). Furthermore, static code optimizations and energy-efficient compilation must be tightly integrated with dynamic schedulers. These orchestrators require advanced particle swarm optimization and deep reinforcement learning techniques to seamlessly ensure workload consolidation, C-state latency balancing, and energyefficient spatiotemporal shifting. Finally, as air cooling reaches its absolute physical limits, liquid-to-chip infrastructure, intelligent DRL HVAC controllers, and waste heat recovery (energy reuse factor) must be standardized, thus bridging the severe gap between 15-year facility lifecycles and 18-month GPU iterations. As is widely documented in the field of sustainable digital infrastructure, addressing the energy crisis requires viewing computing limits as a multidisciplinary continuum. The transition to a net-zero carbon architecture necessitates abandoning the historical criterion of raw performance at all costs. By leveraging renewable sources intrinsically linked to dynamic planning, establishing rigorous transparency to combat the rebound effects of embodied carbon, and directly integrating environmental indicators into global legislation (e.g., the European AI law and CSRD mandates), the research and industry communities can successfully transition
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
to the era of exascale AI while respecting the energy constraints of our time [286].
References [1] International Energy Agency, Co2 emissions associated with electricity generation for data centres, https: //www.iea.org/data-and-statistics/charts/co2-emissions-\ associated-with-electricity-generation-for-data-centres-by-case\ -2020-2035, 2025. IEA dataset/report reference. [2] TOP500 Organization, Top500 supercomputer sites, https://www. top500.org/lists/top500/, 2025. June/November 2025 edition. [3] Meta AI, Introducing meta llama 3, https://ai.meta.com/blog/ meta-llama-3/, 2024. Official blog and technical overview.
[4] NVIDIA, Nvidia blackwell platform arrives to power a new era of computing, https://nvidianews.nvidia.com/news/
nvidia-blackwell-platform-arrives-to-power-a-new-era-of-computing,
2024. Official announcement (whitepaper-style content). [5] AMD, Amd instinct mi300 series accelerators, https://www.amd. com/en/products/accelerators/instinct/mi300.html, 2024. Product whitepaper/specifications. [6] United Nations Framework Convention on Climate Change, Paris agreement, https://unfccc.int/process-and-meetings/ the-paris-agreement/the-paris-agreement, 2015. [7] A. G. Philipo, H. Ning, D. S. Sarwatt, J. A. Mohamed, A. S. Yusufu, F. Shi, S. Urenje, J. Ding, Sustainable ai: Emerging trends, impacts, and future challenges, IEEE Trans. Sustain. Comput. 10 (2025) 1278–1291. [8] W. Lin, J. Lin, Z. Peng, H. Huang, W. Lin, K. Li, A systematic review of green-aware management techniques for sustainable data center, Sustain. Comput. Inf. Syst. 42 (2024) 100989. [9] S. M. Hasan, T. Islam, M. Saifuzzaman, K. R. Ahmed, C.-H. Huang, A. R. Shahid, Carbon emission quantification of machine learning: A review, IEEE Trans. Sustain. Comput. 10 (2025) 1085–1102. [10] S. McGuire, E. Schultz, B. Ayoola, P. Ralph, Sustainability is stratified: Toward a better theory of sustainable software engineering, in: IEEE/ACM 45th Int. Conf. Soft. Eng. (ICSE), 2023, pp. 1996–2008. doi:10.1109/ICSE48619.2023.00169. [11] Google Cloud, Introducing trillium: Google’s 6th generation tpu, https://cloud.google.com/blog/products/compute/ introducing-trillium-6th-gen-tpus, 2024. Accessed: 2026-04-16. [12] G. Raffin, D. Trystram, Dissecting the software-based measurement of cpu energy consumption: A comparative analysis, IEEE Trans. Parallel. Distrib. Syst. 36 (2025) 96–107. [13] Z. Yang, K. Adamek, W. Armour, Accurate and convenient energy measurements for gpus: A detailed study of nvidia gpu’s built-in power sensor, in: SC24: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, 2024, pp. 1–17. doi:10.1109/SC41406.2024. 00028. [14] T. Ilsche, D. Hackenberg, S. Graul, R. Schöne, J. Schuchart, Power measurements for compute nodes: Improving sampling rates, granularity and accuracy, in: IEEE IGSC’15, 2015, pp. 1–8. doi:10.1109/ IGCC.2015.7393710. [15] S. Rajput, T. Widmayer, Z. Shang, M. Kechagia, F. Sarro, T. Sharma, Enhancing energy-awareness in deep learning through fine-grained energy measurement, ACM Trans. Softw. Eng. Methodol. 33 (2024). [16] A. Guldner, R. Bender, C. Calero, G. S. Fernando, M. Funke, J. Gröger, L. M. Hilty, J. Hörnschemeyer, G.-D. Hoffmann, D. Junger, T. Kennes, S. Kreten, P. Lago, F. Mai, I. Malavolta, J. Murach, K. Obergöker, B. Schmidt, A. Tarara, J. P. De Veaugh-Geiss, S. Weber, M. Westing, V. Wohlgemuth, S. Naumann, Development and evaluation of a reference measurement model for assessing the resource and energy efficiency of software products and components—green software measurement model (gsmm), Future Gener. Comput. Syst. 155 (2024) 402–418. [17] J. Zhang, Z. Wang, H. Wang, T. Song, H.-a. Su, R. Chen, Y. Hua, X. Zhou, R. Ma, M. Pan, H. Guan, Ampere: A generic energy
Page 18 of 26
Energy-Aware Computing in 2026 estimation approach for on-device training, SIGMETRICS Perform. Eval. Rev. 53 (2025) 27–32. [18] Z. Zong, R. Ge, Q. Gu, Marcher: A heterogeneous system supporting energy-aware high performance computing and big data analytics, Big Data Research 8 (2017) 27–38. Tutorials on Tools and Methods using High Performance Computing resources for Big Data. [19] M. Weber, J. Dorn, S. Apel, N. Siegmund, Famlem, the fast modular energy meter at code level, in: Proc. 33rd ACM Int. Conf. Found. Soft. Eng., FSE Companion ’25, ACM, 2025, p. 1129–1133. doi:10. 1145/3696630.3728595. [20] J. Treibig, G. Hager, G. Wellein, Likwid: A lightweight performance-oriented tool suite for x86 multicore environments, in: 2010 39th International Conference on Parallel Processing Workshops, 2010, pp. 207–216. [21] S. Rajput, M. Kechagia, F. Sarro, T. Sharma, Greenlight: Highlighting tensorflow apis energy footprint, in: Proc. IEEE/ACM 21st Int. Conf. Min. Softw. Repos. (MSR), 2024, pp. 304–308. [22] M. Paipuri, Ceems: A resource manager agnostic energy and emissions monitoring stack, in: Proc. SC ’24 Workshops, SC-W ’24, IEEE Press, 2025, p. 1862–1866. doi:10.1109/SCW63240.2024.00233. [23] C. Terboven, R. Liem, J. Gracia, K. Haldar, J. Engels, P. Giesselmann, D. Brayford, T. Wilde, C. Simmendinger, M. Marquardt, J. Eitzinger, T. Gruber, Ee-hpc a framework for energy efficient hpc system management, in: Proc. SC ’24 Workshops, 2024, pp. 1878– 1882. doi:10.1109/SCW63240.2024.00236. [24] T. A. Rahmani, G. Belalem, S. A. Mahmoudi, O. R. Merad-Boudia, Real-time ai-powered monitoring for energy-efficient scheduling in multi-node heterogeneous systems, Future Gener. Comput. Syst. 181 (2026) 108428. [25] K. A. Sanni, A. G. Andreou, A historical perspective on hardware ai inference, charge-based computational circuits and an 8 bit chargebased multiply-add core in 16 nm finfet cmos, IEEE J. Emerg. Sel Top. Circ. Syst. 9 (2019) 532–543. [26] K. Wang, H. Zheng, Y. Li, A. Louri, Securenoc: A learning-enabled, high-performance, energy-efficient, and secure on-chip communication framework design, IEEE Trans. Sustain. Comput. 7 (2022) 709– 723. [27] M. H. Sadi, C. Sudarshan, S. R. Nassif, N. Wehn, Efficient deep neural network training with a novel 5.3-bit block floating point data format, IEEE Trans. Circ Syst. A.I. 2 (2025) 150–161. [28] G. Li, S. Ye, C. Chen, Y. Wang, F. Yang, T. Cao, C. Liu, M. M. S. Aly, M. Yang, Lut-dla: Lookup table as efficient extreme low-bit deep learning accelerator, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2025, pp. 671–684. doi:10.1109/HPCA61900.2025. 00057. [29] S. Laskar, P. Majhi, A. Muzahid, E. J. Kim, Supermesh: Energyefficient collective communications for accelerators, in: Proceedings of the 58th IEEE/ACM International Symposium on Microarchitecture, MICRO ’25, ACM, 2025, p. 1640–1655. doi:10.1145/3725843. 3756085. [30] Y. Liu, Y. Xue, N. Crawford, J. Xue, J. Huang, Elk: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques, ACM, 2025, p. 1284–1299. [31] T. Li, H. Zhong, Y. Xu, V. Narayanan, K. Ni, H. Yang, T. Kämpfe, X. Li, Remna: Variation-resilient and energy-efficient mlc fefet computing-in-memory using nand flash-like read and adaptive control, in: Proc. 43rd IEEE/ACM Int. Conf. Computer Aided Des. (ICCAD), ICCAD ’24, ACM, 2025. doi:10.1145/3676536.3676715. [32] D. R. Dipta, B. Gulmezoglu, Mad-en: Microarchitectural attack detection through system-wide energy consumption, IEEE Trans. Inform. Forens. Secur. 18 (2023) 3006–3017. [33] Z. Ebrahimi, A. Kumar, Green: An approximate simd/mimd cgra for energy-efficient processing at the edge, IEEE Trans. Comp.-Aid. Des. Integr. Circ. Syst. 43 (2024) 2874–2887. [34] F. Fayza, C. Demirkiran, S. P. Rao, D. Bunandar, U. Gupta, A. Joshi, Epicarbon: A carbon modeling tool for electro-photonic accelerators, in: 2025 IEEE/ACM Int. Conf. Computer Aided Des. (ICCAD), 2025, pp. 1–9. doi:10.1109/ICCAD66269.2025.11240809.
[35] U. Gupta, M. Elgamal, G. Hills, G.-Y. Wei, H.-H. S. Lee, D. Brooks, C.-J. Wu, Act: designing sustainable computer systems with an architectural carbon modeling tool, in: ACM/IEEE 49th Annu. Int. Symp. Comput. Archit. (ISCA), ISCA ’22, ACM, 2022, p. 784–799. doi:10.1145/3470496.3527408. [36] M. Pashaeifar, M. Kamal, A. Afzali-Kusha, M. Pedram, A theoretical framework for quality estimation and optimization of dsp applications using low-power approximate adders, IEEE Trans. Circ Syst. 66 (2019) 327–340. [37] D. Danopoulos, G. Zervakis, D. Soudris, J. Henkel, Transaxx: Efficient transformers with approximate computing, IEEE Trans. Circ Syst. Art. Intell. 2 (2025) 288–301. [38] A. Golden, M. Elgamal, A. Mahmoud, G. Hills, C.-J. Wu, G.Y. Wei, D. Brooks, Wafer-scale systems: A carbon perspective, SIGENERGY Energy Inform. Rev. 5 (2025) 118–124. [39] J. Kalyanapu, F. Dizani, D. Asher, A. Ghanbari, R. Cammarota, A. Aysu, S. M. Ajorpaz, GateBleed: Exploiting On-Core Accelerator Power Gating for High Performance and Stealthy Attacks on AI, ACM, 2025, p. 308–325. [40] W. Abu Ahmad, A. Bartolini, F. Beneventi, L. Benini, A. Borghesi, M. Cicala, P. Forestieri, C. Gianfreda, D. Gregori, A. Libri, F. Spiga, S. Tinti, Design of an energy aware petaflops class high performance cluster based on power architecture, in: IEEE IPDPS’ 17 Workshops, 2017, pp. 964–973. doi:10.1109/IPDPSW.2017.22. [41] M. Cochet, K. Swaminathan, E. Loscalzo, J. Zuckerman, M. C. D. Santos, D. Giri, A. Buyuktosunoglu, T. Jia, D. Brooks, G.-Y. Wei, K. Shepard, L. P. Carloni, P. Bose, Blitzcoin: Fully decentralized hardware power management for accelerator-rich socs, in: ACM/IEEE 51st Annu. Int. Symp. Comput. Archit. (ISCA), 2024, pp. 801–817. doi:10.1109/ISCA59077.2024.00063. [42] A. M. Panteleaki, K. Balaskas, G. Zervakis, H. Amrouch, I. Anagnostopoulos, Late breaking results: Leveraging approximate computing for carbon-aware dnn accelerators, in: Des. Autom. Test Eur. Conf. Exhib. (DATE), 2025, pp. 1–2. doi:10.23919/DATE64628.2025. 10993191. [43] C. C. Sudarshan, A. Arora, V. A. Chhabria, Greenfpga: Evaluating fpgas as environmentally sustainable computing solutions, in: 2024 61st ACM/IEEE Design Autom. Conf. (DAC), 2024, pp. 1–6. [44] A. Darabi, E. Abiri, Cgi-sram: Memory cell with data-aware write bitline-free and inner readout features for energy-efficient bitwise logic-in-memory operations, Sustain. Comput. Inf. Syst. 50 (2026) 101307. [45] J. Spieck, D. Walter, J. Waschkeit, J. Teich, Co-design of sustainable embedded systems-on-chip, in: Des. Autom. Test Eur. Conf. Exhib. (DATE), 2025, pp. 1–2. doi:10.23919/DATE64628.2025.10992914. [46] K. Shiflett, A. Karanth, R. Bunescu, A. Louri, Albireo: Energyefficient acceleration of convolutional neural networks via silicon photonics, in: ACM/IEEE 48th Annu. Int. Symp. Comput. Archit. (ISCA), 2021, pp. 860–873. doi:10.1109/ISCA52012.2021.00072. [47] J. Li, A. Louri, A. Karanth, R. Bunescu, Gcnax: A flexible and energy-efficient accelerator for graph convolutional neural networks, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2021, pp. 775–788. doi:10.1109/HPCA51647.2021.00070. [48] X. Li, T. Zhou, H. Wang, M. Lin, Energy-efficient computation with dvfs using deep reinforcement learning for multi-task systems in edge computing, IEEE Trans. Sustain. Comput. 10 (2025) 1116– 1127. [49] A. Swaraj, S. Kumar, Is chatgpt-generated code really green?: Evaluating ai-generated solutions for energy-efficient coding practices, IEEE Software 43 (2026) 42–50. [50] S. Georgiou, M. Kechagia, T. Sharma, F. Sarro, Y. Zou, Green ai: do deep learning frameworks have different costs?, in: Proc. 44th Int. Conf. Soft. Eng., ICSE ’22, ACM, 2022, p. 1082–1094. doi:10.1145/3510003.3510221. [51] S. Sikand, V. S. Sharma, V. Kaulgud, S. Podder, Green ai quotient: Assessing greenness of ai-based software and the way forward, in: Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering, ASE ’23, IEEE Press, 2024, p.
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 19 of 26
Energy-Aware Computing in 2026 1828–1833. doi:10.1109/ASE56229.2023.00115. [52] D. Cesarini, A. Bartolini, P. Bonfà, C. Cavazzoni, L. Benini, Countdown: A run-time library for performance-neutral energy saving in mpi applications, IEEE Trans. Comput. 70 (2021) 682–695. [53] V. Stoico, A. C. Dragomir, P. Lago, An empirical study on the performance and energy usage of compiled python code, in: Proc. 29th Int. Conf. Eval. Assess. Softw. Eng. (EASE), EASE ’25, ACM, 2025, p. 46–56. doi:10.1145/3756681.3756972. [54] J. Song, P. Zhu, Y. Zhang, G. Yu, Cloudsimper: Simulating geodistributed datacenters powered by renewable energy mix, IEEE Trans. Parallel. Distrib. Syst. 35 (2024) 531–547. [55] T. Tunzina, M. Mashira, M. M. Islam, Greenai: A comparative analysis of environmental efficiency in llm-generated code, IEEE Access 14 (2026) 18654–18672. [56] V. De Martino, S. Martínez-Fernández, F. Palomba, Do developers adopt green architectural tactics for ml-enabled systems? a mining software repository study, in: IEEE/ACM 47th Int. Conf. Soft. Eng. (ICSE-SEIS), 2025, pp. 135–139. doi:10.1109/ICSE-SEIS66351.2025. 00019. [57] N. R. Shah, M. V. V. S. M. Kumar, D. Baxi, R. Upadrasta, Polyufc: Polyhedral compilation meets roofline analysis for uncore frequency capping, in: 2026 IEEE/ACM International Symposium on Code Generation and Optimization (CGO), 2026, pp. 563–576. doi:10. 1109/CGO68049.2026.11395211. [58] M. Hort, A. Grishina, L. Moonen, An exploratory literature study on sharing and energy use of language models for source code, in: ACM/IEEE Int. Symp. Emp. Soft. Eng Meas. (ESEM), 2023, pp. 1–12. doi:10.1109/ESEM56168.2023.10304803. [59] B. Ayoola, M. Kuutila, R. R. Wehbe, P. Ralph, User personas improve social sustainability by encouraging software developers to deprioritize antisocial features, in: IEEE/ACM 47th In. Conf. Soft. Eng. (ICSE), 2025, pp. 833–845. doi:10.1109/ICSE55347.2025.00135. [60] J. Chen, L. Tan, P. Wu, D. Tao, H. Li, X. Liang, S. Li, R. Ge, L. Bhuyan, Z. Chen, Greenla: Green linear algebra software for gpuaccelerated heterogeneous computing, in: SC ’16: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, 2016, pp. 667–677. doi:10.1109/SC.2016.56. [61] N. v. Kempen, H.-J. Kwon, D. T. Nguyen, E. D. Berger, It’s not easy being green: On the energy efficiency of programming languages, in: 40th IEEE/ACM Int. Conf. Aut. Soft. Eng. (ASE), 2025, pp. 1553– 1565. doi:10.1109/ASE63991.2025.00131. [62] L. Cruz, X. Franch, S. Martínez-Fernández, Innovating for tomorrow: The convergence of software engineering and green ai, ACM Trans. Softw. Eng. Methodol. 34 (2025). [63] V. De Martino, S. Lambiase, F. Pecorelli, W.-J. van den Heuvel, F. Ferrucci, F. Palomba, Sustainability of machine learning-enabled systems: The machine learning practitioner’s perspective, ACM Trans. Softw. Eng. Methodol. (2025). Just Accepted. [64] E. B. Roque, L. Cruz, T. Durieux, Unveiling the Energy Vampires: A Methodology for Debugging Software Energy Consumption, IEEE Press, 2025, p. 2406–2418. [65] J. Maquoi, M. Cauz, B. Vanderose, X. Devroey, Energy codesumption, leveraging test execution for source code energy consumption analysis, in: Proc.33rd ACM Int. Conf. Found. Soft. Eng., FSE Companion ’25, ACM, 2025, p. 1432–1436. doi:10.1145/3696630. 3728707. [66] N. Alizadeh, B. Belchev, N. Saurabh, P. Kelbert, F. Castor, Language models in software development tasks: An experimental analysis of energy and accuracy, in: Proc. IEEE/ACM 22nd Int. Conf. Min. Softw. Repos. (MSR), IEEE Press, 2025, p. 725–736. doi:10.1109/ MSR66628.2025.00109. [67] G. Zanfardino, M. A. Memon, M. Tucci, Enhancing energy efficiency with reusable software ecosystems and persona-based ui/ux, in: Proc.33rd ACM Int. Conf. Found. Soft. Eng., FSE Companion ’25, ACM, 2025, p. 1449–1452. doi:10.1145/3696630.3728710.
[68] A. Tundo, M. Mobilio, S. Ilager, I. Brandić, E. Bartocci, L. Mariani, An energy-aware approach to design self-adaptive ai-based applications on the edge, in: Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering, ASE ’23, IEEE Press, 2024, p. 281–293. doi:10.1109/ASE56229.2023.00046. [69] J. Andrés Larracoechea, P. Roose, S. Ilarri, Energy-aware software design: A hardware and behavioural consumption scoring and labelling algorithm, Expert Systems with Applications 299 (2026) 130046. [70] R. Dhawan, M. Suresh Patel, Carbon-aware ai control plane for devops automation: A reference architecture and next-generation sustainability framework, IEEE Access 14 (2026) 11163–11184. [71] A. Marahatta, Q. Xin, C. Chi, F. Zhang, Z. Liu, Pefs: Ai-driven prediction based energy-aware fault-tolerant scheduling scheme for cloud data center, IEEE Trans. Sustain. Comput. 6 (2021) 655–666. [72] J. S. Shim, B. Han, Y. Kim, J. Kim, Deeppm: Transformer-based power and performance prediction for energy-aware software, in: Des. Autom. Test Eur. Conf. Exhib. (DATE), 2022, pp. 1491–1496. doi:10.23919/DATE54114.2022.9774589. [73] M. A. N. Al-hayanni, A. Rafiev, F. Xia, R. Shafik, A. Romanovsky, A. Yakovlev, Parma: Parallelization-aware run-time management for energy-efficient many-core systems, IEEE Trans. Comput. 69 (2020) 1507–1518. [74] L. Eeckhout, Assessing processor sustainability using the first-order focal carbon model, IEEE Micro 45 (2025) 11–18. [75] Z. Zhou, M. Shojafar, J. Abawajy, H. Yin, H. Lu, Ecms: An edge intelligent energy efficient model in mobile edge computing, IEEE Trans. Green Comm. Network. 6 (2022) 238–247. [76] C. Conficoni, A. Bartolini, A. Tilli, C. Cavazzoni, L. Benini, Integrated energy-aware management of supercomputer hybrid cooling systems, IEEE Trans. Indus. Inform. 12 (2016) 1299–1311. [77] S. Vahabi, F. Righetti, C. Vallati, N. Tonellotto, The impact of prediction models on energy-aware resource management in faas platforms, IEEE Access 13 (2025) 85711–85727. [78] K. Menear, A. Wilkinson, T. Dykes, U.-U. Haus, D. Duplyakin, Energy-aware hpc scheduling with llm-based power prediction, in: Proc. SC ’25 Workshops, 2025, pp. 1997–2006. doi:10.1145/ 3731599.3767563. [79] I. Alan, E. Arslan, T. Kosar, Energy-aware data transfer algorithms, in: SC ’15: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, 2015, pp. 1–12. doi:10.1145/2807591.2807628. [80] A. Venkatesh, A. Vishnu, K. Hamidouche, N. Tallent, D. Panda, D. Kerbyson, A. Hoisie, A case for application-oblivious energyefficient mpi runtime, in: SC ’15: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, 2015, pp. 1–12. doi:10.1145/2807591. 2807658. [81] Y. Xia, M. Zhou, X. Luo, S. Pang, Q. Zhu, A stochastic approach to analysis of energy-aware dvs-enabled cloud datacenters, IEEE Trans. Syst., Man, Cybern.: Syst. 45 (2015) 73–83. [82] M. B. Sikal, J. González-Gómez, H. Khdr, J. Henkel, Contentionaware forecasting of energy efficiency through sequence-based models in modern heterogeneous processors, in: 2025 62nd ACM/IEEE Design Autom. Conf. (DAC), 2025, pp. 1–7. doi:10.1109/DAC63849. 2025.11132825. [83] R. Schoonhoven, B. Veenboer, B. Van Werkhoven, K. J. Batenburg, Going green: optimizing gpus for energy efficiency through modelsteered auto-tuning, in: PMBS’ 22 Workshop, 2022, pp. 48–59. doi:10.1109/PMBS56514.2022.00010. [84] A. K. Kakolyris, D. Masouros, P. Vavaroutsos, S. Xydis, D. Soudris, throttll’em: Predictive gpu throttling for energy efficient llm inference serving, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2025, pp. 1363–1378. doi:10.1109/HPCA61900.2025.00103. [85] Y. Jiang, R. Kanakagiri, R. Basu Roy, D. Tiwari, Thirstyflops: Water footprint modeling and analysis toward sustainable hpc systems, in: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, SC ’25, ACM, 2025, p. 855–869. doi:10.1145/3712285.3759804.
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 20 of 26
Energy-Aware Computing in 2026 [86] Z. Wang, Y. Zhang, F. Wei, B. Wang, Y. Liu, Z. Hu, J. Zhang, X. Xu, J. He, X. Wang, W. Dou, G. Chen, C. Tian, Using analytical performance/power model and fine-grained dvfs to enhance ai accelerator energy efficiency, in: Proc. 30th ACM Int. Conf. Archit. Support Program. Lang. Oper. Syst. (ASPLOS), Vol. 1, ASPLOS ’25, ACM, 2025, p. 1118–1132. doi:10.1145/3669940.3707231. [87] P. S.K., S. Taluri, B. S, L. Karwa, Y. Simmhan, Powertrain: Fast, generalizable time and power prediction models to optimize dnn training on accelerated edges, Future Gener. Comput. Syst. 161 (2024) 329–344. [88] E. Cesario, P. Lindia, F. Lobello, A. Vinci, S. Capalbo, Enhancing energy efficiency in cloud computing through regression models: A data-driven approach with experimental validation, Future Gener. Comput. Syst. 177 (2026) 108245. [89] G. Da Costa, Hardware and application aware performance, power and energy models for modern hpc servers with dvfs, Sustain. Comput. Inf. Syst. 46 (2025) 101106. [90] J. Sun, H. Cho, A lightweight optimal scheduling algorithm for energy-efficient and real-time cloud services, IEEE Access 10 (2022) 5697–5714. [91] U. Wajid, C. Cappiello, P. Plebani, B. Pernici, N. Mehandjiev, M. Vitali, M. Gienger, K. Kavoussanakis, D. Margery, D. G. Perez, P. Sampaio, On achieving energy efficiency and reducing co2 footprint in cloud computing, IEEE Trans. Cloud Comput. 4 (2016) 138–151. [92] K. M. Tarplee, R. Friese, A. A. Maciejewski, H. J. Siegel, E. K. P. Chong, Energy and makespan tradeoffs in heterogeneous computing systems using efficient linear programming techniques, IEEE Trans. Parallel. Distrib. Syst. 27 (2016) 1633–1646. [93] Y. Georgiou, D. Glesser, K. Rzadca, D. Trystram, A schedulerlevel incentive mechanism for energy efficiency in hpc, in: 15th IEEE/ACM Int. Symp. Cluster, Cloud Grid Comput. (CCGRID), 2015, pp. 617–626. doi:10.1109/CCGrid.2015.101. [94] S. Wang, Z. Qian, J. Yuan, I. You, A dvfs based energy-efficient tasks scheduling in a data center, IEEE Access 5 (2017) 13090–13102. [95] Q. Wang, X. Mei, H. Liu, Y.-W. Leung, Z. Li, X. Chu, Energyaware non-preemptive task scheduling with deadline constraint in dvfs-enabled heterogeneous clusters, IEEE Trans. Parallel. Distrib. Syst. 33 (2022) 4083–4099. [96] H. Ali, U. U. Tariq, Y. Zheng, X. Zhai, L. Liu, Contention & energyaware real-time task mapping on noc based heterogeneous mpsocs, IEEE Access 6 (2018) 75110–75123. [97] S. K. Sahoo, A. Dash, S. K. Mishra, D. Puthal, Energy-efficient federated learning scheduler for vm allocation in edge-cloud continuum, IEEE Access 14 (2026) 4274–4291. [98] A. Jayanetti, S. Halgamuge, R. Buyya, Multi-agent deep reinforcement learning framework for renewable energy-aware workflow scheduling on distributed cloud data centers, IEEE Trans. Parallel. Distrib. Syst. 35 (2024) 604–615. [99] C. Xu, K. Wang, P. Li, R. Xia, S. Guo, M. Guo, Renewable energy-aware big data analytics in geo-distributed data centers with reinforcement learning, IEEE Trans. Net. Sci Eng. 7 (2020) 205– 215. [100] V. R. De Carvalho, J. S. Sichman, Toward cost-efficiency and reduced carbon footprint: A multi-armed bandit hyper-heuristic for cloud scheduling problems, IEEE Access 13 (2025) 204656– 204675. [101] B. M. Beena, P. C. Ranga, T. S. S. Manideep, S. Saragadam, G. Karthik, A green cloud-based framework for energy-efficient task scheduling using carbon intensity data for heterogeneous cloud servers, IEEE Access 13 (2025) 73916–73938. [102] X. Tang, F. Liu, D. Xu, J. Jiang, Q. Tang, B. Wang, Q. Wu, C. L. Philip Chen, Llm-assisted reinforcement learning: Leveraging lightweight large language model capabilities for efficient task scheduling in multi-cloud environment, IEEE Trans. Cons. Electr. 71 (2025) 5631–5644. [103] Y. Li, Y. He, J. Lin, Z. Xu, S. Zhang, A reinforcement learningbased population hyper-heuristic for energy-efficient cloud workflow
scheduling problem, IEEE Trans. Serv. Comput. 18 (2025) 2545– 2558. [104] J. Peng, K. Li, J. Chen, K. Li, Reliability/performance-aware scheduling for parallel applications with energy constraints on heterogeneous computing systems, IEEE Trans. Sustain. Comput. 7 (2022) 681–695. [105] W. Guan, Y. K. Zhao, C. Ababei, U-duct: Uncertainty-aware dynamic unified carbon modeling tool for datacenter scheduling, in: IEEE IGSC’24, 2024, pp. 29–34. doi:10.1109/IGSC64514.2024.00014. [106] J. Chen, M. Manivannan, M. Abduljabbar, M. Pericàs, Erase: Energy efficient task mapping and resource management for work stealing runtimes, ACM Trans. Archit. Code Optim. 19 (2022). [107] J. Ding, H. Hoffmann, Dps: Adaptive power management for overprovisioned systems, in: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, SC ’23, ACM, 2023. doi:10.1145/3581784. 3607091. [108] H. Desai, X. Wang, B. Lucia, Energy-aware scheduling and input buffer overflow prevention for energy-harvesting systems, in: Proc. 30th ACM Int. Conf. Archit. Support Program. Lang. Oper. Syst. (ASPLOS), Vol. 2, ASPLOS ’25, ACM, 2025, p. 339–354. doi:10. 1145/3676641.3715995. [109] F. Qiao, Y. Fang, A. Cidon, Energy-aware process scheduling in linux, SIGENERGY Energy Inform. Rev. 4 (2025) 91–97. [110] R. Chen, W. Lin, H. Huang, X. Ye, Z. Peng, Gas-marl: Green-aware job scheduling algorithm for hpc clusters based on multi-action deep reinforcement learning, Future Gener. Comput. Syst. 167 (2025) 107760. [111] B. Qureshi, Profile-based power-aware workflow scheduling framework for energy-efficient data centers, Future Gener. Comput. Syst. 94 (2019) 453–467. [112] H. A. Hassan, S. A. Salem, E. M. Saad, A smart energy and reliability aware scheduling algorithm for workflow execution in dvfs-enabled cloud environment, Future Gener. Comput. Syst. 112 (2020) 431–448. [113] X. Li, X. Jiang, P. Garraghan, Z. Wu, Holistic energy and failure aware workload scheduling in cloud datacenters, Future Gener. Comput. Syst. 78 (2018) 887–900. [114] J. Luo, D. El Baz, R. Xue, J. Hu, Solving the dynamic energy aware job shop scheduling problem with the heterogeneous parallel genetic algorithm, Future Gener. Comput. Syst. 108 (2020) 119–134. [115] Y. Hu, C. Liu, K. Li, X. Chen, K. Li, Slack allocation algorithm for energy minimization in cluster systems, Future Gener. Comput. Syst. 74 (2017) 119–131. [116] D. Ding, X. Fan, Y. Zhao, K. Kang, Q. Yin, J. Zeng, Q-learning based dynamic task scheduling for energy-efficient cloud computing, Future Gener. Comput. Syst. 108 (2020) 361–371. [117] A. Mencaroni, P. Leyman, B. Raa, S. De Vuyst, D. Claeys, Towards net-zero manufacturing: Carbon-aware scheduling for ghg emissions reduction, J. Clean. Product. 529 (2025) 146787. [118] A. Kassab, J.-M. Nicod, L. Philippe, V. Rehn-Sonigo, Green power aware approaches for scheduling independent tasks on a multi-core machine, Sustain. Comput. Inf. Syst. 31 (2021) 100590. [119] Q. Wang, Y. Hao, Y. Xu, T. Ma, X. Zhang, An energy-aware scheduling method for parallel tasks based on an adaptive differential evolution algorithm in a multi-cloud environment, Expert Systems with Applications 296 (2026) 129008. [120] L. Aminoroaya, R. Khorsand, M. Mosleh, I_moaft: An intelligent multi-layered framework for energy-efficient and low-latency task scheduling in fog-cloud computing, Sustain. Comput. Inf. Syst. 50 (2026) 101327. [121] A. K. Singh, S. R. Swain, D. Saxena, C.-N. Lee, A bio-inspired virtual machine placement toward sustainable cloud resource management, IEEE Systems Journal 17 (2023) 3894–3905. [122] S. Singhal, S. Ali, M. Awasthy, D. K. Shukla, R. Tiwari, Rockhyrax: An energy efficient job scheduling using cluster of resources in cloud computing environment, Sustain. Comput. Inf. Syst. 42 (2024) 100985.
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 21 of 26
Energy-Aware Computing in 2026 [123] T. Hidayat, K. Ramli, R. Harwahyu, M. Salman, T. Surya Gunawan, Reinforcement learning-driven hybrid precopy/postcopy vm migration for energy-efficient data centers, IEEE Access 13 (2025) 169521–169533. [124] Z. Li, X. Yu, L. Yu, S. Guo, V. Chang, Energy-efficient and qualityaware vm consolidation method, Future Gener. Comput. Syst. 102 (2020) 789–809. [125] S. Banerjee, S. Roy, S. Khatua, Towards energy and qos aware dynamic vm consolidation in a multi-resource cloud, Future Gener. Comput. Syst. 157 (2024) 376–391. [126] M. Hosseini Shirvani, An energy-efficient topology-aware virtual machine placement in cloud datacenters: A multi-objective discrete jaya optimization, Sustain. Comput. Inf. Syst. 38 (2023) 100856. [127] A. M. Baydoun, A. S. Zekri, Hapso: An aco-initialized, discretization-aware pso for energy- and carbon-efficient vm consolidation in green cloud datacenters, Sustain. Comput. Inf. Syst. 48 (2025) 101258. [128] D. Gu, Y. Zhao, P. Sun, X. Jin, X. Liu, Greenflow: A carbon-efficient scheduler for deep learning workloads, IEEE Trans. Parallel. Distrib. Syst. 36 (2025) 168–184. [129] Y. Chen, Q. Zhang, R. Xing, Y. Li, X. Ma, Y. Zhang, A. Zhou, S. Wang, Slice: Energy-efficient satellite-ground co-inference via layer-wise scheduling optimization, IEEE Trans. Serv. Comput. 18 (2025) 2388–2402. [130] J. Stojkovic, C. Zhang, I. n. Goiri, E. Choukse, H. Qiu, R. Fonseca, J. Torrellas, R. Bianchini, Tapas: Thermal- and power-aware scheduling for llm inference in cloud platforms, in: Proc. 30th ACM Int. Conf. Archit. Support Program. Lang. Oper. Syst. (ASPLOS), Vol. 2, ASPLOS ’25, ACM, 2025, p. 1266–1281. doi:10.1145/ 3676641.3716025. [131] P.-F. Dutot, Y. Georgiou, D. Glesser, L. Lefevre, M. Poquet, I. Rais, Towards energy budget control in hpc, in: 17th IEEE/ACM Int. Symp. Cluster, Cloud Grid Comput. (CCGRID), 2017, pp. 381–390. doi:10.1109/CCGRID.2017.16. [132] H. Shen, H. Wang, J. Gao, R. Buyya, An instability-resilient renewable energy allocation system for a cloud datacenter, IEEE Trans. Parallel. Distrib. Syst. 34 (2023) 1020–1034. [133] A. L. V. Solórzano, K. Sato, K. Yamamoto, F. Shoji, J. M. Brandt, B. Schwaller, S. P. Walton, J. Green, D. Tiwari, Toward sustainable hpc: In-production deployment of incentive-based power efficiency mechanism on the fugaku supercomputer, in: SC24: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, 2024, pp. 1–16. doi:10.1109/SC41406.2024.00030. [134] J. Yang, W. Xiao, C. Jiang, M. S. Hossain, G. Muhammad, S. U. Amin, Ai-powered green cloud and data center, IEEE Access 7 (2019) 4195–4203. [135] Q. Fettes, M. Clark, R. Bunescu, A. Karanth, A. Louri, Dynamic voltage and frequency scaling in nocs with supervised and reinforcement learning techniques, IEEE Trans. Comput. 68 (2019) 375–389. [136] M. Elgamal, D. Carmean, E. Ansari, O. Zed, R. Peri, S. Manne, U. Gupta, G.-Y. Wei, D. Brooks, G. Hills, C.-J. Wu, Cordoba: Carbon-efficient optimization framework for computing systems, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2025, pp. 1289–1303. doi:10.1109/HPCA61900.2025.00098. [137] C. W. Zaw, S. R. Pandey, K. Kim, C. S. Hong, Energy-aware resource management for federated learning in multi-access edge computing systems, IEEE Access 9 (2021) 34938–34950. [138] K. Fan, M. D’Antonio, L. Carpentieri, B. Cosenza, F. Ficarelli, D. Cesarini, Synergy: Fine-grained energy-efficient heterogeneous computing for scalable energy saving, in: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, SC ’23, ACM, 2023. doi:10. 1145/3581784.3607055. [139] A. Singh, G. S. Aujla, R. S. Bali, Container-based load balancing for energy efficiency in software-defined edge computing environment, Sustain. Comput. Inf. Syst. 30 (2021) 100463. [140] Y. Nan, W. Li, W. Bao, F. C. Delicato, P. F. Pires, Y. Dou, A. Y. Zomaya, Adaptive energy-aware computation offloading for cloud of things systems, IEEE Access 5 (2017) 23947–23957.
[141] D. Zhao, J.-t. Zhou, K. Li, Cfws: Drl-based framework for energy cost and carbon footprint optimization in cloud data centers, IEEE Trans. Sustain. Comput. 10 (2025) 95–107. [142] A. Zhong, D. Wu, Z. Zhu, B. Yang, R. Wang, Dual-timescale nonlinear energy optimization for cloud resource provisioning, IEEE Trans. Cloud Comput. 14 (2026) 177–191. [143] H. Huang, W. Lin, J. Lin, K. Li, Power management optimization for data centers: A power supply perspective, IEEE Trans. Sustain. Comput. 10 (2025) 784–803. [144] Y. Son, U. Gupta, A. McCrabb, Y. G. Kim, V. Bertacco, D. Brooks, C.-J. Wu, Greenscale: Carbon optimization for edge computing, IEEE Internet of Things J. 12 (2025) 32379–32393. [145] Z. Zheng, S. Sultanov, M. E. Papka, Z. Lan, Minimizing power waste in heterogenous computing via adaptive uncore scaling, in: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, SC ’25, ACM, 2025, p. 505–518. doi:10.1145/3712285.3759879. [146] W. A. Hanafy, R. Bostandoost, N. Bashir, D. Irwin, M. Hajiesmaili, P. Shenoy, The war of the efficiencies: Understanding the tension between carbon and energy optimization, SIGENERGY Energy Inform. Rev. 4 (2024) 87–93. [147] M. Jayaweera, M. Kong, Y. Wang, D. Kaeli, Energy-aware tile size selection for affine programs on gpus, in: Proceedings of the 2024 IEEE/ACM Int. Symp. Cod. Gen Optim., CGO ’24, IEEE Press, 2024, p. 13–27. doi:10.1109/CGO57630.2024.10444795. [148] A. Krzywaniak, P. Czarnul, J. Proficz, Dynamic gpu power capping with online performance tracing for energy efficient gpu computing using depo tool, Future Gener. Comput. Syst. 145 (2023) 396–414. [149] G. Verma, Sustainable cost-energy aware load balancing in cloud environment using intelligent optimization, Sustain. Comput. Inf. Syst. 46 (2025) 101115. [150] S. MirhoseiniNejad, H. Moazamigoodarzi, G. Badawy, D. G. Down, Joint data center cooling and workload management: A thermalaware approach, Future Gener. Comput. Syst. 104 (2020) 174–186. [151] H. Kahil, S. Sharma, P. Välisuo, M. Elmusrati, Reinforcement learning for data center energy efficiency optimization: A systematic literature review and research roadmap, Applied Energy 389 (2025) 125734. [152] U.S. Environmental Protection Agency, Per- and polyfluoroalkyl substances (pfas), https://www.epa.gov/pfas, 2024. [153] Y. Wu, I. Hua, Y. Ding, Not all water consumption is equal: A water stress weighted metric for sustainable computing, SIGENERGY Energy Inform. Rev. 5 (2025) 84–90. [154] A. Kilian, H. de Meer, G. Schomaker, Energy-optimized supercomputer networks using wind energy, Commun. ACM 68 (2025) 74–79. [155] E. Strubell, A. Ganesh, A. Mccallum, Energy and policy considerations for deep learning in nlp, 2019, pp. 3645–3650. doi:10.18653/ v1/P19-1355. [156] N. Hogade, S. Pasricha, Game-theoretic deep reinforcement learning to minimize carbon emissions and energy costs for ai inference workloads in geo-distributed data centers, IEEE Trans. Sustain. Comput. 10 (2025) 628–641. [157] S. Nasrin, A. Shylendra, N. Darabi, T. Tulabandhula, W. Gomes, A. Chakrabarty, A. R. Trivedi, Enos: Energy-aware network operator search in deep neural networks, IEEE Access 10 (2022) 81447– 81457. [158] Y. Yao, B. Zhu, Y. Xiao, H. Liu, Trade-offs between power consumption and response time in deep learning systems: A queueing model perspective, Sustain. Comput. Inf. Syst. 48 (2025) 101220. [159] P. Patel, E. Choukse, C. Zhang, I. n. Goiri, B. Warrier, N. Mahalingam, R. Bianchini, Characterizing power management opportunities for llms in the cloud, in: Proc. 29th ACM Int. Conf. Archit. Support Program. Lang. Oper. Syst. (ASPLOS), Vol. 3, ASPLOS ’24, ACM, 2024, p. 207–222. doi:10.1145/3620666.3651329. [160] A. Tschand, A. T. R. Rajan, S. Idgunji, A. Ghosh, J. Holleman, C. Kiraly, P. Ambalkar, R. Borkar, R. Chukka, T. Cockrell, O. Curtis, G. Fursin, M. Hodak, H. Kassa, A. Lokhmotov, D. Miskovic, Y. Pan, M. P. Manmathan, L. Raymond, T. S. John, A. Suresh, R. Taubitz, S. Zhan, S. Wasson, D. Kanter, V. J. Reddi, Mlperf
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 22 of 26
Energy-Aware Computing in 2026 power: Benchmarking the energy efficiency of machine learning systems from µwatts to mwatts for sustainable ai, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2025, pp. 1201– 1216. doi:10.1109/HPCA61900.2025.00092. [161] M. M. A. Ifath, I. Haque, Characterizing performance–energy tradeoffs of large language models in multi-request workflows, Proc. ACM Meas. Anal. Comput. Syst. 10 (2026). [162] C. M. John, S. Nassyr, C. Penke, A. Herten, Performance and power: Systematic evaluation of ai workloads on accelerators with caraml, in: Proc. SC ’24 Workshops, SC-W ’24, IEEE Press, 2025, p. 1164–1176. doi:10.1109/SCW63240.2024.00158. [163] Z. Wang, T. Luo, R. S. M. Goh, J. T. Zhou, Edcompress: Energyaware model compression for dataflows, IEEE Trans. Neural Net. Learn. Syst. 35 (2024) 208–220. [164] J. Stojkovic, C. Zhang, I. Goiri, J. Torrellas, E. Choukse, Dynamollm: Designing llm inference clusters for performance and energy efficiency, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2025, pp. 1348–1362. doi:10.1109/HPCA61900.2025.00102. [165] J. Xia, Y. Zhang, Y. Shi, Towards energy-aware federated learning via marl: A dual-selection approach for model and client, in: ACM/IEEE Int. Conf. Computer Aided Des. (ICCAD), 2024, pp. 1–9. doi:10.1145/3676536.3676815. [166] M. Alswaitti, R. Verdecchia, G. Danoy, P. Bouvry, J. E. Pecero, Training green ai models using elite samples, IEEE Trans. Sustain. Comput. 10 (2025) 858–872. [167] E. Barbierato, A. Gatti, Toward green ai: A methodological survey of the scientific literature, IEEE Access 12 (2024) 23989–24013. [168] S. Salehi, A. Schmeink, Data-centric green artificial intelligence: A survey, IEEE Trans. on Art. Intell. 5 (2024) 1973–1989. [169] M. Shen, M. Umar, K. Maeng, G. E. Suh, U. Gupta, Hermes: Algorithm-system co-design for efficient retrieval-augmented generation at-scale, in: ACM/IEEE 52nd Annu. Int. Symp. Comput. Archit. (ISCA), ISCA ’25, ACM, 2025, p. 958–973. doi:10.1145/ 3695053.3731076. [170] J.-W. Chung, Y. Gu, I. Jang, L. Meng, N. Bansal, M. Chowdhury, Reducing energy bloat in large model training, in: Proceedings of the ACM SIGOPS 30th Symp. Oper. Syst. Princ., SOSP ’24, ACM, 2024, p. 144–159. doi:10.1145/3694715.3695970. [171] G. Moro, L. Ragazzi, L. Valgimigli, Carburacy: Summarization models tuning and comparison in eco-sustainable regimes with a novel carbon-aware accuracy, Proceedings of the AAAI Conference on Artificial Intelligence 37 (2023) 14417–14425. [172] Y. Tian, D. Sun, Y. Ding, S. Liu, Cache your prompt when it’s green — carbon-aware caching for large language model serving, Proc. ACM Meas. Anal. Comput. Syst. 10 (2026). [173] K. Chadli, G. Botterweck, T. Saber, Pearl: Performance and energy aware routing for llms, Future Gener. Comput. Syst. 176 (2026) 108218. [174] J. Castaño, J. Bustillo, X. Franch, S. Martínez-Fernández, Datadriven multi-objective optimization of ml inference hardware configurations for energy, performance and cost, Sustain. Comput. Inf. Syst. 50 (2026) 101347. [175] T. Patki, B. Rountree, T. Wilde, A. Bartolini, S. Brink, E. Heiskanen, S. Idgunji, M. Maiterth, J. Rogers, E. Rrapaj, R. Schneider, W. Shin, K. Shoga, C. Simmendinger, N. J. Wright, Z. Zhao, A global perspective on supercomputer power provisioning: Case studies from united states and europe, in: Proceedings of the 39th ACM International Conference on Supercomputing, ICS ’25, ACM, 2025, p. 1034–1051. doi:10.1145/3721145.3734532. [176] F. Banchelli, M. Garcia-Gasulla, F. Mantovani, J. Vinyals, J. Pocurull, D. Vicente, B. Eguzkitza, F. C. Galeazzo, M. C. Acosta, S. Girona, Introducing marenostrum5: A european pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads, Future Gener. Comput. Syst. 176 (2026) 108125. [177] GlobalPetrolPrices, Electricity prices, https://www. globalpetrolprices.com/electricity_prices/, 2026. [178] The Green Grid, Pue: A comprehensive examination of the metric, https://www.thegreengrid.org/en/resources/library-and-tools,
2012. Industry standard metric definition. [179] Uptime Institute, Uptime institute global data center survey,
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 23 of 26
https://datacenter.uptimeinstitute.com/rs/711-RIA-145/images/ 2025.Annual.Survey.Report.pdf, 2025.
[180] ourworldindata.org, Carbon intensity of electricity generation, 2025. URL: https://ourworldindata.org/electricity-mix, accessed: 2026-04-16. [181] S. Laso, P. Rodriguez, J. L. Herrera, J. Berrocal, J. M. Murillo, Energy consumption and workload prediction for edge nodes in the computing continuum, Sustain. Comput. Inf. Syst. 46 (2025) 101088. [182] S. Chen, P. Oikonomou, Z. Hua, N. Tziritas, K. Djemame, N. Zhang, G. Theodoropoulos, Qos-aware placement of interdependent services in energy-harvesting-enabled multi-access edge computing, Future Gener. Comput. Syst. 174 (2026) 108009. [183] Q. Li, S. Wang, C. Xu, X. Ma, M. Xu, A. Zhou, R. Xing, B. Yang, Z. Zhu, Y. Zhang, X. Liu, Exploring real-time satellite computing: From energy and thermal perspectives, in: 2024 IEEE RealTime Systems Symposium (RTSS), 2024, pp. 161–173. doi:10.1109/ RTSS62706.2024.00023. [184] R. M. Hampau, M. Kaptein, R. van Emden, T. Rost, I. Malavolta, An empirical study on the performance and energy consumption of ai containerization strategies for computer-vision tasks on the edge, in: Proc. 26th Int. Conf. Eval. Assess. Softw. Eng. (EASE), EASE ’22, ACM, 2022, p. 50–59. doi:10.1145/3530019.3530025. [185] K. Wang, A. Louri, A. Karanth, R. Bunescu, Intellinoc: A holistic design framework for energy-efficient and reliable on-chip communication for manycores, in: ACM/IEEE 46th Annu. Int. Symp. Comput. Archit. (ISCA), 2019, pp. 1–12. [186] Intel, Intel xeon 6 processors, https://www.intel.com/content/www/ us/en/products/details/processors/xeon.html, 2024. [187] AMD, Amd epyc (turin) processors, https://www.amd.com/en/ products/processors/server/epyc/9005-series.html, 2024. Zen5based EPYC generation; official docs evolving. [188] Amazon Web Services, Aws graviton processors, https://aws. amazon.com/ec2/graviton/, 2023. [189] Y. Xue, J. Huang, Regate: Enabling power gating in neural processing units, in: Proc. 58th IEEE/ACM Int. Symp. Microarch., MICRO ’25, ACM, 2025, p. 1160–1177. doi:10.1145/3725843.3756038. [190] Y. Zhao, K. Wang, A. Louri, Opt-gcn: A unified and scalable chiplet-based accelerator for high-performance and energy-efficient gcn computation, IEEE Trans. Computer-Aided Design of Integr. Circuits and Systems 43 (2024) 4827–4840. [191] B.-S. Han, K. Parekh, W.-C. Lin, T. Paul, A. Gandhi, Z. Liu, Energyefficient gpu sm allocation, SIGMETRICS Perform. Eval. Rev. 53 (2025) 33–38. [192] W. Shin, J. B. White, W. Elwasif, R. F. Da Silva, C. Zimmer, B. Messer, R. Budiardja, A. Georgiadou, V. M. Vergara, J. Lange, M. Maiterth, T. Osborne, L. Huk, J. Holmen, N. Hagerty, A. M. Karimi, T. Naughton, R. Adamson, R. Prout, F. Wang, S. Atchley, K. G. Thach, T. Beck, S. Oral, Towards sustainable post-exascale leadership computing, in: Proc. SC ’24 Workshops, 2024, pp. 1790– 1794. doi:10.1109/SCW63240.2024.00225. [193] CEA, Alice recoque supercomputer, https://www-dcc.extra.cea.fr/ en/AliceRecoque.html, 2024. European HPC infrastructure initiative. [194] C.-Y. Lu, B.-J. Huang, M.-C. Chen, A. Tsai, E. J.-W. Fang, Y. Cho, R. C.-Y. Liu, E. Wang, Y.-M. Tsao, H. Mair, S.-A. Hwang, 8.2: Runtime power management system by on-die power sensor with silicon machine learning-based calibration in a 3nm octa-core cpu, in: 2025 IEEE International Solid-State Circuits Conference (ISSCC), volume 68, 2025, pp. 160–162. doi:10.1109/ISSCC49661.2025.10904564. [195] C. Intel, Running average power limit energy reporting / cve2020-8694 , cve-2020-8695 / intel-sa-00389, 2022. URL: https://www.intel.com/content/www/us/en/developer/articles/ technical/software-security-guidance/advisory-guidance/ running-average-power-limit-energy-reporting.html, accessed:
2026-04-16.
Energy-Aware Computing in 2026 [196] A. Guliani, M. M. Swift, Per-application power delivery, in: Proc. 14th Eur. Conf. Comput. Syst. (EuroSys), EuroSys ’19, ACM, 2019. doi:10.1145/3302424.3303981. [197] Y. Zhang, Q. Wang, Z. Lin, P. Xu, B. Wang, Improving gpu energy efficiency through an application-transparent frequency scaling policy with performance assurance, in: Proc. 19th Eur. Conf. Comput. Syst. (EuroSys), EuroSys ’24, ACM, 2024, p. 769–785. doi:10.1145/3627703.3629584. [198] LLNS, Variorum, 2023. URL: https://variorum.readthedocs.io/en/ latest/, accessed: 2026-05-06. [199] G. Raffin, D. Trystram, O. Richard, Alumet: a Modular Framework to Standardize the Measurement of Energy Consumption, in: PECS 2025 - Workshop on Performance and Energy Efficiency in Concurrent and Distributed Systems, TÜD Dresden University of Technology, Springer, Dresden, Germany, 2025, pp. 1–12. URL: https://hal.science/hal-05246933. [200] T. Gregersen, P. Patel, E. Choukse, Input-dependent power usage in gpus, in: Proc. SC ’24 Workshops, SC-W ’24, IEEE Press, 2025, p. 1872–1877. doi:10.1109/SCW63240.2024.00235. [201] Z. Zhao, B. Austin, E. Rrapaj, N. J. Wright, Understanding vasp power profiles on nvidia a100 gpus, in: Proc. SC ’24 Workshops, SC-W ’24, IEEE Press, 2025, p. 1496–1505. doi:10.1109/SCW63240. 2024.00189. [202] H. S. Lella, R. Chattaraj, S. Chimalakonda, M. Kurra, Towards comprehending energy consumption of database management systems - a tool and empirical study, in: Proc. 28th Int. Conf. Eval. Assess. Softw. Eng. (EASE), EASE ’24, ACM, 2024, p. 272–281. doi:10.1145/3661167.3661174. [203] MLCommons, Mlperf benchmark suite, https://mlcommons.org/, 2025. [204] H. Kim, K.-D. Kang, G. Park, S. Lee, D. Kim, Brokensleep: Remote power timing attack exploiting processor idle states, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2025, pp. 409–422. doi:10.1109/HPCA61900.2025.00040. [205] A. Oliveira-Filho, W. S. de Souza, C. A. V. Sakuyama, S. X. de Souza, Phoeni6: A systematic approach for evaluating the energy consumption of neural networks, Sustain. Comput. Inf. Syst. 47 (2025) 101172. [206] J. You, J.-W. Chung, M. Chowdhury, Zeus: Understanding and optimizing GPU energy consumption of DNN training, in: 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), USENIX Association, Boston, MA, 2023, pp. 119–139. URL: https://www.usenix.org/conference/ nsdi23/presentation/you. [207] F. Chung, H. Kuo, G. Candea, The case for energy clarity, in: Proceedings of the 2025 Workshop on Hot Topics in Operating Systems, HotOS ’25, ACM, 2025, p. 202–209. doi:10.1145/3713082. 3730370. [208] H. Nejati Sharif Aldin, F. Nayebi Pour, From watts to writes: The role of frequency, replication strategies and consistency levels in performance and energy-aware cassandra cluster operations, Future Gener. Comput. Syst. 176 (2026) 108210. [209] Z. Ournani, M. C. Belgaid, R. Rouvoy, P. Rust, J. Penhoat, Evaluating the impact of java virtual machines on energy consumption, in: Proc. 15th ACM/IEEE Int. Symp. Emp. Soft. Eng Meas. (ESEM), ESEM ’21, ACM, 2021. doi:10.1145/3475716.3475774. [210] T. Smejkal, R. Khasanov, J. Castrillon, H. Härtig, Harp: Energyaware and adaptive management of heterogeneous processors, in: Proceedings of the 26th International Middleware Conference, Middleware ’25, ACM, 2025, p. 270–284. doi:10.1145/3721462.3770774. [211] R. N. Tchakoute, C. Tadonki, P. Dokladal, Y. Mesri, Benchmarkbased study of cpu/gpu power-related features through jax and tensorflow, IEEE Access 13 (2025) 184543–184560. [212] D. De Sensi, T. De Matteis, M. Danelutto, Simplifying self-adaptive and power-aware computing with nornir, Future Gener. Comput. Syst. 87 (2018) 136–151. [213] L. Carpentieri, A. De Caro, M. S. Beni, K. Fan, B. Cosenza, Phase-based frequency scaling for energy-efficient heterogeneous
computing, in: IEEE Int. Par Distr. Process. Symp. (IPDPS), 2025, pp. 824–836. doi:10.1109/IPDPS64566.2025.00078. [214] F. Fraternali, A. Bartolini, C. Cavazzoni, L. Benini, Quantifying the impact of variability and heterogeneity on the energy efficiency for a next-generation ultra-green supercomputer, IEEE Trans. Parallel. Distrib. Syst. 29 (2018) 1575–1588. [215] S. Malla, Q. Deng, Z. Ebrahimzadeh, J. Gasperetti, S. Jain, P. Kondety, T. Ortiz, D. Vieira, Coordinated priority-aware charging of distributed batteries in oversubscribed data centers, in: 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2020, pp. 839–851. doi:10.1109/MICRO50266.2020.00073. [216] L. Costero, F. D. Igual, K. Olcoz, Dynamic power budget redistribution under a power cap on multi-application environments, Sustain. Comput. Inf. Syst. 38 (2023) 100865. [217] R. Azimi, C. Jing, S. Reda, Powercoord: Power capping coordination for multi-cpu/gpu servers using reinforcement learning, Sustain. Comput. Inf. Syst. 28 (2020) 100412. [218] M. T. Costa, A. Georgiadou, T. B. White, B. V. Alvarez, J. Polo, W. Shin, P. Olivier, B. Messer, A. F. Lorenzon, Characterizing the impact of gpu power management on an exascale system, in: Proc. SC ’25 Workshops, 2025, pp. 1524–1533. [219] A. Toor, S. ul Islam, N. Sohail, A. Akhunzada, J. Boudjadar, H. A. Khattak, I. U. Din, J. J. Rodrigues, Energy and performance aware fog computing: A case of dvfs and green renewable energy, Future Gener. Comput. Syst. 101 (2019) 1112–1121. [220] X. Ruan, H. Chen, Y. Tian, S. Yin, Virtual machine allocation and migration based on performance-to-power ratio in energy-efficient clouds, Future Gener. Comput. Syst. 100 (2019) 380–394. [221] S. Gonzalo San José, J. M. Marquès, J. Panadero, L. Calvet, Nara: Network-aware resource allocation mechanism for minimizing quality-of-service impact while dealing with energy consumption in volunteer networks, Future Gener. Comput. Syst. 164 (2025) 107593. [222] J. H. Yahya, H. Volos, D. B. Bartolini, G. Antoniou, J. S. Kim, Z. Wang, K. Kalaitzidis, T. Rollet, Z. Chen, Y. Geng, O. Mutlu, Y. Sazeides, Agilewatts: An energy-efficient cpu core idle-state architecture for latency-sensitive server applications, in: 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2022, pp. 835–850. doi:10.1109/MICRO56248.2022.00063. [223] G. Antoniou, H. Volos, D. B. Bartolini, T. Rollet, Y. Sazeides, J. H. Yahya, Agilepkgc: An agile system idle state architecture for energy proportional datacenter servers, in: 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2022, pp. 851–867. doi:10.1109/MICRO56248.2022.00065. [224] J. Bibbens, C. Sigrist, B. Sun, S. Kamali, M. Hajiesmaili, Green bin packing, Proc. ACM Meas. Anal. Comput. Syst. 10 (2026). [225] G. Park, M. Kim, K.-D. Kang, Y. Jeon, S. Kim, D. Kim, Ecocore: Dynamic core management for improving energy efficiency in latency-critical applications, in: Proceedings of the 58th IEEE/ACM International Symposium on Microarchitecture, MICRO ’25, ACM, 2025, p. 1132–1146. doi:10.1145/3725843.3756021. [226] W. Yao, Z. Wang, Y. Hou, X. Zhu, X. Li, Y. Xia, An energy-efficient load balance strategy based on virtual machine consolidation in cloud environment, Future Gener. Comput. Syst. 146 (2023) 222– 233. [227] R. She, Y. Wu, E. Cui, M. Sun, W. Zhao, D. Fu, Energy-efficiency optimization for heterogeneous computing-assisted noma-mec edge ai tasks, Future Gener. Comput. Syst. 162 (2025) 107458. [228] L. Magoula, N. Koursioumpas, I. Stavrakakis, N. Alonistioti, E-split: A hierarchical genetic algorithm for energy-efficient distributed ai services, Computer Communications 246 (2026) 108371. [229] A. Kamatar, M. Gonthier, V. Hayot-Sasson, A. Bauer, M. Copik, R. Castro Fernandez, T. Hoefler, K. Chard, I. Foster, Core hours and carbon credits: Incentivizing sustainability in hpc, in: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, SC ’25, ACM, 2025, p. 870–887. doi:10.1145/3712285.3759858. [230] O. Antepara, Z. Zhao, B. Austin, N. Ding, L. Oliker, N. J. Wright, S. Williams, Benchmark-driven models for energy analysis and
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 24 of 26
Energy-Aware Computing in 2026 attribution of gpu-accelerated supercomputing, in: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, SC ’25, ACM, 2025, p. 888–904. doi:10.1145/3712285.3759815. [231] J. Stojkovic, N. Iliakopoulou, T. Xu, H. Franke, J. Torrellas, Ecofaas: Rethinking the design of serverless environments for energy efficiency, in: ACM/IEEE 51st Annu. Int. Symp. Comput. Archit. (ISCA), 2024, pp. 471–486. doi:10.1109/ISCA59077.2024.00042. [232] R. Basu Roy, T. Patel, B. Li, S. Samsi, V. Gadepally, D. Tiwari, Greenmix: Energy-efficient serverless computing via randomized sketching on asymmetric multi-cores, in: Proc. Int. Conf. High Perform. Comput., Netw., Storage Anal, SC ’25, ACM, 2025, p. 475–489. doi:10.1145/3712285.3759861. [233] Z. Zhou, J. Abawajy, M. Chowdhury, Z. Hu, K. Li, H. Cheng, A. A. Alelaiwi, F. Li, Minimizing sla violation and power consumption in cloud data centers using adaptive energy-aware algorithms, Future Gener. Comput. Syst. 86 (2018) 836–850. [234] E. Khodayarseresht, A. Shameli-Sendi, Q. Fournier, M. Dagenais, Energy and carbon-aware initial vm placement in geographically distributed cloud data centers, Sustain. Comput. Inf. Syst. 39 (2023) 100888. [235] K. West, Y. Moawad, F. Lehmann, V. Bountris, U. Leser, Y. Elkhatib, L. Thamsen, A systematic evaluation of the potential of carbonaware execution for scientific workflows, Future Gener. Comput. Syst. 182 (2026) 108453. [236] P. Wiesner, O. Kao, Moving beyond marginal carbon intensitya poor metric for both carbon accounting and grid flexibility, SIGMETRICS Perform. Eval. Rev. 53 (2025) 108–111. [237] A. Souza, N. Bashir, J. Murillo, W. Hanafy, Q. Liang, D. Irwin, P. Shenoy, Ecovisor: A virtual energy system for carbon-efficient applications, in: Proc. 28th ACM Int. Conf. Archit. Support Program. Lang. Oper. Syst. (ASPLOS), Vol. 2, ASPLOS 2023, ACM, 2023, p. 252–265. doi:10.1145/3575693.3575709. [238] C. Jiang, Y. Wang, D. Ou, Y. Li, J. Zhang, J. Wan, B. Luo, W. Shi, Energy efficiency comparison of hypervisors, Sustain. Comput. Inf. Syst. 22 (2019) 311–321. [239] T. Sukprasert, A. Souza, N. Bashir, D. Irwin, P. Shenoy, On the limitations of carbon-aware temporal and spatial workload shifting in the cloud, in: Proc. 19th Eur. Conf. Comput. Syst. (EuroSys), EuroSys ’24, ACM, 2024, p. 924–941. doi:10.1145/3627703.3650079. [240] C. Qiu, H. Shen, L. Chen, Towards green cloud computing: Demand allocation and pricing policies for cloud service brokerage, IEEE Trans. Big Data 5 (2019) 238–251. [241] Z. Ournani, R. Rouvoy, P. Rust, J. Penhoat, On reducing the energy consumption of software: From hurdles to requirements, in: Proc. 14th ACM/IEEE Int. Symp. Emp. Soft. Eng Meas. (ESEM), ESEM ’20, ACM, 2020. doi:10.1145/3382494.3410678. [242] M. Weber, C. Kaltenecker, F. Sattler, S. Apel, N. Siegmund, Twins or false friends? a study on energy consumption and performance of configurable software, in: Proc.45th Int. Conf. Soft. Eng., ICSE ’23, IEEE Press, 2023, p. 2098–2110. doi:10.1109/ICSE48619.2023.00177. [243] Y. Sharma, W. Si, D. Sun, B. Javadi, Failure-aware energy-efficient vm consolidation in cloud computing systems, Future Gener. Comput. Syst. 94 (2019) 620–633. [244] P. Rodríguez, J. Mateos-Bravo, S. Laso, J. L. Herrera, J. Berrocal, Predictively controlling the computing continuum with distributed energy-aware orchestration, J. Syst. Arch. 175 (2026) 103752. [245] OpenAI, Ai and compute, https://openai.com/index/ ai-and-compute/, 2018. [246] D. Zhao, N. C. Frey, J. McDonald, M. Hubbell, D. Bestor, M. Jones, A. Prout, V. Gadepally, S. Samsi, A green(er) world for a.i., in: IEEE IPDPS ’22 Workshops, 2022, pp. 742–750. doi:10.1109/ IPDPSW55747.2022.00126. [247] H. Järvenpää, P. Lago, J. Bogner, G. Lewis, H. Muccini, I. Ozkaya, A synthesis of green architectural tactics for ml-enabled systems, in: IEEE/ACM 46th Int. Conf. Soft. Eng. (ICSE-SEIS), 2024, pp. 130–141. [248] M. Shoeybi, et al., Megatron-lm: Training multi-billion parameter language models, arXiv preprint arXiv:1909.08053 (2019).
[249] W. A. Hanafy, Q. Liang, N. Bashir, D. Irwin, P. Shenoy, Carbonscaler: Leveraging cloud workload elasticity for optimizing carbonefficiency, Proc. ACM Meas. Anal. Comput. Syst. 7 (2023). [250] G. Spillo, A. De Filippo, C. Musto, M. Milano, G. Semeraro, Balancing carbon footprint and algorithm performance in recommender systems: A comprehensive benchmark, Sustain. Comput. Inf. Syst. 49 (2026) 101286. [251] J. Yu, J. Kim, E. Seo, Know your enemy to save cloud energy: Energy-performance characterization of machine learning serving, in: IEEE Int. Symp. High-Perform. Comput. Archit. (HPCA), 2023, pp. 842–854. doi:10.1109/HPCA56546.2023.10070943. [252] D. Dong, H. Jiang, X. Wei, Y. Song, X. Zhuang, J. Wang, Etnas: An energy consumption task-driven neural architecture search, Sustain. Comput. Inf. Syst. 40 (2023) 100926. [253] S. Lambert, E. Caron, L. Lefevre, R. Grivel, Consolidation of virtual machines to reduce energy consumption of data centers by using ballooning, sharing and swapping mechanisms, Future Gener. Comput. Syst. 174 (2026) 107968. [254] M. Belghachi, Adaptive federated learning for energy-aware wireless edge networks in 6g environments, Computer Communications 250 (2026) 108452. [255] P. Fusco, G. Rimoli, A. Guerriero, F. Palmieri, M. Ficco, On-device training and pruning for energy saving and continuous learning in resource-constrained mcus, Future Gener. Comput. Syst. 176 (2026) 108194. [256] Z. Xu, L. Zhao, W. Liang, O. F. Rana, P. Zhou, Q. Xia, W. Xu, G. Wu, Energy-aware inference offloading for dnn-driven applications in mobile edge clouds, IEEE Trans. Parallel. Distrib. Syst. 32 (2021) 799–814. [257] K. Gai, X. Qin, L. Zhu, An energy-aware high performance task allocation strategy in heterogeneous fog computing environments, IEEE Trans. Comput. 70 (2021) 626–639. [258] P. Li, M. J. Islam, S. Ren, A case study of environmental footprints for generative ai inference: Cloud versus edge, SIGMETRICS Perform. Eval. Rev. 53 (2025) 21–26. [259] W. Luo, T. Yu, S. Dai, B. Guo, X. Lin, Y. Pu, Semi-asynchronous energy-efficient federated prototype learning for end-edge-cloud architectures, Future Gener. Comput. Syst. 179 (2026) 108351. [260] C. Cho, Y. Son, S. Park, Y. G. Kim, Clover: Carbon optimization of federated learning over heterogeneous clients, in: Proc. 29th ACM/IEEE Int. Symp. Low Power Electr. Des., ISLPED ’24, ACM, 2024, p. 1–6. doi:10.1145/3665314.3670829. [261] Z. Yang, S. Zhang, C. Li, M. Wang, J. Yang, M. Zhang, Joint heterogeneity-aware personalized federated search for energy efficient battery-powered edge computing, Future Gener. Comput. Syst. 146 (2023) 178–194. [262] L. Wu, W. A. Hanafy, A. Souza, K. Nguyen, J. Harkes, D. Irwin, M. Satyanarayanan, P. Shenoy, Carbonedge: Leveraging mesoscale spatial carbon-intensity variations for low carbon edge computing, in: Proc. 34th Int. Symp. High-Perf. Par Distr.Comput., HPDC ’25, ACM, 2025. doi:10.1145/3731545.3731576. [263] B. Acun, B. Lee, F. Kazhamiaka, K. Maeng, U. Gupta, M. Chakkaravarthy, D. Brooks, C.-J. Wu, Carbon explorer: A holistic framework for designing carbon aware datacenters, in: Proc. 28th ACM Int. Conf. Archit. Support Program. Lang. Oper. Syst. (ASPLOS), Vol. 2, ASPLOS 2023, ACM, 2023, p. 118–132. doi:10.1145/3575693. 3575754. [264] J. Chen, Z. Zhong, K. Sun, C. Ma, R. Mao, Y. Wang, Lift: Exploiting hybrid stacked memory for energy-efficient processing of graph convolutional networks, in: 2023 60th ACM/IEEE Design Autom. Conf. (DAC), 2023, pp. 1–6. doi:10.1109/DAC56929.2023.10247871. [265] W.-T. Chang, C.-F. Wu, Y.-C. Lo, P-dac: Power-efficient photonic accelerators for llm inference, in: 2025 62nd ACM/IEEE Design Automation Conference (DAC), 2025, pp. 1–7. doi:10.1109/DAC63849. 2025.11132618. [266] J. Yang, S. Dong, Z. Fu, H. Shang, A. Basu, High energy-efficiency and low latency in-memory computing using analog accumulator and in-memory adc with shared, in: Proceedings of the 62nd Annual
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 25 of 26
Energy-Aware Computing in 2026 ACM/IEEE Design Autom. Conf., DAC ’25, IEEE Press, 2025. doi:10.1109/DAC63849.2025.11133014. [267] C. C. Sudarshan, N. Matkar, S. Vrudhula, S. S. Sapatnekar, V. A. Chhabria, Eco-chip: Estimation of carbon footprint of chipletbased architectures for sustainable vlsi, in: IEEE Int. Symp. HighPerform. Comput. Archit. (HPCA), 2024, pp. 671–685. doi:10.1109/ HPCA57654.2024.00058. [268] Y. Zhao, Y. K. Zhao, C. Wan, Y. C. Lin, 3d-carbon: An analytical carbon modeling tool for 3d and 2.5d integrated circuits, in: 2024 61st ACM/IEEE Design Autom. Conf. (DAC), 2024, pp. 1–6. [269] H. Zhang, H. Zhang, C. Xia, Z. Huang, Y. Chen, A. Barnard, Rocket: An rns-based photonic accelerator for high-precision and energy-efficient dnn training, in: Proc. of the ACM Int. Conf. on Supercomputing, ICS ’25, ACM, New York, NY, USA, 2025, p. 1020–1033. doi:10.1145/3721145.3734529. [270] C.-J. Wu, B. Acun, R. Raghavendra, K. Hazelwood, Beyond efficiency: Scaling ai sustainably, IEEE Micro 44 (2024) 37–46. [271] L. Eeckhout, Kaya for computer architects: Toward sustainable computer systems, IEEE Micro 43 (2023) 9–18. [272] T. Eilam, P. Bose, L. P. Carloni, A. Cidon, H. Franke, M. A. Kim, E. K. Lee, M. Naghshineh, P. Parida, C. S. Stein, A. N. Tantawi, Reducing datacenter compute carbon footprint by harnessing the power of specialization: Principles, metrics, challenges and opportunities, IEEE Trans. Semic. Manuf. 37 (2024) 481–488. [273] G. Mersy, S. Krishnan, Toward a life cycle assessment for the carbon footprint of data, SIGENERGY Energy Inform. Rev. 4 (2025) 25–33. [274] J. Wang, D. S. Berger, F. Kazhamiaka, C. Irvene, C. Zhang, E. Choukse, K. Frost, R. Fonseca, B. Warrier, C. Bansal, J. Stern, R. Bianchini, A. Sriraman, Designing cloud servers for lower carbon, in: ACM/IEEE 51st Annu. Int. Symp. Comput. Archit. (ISCA), ISCA ’24, IEEE Press, 2025, p. 452–470. doi:10.1109/ISCA59077. 2024.00041. [275] D. Halder, D. Banerjee, A. Mani, A. Gandhi, E. Zadok, How carbon metrics impact device selection, SIGMETRICS Perform. Eval. Rev. 53 (2025) 104–107. [276] J. Roelandts, A. Naithani, L. Eeckhout, The architectural sustainability indicator, IEEE Comp. Arch. Lett. 24 (2025) 205–208. [277] European Union, Corporate sustainability reporting directive (csrd), https://eur-lex.europa.eu/eli/dir/2022/2464, 2022. [278] C. C. Sudarshan, A. Arora, V. A. Chhabria, Beyond the surface: The necessity for detailed metrics in corporate sustainability reports, in: IGSC’24, 2024, pp. 145–150. doi:10.1109/IGSC64514.2024.00035. [279] E. Peltonen, S. Bayhan, D. Bermbach, S. Buschjäger, V. Degeler, A. Y. Ding, O. D. Incel, D. Katare, M. B. Kjærgaard, S. Leroux, T. Mahmoodi, Z. A. Mann, N. Meratnia, A. D. Pimentel, J. S. Rellermeyer, E. Rivière, D. Sapra, G. Solmaz, B. van der Waaij, Rethinking computing systems in the era of climate crisis: A call for a sustainable computing continuum, IEEE Internet Computing 29 (2025) 8–18. [280] European Union, Artificial intelligence act, https://eur-lex.europa. eu/eli/reg/2024/1689, 2024. [281] P. Migliarini, M. Autili, P. Inverardi, P. Pelliccione, Ethical prompt engineering for ai-driven se: Evidence-informed interaction-time governance roadmap to 2030, ACM Trans. Softw. Eng. Methodol. (2026). Just Accepted. [282] Z. Li, J. Ge, C. Li, H. Yang, H. Hu, B. Luo, V. Chang, Energy cost minimization with job security guarantee in internet data center, Future Gener. Comput. Syst. 73 (2017) 63–78. [283] A. Kumar, N. M. Pindoriya, Toward sustainable data center operation: A review on existing infrastructures, integrated smart energy management frameworks, and future perspectives, Renewable and Sustainable Energy Reviews 230 (2026) 116664. [284] A. Jahanshahi, N. Yu, D. Wong, Powermorph: Qos-aware server power reshaping for data center regulation service, ACM Trans. Archit. Code Optim. 19 (2022). [285] I. Chrysakis, E. Agorogiannis, N. Tsampanaki, M. Vourtzoumis, E. Chondrodima, Y. Theodoridis, D. Mongus, B. Capper, M. Wagner, A. Sotiropoulos, F. A. Coelho, C. V. Brito, P. Protopapas,
D. Brasinika, I. Fergadiotou, C. Doulkeridis, Multi-partner project: Green.dat.ai: A data spaces architecture for enhancing green ai services, in: Des. Autom. Test Eur. Conf. Exhib. (DATE), 2025, pp. 1–7. doi:10.23919/DATE64628.2025.10992729. [286] R. Verdecchia, P. Lago, C. de Vries, The future of sustainable digital infrastructures: A landscape of solutions, adoption factors, impediments, open problems, and scenarios, Sustain. Comput. Inf. Syst. 35 (2022) 100767.
R.N. Tchakoute and C. Tadonki: Preprint submitted to Elsevier
Page 26 of 26