arXiv:2609.08688v1 [cs.DC] 8 Sep 2026
Measuring Sustainability in Multi-Scale High-Performance Computing 1st Carlos J. Barrios H.
2nd Frédéric Le Mouël
3rd Yves Denneulin
Universidad Industrial de Santander SC3UIS, CAGE Bucaramanga, Colombia LIG/INRIA Grenoble Grenoble, France INSA Lyon, INRIA CITI Laboratory Villeurbanne, France 0000-0002-3227-8651
INSA Lyon, INRIA CITI Laboratory Villeurbanne, France 0000-0002-7323-4057
LIG/INRIA Grenoble Université Grenoble-Alpes, Grenoble-INP Grenoble, France 0000-0002-0340-2094
Abstract—The transition from traditional High Performance Computing (HPC) to the Computing Continuum emphasizes efficient resource management and sustainable practices across Multi-Scale hybrid architectures. This paper introduces a multidimensional metric framework to characterize these systems and guide deployment strategies for modern workloads. The framework combines Architectural Performance metrics (such as Throughput, Latency, Scalability), System Utilization, and key Sustainability and Accuracy indicators (such as Energy Efficiency and Power Consumption). Using a modular hybrid testbed, experiments reveal complex relationships among metrics, especially the trade-offs between accuracy and energy, and the efficiency of hybrid nodes. The guidelines help identify optimal operating points and lay the groundwork for improving orchestrators and schedulers (e.g., Kubernetes) to assign demanding applications, including AI and Quantum Computing, to suitable system modules, ensuring high performance and sustainability. Index Terms—Multi-Scale HPC, Sustainability, Performance Analysis
I. I NTRODUCTION Transitioning from High Performance Computing (HPC) to Advanced Computing within the Computing Continuum requires an understanding of how to integrate multi-scale systems. This involves utilizing parallelization and distribution by combining diverse computing resources and multiscale workloads across multiple systems, ensuring sustainable performance while effectively managing complexity and scalability. Based on functional requirements, different workloads are allocated among connected computing resources for parallel processing. When viewed horizontally, these resources are equally diverse and integrated into hybrid systems and heterogeneous architectures, such as CPU-GPU based systems. Looking at it vertically, we observe different levels or layers, depending on what we wish to identify; for example, in a continuous computing architecture consisting of various interconnected levels [1]. The integration and convergence of different computational systems and architectures aim to enhance performance and improve computational efficiency
for complex tasks across multiple scales. This methodology, known as hybrid HPC or heterogeneous HPC, leverages the strengths of various computer architectures to handle diverse workloads effectively. To grasp these concepts, one must analyze architectural diversity through the lens of heterogeneity and scaling through the lens of hybridity. The specialized community has long been developing a comprehensive definition of heterogeneous computing that encompasses the smooth and coordinated deployment of various high-performance machines, including parallel systems, to achieve ultra-fast processing for demanding tasks with diverse computational needs [2]. In terms of systems, the hybrid computing approach represents the intersection of three broad paradigms for computing infrastructure and use: (1) Ownercentric (traditional) HPC; (2) Grid computing (resource sharing); (3) Cloud computing (on-demand resource/service provisioning) [3]. Definitions based on theoretical assumptions face technical implementations of hybrid computing date back to the sixties, mixing both architectural diversity and the increase in parallel processes and data scale, as noted in [4]. These paradigms are important because they allow us to identify the attributes necessary to support different use cases that traditional HPC systems must contain, guiding us precisely towards Multi-Scale HPC systems. Figure 1 shows the correlation between workloads and scale within the framework of High Performance Computing (HPC). The vertical axes illustrate the performance, specifically the peak performance measured in FLOPS, alongside the data capacity and bandwidth. The classification of largescale computation depends on the magnitude of operations and the volume of data. Workloads, classified by application type as intensive or massive, are supported by various computer architecture organizations that inherently foster parallelism and High Throughput Computing (HTC). HPC systems for execution and deployment are tailored to specific workload types, taking into account the technological infrastructure and software platforms used for application
Fig. 1. Resources and Workloads Scale
development, deployment, and execution. The organization shown in Figure 1 is directly associated with the processing units in that environment. For instance, the first three categories, single-core, quad-core, and multi-core, reflect a similar architecture that suits monolithic or owner-centric computing. In contrast, many-core and supercomputer clusters scale up, sharing resources and forming interconnected components, underscoring the significance of network capabilities at this scale. This journey begins with Warehouse-Scale Computing (WSC) and extends to Grid, Cloud, and Modular Supercomputing Systems [5], all of which imply scalable distributed computing environments. We can characterize computation at scale by examining workloads, HPC systems, performance, and scalability. This includes computing related to Edge, Milliclusters, IoT, or any embedded systems that support parallelism (or HPC@Pocket systems [6]), as well as Exascale and PostExascale computing [7], hybrid and reconfigurable systems [8], and even non-von Neumann architectures that integrate Quantum Computing into these Multi-Scale HPC systems [9] [10]. Additionally, it encompasses quantum computing, particularly from a hybrid-computing perspective that integrates classical computer architectures. Adopting a holistic view provides a thorough understanding of computing systems, marking the shift from traditional HPC to more intricate Multi-Scale environments. This approach facilitates the creation of a performance metric framework. Recognizing the ability to navigate complex, interconnected computational landscapes is essential for adopting a unified approach across various computing paradigms. The heterogeneity and hybrid nature of these systems help clarify behaviors, trade-offs, sustainability, and energy consumption metrics of different features. By exploring expanding scales, applications, and multi-level behaviors, one can outline a Multi-Scale High Performance Computing (Multi-Scale HPC) system that ensures efficient resource allocation and highlights key metrics for system monitoring and design. In other words, we define Multi-Scale HPC as the capability of a system and its software framework to manage workloads across multiple interconnected dimensions of scale: (i) Hardware Scale (from heterogeneous cores to accelerated nodes),
(ii) Algorithmic/Computational Scale (from mixed/variableprecision operations to massive multiphysics simulations), and (iii) Temporal/Energy Scale (energy consumption dynamics based on the application phase). This contribution presents application metrics and multilevel behaviors, aiming to create a metric framework that leverages workload analysis to support monitoring techniques. It also lays the groundwork for designing sustainable MultiScale hybrid systems. The discussion covers components for understanding and integrating traditional HPC architectures into distributed, multi-scale systems, as well as guidelines for characterizing these architectures using sustainability and scalability metrics. Addresses diverse workloads for effective deployment, especially in hybrid and heterogeneous architectures. This paper is structured as follows: after this introduction, Section II highlights the essential components and characteristics that shape our guidelines. These guidelines are examined in Section III, where we emphasize the connection between resource utilization and performance, allowing us to identify sustainability and precision metrics. The relevant analysis results are provided in Section IV, followed by a discussion of our further work in Section V. Finally, Section VI concludes our findings. II. K EY C ONCEPTUAL E LEMENTS AND C HARACTERISTICS When examining all interconnected components: hardware, software, algorithms, and data, within a framework supported by a Multi-Scale HPC system, it is crucial to effectively integrate and define architectures and environments. MultiScale HPC systems require incorporating models that operate at varying scales. For example, in exascale computing environments, Multi-Scale applications must efficiently scale across thousands of nodes, which requires meticulous coordination of computational resources and algorithms. Furthermore, this demands malleability [11] to accommodate diverse workload profiles, ensuring effective utilization of system resources and maximizing performance and throughput. Integrated architectures in Multi-Scale HPC systems are complex, involving customizable, heterogeneous hardware such as FPGAs, hybrid GPU-CPU systems, and networked clusters. Optimizing these architectures requires a comprehensive approach to components and interactions, managing interdependencies. Balancing computational loads and ensuring efficient data transfer between storage and processing units are vital to improving performance. Three key conceptual elements can be delineated: first, architectural integration, which presents a methodology for uniting diverse computing paradigms; Secondly, the distributed organization of computing resources, characterized as a distributed HPC system; and finally, Multi-Scale computing, aimed at addressing complexity and scalability 1 [12]. 1 In effect, Multi-Scale computing and traditional computing differ fundamentally in their approach to problem-solving, particularly in terms of complexity and scalability.
Figure 2 displays the layout of the computational architecture for the servers within a Multi-GPU-CPU HPC cluster node. These nodes are designed to optimize computational performance by leveraging the combined power of CPUs and GPUs. They are essential for addressing complex tasks in AI, scientific research, and data analytics, making them vital components of modern HPC systems. When each node is outfitted with this hybrid setup, the system can effectively handle increasing workload complexity and data volume. For instance, applications that use Massive Parallel Processing (MPP) benefit from this configuration, which requires multilevel parallelism to achieve scalability.
systems [13]. Nevertheless, challenges related to resource allocation and efficient execution persist, especially due to the diversity of required workloads and resources; for instance, the need for effective communication strategies, job scheduling techniques [14], and optimization strategies [15]. Technically, it is possible to observe a Multi-Scale HPC system that defines architectural layers, identifies roles and interactions according to functionality, and uses levels as an abstraction to describe the hierarchy within the multitier architecture. Figure 4 shows an abstraction that identifies different components of a Multi-Scale HPC system.
Fig. 4. Multilevel and Multi-Layer HPC Architecture
Fig. 2. Organization of a Multi-GPU-CPU HPC Cluster Node
Multilevel parallelism is the ability to exploit parallelism at multiple levels within a computational task. This concept is especially valuable for improving the efficiency and speed of computations by leveraging different forms of parallelism, including coarse-grained, fine-grained, and instruction-level parallelism. In a complex scientific application, fine-grained computing is directed to the processing cores, while coarsegrained components are allocated to processors or nodes. This allocation increases committed infrastructure resources based on the specific computation requirements. Additionally, this observed behavior improves scalability, allowing MultiScale HPC systems to accommodate parallelism as needed and support more complex applications effectively. Figure 3 illustrates a well-known pipeline for exploiting multilevel parallelism, noting that varying granularity introduces different architectural elements.
Fig. 3. Multilevel Parallelism
This principle has been recognized and utilized for some time in various scale-up workloads, known as ultrascale
A Multi-Scale HPC system can comprise various types of platforms. Figure 4 shows four categories of infrastructure. Horizontally, these include high-performance pure CPUs, hybrid CPU-GPU systems, FPGA-based reconfigurable platforms, heterogeneous systems with varying processing units per node, and small pure CPU architectures. These emphasize both local computing and storage. Vertically, each layer represents a specific functional abstraction, starting from computing in the core layer, moving through the fog/exchange layer, which handles intermediate resources as scale and complexity increase, and finally reaching the interaction/interface layer at the cloud level. These two dimensions integrate architectural designs and HPC configurations into distributed Multi-Scale (HPC) systems by analyzing feature overlap across vertically and horizontally defined parameters and metrics. For instance, when examining horizontal factors, it is important to consider the number of processing cores or accelerators, the level of process support, and the grain types supported. Vertically, one must assess the algorithm’s complexity, the tasks or jobs to be performed, and scalability. This method promotes alignment between scalability, performance, interoperability, and resource management, aiding in the development of workload profiles. This paper does not define I/O performance, as it was held constant to isolate the relationship between compute precision and energy. However, tools that use alternative methods can help in understanding complex workloads. Evaluating these workloads improves characterization, enabling storage systems to optimize I/O performance for specific HPC tasks. This approach also applies to modular and Multi-Scale HPC systems [16].
III. M ULTIDIMENSIONAL M ETRIC F RAMEWORK Key metrics are vital for evaluating HPC system performance, especially as HPC moves toward Multi-Scale distributed systems. They assess system performance and resource usage, support performance evaluation through holistic, monitoring-based analysis and visualization tools [17], optimize current systems, and guide future architectural design for complex Multi-Scale computations. We identify two main metric categories: architectural capabilities and environmental behavior, with additional metrics, such as energy efficiency, at their intersection. We propose guidelines for applying these metrics to ensure smooth deployment across hybrid architectures and multilevel functions. Rather than evaluating standard isolated metrics (e.g., peak FLOPS, raw throughput, or network latency) that fail to capture energy and precision tradeoffs in heterogeneous environments, we formulate composite operational indicators. We map execution attributes into a multi-scale vector: Θ = ⟨EERC, CF, RA⟩
(1)
Where: Energy-Efficiency Resource Conversion (EERC): Measures effective computational output per watt-hour, considering power caps. Algorithmic Complexity Factor (CF ): Evaluates runtime scaling based on overhead O(f (n)), especially under mixed-precision regimes, and Required Accuracy (RA): Penalizes overkill by scaling energy efficiency against error tolerance limits. These metrics influence the Characteristic Ratio (CR) in Equation 2, going beyond static benchmarks. A. Metric Framework Model Before displaying performance metrics, users can define the categories in the visual model depicted in Figure 5. This model serves as a conceptual framework for understanding how critical metric categories converge and aids in system characterization.
Fig. 5. Multi-Dimensional HPC Metric Framework
Figure 5 emphasizes that evaluating systems requires moving beyond individual measures to an analysis of their convergence: 1) Architectural Performance (e.g., Throughput, Scalability) focuses on the system’s raw capability and speed. 2) System Utilization (e.g., Processor, Memory, Network Bandwidth) measures how efficiently internal resources are being used. 3) Sustainability and Accuracy (e.g., Energy Efficiency, Accuracy) capture the integrity of the results and the system’s environmental cost. The central convergence element characterizes the optimal system by linking metrics, such as the Energy Efficiency Ratio, derived from performance and power data. System characterization uses operational data to identify the optimal operating point, guiding scheduling to improve hardware efficiency and performance. B. Architectural Performance Metrics Our holistic approach assesses architectural performance metrics like latency, throughput, and memory bandwidth to improve system performance. These metrics help in designing efficient, scalable, and reliable systems. We propose using the following metrics : • Throughput: This metric quantifies the volume of workload supported as efficient processing capabilities. • Latency: The time required to complete a single operation or task. • Scalability: This metric measures a system’s ability to handle increased workloads by adding resources. It includes linear scalability, where performance increases proportionally, and sub-linear scalability, where gains lessen with more resources. Understanding this helps develop strategies like strengthening large-scale applications within the broader HPC and data analytics convergence in Multi-Scale HPC architectures [18]. Latency and throughput are key system performance indicators. Low latency ensures quick responses, while high throughput enables handling multiple requests simultaneously. Workloads impact both; increasing workloads require scalable systems that adapt without performance loss. Scalability is vital for robust design. C. System Utilization Metrics The utilization metrics show the impact of the resource. System metrics reveal how efficiently resources are used, identify bottlenecks, optimize performance, and prevent under or overuse. The following are the key metrics of this work. • Processor Utilization: This shows the percentage of Processor capacity used during computations. High utilization indicates efficient resource use. • Memory Usage: Monitoring memory usage helps ensure applications do not exceed available memory, which can lead to performance degradation or failures. • Network Bandwidth: In distributed systems, the volume of data transmitted over a given time frame is crucial.
High bandwidth ensures effective data exchange among nodes. As with architectural metrics, there are interdependencies among performance indicators. For instance, high Processor Utilization often coincides with elevated memory usage in memory-intensive applications. However, if memory bandwidth reaches its limit, the Processor Unit may be underutilized while waiting for data. Similarly, Processor Unit utilization can increase in network-intensive applications due to packet processing overhead. Monitoring these metrics together reveals inefficiencies. Applications that transfer large amounts of data over the network also require significant memory for buffering and caching [19]. D. Sustainability and Accuracy Metrics Sustainability and Accuracy metrics for Multi-Scale HPC systems are used to measure and optimize resource efficiency, energy consumption, and overall sustainability, particularly to support HPC systems for Artificial Intelligence and other massive, intensive workloads. • Energy Efficiency: This metric evaluates the total energy consumed to complete a specific computational task. It is particularly useful for comparing the efficiency of different HPC systems, application executions, or configurations. • Power Consumption: As energy efficiency becomes increasingly important, measuring HPC systems’ power consumption helps assess their environmental impact and operational costs. • Accuracy: Ensuring that the results produced by the system meet the required precision and correctness standards is vital, especially in scientific and Artificial Intelligence (AI) or Deep Learning (DL) computations. This research emphasizes the importance of refining workload-specific accuracy standards, as different Multi-Scale HPC tasks require varying levels of precision. Scientific simulations, for instance, demand extremely high precision, such as double-precision floating-point, to prevent errors and ensure valid results. In contrast, AI and deep learning often use lower-precision formats, such as F P 16 or lower, to improve performance and reduce energy use with little impact on accuracy. Quantum computing precision hinges on maintaining coherence and minimizing errors for reliable calculations. Additionally, a thorough, multi-level analysis is essential, assessing accuracy at various stages: instructionlevel, core-level, node-level, and inter-node communication, rather than only at the final output. While current evaluations focus on the minimal impact on heterogeneous elements of a Multi-Scale system, it is also vital to recognize that convergence depends on other metrics; a system should be characterized not only by speed or efficiency, but also by the accuracy it provides at those levels. For example, the Optimal Operation Point (OOP) [20] can be defined by a function OOP = f (T hroughput, Scalability, EE, Accuracy), where EE represents energy efficiency. Incorporating these
factors allows precision to serve as a comprehensive metric for complex and diverse workloads. Characterizing the performance and sustainability of MultiScale HPC systems is important for ensuring the accuracy and energy efficiency of artificial intelligence applications. Consequently, all metrics must be harmonized within this multi-level framework, emphasizing the convergence of measurements to define and delineate these systems. However, it is important to formalize the different metrics and associate them with quantifiable ratios. E. Characteristic Ratio Taking into account the different dimensions of the metrics, it is possible to introduce a formalized structure for calculating and weighting the characteristics of the system, defining a Characteristic Ratio (CR) that facilitates the transition from guidelines to a framework. Then, we propose the use of the formula: CR =
A · EERC CF · RA
(2)
P
Where A = i wi Mi represents the composite Architectural Performance index derived from normalized throughput, latency, and scalability; EERC is the Energy-Efficiency Resource Conversion ratio; CF denotes the algorithmic complexity factor; and RA specifies the workload’s required accuracy. It is important to note that EERC as Energy Efficiency Ratio is explicitly defined as T hroughput , and taking into account P ower that the Complexity Factor (CF) is a variable based on the difficulty or interdependence of the workload in execution, a dynamic, dimensionless measure reflecting the computational difficulty and interdependency within a workload. For this approach, CF is assigned based on algorithmic complexity: O(n) workloads are assigned a CF of 1.0 while O(n2 ) or higher interdependencies are assigned a CF of 1.5. Additionally, the Required Accuracy RA must specify the level of numerical precision and correctness that the workload demands of the system. This is useful for balancing performance and sustainability. Evidently, the RA in the denominator of the equation 2 penalizes when the workload demands a very high level of accuracy. This approach to applying convergence across multidimensional elements is often used for Multi-Scale (and large-scale) systems because simple maximum-speed metrics cannot effectively capture the ’extreme heterogeneity.’ Therefore,Pits numerator should be the sum of architectural metrics ( ArchitecturalM etrics), and the metric should focus on the system’s internal features rather than on raw performance levels alone. IV. M ULTI -S CALE HPC S YSTEMS C HARACTERIZATION According to our modular architecture guidelines for a Multi-Scale HPC system, the three key stages help us understand architecture, performance metrics related to system utilization and sustainability, and accuracy metrics for managing workloads. Architectural complexity arises from heterogeneous computing resources, primarily driving horizontal
scaling that enhances support by adding more nodes. However, as shown in the proposed testbed, four distinct processing capabilities and configurations influence both individual and module performance. As mentioned previously, various infrastructure configurations enable observation from elevated platforms. For instance, some platforms rely solely on CPUs with fixed core counts and processor frequencies, while others integrate CPU-GPU combinations with different core counts and GPU families on interconnected nodes. The data collected comes from four modules: the first module includes nodes 1 to 6 with lowprofile CPU-only configurations; the second module, covering nodes 7 to 20, employs a CPU-GPU setup with four mid-range CPUs and four mid-range GPUs; the third module consists of nodes 21 to 28, which include two high-range CPUs and two high-range GPUs; and the final module, containing nodes 29 and 30, features a different family of two high-range CPUs and two high-range GPUs per node. This configuration is available across various frameworks. For example, in extensive systems like Grid50002 or hybrid platforms presented in the SCALAC system3 , we can identify distinct modules: a basic configuration with only a CPU module comprising two Intel Xeon CPUs operating at 2.4 GHz with 24 cores, providing 102 GB of RAM. Another set of CPU-GPU nodes features Intel Xeon CPUs running at 2.67 GHz (16 cores) and another at 2.4 GHz (64 cores), along with 2 NVIDIA K20 and K80 GPUs and 320 GB of RAM. Furthermore, CPU-GPU nodes employ AMD EPYC 9534 and 9554 processors operating at 2.5 GHz, paired with 2 AMD Instinct MI210. All modules are interconnected through high-bandwidth Gigabit Ethernet and InfiniBand networks simultaneously. The evaluation of workloads ranging from 5000M B to 30000M B is conducted across a spectrum of 2 to 30 diverse nodes, achieving average processor utilization between 89.7 and 97.3 percent. Throughput is determined by dividing the workload by the total time, adjusting for the number of nodes, and then refining it using a logarithmic efficiency factor [21]. Figure 6 illustrates the relationship between throughput and scalability, detailing how throughput varies with the number of nodes within the hybrid CPU-GPU cluster. This expected behavior is observed across different workloads with high memory usage, even on similar platforms. Various factors can lead to high memory utilization, including process types, memory management strategies, and system architecture. However, in our tests, it is mainly due to the process itself. The system’s heterogeneous and hybrid architecture helps manage these trade-offs, allowing for resource optimization. For example, tasks requiring low-to-moderate accuracy can be assigned to energy-efficient components such as GPUs or specialized low-power cores. In contrast, highprecision tasks are reserved for traditional CPU nodes, accepting increased energy use to maintain accuracy. However, following our established guidelines, it is crucial 2 https://www.grid5000.fr 3 https://scalac.redclara.net
Fig. 6. Throughput
to examine the relationships among throughput, scalability, and the number of nodes while considering the deployed workloads, as illustrated in Figure 7.
Fig. 7. Throughput and Scalability
Throughput and scalability are key metrics in computing, particularly in the realm of Multi-Scale HPC. High throughput facilitates efficient handling of complex computations, while scalability ensures that systems can expand and adapt to growing demands without compromising performance. Together, these factors are essential for enhancing computational capabilities. Figure 7 presents the optimal workload that maximizes throughput given the platform’s characteristics. However, there is significant overloading for workloads ranging from 20000 to 30000 MB, as throughput is limited by the slowest component, particularly noticeable during transitions between modules 1
and 2 (only CPU and CPU-GPU modules). In contrast, performance remains stable when shifting from module 3 to module 4 (CPU-GPUs Module but with two different processors and GPU families). By analyzing this behavior, insights can be gained regarding the system’s capacity to manage increasing workloads effectively. At the same time, latency varies with workload size; specifically, high latency is observed on nodes 1 to 6, while low latency is observed in the range of 8 to 30 nodes, with a slight increase as workloads grow. Consequently, due to the architectural configuration, high bandwidth is provided to modules with the most resources, as shown in Figure 8.
Fig. 8. Latency
Figure 9 illustrates the memory consumption of both the CPUs and the GPUs, encompassing elements of both the host and the device throughout the trials carried out. A clear depiction emphasizes the efficiency and interaction between the host and device, highlighting the potential architectural variability that may arise from inefficient interactions between accelerators. The involvement of the CPU and GPU demonstrates the impacts on these resources when diverse modules participate in the overall execution.
node. The energy efficiency ratio is determined by comparing the throughput to the overall power consumption. We examined this correlation by analyzing the total energy used to execute all tasks on the active nodes. The differences illustrated in Figure 10 stem from variations in power consumption among the modules, with CPU-GPU nodes reflecting the aggregate of both types. When combined within CPUGPU nodes, their total power usage accurately represents the sum of their individual inputs, significantly affecting energy efficiency, performance, and sustainability. Then, by leveraging the strengths of each module and optimizing their combined usage in CPU-GPU nodes, it is possible to achieve a balance between performance and energy efficiency, which is increasingly important for Multi-Scale HPC systems as will be discussed in Section V. In this paper, we introduce the relationship between accuracy and energy consumption, enabling the use of different analysis mechanisms, such as trade-off curve analysis and architectural influence analysis. In Multi-Scale hybrid HPC architectures, balancing precision and energy consumption is crucial because these systems combine quantum, analog, and classical components. Higher accuracy usually consumes more energy, which is challenging in low-power settings. Techniques such as reducing precision or selective processing can lower energy consumption while supporting tasks like machine learning inference or real-time analytics. This tradeoff is complex in hybrid systems where subsystems differ in speed, accuracy, and energy needs. Managing this balance is essential for optimal performance and sustainability without sacrificing key results.
Fig. 9. Architecture Performance
Fig. 10. Energy Efficiency Ratio
When assessing architectural features, CPU nodes typically have higher power consumption, even though they run at reduced frequencies. In contrast, CPU-GPU nodes demonstrate superior energy efficiency, averaging around 100 Watts per
Figure 10 is depicted in three dimensions to illustrate the system and incorporate the metrics discussed above. While efficiency improves with increased processing support, energy consumption also increases as more nodes are added. Addi-
tionally, understanding the relationship between accuracy and energy is vital for a complete analysis. To facilitate this, we used a simple backpropagation algorithm [22], which can be characterized in terms of matrix multiplication. While backpropagation serves as a proxy for the matrix-heavy operations used in AI, subsequent validation with the MLPerf benchmark suite is planned to further demonstrate the framework’s versatility across diverse production workloads. The initial accuracy is depicted in Figure 11.
Fig. 11. Accuracy
More accuracy is achieved with increasing workloads and node counts, and GPU nodes deliver better performance. By analyzing the experience, we find that both accuracy and energy consumption depend on workload size and node configurations. Energy consumption increases with workload size and total node power, as illustrated in the figure. In Figure 12, energy consumption is measured in joules, and the accuracy is limited to a maximum value of 1.0.
approach to resource selection by penalizing Accuracy Overkill in simple tasks. Accuracy standards depend on hardware features, but this paper relies on manufacturer data without in-depth analysis. I/O performance, which is significant in HPC workloads, is not included but will be in future CR versions, with an I/O Wait-Time coefficient. I/O bottlenecks are discussed in upcoming research. The convergence analysis in Figure 12 highlights the trade-off between quality and environmental cost, supporting the inclusion of RA in the denominator of CR. Data show saturation [24], where small accuracy reductions (e.g., 98% vs. 99.5%) significantly cut energy (25%) with minimal quality loss, but higher fidelity costs more energy for little gain. Setting an optimal RA based on this trade-off helps CR penalize over-resourced accuracy, promoting sustainability without sacrificing quality. Results show that effective characterization indicates potential support for Multi-Scale HPC systems and helps inform deployment policies and the development of new metrics. A. Extending the CR Metric to Quantum and Edge Domains To avoid overgeneralization, our current experimental validation focuses on classical heterogeneous HPC nodes. However, the mathematical formulation of the Characteristic Ratio (CR) is designed to accommodate Quantum and Edge paradigms by mapping their technology-specific parameters into our multidimensional vector Θ = ⟨EERC, CF, RA⟩: • Quantum Computing Formulation: For quantumclassical hybrid execution, the accuracy parameter RA is redefined as a function of Quantum Gate Fidelity (F ) and coherence time limits (T1 , T2 ), such that RAQ = f (F, T1 ). The algorithmic complexity CF incorporates the circuit depth (D) and the qubit width (Q). The energy efficiency ratio EERC accounts for the total cryogenic cooling overhead per executed circuit. • Edge/IoT Constrained Systems: In edge environments, EERC is adjusted to penalize network transmission overhead (∆tnet ) alongside computational power consumption (Pedge ), reflecting tight energy budgets and bandwidth constraints: EERCedge =
Fig. 12. Accuracy vs Energy Consumption
The relationship between Accuracy and Energy Consumption motivates introducing a Required Accuracy (RA), a performance indicator that is part of the CR in formula 2. Unlike the Energy Delay Product (EDP) [23], which focuses on time-energy trade-offs, our CR metric offers a more nuanced
Throughput Pedge + β · Latencynet
(3)
This formulation describes EERCedge , a refined metric of energy efficiency tailored for distributed edge environments. The numerator reflects the effective computational output, while the denominator expands local power dissipation (Pedge ) to include network communication overhead (Latencynet ). The parameter β serves as an application-specific penalty coefficient that balances offload latency and transmission energy costs, preventing the framework from favoring nodes with limited energy or high delays. Extending the CR formulation to quantum and edge regimes confirms its viability as a generalizable metric and provides the necessary theoretical foundation for addressing sustainability and deployment trade-offs.
V. D ISCUSSION AND F URTHER W ORK In high-workload Multi-Scale HPC systems, especially those engaged in intensive processing such as AI and Quantum Computing, monitoring sustainability metrics requires a hybrid strategy to ensure optimal deployment. Surpassing energy efficiency thresholds can reduce effectiveness, due to physical limitations [25], and limited performance benefits of higher energy use. This raises questions about practicality, environmental responsibility, and cost. These systems aim to provide tailored capabilities within each module by deploying workloads to improve performance. For example, in CPUand-GPU-combined nodes, power consumption depends on the workload and module efficiency. Integrated GPUs that share memory with CPUs can reduce power and heat, making them suitable for energy-constrained settings. Naturally, accuracy and correctness standards must be taken into account. Hardware components, such as support for various floating-point formats, algorithms like error-correction codes, and the application’s tolerance for approximation, all affect the system. However, this document primarily relies on manufacturers’ standard values and metrics without delving into them. Additionally, although I/O performance is crucial for understanding complex workloads and optimizing storage systems in HPC environments, it is not explicitly defined here. Future iterations of the CR formula will integrate an I/O WaitTime coefficient to account for data-intensive bottlenecks. I/O often acts as a significant bottleneck in many modern HPC and data-intensive applications. These aspects are explored in further research initiatives. To implement monitoring systems that analyze metrics and facilitate analysis, especially considering heterogeneity, a factor w can be introduced into equation 2. This factor, known as the weighting factor wm,W [26], transforms generic hardware performance data into a workload-aware metric. It emphasizes architectural features critical to the specific job, such as assigning a higher weight to Memory Bandwidth for data-intensive workloads. However, applying regression coefficients (weights) to low-level hardware counters (architectural metrics) to predict a higher-level composite outcome (e.g., performance or energy) is a standard practice in performance and power modeling and could make the model more complex or introduce biases in the correlations within a holistic approach. The evolution of applications, along with the complexity and behavior of workloads in Multi-Scale HPC systems, requires new paradigms and guidelines for analyzing and evaluating performance and sustainability that extend beyond traditional metrics. These new frameworks should consider energy efficiency, workload-specific performance, scalability, resilience, carbon footprint, and more. These guidelines are designed to empower developers rather than impose control, equipping them with vital tools to enhance the deployment and performance of algorithms. They acknowledge the distinct strengths of each component within a Multi-Scale HPC system. Using multidimensional analysis and a comprehensive strategy, we can establish measurable sustainability metrics,
observing all elements of the system. Additionally, these guidelines clarify the configurations for data and application programming, such as setting up a Kubernetes infrastructure that aligns with specific metrics and limitations. This approach helps characterize both multi-HPC systems and applications, ensuring a sufficient standard of performance accuracy. Ongoing efforts focus on meeting the rapidly growing data needs of fields such as quantum computing and large-scale AI. This involves using Multi-Scale HPC systems with specialized modules—such as those for quantum simulators, rather than real quantum computers. These developments address not only data volume but also the management of large datasets, storage issues, speed enhancements, and data integration complexities. As data grows, AI algorithms tend to improve in accuracy, quantum computing can better maintain coherence and reduce errors, and scientific fields benefit from increased scalability, leading to better simulation results and more precise calculations. Clearly addressing these factors is essential, as they directly impact sustainability metrics. VI. C ONCLUSION The results presented in this paper showcase their relevance across a range of workloads, primarily low-interaction types similar to those encountered in scientific applications, and high-interaction varieties, such as those found in AI applications. Subsequently, metrics were proposed from a multidimensional perspective to facilitate a comprehensive understanding of the system. These guidelines aid in identifying behavioral patterns by examining sustainability metrics and in characterizing a ratio that outlines Multi-Scale HPC systems, thereby providing insights for behavior prediction and configuration enhancement. The Characteristic Ratio (CR) framework deliberately excludes a utility-based Error Tolerance Cost for two main reasons. First, such a cost (e.g., X lost per 1% decrease in accuracy) is subjective, highly dependent on the application owner, and would undermine the CR’s goal of being an objective metric based solely on physical and architectural data. Second, this cost is already implicitly addressed by the denominator through the Required Accuracy (RA) constraint. Observing the last Figure 12, it demonstrates that the system must meet the RA requirement to achieve a meaningful score; if a module fails to reach this minimum accuracy, its CR score will be disproportionately low, making it unsuitable for allocation regardless of the module’s energy efficiency. This is crucial for our proposal because it ensures that the sustainability guidelines derived from the CR are universally applicable across financial models, providing a hardwarecentric, physically grounded guide for optimal resource selection [27] [28]. The proposed metrics and guidelines indicate that monitoring loses effectiveness once the application’s key features are identified, such as in well-known applications or pre-profiled containerized environments. Instead, it is essential to develop policies and strategies that promote execution and optimization within modular systems, especially in modern Multi-Scale
HPC systems. These strategies aim to enhance configurations involving orchestrators, schedulers, and deployers to boost computational efficiency [29]. For example, Kubernetes cluster setups can be recommended for large-scale systems, directing traffic to specific modules based on workload and energy efficiency goals. However, because workloads are diverse and complex, and due to the modular structure of Multi-Scale HPC systems, assessing individual metrics alone is inadequate. Our advice emphasizes that effective characterization depends on linking various metrics to the specific analysis context, allowing accurate profiling and characterization of system components and execution types. In conclusion, our study emphasizes the urgent need for effective Resource Allocation and Sustainability Guidelines in Multi-Scale Hybrid HPC Architectures. We propose the Characteristic Ratio (CR), a detailed, multi-dimensional metric that goes beyond basic energy-per-operation measurements. The CR assesses trade-offs among architectural performance, utilization, and energy efficiency, while also considering Required accuracy (RA), linking sustainability to workload quality. Unlike subjective or cost-based models, our framework provides a physically justified method for identifying the optimal operating point, thereby maximizing computational efficiency and reducing environmental impact. These guidelines enable HPC administrators to make informed deployment decisions, fostering resource allocation that balances speed with sustainable, high-quality performance across the Computing Continuum.
ACKNOWLEDGMENTS The authors thank L. A. Torres, S. Gelvez, F. Mejia and P. Rojas for data on their experiences on various platforms, particularly the Grid5000 platform supported by INRIA and its scientific interest group, including CNRS, RENATER, multiple universities and organizations, and the SC3UIS Center 4 . R EFERENCES [1] I. Mongiardo, L. Massari, M. Carla Calzarossa, B. Bermejo, and D. Tessera, “Application placement in the cloud continuum with resource overbooking,” IEEE Access, vol. 13, pp. 56 471–56 484, 2025. [2] A. Khokhar, V. Prasanna, M. Shaaban, and C.-L. Wang, “Heterogeneous computing: challenges and opportunities,” Computer, vol. 26, no. 6, pp. 18–27, 1993. [3] G. Mateescu, W. Gentzsch, and C. J. Ribbens, “Hybrid computingwhere hpc meets grid and cloud computing,” Future Gener. Comput. Syst., vol. 27, no. 5, p. 440–453, May 2011. [Online]. Available: https://doi.org/10.1016/j.future.2010.11.003 [4] T. D. Truitt, “Hybrid computation: What is it? who needs it?” IEEE Spectrum, vol. 1, no. 6, pp. 132–146, 1964. [5] E. Suarez, N. Eicker, T. Moschny, S. Pickartz, C. Clauss, V. Plugaru, A. Herten, K. Michielsen, and T. Lippert, “Modular supercomputing architecture,” May 2022. [Online]. Available: https://doi.org/10.5281/ zenodo.6508394 [6] H. Bauer, Y. Goh, S. Schlink, and C. Thomas, “The supercomputer in your pocket,” McKinsey on Semiconductors, vol. 112, pp. 14–27, 2012. 4 https://www.sc3.uis.edu.co
[7] X.-K. Liao, K. Lu, C.-Q. Yang, J.-W. Li, Y. Yuan, M.-C. Lai, L.B. Huang, P.-J. Lu, J.-B. Fang, J. Ren, and J. Shen, “Moving from exascale to zettascale computing: Challenges and techniques,” Frontiers of Information Technology & Electronic Engineering, vol. 19, no. 10, pp. 1236–1244, 2018. [8] M. Lyashov, A. Bereza, J. Alekseenko, and L. Blanco, “The hybrid reconfigurable system for high-performance computing,” in 2015 9th International Conference on Application of Information and Communication Technologies (AICT), 2015, pp. 258–262. [9] P. Naayini, C. Bura, and A. K. Jonnalagadda, “The convergence of distributed computing and quantum computing: A paradigm shift in computational power,” International Journal of Scientific Advances (IJSCIA), vol. 6, no. 2, pp. 265–275, 2025. [10] M. Suchara, Y. Alexeev, F. T. Chong, H. Finkel, H. Hoffmann, J. Larson, J. C. Osborn, and G. Smith, “Hybrid quantum-classical computing architectures,” in Proceedings of the 3rd International Workshop on Post-Moore Era Supercomputing, Dallas, TX (11/2018), url=https://api.semanticscholar.org/CorpusID:202593750, 2018. [11] A. Tarraf, M. Schreiber, A. Cascajo, J.-B. Besnard, M.-A. Vef, D. Huber, S. Happ, A. Brinkmann, D. E. Singh, H.-C. Hoppe, A. Miranda, A. J. Peña, R. Machado, M. Garcia-Gasulla, M. Schulz, P. Carpenter, S. Pickartz, T. Rotaru, S. Iserte, V. Lopez, J. Ejarque, H. Sirwani, J. Carretero, and F. Wolf, “Malleability in modern hpc systems: Current experiences, challenges, and future opportunities,” IEEE Transactions on Parallel and Distributed Systems, vol. 35, no. 9, pp. 1551–1564, 2024. [12] S. Alowayyed, D. Groen, P. V. Coveney, and A. G. Hoekstra, “Multiscale computing in the exascale era,” Journal of Computational Science, vol. 22, pp. 15–25, 2017. [Online]. Available: https: //www.sciencedirect.com/science/article/pii/S1877750316302988 [13] J. Carretero, J. Garcia-Blas, and R. Ciegis, “Introduction to sustainable ultrascale computing systems and applications,” The Journal of Supercomputing, vol. 72, no. 11, pp. 4043–4046, 2016. [14] Y. Dai, Y. Dong, K. Lu, R. Wang, W. Zhang, J. Chen, M. Shao, and Z. Wang, “Towards scalable resource management for supercomputers,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis, 2022, pp. 1–15. [15] A. Mehta, “High-performance computing in big data analytics: Architectures, scalability, and optimization strategies,” International Journal of Artificial Intelligence, Data Science, and Machine Learning, vol. 2, no. 2, p. 22–29, Jun. 2021. [Online]. Available: https://ijaidsml.org/index.php/ijaidsml/article/view/26 [16] H. Devarajan and K. Mohror, “Extracting and characterizing i/o behavior of hpc workloads,” in 2022 IEEE International Conference on Cluster Computing (CLUSTER), 2022, pp. 243–255. [17] S. Shilpika, B. Lusch, M. Emani, F. Simini, V. Vishwanath, M. E. Papka, and K.-L. Ma, “A multi-level, multi-scale visual analytics approach to assessment of multifidelity hpc systems,” in 2024 IEEE 24th International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2024, pp. 478–488. [18] D. Elia, S. Fiore, and G. Aloisio, “Towards hpc and big data analytics convergence: Design and experimental evaluation of a hpda framework for escience at scale,” IEEE Access, vol. 9, pp. 73 307–73 326, 2021. [19] J. Ravi, S. Byna, Q. Koziol, H. Tang, and M. Becchi, “Evaluating asynchronous parallel i/o on hpc systems,” in 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2023, pp. 211– 221. [20] C. Reuter, J.-P. Prote, and C. Witthohn, “Global production networks – an approach to find the optimal operating point in the conflict between risk- and cost-minimization,” Procedia CIRP, vol. 41, pp. 532–537, 2016, research and Innovation in Manufacturing: Key Enabling Technologies for the Factories of the Future - Proceedings of the 48th CIRP Conference on Manufacturing Systems. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S2212827115010896 [21] J.-A. Pineiro, M. Ercegovac, and J. Bruguera, “Algorithm and architecture for logarithm, exponential, and powering computation,” IEEE Transactions on Computers, vol. 53, no. 9, pp. 1085–1096, 2004. [22] R. Rojas, The Backpropagation Algorithm. Berlin, Heidelberg: Springer Berlin Heidelberg, 1996, pp. 149–182. [Online]. Available: https://doi.org/10.1007/978-3-642-61068-4 7 [23] J. H. Laros III, K. Pedretti, S. M. Kelly, W. Shu, K. Ferreira, J. Vandyke, and C. Vaughan, Energy Delay Product. London: Springer London, 2013, pp. 51–55. [Online]. Available: https: //doi.org/10.1007/978-1-4471-4492-2 8
[24] M. Brand, F. Hannig, O. Keszocze, and J. Teich, “Precision- and accuracy-reconfigurable processor architectures—an overview,” IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 69, no. 6, pp. 2661–2666, 2022. [25] N. G. Anderson, “Energy efficiency limits in approximate computing: A fundamental physical perspective,” 2016 50th Asilomar Conference on Signals, Systems and Computers, pp. 1653–1657, 2016. [Online]. Available: https://api.semanticscholar.org/CorpusID:2732491 [26] A. Shahid, M. Fahad, R. R. Manumachu, and A. Lastovetsky, “Improving the accuracy of energy predictive models for multicore cpus by combining utilization and performance events model variables,” Journal of Parallel and Distributed Computing, vol. 151, pp. 38– 51, 2021. [Online]. Available: https://www.sciencedirect.com/science/ article/pii/S0743731521000137 [27] C. Cai, L. Wang, S. U. Khan, and J. Tao, “Energy-aware high performance computing: A taxonomy study,” in 2011 IEEE 17th International Conference on Parallel and Distributed Systems, 2011, pp. 953–958. [28] G. Guo, Y. Qi, S. Lai, Z. Liu, and J. Yen, “The latency accuracy trade-off and optimization in implied volatility-based trading systems,” Expert Syst. Appl., vol. 221, no. C, Jul. 2023. [Online]. Available: https://doi.org/10.1016/j.eswa.2023.119714 [29] J. Bader, J. Belak, M. Bement, M. Berry, R. Carson, D. Cassol, S. Chan, J. Coleman, K. Day, A. Duque, K. Fagnan, J. Froula, S. Jha, D. S. Katz, P. Kica, V. Kindratenko, E. Kirton, R. Kothadia, D. Laney, F. Lehmann, U. Leser, S. Lichołai, M. Malawski, M. Melara, E. Player, M. Rolchigo, S. Sarrafan, S.-J. Sul, A. Syed, L. Thamsen, M. Titov, M. Turilli, S. Caino-Lores, and A. Mandal, “Novel approaches toward scalable composable workflows in hyper-heterogeneous computing environments,” in Proceedings of the SC ’23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, ser. SC-W ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 2097–2108. [Online]. Available: https://doi.org/10.1145/3624062.3626283