ConceptioArchivearXiv CS
arXiv CSopen access

DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute Continuum

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute Continuum Aasish Kumar Sharma1 , Felix Stein1 , Mirac Aydin2 , Michael Bidollahkhani1 Sachin P. Nanavati3 , Mohsen Seyedkazemi Ardebili4 , Giorgi Mamulashvili1 , Mojtaba Akbari1 Jonathan Decker1 , Zoya Masih2 , Julian M. Kunkel1,2

arXiv:2605.25292v1 [cs.DC] 24 May 2026

1

Georg-August-Universität Göttingen, 2 GWDG mbH, Göttingen, Germany, 3 NAG, Oxford, UK, 4 Università di Bologna, Italy

Abstract—This paper presents the DECICE project (Device Edge Cloud Intelligent Collaboration framEwork), a Horizon Europe Research and Innovation Action (Grant No. 101092582, December 2022 to November 2025) that developed an opensource framework for intelligent workload scheduling across the cloud-HPC-edge compute continuum. A consortium of 12 partners across 6 European countries organized the work into six work packages covering AI-driven scheduling, digital twin infrastructure, system architecture and integration, monitoring, use case validation, and dissemination. The two core technical contributions are an Integrated AI Scheduler (IAIS) employing RNN-based prediction and formal workflow modeling for constraint-aware workload mapping, and a Digital Twin aggregating real-time metrics with carbon intensity and anomaly prediction for energy-aware scheduling. The framework operates within Kubernetes environments, supports unified workflow ingestion from multiple formats, and bridges cloud-native and HPC orchestration through a Slurm integration layer. We present the project vision, the overall architecture, contributions from each work package, quantitative evaluation results, and the opensource release. Index Terms—compute continuum, Heterogeneous HPC Systems, AI-Driven Scheduling, Workflow Scheduling, digital twin, Kubernetes, energy-aware computing, Horizon Europe

I. P ROJECT OVERVIEW A. Project Details DECICE (Device Edge Cloud Intelligent Collaboration framEwork) is a Horizon Europe Research and Innovation Action funded under programme HORIZON-CL4-2022-DATA01-02, Grant Agreement No. 101092582. The project ran from December 2022 to November 2025 and is currently in its dissemination and community adoption phase. The consortium comprised 12 active partners across 6 countries: Georg-August-Universität Göttingen (UGOE, coordinator) and GWDG (Germany); E4 Computer Engineering and Consorzio TOP-IX (Italy); KTH Royal Institute of Technology (Sweden); University of Stuttgart / HLRS (USTUTT, Germany); Huawei Technologies Düsseldorf (HWDU, Germany); SYNYO GmbH (Austria); Marmara University (MARUN, Turkey) and BIGTRI (Turkey); Università di Bologna (UNIBO, Italy); and the Numerical Algorithms Group (NAG, UK).

B. Motivation and Objectives Modern computing workloads span heterogeneous infrastructure ranging from IoT devices and edge nodes through cloud platforms to HPC clusters. This compute continuum introduces scheduling challenges: workloads must be matched to resources satisfying computational, memory, and data locality constraints while meeting energy, latency, and cost requirements. DECICE addressed six objectives: (O1) leverage the full cloud-HPC-edge continuum; (O2) build an AI scheduler with dynamic load balancing, energy-efficiency, and green energy awareness; (O3) design APIs increasing control over network, computing, and data resources; (O4) implement a Dynamic Digital Twin with AI-based prediction; (O5) demonstrate usability through real-life use cases; and (O6) enable trustworthy and security-compliant service deployment. C. Relevance to COMPSAC Themes DECICE addresses COMPSAC symposia themes including Applied Artificial Intelligence (AI-driven scheduling), Autonomous Systems (adaptive workload placement), Computer Architecture and Platforms (Kubernetes federation across HPC, cloud, and edge), Smart IoT Systems (Digital Twin and edge monitoring), and Security, Privacy and Trust (trustworthy deployment under O6). II. W ORK PACKAGE S TRUCTURE AND C ONTRIBUTIONS The project was organized into six work packages. Fig. 1 shows the overall framework architecture that integrates contributions from all WPs. WP1: Project Management (Lead: UGOE) coordinated consortium activities, technical steering, and milestone tracking across all partners and deliverables. WP2: AI Scheduler for Optimization and Adaptation (Lead: UGOE, after transfer from FBU; key contributors: GWDG, NAG, UNIBO, MARUN, BIGTRI) developed the two core technical components. The Integrated AI Scheduler (IAIS) optimizes job placement across the compute continuum using multiple strategies. HOSHMAND [2], an RNN-based scheduler, learns from historical job-resource state pairs to predict optimal allocations and eliminate redundant scheduling computations. GrapheonRL [3], a Graph Neural Network and

Fig. 1. DECICE framework architecture integrating contributions from all work packages. Upper layer: DECICE Framework (WP2 AI Scheduler and Digital Twin, WP3 Control Manager and APIs). Lower layer: Kubernetesbased Compute Plane (WP4 monitoring, WP5 deployment) federating HPC, cloud, and edge resources.

Fig. 2. WP2: IAIS data flow. Heterogeneous node and job attributes enter through ETL, pass through the AI scheduler’s RNN model and LB-KAIROS simulation, producing Kubernetes-compatible placement decisions.

Reinforcement Learning framework, models workflows as dependency-aware graphs enabling constraint-aware scheduling without mathematical reformulation. The workflow-driven modeling framework [1] formally decomposes heterogeneous HPC scheduling into task-node mapping and schedule derivation, benchmarking exact solvers (MILP via PuLP, SCIP, Gurobi, OR-Tools), constraint programming (CP-SAT), and heuristics (GA, PSO, ACO, SA, HEFT, OLB) across progressively complex workflow scenarios. Fig. 2 illustrates the IAIS data flow, and Fig. 3 shows the Digital Twin architecture developed by UNIBO, UGOE, and GWDG, which provides node-level power consumption metrics, carbon intensity prediction [9], and anomaly detection [8] to enable energy-aware scheduling decisions. WP3: Open Framework and Virtual Training Environment (Lead: GWDG; key contributors: UGOE, E4, USTUTT, HWDU, TOP-IX) designed the overall system architecture, the DECICE Control Manager, and the API layer. WP3 also developed the Synthetic Test Environment for training and validating AI scheduler models, the unified workflow ingestion layer accepting Kubernetes, Snakemake, and Argo Workflows formats, the Kubernetes-Slurm Integration (KSI) framework [5] enabling Kubernetes workloads on rootless

Fig. 3. WP2: Digital Twin architecture. The Digital Twin Core (DTC) maintains cluster state via InfluxDB. ML models predict carbon intensity, detect node anomalies, and forecast energy consumption. DT Clients integrate with the scheduler, control manager, and virtual test environment.

HPC systems, and ephemeral cluster management [10]. The final architecture (D3.2) defines how all components interact as modular, portable services across the federated infrastructure. WP4: Cloud Management Framework Integration (Lead: USTUTT; key contributors: GWDG, UGOE, KTH, HWDU) implemented the Prometheus-based monitoring stack that collects metrics from HPC, cloud, and edge nodes, including CPU/GPU utilization, memory pressure, network throughput, and storage I/O. The monitoring data feeds both the Digital Twin and the AI Scheduler’s decision pipeline. Fault-tolerance mechanisms for Kubernetes and Slurm in HPC environments were also investigated [6]. WP5: Deployment, Validation and Performance Assessment (Lead: E4; key contributors: UGOE, GWDG, USTUTT, TOP-IX, MARUN, BIGTRI, UNIBO) specified the development environment, deployed the integrated framework across two project phases, and conducted performance evaluation across real-world use cases including cloud robotics (Marmara/BIGTRI), scientific computing (KTH/PDC), and edge inference. The performance evaluation report (D5.5) validated the complete DECICE stack under production-like conditions. WP6: Dissemination and Exploitation (Lead: SYNYO) managed community engagement, with tutorials at ISC-HPC 2024 and 2025, the open-source release on GitHub, and an exploitation strategy (D6.2) outlining sustainability pathways for the framework beyond the funded period. III. E VALUATION R ESULTS Fig. 4 presents the IAIS scalability evaluation, measuring training and validation times from 10 jobs on 10 nodes to 5000 jobs on 5000 nodes at varying resource utilization levels. Even at the largest scale with 100% utilization, execution completes within 300 seconds, demonstrating practical applicability to production-sized clusters. Fig. 5 compares 10 scheduling tools across three workflow scenarios of increasing complexity. Exact MILP solvers pro-

Fig. 4. IAIS scalability: training and validation times across problem sizes up to 5000 jobs × 5000 nodes at 50% and 100% resource utilization.

compute continuum. Through six work packages and 12 consortium partners, the project produced an Integrated AI Scheduler with RNN-based prediction and formal workflow modeling, a Digital Twin with energy-aware and carbonaware prediction, a unified Kubernetes-based orchestration layer bridging cloud and HPC environments, a comprehensive monitoring stack, and validated use cases in cloud robotics, scientific computing, and edge inference. The modular architecture enables independent adoption of individual components, and the open-source release ensures continued community development beyond the funded period. Future directions include scaling to production workloads with thousands of tasks, extending GrapheonRL [3] to multi-objective scheduling across larger compute continuum deployments, and longitudinal energy-aware scheduling using real carbon intensity data.

25.46

30

27.6528 27.65

ACKNOWLEDGMENT Workflow 1: Without Data Transfer Time Workflow 2: With Data Transfer Time Workflow 3: With Higher Number of Dependencies

20

MILP Tools (With Exact Solution)

PSO

ACO

SA

0.01 0.001 0.001 0.01 0.001 0.01

GA

0 0.001 0.01 0.002 0.02

3.92 4.71 4.65 0.7 0.81 0.81

3.89 4.59 4.13

Gurobi OR-Tools

0.07 0.08 0.1

0.06 0.05 1.41

SCIP

0 PuLP

1.7103 1.71 1.39 1.3862

0.51 0.64

5.41

10

5

IEEE Computers, Software, and Applications Conference (COMPSAC 2026), Madrid, Spain, July 7–10, 2026.

This work was funded by the European Union’s Horizon Europe programme, Grant Agreement No. 101092582 (DECICE). The authors thank all 12 consortium partners for their contributions across six work packages. Website: https://www. decice.eu; source code: https://github.com/DECICE-project.

15

3.04 3.29

Run-Time (seconds)

25

© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses. Accepted at the 50th

HEFT

OLB

H/MH Techniques (With Approximate Solutions)

Different Tools and Techniques

Fig. 5. Runtime comparison of 10 scheduling tools across three workflow scenarios. PuLP and OR-Tools scale steeply under high dependency complexity (Workflow 3); Gurobi’s internal presolving maintains efficiency; heuristics remain sub-second throughout.

vide optimal solutions but scale steeply under high dependency counts: PuLP reaches 25.46s and OR-Tools 27.65s on Workflow 3, while Gurobi’s internal presolving keeps it at 1.41s. Heuristics (GA, PSO, ACO, SA, HEFT, OLB) maintain sub-second runtimes throughout, with HEFT at 0.01 to 0.02s. This benchmarking [1] guides scheduling strategy selection within the IAIS based on workload characteristics and time constraints. The project produced over 15 peer-reviewed publications consortium-wide, spanning IEEE COMPSAC [1]–[3], ACM Computing Frontiers [4], Future Generation Computer Systems [8], [9], Journal of Supercomputing [10], IARIA Intelligent Systems [5], ADVCOMP [6], and Springer FTC [7]. All software components were released under open-source licenses. Tutorials reached over 100 participants at ISC-HPC 2024 and 2025. IV. C ONCLUSION DECICE delivered a validated, open-source framework for AI-driven workload scheduling across the cloud-HPC-edge

R EFERENCES [1] A. K. Sharma, C. Boehme, P. Gelß, R. Yahyapour, and J. Kunkel, “Workflow-driven modeling for the compute continuum: An optimization approach to automated system and workload scheduling,” in Proc. IEEE COMPSAC, 2025. doi: https://doi.org/10.1109/COMPSAC65507. 2025.00343 [2] M. Bidollahkhani, A. K. Sharma, and J. Kunkel, “HOSHMAND: Accelerated AI-driven scheduler emulating conventional task distribution techniques for cloud workloads,” in Proc. IEEE COMPSAC, 2024. doi: https://doi.org/10.1109/COMPSAC61105.2024.00372 [3] A. K. Sharma and J. M. Kunkel, “Grapheon RL: A Graph Neural Network and Reinforcement Learning Framework for Constraint and Data-Aware Workflow Mapping and Scheduling in Heterogeneous HPC Systems,” in Proc. IEEE COMPSAC, 2025. doi: https://doi.org/10.1109/ COMPSAC65507.2025.00341 [4] J. Kunkel et al., “DECICE: Device-Edge-Cloud Intelligent Collaboration Framework,” in Proc. ACM Computing Frontiers (CF’23), 2023. doi: https://doi.org/10.1145/3587135.3592179 [5] J. Decker et al., “Enabling Kubernetes workload execution on rootless HPC systems with KSI,” Int. J. Advances in Intelligent Systems, vol. 18, no. 3&4, pp. 126–136, 2025. [6] M. Aydin, M. Bidollahkhani, and J. Kunkel, “Comparing fault-tolerance in Kubernetes and Slurm in HPC infrastructure,” in Proc. ADVCOMP, ISBN 978-1-68558-184-8, 2024, pp. 40–48. [7] M. Bidollahkhani et al., “Design and implementation of integrated AI scheduler for dynamic cloud workloads allocation in Kubernetes environments,” in Proc. FTC, Springer LNNS, 2025. doi: https://doi. org/10.1007/978-3-032-07986-2 25 [8] M. Molan et al., “GRAAFE: GRaph Anomaly Anticipation Framework for Exascale HPC systems,” Future Gener. Comput. Syst., vol. 160, 2024. doi: https://doi.org/10.1016/j.future.2024.06.032 [9] M. S. Ardebili, A. Acquaviva, L. Benini, and A. Bartolini, “HazardNet: A thermal hazard prediction framework for datacenters,” Future Gener. Comput. Syst., vol. 155, pp. 340–353, 2024. doi: https://doi.org/10.1016/ j.future.2024.01.031 [10] J. Decker and J. M. Kunkel, “Ephemeral Kubernetes: dynamically deleting and recreating clusters using Warewulf,” J. Supercomput., vol. 81, 2025. doi: https://doi.org/10.1007/s11227-025-07668-y

Record · ID 224471 · SHA-256 8e0d862cd2c20923
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.