Toward System-of-Systems Integration for Composable Cloud-HPC-Edge AI Platforms Sumit Rakesh
[email protected] Luleå tekniska universitet Luleå, Sweden
arXiv:2609.33427v1 [cs.DC] 27 Sep 2026
Abstract Modern AI platforms increasingly combine infrastructure stacks and operating models designed around different assumptions, including cloud-style service platforms, HPC workload-management systems, cloud-native orchestration, data and artifact systems, managed connectivity, observability, and edge or cyber-physical environments. Existing work demonstrates effective bridges between selected stacks, but a general way to reason about composition across independently controlled systems remains underdeveloped. We argue that such platforms can be usefully viewed as systems of systems (SoS) when independently useful systems retain their own control, management, lifecycles, policies, and failure semantics while contributing to a higher-level AI platform capability. We frame composable integration as an approach to cross-system coordination based on interfaces, contracts, mappings, references, policy context, and operational evidence, while preserving native control planes and avoiding dependence on a single topology or orchestration stack. The resulting direction is converged in use and federated in control. The paper presents a peer constituent-system view, a boundary test for distinguishing constituent systems from components, local dependencies, and independently useful systems outside the current SoS boundary, a local, shared, and scoped responsibility model, seven integration surfaces, and a representative cross-system workflow. It concludes with evidence classes and research questions for evaluating interoperability, governance, observability, fault containment, evolution, and reuse.
CCS Concepts • Computer systems organization → Distributed architectures; Cloud computing; Embedded and cyber-physical systems; Heterogeneous (hybrid) systems.
Keywords Cloud-HPC-Edge AI platforms, computing continuum, system-ofsystems, composable integration, AI infrastructure, cloud-native orchestration, interoperability
1
Introduction
AI infrastructure increasingly spans operating domains designed around different assumptions. A contemporary AI platform may combine large-scale training and fine-tuning, interactive development, online and batch inference, governed data and artifact management, accelerator-aware execution, cross-system observability, and interaction with sensors, cameras, robots, immersive clients, or industrial gateways. Cloud-style service platforms emphasize accessible services and resource abstraction [18]. HPC systems emphasize managed execution, accelerator locality, high-performance
Rajkumar Saini
[email protected] Luleå tekniska universitet Luleå, Sweden communication, and accounting [26]. Cloud-native orchestration emphasizes declarative control, automation, and repeatable deployment [4, 30]. Edge and cyber-physical systems emphasize locality, latency, device ownership, intermittent connectivity, and safety or privacy constraints [1, 8, 25]. The integration challenge emerges when these systems must participate in shared AI workflows while retaining their distinct control and operating models. Whether a Cloud-HPC-Edge AI platform should be treated as a system of systems (SoS) depends on the degree of independence among its participating systems. The SoS framing becomes appropriate when independently useful systems retain their own control, management, lifecycle, and failure semantics while contributing to a higher-level capability that no single system can provide alone. Although a multi-site federation makes the SoS character readily apparent, a single-site deployment may exhibit the same properties when independently managed systems retain distinct operators, lifecycles, policies, and control planes. Maier identifies operational and managerial independence as defining SoS characteristics [17], while ISO/IEC/IEEE guidance treats constituent systems as systems with their own lifecycle concerns that also participate in a SoS [13, 14]. These properties provide a stronger basis for defining constituent-system boundaries than grouping technologies within the same deployment. We present composable integration as the corresponding approach to SoS integration. It is distinct from composable infrastructure in the hardware-disaggregation sense and from applicationservice composition through a service mesh. Prior composablesystem work focuses on assembling compute, memory, storage, or accelerator resources under an infrastructure control plane [5], while service-mesh guidance addresses communication and policy among application services [2]. Our concern is platform-level composition across independently controlled systems. Such a composition relies on explicit interfaces, workflow contracts, identity and policy mappings, dataset and artifact references, deployment specifications, telemetry, and recovery expectations rather than on replacing native control planes or imposing a single orchestration stack. This paper makes three conceptual contributions. 1. First, it introduces a peer constituent-system integration model and a boundary test that distinguishes constituent systems from components, local dependencies, and independently useful systems outside the current SoS boundary. 2. Second, it defines local, shared, and scoped responsibilities together with seven integration surfaces that preserve native authority while making cross-system obligations explicit. 3. Third, it maps a representative remotely initiated edge-totraining-to-serving workflow across these views and derives testable
SoCC ’26, November 18–20, 2026, Singapore
evidence and research questions concerning interoperability, governance, fault containment, evolution, and reuse.
2
Cloud-HPC-Edge Convergence and the Architectural Gap
Three trends make the problem immediate. First, AI workload and accelerator diversity are widening. Training and inference suites cover workloads with different latency, data, memory, accuracy, and throughput objectives [19, 23]. Hardware surveys likewise span GPUs, TPUs, FPGAs, ASICs, NPUs, RISC-V accelerators, nearmemory designs, and other co-processors [29]. Placement, therefore, depends on workload stage, data locality, latency, accelerator capability, energy envelope, trust boundary, and failure risk—not only available compute. Second, the Cloud-HPC divide is becoming operationally visible. AI-factory work argues for a dual-stack direction that combines HPC performance with cloud-native usability and service-facing interfaces [10]. Current AI-factory reference architectures likewise separate compute, storage, external access, support, and out-ofband management network roles, illustrating the distinct operational paths present even within tightly integrated deployments [20]. Systems research has explored container orchestration on HPC [36], cloud-native workloads on HPC resources while preserving HPC accounting [3], multi-tenant RDMA for Kubernetes on HPC fabrics [9], and Kubernetes-Slurm integration for acceleratorbacked LLM serving [32]. These results show that selected bridges are feasible; they do not, by themselves, define ownership, policy, storage exposure, telemetry, or recovery across the complete platform. Third, AI increasingly spans the edge-to-cloud continuum. Surveys emphasize decentralized resources, low-latency processing, IoT-generated data, heterogeneous deployment models, and reproducibility challenges [1, 25]. DECICE adds scheduling, monitoring, and digital-twin coordination across cloud, HPC, and edge environments [28]. Research infrastructures such as the National Research Platform and CHI@Edge demonstrate multi-site Kubernetes, heterogeneous GPU/CPU resources, portals, edge experimentation, and operational monitoring [16, 34]. EuroHPC coordination of AI Factories and Antennas similarly targets technical and procedural interoperability and the sharing of data, applications, services, and knowledge across a distributed AI ecosystem [6]. Scientific cyberinfrastructure work documents convergence between AI and HPC, similarly at scale [12]. Existing references clarify important slices of the problem. HPC security guidance defines specialized threat and architecture concerns [11], while Zero Trust makes access decisions resource-centric rather than location-centric [24]. Kubernetes multi-tenancy guidance distinguishes control-plane and data-plane isolation [31]. Yet none of these viewpoints alone answer four cross-system questions: which elements are constituents, what remains locally controlled, which contracts are shared or scoped, and what evidence proves that a workflow is intentionally composed rather than accidentally wired together. That gap motivates a platform-level SoS model rather than another universal scheduler or reference topology.
Rakesh et al.
3
Requirements and Design Principles
The proposed SoS lens begins with system boundaries, responsibilities, and workflow intent rather than a specific implementation, technology stack, topology, or prescriptive deployment blueprint. The following statements consolidate the requirements and design principles that guide composable integration across independently controlled systems. They draw on SoS engineering, Cloud-HPC convergence, edge-continuum research, security guidance, artifact management, and interoperability work [1, 11, 13, 14, 17, 24, 25, 27, 31, 33]. D1: Preserve constituent-system autonomy. Native schedulers, orchestrators, data systems, managed fabrics, identity services, observability platforms, and edge controllers retain the control that makes them independently useful. D2: Make boundaries and ownership explicit. The SoS view distinguishes components, subsystems, constituent systems, and shared integration services. It also identifies ownership of scheduling, metadata, deployment, identity, telemetry, and recovery. D3: Compose through interfaces and contracts. Cross-system behavior is coordinated through APIs, events, references, policy mappings, deployment specifications, telemetry, and operational evidence rather than by replacing native control planes. D4: Support multiple execution mappings. Batch and distributed training, serving, batch inference, stream ingestion, interactive sessions, and edge inference or control may be assigned to, or span, different systems. Training and inference are lifecycle stages rather than fixed constituent-system classes. D5: Treat data, artifacts, provenance, and policy as integration objects. Dataset references, checkpoints, model artifacts, registry entries, lineage, retention, access context, and transfer paths are explicit rather than hidden inside scripts or job directories [27]. D6: Make tenancy, observability, and recovery cross-system properties. Isolation should reflect trust relationships and data sensitivity. Telemetry and audit evidence should support diagnosis, accountability, fault containment, and recovery [11, 21, 22, 24, 31]. D7: Remain topology-neutral and evolvable. Constituent-system boundaries and cross-system contracts should remain independent of a particular physical deployment, virtualization model, network topology, orchestration stack, or vendor implementation. Constituent systems should be added, replaced, or reconfigured through defined interfaces and contracts without redesigning the overall integration model. D8: Target selective workflow interoperability. The objective is sufficient technical, syntactic, semantic, and pragmatic interoperability to carry project, identity, dataset, job, artifact, endpoint, telemetry, quota, and failure context for a selected workflow, rather than universal interoperability among all systems [33]. Together, D1–D8 define complementary views for reasoning about peer constituent-system boundaries, constituent qualification, local, shared, and scoped responsibilities, integration surfaces, workflow realization, and system evolution. These views separate the principal integration concerns while allowing implementations, technologies, and topology to change without altering the underlying SoS reasoning model [15].
Toward System-of-Systems Integration for Composable Cloud-HPC-Edge AI Platforms
SoCC ’26, November 18–20, 2026, Singapore Composable Cloud-HPC-Edge AI Platform
4 System-of-Systems Integration Model 4.1 Peer Constituent-System View Figure 1 uses four peer constituent-system classes in the base view: an HPC workload-management system, a cloud-native orchestration system, a data/artifact system, and an edge/CPS environment. The classes are peers because each may own a distinct control boundary, operator group, lifecycle, policy model, and failure domain. They represent persistent operating and management boundaries rather than steps in an execution sequence. Training, serving, data preparation, and evaluation may be assigned to one constituent-system class or span several classes. The peer view, therefore, avoids treating infrastructure as a homogeneous resource pool. The HPC system retains authority over queueing, accounting, allocation, and locality. The cloud-native system retains authority over API state, controllers, namespaces, and service lifecycle. The data/artifact system controls metadata, storage layout, versioning, retention, and lineage. The edge/CPS environment controls devices, gateways, local capture, safety or privacy enforcement, and disconnected operation. The higher-level AI platform capability emerges from their coordinated contribution rather than from transferring these responsibilities to a new controller. The composable SoS integration layer is therefore a binding role rather than another execution system. It defines adapters, contracts, mappings, references, and evidence needed at system boundaries. It may be implemented by several services or mechanisms and need not lie on every data path. A possible common service/access layer can expose portals, APIs, catalogs, and tenant-aware entry without becoming a constituent by default. A managed connectivity fabric is included only when connectivity has independent management, policy, telemetry, lifecycle, and failure behavior. Cross-cutting concerns such as identity, governance, observability, audit, recovery, and evolution may remain local, be provided through shared services, or be handled by independent systems that satisfy the boundary test themselves. This separation underlies the principle: converged in use, federated in control. Users may experience a common platform entry point and workflow-status view, while each participating system remains responsible for its local decisions. Shared interfaces, contracts, and operational evidence provide enough context to coordinate workflows and diagnose failures without requiring complete state replication or global control. Figure 1 presents a structural view of operating and management boundaries and the integration relationships among them. The four peer constituent-system classes retain their native interfaces and control planes, while the composable SoS integration layer coordinates the cross-system exchange of workflow intent, policy context, dataset and artifact references, deployment specifications, telemetry, and recovery expectations. A possible service/access layer may provide a shared entry point but is not a constituent system by default. Managed connectivity and other supporting or cross-cutting capabilities are classified as constituent systems only when they satisfy the boundary test in Section 4.2.
SoS-level capability: converged in use, federated in control
Possible common service / access layer Portal, API, catalog, and tenant-aware entry workflow intent • access context • status
HPC workloadmanagement system Cloud-native orchestration system
Composable SoS integration layer interface adapters and bindings workflow contracts policy mappings, data/artifact references, telemetry, recovery assumptions not a master control plane connectivity contracts • path policy fabric telemetry • status
Data / artifact system
Edge / cyber-physical environment
Optional managed connectivity fabric QoS, segmentation, path policy, telemetry, fabric management
Cross-cutting systems and concerns identity, policy, tenancy, governance, observability, audit, recovery, evolution, operational evidence
Figure 1: Peer constituent-system classes integrated through a composable SoS integration layer, with a possible common service/access layer, an optional managed connectivity fabric, and cross-cutting systems and concerns. The constituentsystem blocks represent distinct operating and management boundaries rather than fixed workload destinations. The service/access layer and managed connectivity fabric represent optional roles and are not assumed to be constituent systems by default. Constituent systems may interoperate through shared interfaces and contracts while retaining local control.
4.2
Constituent-System Boundary Test
Constituent-system classification begins by determining which platform elements have both local independence and an explicit role in the higher-level AI platform capability. Workload managers, orchestration environments, data and artifact systems, connectivity fabrics, identity services, observability platforms, and edge/CPS environments are not constituent systems merely because they are technically complex. A candidate exhibits local independence when it provides a meaningful capability of its own, retains native control, has an independently managed lifecycle, and has a recognizable failure and recovery model. It has an integration role when it exposes interfaces or contracts through which it contributes to the higher-level AI platform capability. Classification is specific to the SoS scope being examined rather than to a technology category. An HPC workload-management system may qualify when it controls queues, resource allocation, accounting, operational policy, and recovery. A cloud-native orchestration system may qualify when it controls its API, desiredstate mechanisms, namespaces, service lifecycle, and upgrades. An edge/CPS environment may qualify when it has a local operational purpose, device or gateway control, privacy or safety constraints, and recognizable disconnected or degraded behavior. Capabilities such as data/artifact management, connectivity, identity, and observability are deployment-dependent because they may be operated as part of a larger system or as independent systems. The boundary test shown in Figure 2 and operationalized by the questions in Table 1 distinguishes these cases. A capability remains a component, subsystem, or local dependency when it lacks local independence: it does not provide a useful capability of its own or does not retain its own control, lifecycle, and recognizable failure and recovery model. Its scale or technical complexity does not
SoCC ’26, November 18–20, 2026, Singapore
Rakesh et al.
Table 1: Operational questions for the constituent-system boundary test.
Constituent-System Boundary Test Candidate platform element Not a constituent system in this SoS view
Component, subsystem, or local dependency
No
Boundary question Useful purpose
Operational interpretation If platform integration is removed, does the element still provide a meaningful capability? If not, it is probably a component or subsystem. Does it have its own scheduler, controller, fabric manNative control ager, identity provider, metadata service, or management plane? Independent lifecycle Can it be owned, operated, upgraded, staffed, or budgeted separately from the rest of the platform? Failure and recovery Does its failure create a recognizable operational event model with an independently defined recovery procedure? Interfaces/contracts Does it expose APIs, paths, events, policies, telemetry, or service endpoints through which other systems interact with it? Higher-level contribu- Does it contribute to the higher-level AI platform capation bility that no single participating system can provide alone?
Local independence? Useful purpose • Native control • Independent lifecycle • Failure and recovery model
Yes No
Independent system, but outside the current SoS boundary
Integration role? Interfaces/contracts • Contribution to the higher-level AI platform capability
Yes
Constituent system in the SoS model
Figure 2: Constituent-system boundary test. A candidate platform element is treated as a constituent system in the current SoS view when it has both local independence and an explicit integration role in the higher-level AI platform capability. An element without local independence is treated as a component, subsystem, or local dependency; an independently useful system without an integration role remains outside the current SoS boundary. change that classification. For example, a storage pool, network fabric, identity mechanism, or monitoring service operated as part of a larger system would not qualify as a constituent system solely because it is large or technically complex. Local independence alone is insufficient. An independently useful system may remain outside the current SoS boundary when it does not participate in the higher-level capability. Participation is established through explicit interfaces or contracts, such as job specifications, deployment requests, dataset and artifact references, policy mappings, telemetry context, or recovery expectations. The boundary test, therefore, avoids treating every complex subsystem as a constituent system while recognizing independently controlled systems whose participation materially shapes the operation, governance, recovery, or evolution of the integrated AI platform. Figure 2 separates three outcomes. A candidate lacking local independence is treated as a component, subsystem, or local dependency rather than a constituent system in the current SoS view. An independently useful candidate without an explicit integration role in the higher-level AI platform capability remains outside the current SoS boundary. Only a candidate satisfying both conditions is classified as a constituent system. Technical complexity alone therefore does not determine SoS status. Classification remains context-dependent. A local storage pool, internally managed network, embedded monitoring capability, or system-specific identity mapping is normally treated as a component or subsystem when it is operated as part of a larger system. By contrast, a storage platform, managed WAN or RDMA fabric, identity provider, or observability service may qualify as a constituent system when it has an independent operational purpose, operators, management plane, lifecycle, interfaces, and recovery procedures.
4.3
Local, Shared, and Scoped Responsibilities
The boundary test identifies which systems participate in the SoS; the responsibility model clarifies how authority and coordination
Table 2: Local, shared, and scoped responsibilities. Area
Local authority of con- Shared through compos- Scoped or isolated stituent systems able integration when required Control and Native scheduling, Submission and deploy- Project queues or parexecution desired-state control, ment interfaces, workflow titions, namespaces, placement, execution, identifiers and state, hand- reservations, virtual accounting, and local off events, placement con- clusters or control recovery. straints, and quota or allo- planes, and dedicated cation summaries. edge-control domains. Data and arti- Storage layout, metadata Dataset and artifact iden- Tenant buckets or facts services, file and object tifiers, locations, versions, volumes, encryption dostores, checkpoints, provenance, access context, mains, quotas, retention model registries, edge and transfer rules. policies, and regulatedbuffers, and retention data boundaries. enforcement. and fabric Reachability requirements, Tenant VLANs or VRFs, Connectivity Routing control, segmentation permitted paths, latency RDMA domains, dedienforcement, WAN or or QoS intent, exposure cated tunnels or WAN VPN operation, failure do- rules, connectivity status, paths, and isolated netmains, and local network and fabric telemetry. work segments. telemetry. Identity and Native identities, service Identity federation, subject Tenant identity policy accounts, authorization and role mapping, policy providers, projectmodels, and local policy translation, trust context, specific roles and service enforcement. and audit-correlation iden- accounts, external trust tifiers. domains, and workflowspecific data-use policies. Observability Local metrics, logs, traces, Cross-system correlation, Restricted dashboards and recovery health checks, alerts, run- workflow telemetry, SLO and log views, tenant books, and recovery mech- context, audit trails, and or project SLOs, reguanisms. failure or recovery status. lated audit retention, and workflow-specific recovery objectives.
are divided among them. Local authority remains with native constituent systems. Shared responsibilities are limited to the contracts and context required for cross-system coordination. Scoped controls apply additional isolation, reservation, or policy to a project, tenant, organization, site, workflow, or risk class. Table 2 applies this distinction across the principal platform areas.
5
Composable Integration Surfaces
Composable integration operates at the platform/system-integration layer. It does not pool hardware under a single infrastructure controller, nor does it turn independently managed systems into microservices. Its purpose is to expose enough stable structure for a workflow to cross boundaries while native systems retain authority.
Toward System-of-Systems Integration for Composable Cloud-HPC-Edge AI Platforms
We identify seven primary integration surfaces. The access surface defines how users, tenants, and external services submit workflow intent, discover capabilities, and obtain status. It may be realized through a common portal, APIs, command-line or notebook gateways, catalogs, native system entry points, or a combination of these. The control-plane interface surface exposes native capabilities through scheduler interfaces, orchestration APIs, storage APIs, fabric telemetry, edge-deployment interfaces, and policy endpoints. The execution surface maps requested activities to batch jobs, containers, services, notebooks, pipelines, or edge inference and control tasks. The data and artifact surface carries dataset and artifact references, object or file paths, checkpoints, registry entries, model versions, provenance, and transfer rules [27]. The observability surface defines how metrics, logs, traces, alerts, SLO context, audit events, and incident identifiers are exposed and correlated across system boundaries. OpenTelemetry and Prometheus provide useful mechanisms [21, 22], but the architectural requirement is to preserve workflow, tenant, and provenance context as operational evidence crosses systems. The policy and governance surface defines how identity, roles, quotas, tenant isolation, dataaccess rules, placement constraints, and compliance obligations are translated across local control domains, while enforcement remains with the native systems. The recovery and evolution surface makes explicit the failure domains, recovery responsibilities, fallback paths, compatibility constraints, upgrade boundaries, and conditions under which a constituent system may be replaced or extended. Each integration surface should expose only the information and capabilities required by the workflow. Native constituent systems retain authority over their internal operations: the HPC workload manager controls job scheduling and accounting; the cloud-native orchestrator controls services and desired state; storage systems control metadata and data layout; managed connectivity fabrics, when present, control network behavior; and edge/CPS systems control local execution. The composable integration layer exchanges workflow intent, references, policy context, and operational evidence without taking over these local decisions. This separation of cross-system coordination from native control distinguishes composable SoS integration from a monolithic platform control plane.
6
Representative Cross-System Workflow
Figure 3 illustrates how the integration surfaces operate in a remotely initiated edge-to-training-to-serving feedback loop. The example is illustrative rather than prescriptive. It does not define a mandatory pipeline or permanently assign lifecycle stages to particular constituent-system classes. Instead, it shows how activities remain under native control while cross-system handoffs carry explicit workflow intent, references, policy context, and operational evidence. In other deployments, edge capture may operate autonomously and enter the workflow at dataset registration or another coordination point.
SoCC ’26, November 18–20, 2026, Singapore
6.1
Remotely Initiated Edge - to - Training - to Serving Feedback Loop
The proposed integration model can support multiple cross-system AI workflows. We use a remotely initiated edge-to-training-toserving feedback loop because it spans the principal constituentsystem classes shown in Figure 1 and exercises several distinct integration needs. These include workflow intent, edge-side data selection, dataset registration, HPC-managed training or evaluation, artifact publication, cloud-native serving, edge deployment, telemetry, policy context, and feedback for retraining. Other workflows may omit, reorder, replace, or add activities, participating systems, and handoffs. A user or tenant initiates the workflow through a possible common service/access layer. The submitted workflow or control intent may configure edge-side capture and local filtering, while the edge/CPS environment retains authority over device operation, privacy enforcement, and local data selection. Selected data is then registered in the data/artifact system together with the metadata, policy context, and access information required by downstream activities. The data/artifact system exposes the registered dataset through an explicit reference, which is combined with a job specification and submitted to the HPC workload-management system. The HPC system applies its native queueing, placement, accounting, and failure-handling policies. Training, fine-tuning, or evaluation produces a model artifact, an evaluation result, and associated lineage information. These outputs are published through the data/artifact system, which provides the versioned references required by downstream consumers. The published artifact may be used by the cloud-native orchestration system to create a serving endpoint, by the edge/CPS environment for edge deployment, or by both. These are independent downstream mappings rather than sequential stages. The solid arrows in Figure 3 denote governed workflow handoffs. C1 carries workflow or control intent. C2 carries selected data and policy context. C3 carries the dataset reference and job specification. C4 carries the model artifact and evaluation result. C5-S and C5-E carry the serving specification and edge rollout profile, respectively. The dashed C6 arrows carry telemetry, audit records, evaluation results, and failure evidence into the cross-system evidence view. The dashed C7 path carries status, evaluation outcomes, or an optional retraining trigger back to the service/access layer. Identity and policy, tenancy and governance, provenance, telemetry and audit, and recovery context are attached only to the handoffs for which they are relevant.
6.2
Alternative Workflow and Tenancy Mappings
The workflow mapping in Figure 3 is illustrative rather than fixed. For example, a cloud-native service endpoint may use accelerators allocated by an HPC workload-management system. Prior work integrating Kubernetes, Slurm, and vLLM demonstrates the feasibility of this execution pattern [32]. From the SoS perspective, the relevant requirement is that the cross-system interface preserves HPC allocation, accounting, observability, and failure semantics while the service-facing system retains responsibility for endpoint lifecycle.
SoCC ’26, November 18–20, 2026, Singapore
Rakesh et al.
Representative Cross-System Workflow
Illustrative realization across peer constituent-system boundaries; local control remains native Submit intent
tenant context • workflow request
Status / retraining decision
Possible common service / access layer
Shared workflow entry, status, and tenant-aware access
policy • operator • automation
C1
Edge / CPS environment
Capture + local filter privacy / local selection
Data / artifact system HPC workload-management system
Deploy to edge
rollout profile • constraints
C2
C5-E Register dataset
Publish artifact
metadata • policy reference
C3 Train / fine-tune / evaluate
version • lineage
C4 C5-S
queue • placement • accounting
Serve endpoint
Cloud-native orchestration system
handoff
deployment specification
telemetry / feedback
C6
Cross-system evidence
telemetry • audit • evaluation • failure status
C7
Cross-cutting context across relevant handoffs: identity/policy • tenancy/governance • provenance • telemetry/audit • recovery expectations
C1 workflow intent • C2 selected data + policy • C3 dataset reference + job specification • C4 trained model + evaluation • C5-S serving specification • C5-E edge rollout profile • C6 telemetry and audit • C7 evaluation or retraining trigger
Figure 3: Representative cross-system AI workflow across peer constituent-system boundaries. A possible common service/access layer initiates the remotely triggered workflow, while activities execute under the native control of the edge/CPS, data/artifact, HPC workload-management, and cloud-native orchestration systems. Solid arrows denote governed handoffs; dashed arrows denote telemetry, operational evidence, and feedback. After artifact publication, cloudnative serving and edge deployment are independent downstream destinations. Cross-cutting context accompanies relevant handoffs rather than forming a separate workflow stage.
The same logical workflow may also be scoped differently for different tenants or trust relationships. A multi-tenant research platform may expose common training, storage, serving, and edge capabilities through distinct namespaces, identity mappings, storage domains, network segments, audit views, reservations, or virtual control planes. Kubernetes multi-tenancy guidance and virtual-cluster research show that control-plane and data-plane isolation are architectural choices rather than minor deployment details [31, 35]. These variants reinforce the central principle: workflow stages may be assigned to or span different constituent-system classes, while native systems retain local authority, and cross-system behavior is coordinated through explicit interfaces and contracts.
7
Feasibility Evidence and Research Agenda
A credible SoS integration approach must be testable at the level of system boundaries, workflow handoffs, operations, tenancy, evolution, and cost. We therefore identify six classes of evidence. Boundary and ownership evidence tests whether operators can consistently distinguish components, constituent systems, and external systems, and whether ownership, interfaces, and failure domains are explicit. Workflow and interoperability evidence tests whether each cross-system handoff can be described at the technical, syntactic, semantic, and pragmatic levels [33]. Operational evidence tests whether telemetry and audit records preserve workflow context, support diagnosis, and demonstrate fault containment. Tenancy and policy evidence tests whether the same workflow contracts can be safely scoped for internal, external, or regulated users. Evolution and replacement evidence tests whether a scheduler, data system, edge site, serving runtime, identity provider, or observability backend can be added or replaced without redesigning the full platform.
Performance and operational-cost evidence tests whether the benefits of composition justify its overhead in latency, throughput, integration effort, diagnosis time, policy administration, and workflow reuse. These evidence classes lead to five research questions. RQ1: How should operators distinguish a constituent system from a component, subsystem, or external system? RQ2: What is the minimum interface and contract description required for independently controlled systems to participate in reusable workflows? RQ3: How should shared concepts such as project, job, dataset, model artifact, tenant, quota, SLO, provenance, and failure domain be mapped across control planes? RQ4: How should placement and handoff decisions account for locality, latency, accelerator type, energy, isolation, and recovery, building on continuum-scheduling work such as DECICE [28]? RQ5: How can operators diagnose cross-system failures, verify containment, and replace a constituent system while preserving workflow contracts and operational evidence? Evaluation can proceed incrementally. An initial prototype should classify the participating systems, execute one governed cross-system workflow, and correlate telemetry across its handoffs. A stronger evaluation should add fault injection, tenant-specific policies, constituent-system replacement, and measurements of interface overhead, diagnosis effort, and workflow reuse.
8
Conclusion
This paper advances a system-of-systems framing for composable Cloud-HPC-Edge AI platforms. It makes constituent-system boundaries, ownership, interfaces, contracts, and operational evidence explicit. Consistent with the architecture-description perspective of ISO/IEC/IEEE 42010 [15], the framing is expressed through complementary views of peer constituent-system classes, constituent qualification, local, shared, and scoped responsibilities, integration surfaces, workflow realization, and system evolution. Cloud-HPC-Edge AI platforms become difficult to reuse when portals, schedulers, orchestrators, data services, connectivity fabrics, observability mechanisms, and edge environments are connected through deployment-specific scripts, implicit mappings, and undocumented failure assumptions. A system-of-systems perspective provides a more disciplined alternative. Composable integration preserves native authority while defining reusable contracts for access, control-plane interaction, execution, data and artifacts, policy and governance, observability, recovery, and evolution. The proposed direction is converged in use, federated in control, and composed through explicit interfaces and evidence. This framing can serve as a basis for reproducible architecture descriptions, testable evaluation criteria, incremental constituent-system replacement, and future federation across sites and organizations. It may also extend to additional independently managed systems as computing platforms evolve. For example, emerging hybrid quantumHPC environments introduce distinct execution models, control mechanisms, interfaces, and lifecycles that could be incorporated using the same boundary and integration principles [7].
Disclosure Generative AI tools assisted with language editing, graphics enrichment, and reference checking. The authors verified the claims, citations, and final wording and take responsibility for the content.
Toward System-of-Systems Integration for Composable Cloud-HPC-Edge AI Platforms
References [1] Loris Belcastro, Fabrizio Marozzo, Alessio Orsino, Domenico Talia, and Paolo Trunfio. 2026. Navigating the Edge-Cloud Continuum: A State-of-Practice Survey. IEEE Access 14 (2026), 40622–40647. doi:10.1109/ACCESS.2026.3673012 [2] Ramaswamy Chandramouli and Zack Butcher. 2020. Building Secure MicroservicesBased Applications Using Service-Mesh Architecture. NIST Special Publication 800-204A. National Institute of Standards and Technology, Gaithersburg, MD, USA. doi:10.6028/NIST.SP.800-204A [3] Antonis Chazapis, Evangelos Maliaroudakis, Fotis Nikolaidis, Manolis Marazakis, and Angelos Bilas. 2024. Running Cloud-Native Workloads on HPC with High-Performance Kubernetes. arXiv preprint arXiv:2409.16919 (2024). arXiv:2409.16919 [cs.DC] doi:10.48550/arXiv.2409.16919 [4] Cloud Native Computing Foundation. 2024. CNCF Cloud Native Definition v1.1. https://github.com/cncf/toc/blob/main/DEFINITION.md. Approved by the CNCF Technical Oversight Committee and Governing Board on February 26, 2024; accessed: 2026-07-14. [5] Kaoutar El Maghraoui, Lorraine M. Herger, Chekuri Choudary, Kim Tran, Todd Deshane, and David Hanson. 2021. Performance Analysis of Deep Learning Workloads on a Composable System. In 2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE, 951–954. https: //ieeexplore.ieee.org/document/9460648 [6] European High Performance Computing Joint Undertaking (EuroHPC JU). 2026. Cooperation of Artificial Intelligence Factories and Factories Antennas. https://www.eurohpc-ju.europa.eu/cooperation-artificial-intelligencefactories-and-factories-antennas-0_en. Published April 28, 2026; accessed: 202607-14. [7] European High Performance Computing Joint Undertaking (EuroHPC JU). 2026. Quantum Computing & Access. https://www.eurohpc-ju.europa.eu/quantumtechnologies/quantum-computing-access_en. Accessed: 2026-07-14. [8] European Telecommunications Standards Institute. 2025. Multi-access Edge Computing (MEC); Framework and Reference Architecture. ETSI Group Specification ETSI GS MEC 003 V4.1.1. European Telecommunications Standards Institute. https://www.etsi.org/deliver/etsi_gs/mec/001_099/003/04.01.01_60/gs_ mec003v040101p.pdf [9] Philipp A. Friese, Ahmed Eleliemy, Utz-Uwe Haus, and Martin Schulz. 2025. Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes. In 2025 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, 1–10. doi:10.1109/CLUSTER59342.2025.11186471 [10] Pedro Garcia Lopez, Daniel Barcelona Pons, Marcin Copik, Torsten Hoefler, Eduardo Quiñones, Maciej Malawski, Peter Pietzuch, Alberto Marti, Thomas Ohlson Timoudas, and Aleksander Slominski. 2025. AI Factories: It’s Time to Rethink the Cloud–HPC Divide. arXiv preprint arXiv:2509.12849 (2025). arXiv:2509.12849 [cs.DC] doi:10.48550/arXiv.2509.12849 [11] Yang Guo, Ramaswamy Chandramouli, Lowell Wofford, Rickey Gregg, Gary Key, Antwan Clark, Catherine Hinton, Andrew Prout, Albert Reuther, Ryan Adamson, Aron Warren, Purushotham Bangalore, Erik Deumens, and Csilla Farkas. 2024. High-Performance Computing Security Architecture, Threat Analysis, and Security Posture. NIST Special Publication 800-223. National Institute of Standards and Technology, Gaithersburg, MD, USA. doi:10.6028/NIST.SP.800-223 [12] E. A. Huerta, Asad Khan, Edward Davis, Colleen Bushell, William D. Gropp, Daniel S. Katz, Volodymyr Kindratenko, Seid Koric, William T. C. Kramer, Brendan McGinty, Kenton McHenry, and Aaron Saxton. 2020. Convergence of Artificial Intelligence and High Performance Computing on NSF-Supported Cyberinfrastructure. Journal of Big Data 7, 1, Article 88 (2020). doi:10.1186/s40537-020-00361-2 [13] ISO/IEC/IEEE. 2019. Systems and Software Engineering—Guidelines for the Utilization of ISO/IEC/IEEE 15288 in the Context of System of Systems (SoS) (1 ed.). International Standard ISO/IEC/IEEE 21840:2019. International Organization for Standardization. https://www.iso.org/standard/71956.html [14] ISO/IEC/IEEE. 2019. Systems and Software Engineering—System of Systems (SoS) Considerations in Life Cycle Stages of a System (1 ed.). International Standard ISO/IEC/IEEE 21839:2019. International Organization for Standardization. https: //www.iso.org/standard/71955.html [15] ISO/IEC/IEEE 42010:2022. 2022. Software, Systems and Enterprise—Architecture Description. International Standard ISO/IEC/IEEE 42010:2022. International Organization for Standardization. https://www.iso.org/standard/74393.html [16] Kate Keahey, Michael Sherman, Jason Anderson, and Mark Powers. 2025. CHI@Edge: Supporting Experimentation in the Edge to Cloud Continuum. In Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration (PEARC ’25). Association for Computing Machinery, New York, NY, USA, Article 11, 8 pages. doi:10.1145/3708035.3736014 [17] Mark W Maier. 1998. Architecting principles for systems-of-systems. Systems Engineering: The Journal of the International Council on Systems Engineering 1, 4 (1998), 267–284. [18] Peter M. Mell and Timothy Grance. 2011. The NIST Definition of Cloud Computing. NIST Special Publication 800-145. National Institute of Standards and Technology, Gaithersburg, MD, USA. doi:10.6028/NIST.SP.800-145
SoCC ’26, November 18–20, 2026, Singapore
[19] MLCommons. [n. d.]. MLPerf Training. https://mlcommons.org/benchmarks/ training/. Accessed: 2026-07-14. [20] NVIDIA Corporation. 2026. Networking Logical Architecture. NVIDIA HGX AI Factory Enterprise Reference Architecture. https://docs.nvidia.com/enterprisereference-architectures/hgx-ai-factory/latest/network-logical-architecture. html Last updated May 18, 2026; accessed: 2026-07-14. [21] OpenTelemetry Authors. 2026. What Is OpenTelemetry? https://opentelemetry. io/docs/what-is-opentelemetry/. Accessed: 2026-07-14. [22] Prometheus Authors. [n. d.]. Overview: What is Prometheus? https://prometheus. io/docs/introduction/overview/. Accessed: 2026-07-14. [23] Vijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Brian Anderson, Maximilien Breughe, Mark Charlebois, William Chou, Ramesh Chukka, Cody Coleman, Sam Davis, Pan Deng, Greg Diamos, Jared Duke, David Fick, Jason Scott Gardner, Itay Hubara, Suryanarayana Idgunji, Thomas B. Jablin, Jeff Jiao, Tom St. John, Pankaj Kanwar, David Lee, Jeffery Liao, Anton Lokhmotov, Francisco Massa, Peng Meng, Paulius Micikevicius, Colin Osborne, Gennady Pekhimenko, Arun Tejusve Rajan, Dilip Sequeira, Ashish Sirasao, Fei Sun, Hanlin Tang, Michael Thomson, Frank Wei, Ephrem Wu, Lingjie Xu, Koichi Yamada, Bing Yu, George Yuan, Aaron Zhong, and Peizhao Zhang. 2020. MLPerf Inference Benchmark. In Proceedings of the 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA ’20). IEEE, 446–459. doi:10.1109/ISCA45697.2020.00045 [24] Scott W. Rose, Oliver Borchert, Stuart Mitchell, and Sean Connelly. 2020. Zero Trust Architecture. NIST Special Publication 800-207. National Institute of Standards and Technology, Gaithersburg, MD, USA. doi:10.6028/NIST.SP.800-207 [25] Daniel Rosendo, Alexandru Costan, Patrick Valduriez, and Gabriel Antoniu. 2022. Distributed Intelligence on the Edge-to-Cloud Continuum: A Systematic Literature Review. J. Parallel and Distrib. Comput. 166 (2022), 71–94. doi:10.1016/ j.jpdc.2022.04.004 [26] SchedMD. [n. d.]. Slurm Workload Manager Overview. https://slurm.schedmd. com/overview.html. Accessed: 2026-07-14. [27] Marius Schlegel and Kai-Uwe Sattler. 2023. Management of Machine Learning Lifecycle Artifacts: A Survey. ACM SIGMOD Record 51, 4 (2023), 18–35. doi:10. 1145/3582302.3582306 [28] Aasish Kumar Sharma, Felix Stein, Mirac Aydin, Michael Bidollahkhani, Sachin P. Nanavati, Mohsen Seyedkazemi Ardebili, Giorgi Mamulashvili, Mojtaba Akbari, Jonathan Decker, Zoya Masih, and Julian M. Kunkel. 2026. DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud–HPC–Edge Compute Continuum. arXiv preprint arXiv:2605.25292 (2026). arXiv:2605.25292 [cs.DC] doi:10.48550/arXiv.2605.25292 [29] Cristina Silvano, Daniele Ielmini, Fabrizio Ferrandi, Leandro Fiorin, Serena Curzel, Luca Benini, Francesco Conti, Angelo Garofalo, Cristian Zambelli, Enrico Calore, Sebastiano Fabio Schifano, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Nicola Petra, Davide De Caro, Luciano Lavagno, Teodoro Urso, Valeria Cardellini, Gian Carlo Cardarilli, Robert Birke, and Stefania Perri. 2025. A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms. Comput. Surveys 57, 11, Article 286 (2025). doi:10.1145/3729215 [30] The Kubernetes Authors. 2026. Kubernetes Components. https://kubernetes.io/ docs/concepts/overview/components/. Last modified May 30, 2026; accessed: 2026-07-14. [31] The Kubernetes Authors. 2026. Multi-tenancy. https://kubernetes.io/docs/ concepts/security/multi-tenancy/. Accessed: 2026-07-14. [32] Tim Trappen, Robert Keßler, Roland Pabel, Viktor Achter, and Stefan Wesner. 2025. Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM. In Proceedings of the 2025 Workshop on Middleware for Next Generation Data-Intensive Applications (MIND ’25). Association for Computing Machinery, New York, NY, USA, 13–18. doi:10.1145/3774902.3776632 [33] Wenguang Wang, Andreas Tolk, and Weiping Wang. 2009. The Levels of Conceptual Interoperability Model: Applying Systems Engineering Principles to M&S. In Proceedings of the 2009 Spring Simulation Multiconference (SpringSim ’09). Society for Computer Simulation International, San Diego, CA, USA, Article 168, 9 pages. doi:10.5555/1639809.1655398 [34] Derek Weitzel, Ashton Graves, Sam Albin, Huijun Zhu, Frank Wuerthwein, Mahidhar Tatineni, Dmitry Mishin, Elham Khoda, Mohammad Sada, Larry Smarr, Thomas DeFanti, and John Graham. 2025. The National Research Platform: Stretched, Multi-Tenant, Scientific Kubernetes Cluster. In Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration (PEARC ’25). Association for Computing Machinery, New York, NY, USA, Article 69, 5 pages. doi:10.1145/3708035.3736060 [35] Chao Zheng, Qinghui Zhuang, and Fei Guo. 2021. A Multi-Tenant Framework for Cloud Container Services. In 2021 IEEE 41st International Conference on Distributed Computing Systems (ICDCS). IEEE, 359–369. doi:10.1109/ICDCS51616.2021.00042 [36] Naweiluo Zhou, Yiannis Georgiou, Li Zhong, Huan Zhou, and Marcin Pospieszny. 2020. Container Orchestration on HPC Systems. In 2020 IEEE 13th International Conference on Cloud Computing (CLOUD). IEEE, 34–36. doi:10.1109/CLOUD49709. 2020.00017