Clustered Edge Intelligence: Beyond Just Convergence of Edge Computing and AI Chinmaya Kumar Dehurya , Boris Sedlakb , Alaa Salehc , Ilir Murturid , Lauri Lovéne , Satish Narayana Sriramaf , Praveen Kumar Dontag,∗ a Department of PComputer Science, IISER Berhampur, Odisha, India b Distributed Systems Group, TU Wien, 1040 Vienna, Austria. c Department of Computer Science, University of Helsinki, 00100 Helsinki, Finland d Department of Mechatronics, University of Prishtina, Prishtina 10000, Kosova. e Center for Applied Computing, University of Oulu, 90100 Oulu, Finland f School of Computer and Information Sciences, University of Hyderabad, India. g Department of Computer and Systems Sciences, Stockholm University, 164 25 Stockholm, Sweden.
Abstract
arXiv:2607.20937v1 [cs.AI] 23 Jul 2026
We are moving from an information age to the age of intelligence. A decade, or possibly less than that, data will not be the gold anymore rather the derived intelligence out of the data and the information we posses from the edge of the network. Existing Edge Intelligence research focuses mainly on two directions: using AI for edge resource management and deploying lightweight AI models on edge devices. However, existing edge-computing research lacks an intelligence-centric framework in which derived intelligence is treated as a first-class, independently manageable entity that can be described, discovered, observed, shared, reused, and dynamically clustered across heterogeneous edge devices and applications. To address these research gaps, we introduced Clustered Edge Intelligence (CEI), a visionary intelligence-centric approach. The aim of CEI is to make intelligence a shareable and reusable first-class entity that can be independently represented, discovered, observed, exchanged, and managed across the distributed edge–cloud continuum. We present a three-layer CEI architecture and examine enabling technologies and research dimensions, including intelligence inventories, semantic knowledge representation, communication, discoverability, observability, lifecycle automation, clustering mechanisms, marketplaces, interoperability, and standardization. Keywords: Clustered Edge Intelligence, intelligence-centric clustering, edge agent, Edge Computing, AI understand or to deal with new or trying situations. In computer science, since decades a portion of the research We are transitioning from the Information Age to the revolving around embedding high-degree of intelligence onto Age of Intelligence, where the saying “data is the new gold” the servers with high-end resource configuration. However, will soon be a thing of the past. The spotlight will shift the research is slowly moving towards embedding certain from merely managing and analyzing vast quantities of level of intelligence to resource-constraints devices, e.g. the data—say zettabytes to exabytes to yottabytes—to manedge devices. aging and utilizing the derived intelligence. The process of In the context of edge infrastructure, consisting of biltransforming raw data into actionable intelligence involves lions of edge devices, lets take an example and understand a series of steps where data is contextualized into informathe transformation of data into intelligence, as shown in tion, understood as knowledge, and ultimately applied to Figure 1. In this example, data is first collected from make informed decisions. two sensors, say numerical value 45 from a temperature sensor and 97 from a PM2.5 sensor. Let’s assume that Knowledge & Intelligence. In the domain of computer sciboth sensors are connected to one edge device. These raw ence, knowledge refers to the information, facts, or undata points, however, are devoid of context until they are derstanding of the surrounding environment, including all given meaning, transforming them into information. For tangible and intangible entities that we have acquired through instance, the value 45 is contextualized as 45◦ C temp, and a set of dedicated resource-constraint devices and networks 97 is understood as 97µg/m3 PM2.5, providing a clear un[1]. On the other hand, intelligence encompasses the cogderstanding of what these numbers represent in the real nitive abilities, reasoning, and learning capacity that make world. This information is then processed further, where use of the existing knowledge [2]. Based on the Merriamthe device recognizes or becomes aware of the current en1 Webster dictionary, intelligence is the ability to learn or vironmental conditions, thus converting this information into knowledge. Knowledge in this context means that a ∗ Corresponding author system is not only aware of the current temperature and 1 1. Introduction
https://www.merriam-webster.com/dictionary/intelligence
97 PM2.5 sensor
Data
Give a meaning
PM2.5
Information
Aware of Use of current knowledge temp
Realization of having information
Aware of current Use of PM2.5 knowledge value
Knowledge
Categorize
Very cold Cold
Take action
...
Temp sensor
temp
Realization of having information
Very hot
Categorize
Very good good ...
45
Give a meaning
Take action
worse air quality
Intelligence
Edge Device Figure 1: The transformation from Data to Intelligence: an example in Edge device
air quality levels but also understands their significance. Finally, this knowledge is used to categorize the environmental conditions into states such as Very Hot or Good air quality. Based on environmental conditions, the edge device may take appropriate action. For instance, in a smart home scenario, the edge device may directly communicate with the air purifier to turn on when the air quality is low. It may also directly instruct the air conditioning system to cool the environment. This final stage, where the system uses its knowledge to make decisions and act, represents the transformation of knowledge into intelligence. This demonstrates a clear distinction between knowledge and intelligence within the context of an edge infrastructure. Can we refer intelligence available at the edge of the network as “Edge Intelligence”? More onto this is discussed in Section 2.
frastructure, where the number of edge devices may scale beyond few millions or billions, management of those intelligence is a cumbersome task. The current research is mostly progressing through clustering resource-constraint IoT devices [7], where clusters are formed mainly based on the geographic proximity, communication distance, sensing coverage, or proximity to an edge server [8, 9, 10]. Such clusters are mainly responsible for performing a set of tasks with a single objective to achieve. This approach itself becoming a hurdle for implementing the edge intelligence that may involves a large number of IoT devices deployed in different geographical locations or of different types. For instance, with a conventional location-based clustering approach [11], it may not be possible to form a cluster of sensors that are available in the city buses. We may face the similar difficulties in case of clustering the sensors deployed in different solar panels in a smart city. To address the above limitations, this paper introduces Clustered Edge Intelligence (CEI), an intelligence-centric vision where intelligence is considered as a first-class entity for its efficient management across distributed computing continuum. The major contributions of this paper are summarized as follows.
Cohesion & Coherence . In general, the term Cohesion can be defined as binding different elements/components to create a singular unit [3]. In the context of edge intelligence, Cohesion would refer to the integration and binding of intelligence available at different types of edge devices [4]. On the other hand, the term Coherence can be defined as establishing a logical connection between diverse elements/components. In the context of edge intelligence, coherence can be seen as the connectedness/consistency that exists among intelligence possessed by different edge devices. For instance, a logical connection can be established between different sensors in the smart home scenario. The luminosity sensor and the proximity sensor are logically connected to the smart bulb. The level of CO2 is logically connected with traffic congestion and noise level. Higher noise levels on the road might indicate a higher value of CO2, and in a similar way, a higher value of CO2 may indicate higher noise pollution and more traffic congestion [5, 6]. This shows a coherence or a logical relationship among different types of intelligence available at the edge of the network. In a very large-scale of edge in-
• To address the above limitations, this paper introduces Clustered Edge Intelligence (CEI), an intelligencecentric vision where intelliegnce is considered as a first-class entity for its efficient management across distributed computing continuum. • We revisited the meaning of Edge Intelligence and classify it into three forms: AI for edge resource management, AI deployed at the edge, and intelligence available at the edge (our primary focus). • CEI as a new paradigm is introduced where intelligence can be represented, discovered, observed, shared, reused, and clustered across heterogeneous edge devices. Following this, 2
• we compared device-centric and intelligence-centric clustering and show how CEI decouples intelligence from the underlying hardware.
Table 1: List of acronyms used within this paper. Acronym Description EI Edge Intelligence EI-1F Edge Intelligence – First Form (AI for Edge Resource Management) EI-2F Edge Intelligence – Second Form (AI at the Edge Infrastructure) EI-3F Edge Intelligence – Third Form (Intelligence Available at the Edge) ERM Edge Resource Management EC Edge Computing CEI Clustered Edge Intelligence EA Edge Agent DCC Device-Centric Clustering ICC Intelligence-Centric Clustering FL Federated Learning RL Reinforcement Learning IoT Internet of Things API Application Programming Interface REST Representational State Transfer Pub-Sub Publish–Subscribe Communication Model MQTT Message Queuing Telemetry Transport HTTP Hypertext Transfer Protocol IP intelligence provider HAL Hardware Abstraction Layer MPU Microprocessor Unit EIM Edge Intelligence Marketplace EIP Edge Intelligence Producer ICC Intelligence Cluster DD Device Discovery D2D Device-to-Device SD Service Discovery SR Service Registry
• A three-layer CEI architecture is proposed consisting of edge devices, edge controllers, and cloud environments. • We identified key baseline technologies required for design and implementation of CEI, including abstraction, knowledge graphs, communication, discoverability, observability, clustering, and lifecycle-support mechanisms. • Finally, we discussed major research dimensions and representative use cases to guide future research on intelligence-centric edge systems. The remainder of this paper is organized as follows. In addition, for a quick reference, the overall outline of the manuscript is presented in Figure 2. Section 2 discusses the different forms of Edge Intelligence, including AI for edge resource management, AI deployed at the edge, and intelligence available at the edge. Section 3 introduces the concept of Clustered Edge Intelligence, explains the motivation behind clustering intelligence, and compares device-centric and intelligence-centric clustering mechanisms. Section 4 presents the proposed CEI architecture, including the edge device layer, edge controller layer, cloud environment layer, and key system characteristics. Section 5 discusses the baseline technologies required to realize CEI, such as abstraction, knowledge graphs, communication mechanisms, discoverability, observability, clustering mechanisms, and other supporting technologies. Section 6 highlights the major research dimensions, including intelligence marketplaces, intelligence lifecycle automation, clustering mechanisms, the use of language models, interoperability, and standardization. Section 7 presents representative use cases of CEI in Industry 4.0, smart city and traffic management, smart agriculture, and other application domains. Section 8 provides the survey of existing research works. Finally, Section 9 concludes the paper and outlines future research directions. Table 1 provides the list of frequently used acronyms in this article.
Edge Tensor Processing Unit (EdgeTPU)-optimized models, YOLOv5-Nano 2 , ShuffleNet [16], etc) deployed at the edge of the network [17], as shown in Figure 3. In this article, this definition is referred to as 2nd Form (acronym EI-2F) and used henceforth. However, in this research, it is believed that the term EI may also refereed to the intelligence available at the edge of the network, as shown in Figure 3, considered as 3rd Form (acronym EI-3F) and used henceforth. Such intelligence can be obtained using any AI, non-AI algorithms or any custom script. This can be realized in the context of smart city, supply chain, smart agriculture, and industry 4.0 real-world scenarios [18]. Detailed explanation of all possible forms of EI and their comparisons are provided below.
2. Different forms of “Edge Intelligence” 2.1. EI-1F: AI for edge resource management Unlike the conventional approach of manual and bruteforce approach of managing the edge resource management, the recent advancement of AI/ML is significantly integrated for the same purpose. AI is used in handling wide-range of tasks related to edge resource management, including network allocation [19], energy-aware scheduling [20], task migration [21], compute allocation [22], upload scheduling and many more.
The term “Edge Intelligence (EI)” may have different meaning in different context and situations [12]. This can be referred to the design and deployment of AI algorithm for management of the edge computing, storage and network resources [13]. This definition is referred to as 1st Form (acronym EI-1F) in this article and used henceforth. On the other hand, in recent days the term “Edge Intelligence (EI)” is also refers to the minimalist version of an AI model (e.g. MobileNetV2 [14], SqueezeNet [15],
2 https://docs.ultralytics.com/models/yolov5
3
2. Different forms of “Edge Intelligence”
3. Clustered Edge Intelligence
2.1. EI-1F: AI for edge resource management
3.2. An example of CEI
4.1. Edge device layer
3.3. Device- vs Intelligencecentric clustering mechanism 3.3.1. Device-centric clustering Approach 3.3.2. Intelligencecentric clustering Approach
3.3.3. Comparision
2.3. EI-3F: Intelligence available at edge
4. CEI Architecture
3.1. Why “Clustered”? 2.4. Comparision
2.2. EI-2F: AI at edge infrastructure
9. Conclusions and Future Works
CEI Manuscript Organisation
1. Introduction
4.2. Edge controller layer 4.3. Cloud Environment layer 4.4. Characteristic of CEI
5. Baseline Technologies 5.1. Abstraction
5.3.1. REST
5.2. Knowledge Graph
5.3.2. Pub-Sub
5.3. Communication
5.3.3. 5G/6G and LoRaWAN
5.4. Discoverability 5.5. Observability 5.6. Clustering mechanisms 5.7. Other baseline technologies
6. Research Dimensions 6.1. Intelligence marketplace 6.2. Intelligence discoverability & observability 6.3. Intelligence lifecycle automation
5.4.1. Service Discovery
6.4. Clustering mechanism
5.4.2. Device Discovery
6.5. Use of language models
5.4.3. Network Discovery
6.6. Interoperability and Standardization
7. Use cases
7.1. Industry 4.0
7.2. Smart City & Traffic Management System 7.3. Smart Agricultural
8. Literature Survey 8.1. Paradigms and Taxonomy 8.2. Semantic Representation & Abstraction 8.3. From DeviceCentric to Clustering-Centric
7.4. Other use cases
Figure 2: Organization of the manuscript
For example, edge devices such as roadside CCTV analytics units or Jetson-based video processors often undergo thermal stress during continuous high-rate inference [23]. A TinyML-based regression model can be embedded locally to predict short-term thermal load (e.g., 30–60 seconds ahead) based on current CPU/GPU utilization, ambient temperature, and inference rate. When the predicted temperature exceeds a critical threshold (e.g., 75°C), the system can autonomously downscale computation by lowering Frames Per Second (FPS), switching to a lighter AI model [24], or reducing hardware clock speeds [25]. Another example can be seen in network bandwidth allocation [19]. Smart traffic poles frequently execute multiple competing tasks-such as license plate recognition, pedestrian detection, and environmental sensingover bandwidth-limited 4G/5G uplinks. A lightweight reinforcement learning (RL) agent deployed within the pole can dynamically prioritize bandwidth allocation based on contextual demand [26, 27]. For example, under high traffic load, more bandwidth is allocated to CCTV analytics. Further, in remote isolated areas, solar-powered IoT nodes must preserve energy while performing the task assigned to it. A TinyLSTM (Tiny Long Short-Term Memory) model3 can predict future battery levels and solar input over the upcoming operational cycle. Based on these predictions, the node autonomously reduces sensor sampling frequency, disables non-critical sensing modules, or defers communication-intensive tasks when low-energy conditions are forecasted [28]. Through predictive analytics, AI can forecast edge and cloud resource failures, ensuring timely maintenance and reducing downtimes. Additionally, AI-driven security protocols can detect and mitigate potential threats, enhancing the robustness of the edge network. Furthermore, AI facilitates dynamic resource allocation, optimizing bandwidth and storage based on real-time data processing needs. In
essence, the integration of AI into edge-cloud infrastructure management promises enhanced efficiency, security, and adaptability. Design of AI for edge device management is also referred to as the 1st form of EI in this article. 2.2. EI-2F: AI at edge infrastructure AI at the edge infrastructure, refers to the deployment of AI algorithms, especially the tiny version of AI, and minimalist models on the resource-constraint edge devices, close to the source of data, rather than in a centralized cloud-based system. This approach offers several advantages, including real-time processing, efficiency, speed, and enable autonomy. Lightweight AI models are deployed at the edge in a wide range of real-world applications. Albanese et al. [29] highlighted the potential of edge AI in automating pest detection within fruit orchards, allowing for real-time monitoring and timely interventions. This enhances the overall efficiency of smart agriculture practices. Edge intelligence and deep learning are also applied in autonomous navigation in vineyards by Aghi et al. [30]. The authors present a novel control mechanism that uses a custom-trained segmentation network and a low-range RGB-D camera to produce smooth trajectories and stable control in various vineyard scenarios. Radanliev et al. in [31] discuss the integration of AI and edge computing in supply chains, emphasizing the mitigation of cyber risks, offering predictive cyber risk analytics with real-time intelligence from IoT networks at the edge. Recently Large Language Models (LLMs) are also being deployed on these edge devices providing intelligence for performing various natural language processing (NLP) tasks such as translation, question answering, text summarization, and generating clinical summaries from multimodal data such as text, sensor data, images and audio [32]. This version of EI, also referred to as the 2nd form of EI, in this article.
3 https://github.com/cpllab/tinylstm
4
EI-2F: AI at Edge Input
Edge Device -1 AI model
Output (Predicted)
AI for ERM Input
Edge Device -2 AI model
Edge Gateway
Output (Predicted)
AI for ERM Input
Edge Device - 3 Custom script
Output
Edge Gateway
AI for ERM Input
Edge Device - 4 non-AI algorithm
Cloud environment
Output
Input
Edge Gateway
...
...
AI for ERM
Edge Device if-else rule
Output
EI-3F: Intelligence Available at Edge of the network
AI for ERM
EI-1F: AI for edge * AI for ERM: AI for Edge Resource Management * EI-1F : 1st Form of EI ; EI-2F : 2nd Form of EI ; EI-3F : 3rd Form of EI
Figure 3: Different forms of Edge Intelligence.
making. However, achieving this goal introduces several challenges, including intelligence synchronization, distributed learning, intelligence discoverability & observability, and maintaining consistency across devices. The deployment scenarios of EI-3F are particularly suited for collaborative learning, where intelligence sharing among edge devices can mutually benefit multiple systems. A key example includes interconnected vehicles in a fleet, where each vehicle shares its intelligence to enhance navigation, safety, and traffic management in real-time [33].
2.3. EI-3F: Intelligence available at edge However, unlike the above-mentioned meanings of “edge intelligence”, it can also be referred to as the “intelligence available at the edge of the network ”, also known as 3rd form of EI. Some of the examples include intelligence related to energy production by Photovoltaic panels deployed atop city buses, intelligence related to trees, wildlife, and animals in the forest, and intelligence connected to soil conditions in the smart agriculture context. In the context of smart traffic, the camera not only captures edge knowledge (e.g. number of vehicles, their speed, traffic conditions, etc.) but also processes that knowledge onthe-spot to produce and share those intelligence (e.g. if a traffic is highly congested, if a city bus is delayed, if a vehicle is going beyond speed limit in a specific area, etc.) among other edge devices. Edge intelligence helps smarter decision making and can be used to enable predictive analytics and smarter services. In this article, our focus is to explore this version of EI, including the novelty, challenges, potential solutions, and baseline technologies or concepts. More on this is discussed on the later sections.
3. Clustered Edge Intelligence: Beyond EC and AI convergence Edge Intelligence (EI) in general refers to the convergence of edge computing and artificial intelligence [34, 35, 36, 37, 38, 39], be it AI for edge or AI on edge. In conventional EI, the AI/ML algorithms are deployed on edge computing node, which complement its counterpart cloud deployment to provide realtime decision-making and predictive analytics service, as also discussed in Section 2.1 and 2.2. For example in a modern traffic control system, the identification and blurring of a person’s face can be done at the edge using light-weight AI algorithms [40]. One of the major advantages of EI is to perform AI/ML tasks and data analysis locally on the device, instead of sending all data to a central server, which can reduce latency, maintain privacy, and reduce network bandwidth consumption. Some of the applications of conventional EI includes:
2.4. Comparison of three forms Table 2 provides a comparative analysis of the three forms of Edge Intelligence (EI-1F, EI-2F, and EI-3F), taking some of the key parameters into account, such as hosting environment, purpose, challenges, development focus, deployment scenarios, complexity, interdependency, scalability, and security impact. Unlike EI-1F and EI-2F, the primary goal of EI-3F (Intelligence available at Edge) is to obtain and share intelligence across devices to enable collaborative decision-
• Smart Home: The devices such as smart speakers, 5
Parameter Hosting environment Purpose Challenges
Development focus Deployment scenario
Complexity
Scalability
Security pact
im-
Table 2: Comparison of different forms of Edge Intelligence EI-1F: AI for ERM EI-2F: AI@Edge EI-3F: Intelligence available at Edge Edge Device, edge-cloud continEdge device Edge Device, edge-cloud continuum, netuum, network devices work devices Edge computing and network reprocessing raw data available at Obtaining & sharing intelligence, collborasource management edge tive decision-making Making AI model tiny, limited com- Limited computational resources, Intelligence synchronization, distributed puting resource, security, privacy & model optimization & compression, learning, intelligence discoverability, inTrust in AI-driven decisions lack of standardization, deployment telligence observability, consistency across complexity, hardware heterogeneity devices - system optimization - edge infras- model design, pruning, quantiza- distributed learning - peer-to-peer inteltructure management - edge secu- tion, and other techniques to make ligence exchange - collaborative learning. rity AI models compact This is required when (a)- dynamic where immediate action is required - suitable for collaborative learning - when edge resource allocation leads to based on data - e.g. autonomous one edge device can be benifited from othperformance gains (b)- dynamically vehicles, health monitoring, or iners - e.g. interconnected vehicles in a fleet securing edge device dustrial automation. Complex because of device and reComplexity lies in making the complexity lies in - communication protosource heterogenity model compact without compro- cols - distributed system - heterogeneous mising accuracy network Not very challenging as the AI not very challenging, as the AI very challenging needs to manage single edge de- models only concern about data vice’s resources generated by that device only The entire edge infrastructure users’ applications users’ applications
thermostats, and security cameras can use EI to analyze smart home sensor data and make decisions locally, such as adjusting temperature and indoor luminosity, notifying wastage of electricity or alerting homeowner the potential security threats, etc [41].
security, lower network bandwidth consumption, reduced service cost, and improved service quality. However, it does not fully address the research question of how edge devices with similar characteristics, capabilities, or configurations can directly share their learning or derived intelligence without relying on a central system. Conventional Edge Intelligence research therefore provides limited support for peer-level collaboration, transfer of learning and knowledge, and collaborative reasoning. To address this gap, we revisit the concept of Edge Intelligence from a different perspective and interpret it as “Intelligence Available at the Edge”, the 3rd form of EI (EI-3F), as briefed in Section 2.3. In EI-3F, intelligence is treated as a first-class entity that can be uniquely identified, discovered, observed, shared, and reused. Edge agents (EAs) deployed on edge devices are responsible for deriving intelligence from locally available data and knowledge. An edge agent is an autonomous software entity operating on an edge device that processes local data and knowledge, derives and publishes intelligence, discovers and consumes external intelligence, and collaborates with other edge agents to support contextaware decision-making. An edge agent consists of three primary components: an AI model or algorithm, metadata, and business logic. The metadata may describe the type of data processed, location, data-generation frequency, communication requirements, and other contextual properties. The business logic defines how the data is used, how the agent communicates with other edge agents, and how contextual awareness is incorporated into decision-making. In conventional approaches, these components are often hard-coded and tightly coupled with the underlying host device. In the proposed approach, the edge agent encapsulates these com-
• Industry 4.0: EI can be used in modern manufacturing industries to analyze robot and machinery sensor data to detect anomalies, predict failure and maintenance needs, optimize production pipeline, etc [42]. • Smart Traffic and Vehicular network: With EI a wide range of traffic management tasks can be performed, including CCTV footage analysis, predict traffic flow, congestion, and accidents [43]. With EI vehicles can exchange information about road conditions and traffic. In autonomous driving, EI enable sensors data analysis in real-time manner, make decisions about steering, braking, and acceleration. This can play an important role in improving the safety and reliability of autonomous driving [44]. • Retail: EI allows the retailer to analyze data from sensors and cameras to implement smart shelves, track customer behavior, optimize checkout process, detect fraud, and improve inventory management. Under the current approach, edge devices remain highly dependent on a central system, with limited or negligible opportunities to share their learning directly with peer edge devices. For instance, in Federated Learning (FL), multiple edge devices contribute to improving and updating a central machine-learning model while keeping their training data locally [45, 46]. This approach helps achieve several important objectives, including reduced latency, improved reliability and scalability, enhanced privacy and 6
ponents and decouples them from the host device, thereby improving portability, reusability, and manageability. Edge agents are also responsible for producing intelligence and managing its lifecycle, including publication, update, monitoring, and retirement. Multiple edge agents may be colocated on the same edge device; however, a single edgeagent instance is assumed to operate on one edge device at a given time [24]. This raises an important question: if intelligence is treated as a first-class entity, why is there still a need to cluster multiple intelligences?
external intelligence may come from other edge devices. For instance, the EA deployed on Edge Device - 2 may provide environmental intelligence such as recent change in CO2 levels, temperature, and humidity. The EA on Edge Device - 3 may provide intelligence if there is any recent significant change in Air Quality. By comparing the footage-based fire detection with real-time external environmental intelligence, the agent on Edge Device -1 can more accurately and reliably confirm whether a fire is actually occurring. Only after this validation process is complete the system raise an alert. This approach significantly reduces the likelihood of false alarms and ensures a more reliable response to potential fire hazards. Through the CEI (Clustered Edge Intelligence) approach, we can establish a cluster of related intelligence produced by various agents. As illustrated in Figure 4(c), this cluster integrates intelligence such as recent changes in CO2 levels, temperature, and humidity from Edge Device-2, along with significant recent changes in air quality from Edge Device-3. It’s important to note that these edge devices do not share raw data; instead, the deployed EAs share derived intelligence based on the data available to each respective device. This shared intelligence is collectively analyzed to assess the likelihood of a fire, enhancing the system’s ability to detect and respond to potential threats effectively. At this point, one might wonder how intelligence clustering is different or similar to the clustering of underlying edge devices. This also raises another question: why to form clusters based on the relativity of intelligence rather than the devices that provide it? We will address both questions in detail in the following subsection.
3.1. Why “Clustered”? Traditionally, intelligence is derived based on specific information, mostly in the cloud computing environment. For example, whether a road is congested or not is typically determined in a cloud environment after receiving location information from multiple vehicles or analyzing video footage. These are classic examples where intelligence is derived from the same type of data. However, we believe that multiple types of information can be used to derive more reliable intelligence. For example, location data from vehicles, CCTV footage, and CO2 levels in a specific road segment can be combined to more accurately and reliably determine traffic congestion. Extending this to a more practical scenario, we propose that multiple types of intelligence, produced by edge agents deployed on edge devices, can be semantically clustered together to generate more complex intelligence. The following subsection provides a more concrete example to simplify the concept behind Clustered Edge Intelligence (CEI). 3.2. Clustering Edge Intelligence : An Example Let’s explore the benefits of Clustered Edge Intelligence (CEI) with an example of a remote fire detection and alert system, as shown in Figure 4. In the conventional approach, as in Figure 4(a), the setup relies solely on one type of edge sensor, combined with onboard data analytics, say CCTV camera sensor combined with on board footage analytics. Onboard footage analytics (which could be a tiny AI model or specialized computer-vision algorithm) continuously captures the video footage from CCTV camera to detect visual indicators of fire within the view-point of the camera. If a fire is detected, the system immediately triggers an alert, notifying the relevant authorities or systems to respond. This approach is straightforward but has limitations, primarily because it depends entirely on the visual detection of fire. As a result, it may be prone to false alarms or could miss fires that do not have clear visual signs, such as those covered by obstacles or occurring in areas not covered by the camera. However, with the CEI approach (Figure 4(b)), the fire detection mechanism can be further enhanced by integrating additional layers of intelligence and data validation before raising an alert. Instead of raising an alert immediately upon detecting fire, the edge device can cross-check and validate the detection with external intelligence. This
3.3. Device- vs Intelligence-centric clustering mechanism In continuation to explore the concept of clustered edge intelligence and our previous work on clustered and cohesive edge intelligence [47], in this subsection, we compare two different clustering approaches: Device-centric clustering and Intelligence-centric clustering approaches. Figure 5 presents a comparative analysis of two different clustering approaches. 3.3.1. Device-centric clustering (DCC) Approach: In device-centric clustering approach (Figure 5(a)), clusters of physical edge devices are formed based on their physical characteristics, proximity, or functional roles, rather than the specific intelligence they generate. Each cluster contains a group of devices that are managed by an edge controller, which communicates with a fog node connected to cloud computing resources. The edge controller maintains a device inventory, mapping each device (e.g., D1, D2) to the specific type of intelligence (e.g., i1, i2) it generates. For instance, devices D1 and D2 generate intelligences i1 and i2, respectively. This approach organizes the devices into clusters without considering the specific type of intelligence they produce, instead of primarily focusing on the management of the devices themselves. 7
Traditional Approach CCTV
New Approach CCTV
Fire Dection Cluster
On-board Footage Analytics
Recent increase in CO2 value
On-board Footage Analytics Edge Device-2 Detected Fire
Raise Alert Edge Device
(a)
Cross-check & validate with external intelligence
External Intelligence
Detected Fire
Recent increase in CO2 value Recent increase in Temp value Recent significant change in Humidity
On-board Footage Analytics Recent increase in Temp value
Recent significant change in Humidity
Recent change in Air Quality (e.g. PM2.5)
Recent change in Air Quality (e.g. PM2.5)
Raise Alert Edge Device-1
Edge Device-3
(b)
(c)
Figure 4: Comparison between traditional and CEI-enabled fire-detection mechanisms (an example): (a) single-source detection, (b) intelligence-assisted validation, and (c) logical clustering of related intelligence for reliable fire detection.
Each device within the cluster is usually identified by a unique identifier. The intelligence could either be stored directly in the inventory or queried from the device in real-time when needed. Accessing the intelligence in a device-centric clustering approach requires the user to have knowledge of specific device identifiers. This means that before any intelligence (such as “What is the percentage of changes in CO2 value in last 10mins in location X?”) can be accessed, the user must first ascertain whether the desired device is present within the network. The process begins with the user sending a query to the edge gateway, which serves as the management hub for the device cluster. The query typically asks whether a particular device is available in the inventory. If the device is present, the edge gateway retrieves the surrounding intelligence from that specific device. Further, in a traditional approach, the intelligence is computed/derived on-demand either on the source edge device or computed on the edge device where the request is generated. An important aspect of this approach is the tight coupling between devices and the intelligence they generate. This tight coupling means that the device-centric approach emphasizes the management and availability of physical devices. To access a particular piece of intelligence, it is crucial first to ensure that the device capable of providing that intelligence is available in the network. This process is repeated for all devices connected to the edge gateway, making the device the primary focus of the clustering and intelligence retrieval processes. It reflects a more hardware-oriented view of edge computing, where the infrastructure and devices themselves are central, and the intelligence they provide is secondary. This approach can be advantageous in scenarios where specific devices are known to generate unique and critical intelligence. The primary challenge with this approach is
the risk of vendor lock-in. Different vendors may use distinct data models to provide data or intelligence, leading to compatibility issues. Furthermore, if multiple devices offer the same type of intelligence, it becomes difficult to seamlessly substitute one device with another from a different vendor in case of failure or unavailability. This lack of interoperability can restrict the effective use of alternative devices, limiting flexibility and resilience in the system. 3.3.2. Intelligence-centric clustering (ICC) Approach: On the contrary, as proposed in this research work, in the intelligence-centric clustering approach (Figure 5(b)), the fundamental difference is that the focus is not primarily on device management but on the management of intelligence available. There is a shift from focusing on the physical devices to prioritizing the intelligence they generate. In this approach, the edge controller maintains an intelligence inventory, mapping each type of intelligence (e.g., i1, i2) to the corresponding devices that provide it. For example, intelligence i1 might be derived from device D1, while intelligence i2 could be generated by one or more number of devices from multiple devices such as D2, D3, and D4. In this approach, clusters (Cluster-1 and Cluster2) are formed by grouping together intelligences rather than devices. It is to be noted that, the inventory may include intelligence for which there are currently no connected devices, highlighting the system’s flexibility and its decoupling of intelligence from specific hardware. When a user seeks to obtain certain environmental intelligence (e.g., last 30min average humidity on location A, Air quality (P M 2.5 value) improved by X% in last 24hours, etc.), they send a query to the edge controller. The controller response process is streamlined: it checks its intelligence inventory to determine if the requested intelligence is within its scope. If the intelligence is available, 8
Cloud Computing
Cloud Computing
Devices : D1, D2, D3, ..... Intelligence : i1, i2, i3, .....
Fog Node
Device Inventory D1 -> i1 D2 -> i2 ...
Edge Controller
Device Inventory D6 -> i1 D7 -> i2 ...
Edge Controller
D1
D2
D6
D7
i1
i2
i1
i2
D3 i3
Fog Node
D5 i2
D4
D8
D9
i3
i3
Intelligence Inventory
Edge Controller
i1
i1 -> D1 i2 -> D2, D3, ... ...
i2
i3
Intelligence Inventory
Edge Controller
i1
Cluster - 1
i1 -> D6 i2 -> D7 ...
i2
i3
Cluster - 2
i2
Cluster - 1
D1
Cluster - 2
(a) - Device-centric clustering approach
D2
D3
D4
D5
D6
D7
D8
D9
(b) - Intelligence-centric clustering approach
Figure 5: A pictorial representation of (a) Device-centric clustering approach, and (b) intelligence-centric clustering approach.
approaches, taking several aspects into account, as summarized below.
the controller queries the appropriate device and forwards the resulting data back to other edge devices or even the end-user. It is not required to have any knowledge of the specific devices providing the intelligence. The management of devices is kept abstract, meaning, details such as the device’s age, vendor, operating frequency, or other hardware-related characteristics are abstracted out from other edge devices or the end-user. This abstraction simplifies the user experience, allowing them to focus solely on the intelligence they need rather than the complications of the devices that produce it. Another important aspect of this approach is the flexibility it offers. Users querying the same edge controller for the same intelligence at different times might receive data from different devices, depending on which devices are available and capable of providing the intelligence at that moment. This ability to source the same intelligence from multiple devices reduces dependency on any single device and enhances the system’s resilience. If a user receives no response to their intelligence query, it could indicate that the relevant device is unavailable, or it might be due to security restrictions that prevent access to the intelligence. The intelligence-centric approach is characterized by a loose coupling between devices and the intelligence they generate. This decoupling allows for greater flexibility and scalability, as intelligence can be accessed from any suitable device without the user needing to manage or even be aware of the underlying hardware.
Table 3: Comparison of Device-centric clustering (DCC) & intelligence-centric clustering (ICC) approach Parameters DCC ICC Device Hardware very tightly coupled very loosely couand Intelligence pled Coupling Vendor lock-in hard to avoid can be easily issue avoided Management Focus on management Focus on manageof devices. Intelligence ment of intelliwil not be affected. gence. Devices wil not be affected. Clustering strat- implemented at device implemented at egy implementa- level Edge gateway tion level Discoverability focus is on edge devices Focus is on edge & observability intelligence Security & priDoes not help in en- Helps in enhancvacy hancing security and ing security and privacy privacy
• Device Hardware and Intelligence Coupling: DCC shows a “very tightly coupled” relationship between device hardware and the intelligence derived, meaning that the intelligence is often inseparable from the device that produces it [48]. On the other hand, in ICC, intelligence is abstracted from the hardware, making it very loosely coupled. • Vendor Lock-in Issue: In DCC, the vendor lock-in issue is “hard to avoid” due to the tight coupling between devices and the intelligence they generate. Each vendor might use proprietary data models and standards, making it difficult to integrate devices
3.3.3. Comparison of device-centric and intelligence-centric clustering Table 3 provides a comparison of Device-Centric Clustering (DCC) and Intelligence-Centric Clustering (ICC) 9
from different vendors [49, 50]. In contrast, in ICC this issue can be easily mitigated because intelligence is loosely coupled with devices, which makes it interoperable across different vendor devices.
sensors are either embedded on-board or wirelessly connected via technologies like Long Range Wide Area Network (LoRaWAN), Wi-Fi, or Bluetooth. The edge device receives the data, processes, interprets, and contextualizes each data point (e.g., identifying whether the value “40” refers to temperature, humidity, or another metric). It is possible that an edge device can be connected to multiple sensors. In such a case, the edge device will be equipped with different varieties of knowledge sets. Further, as discussed before, limited computing resources are available on each edge device. With such limited computing resource, edge devices are capable of performing several tasks, including processing specific data types, acting on data, executing cloud-based instructions, communicating with other devices, and discovering peers. Conversely, edge devices have “Requirements,” such as accessing core intelligence from other devices, retrieving intelligence from shared Intelligence Inventories (discussed later), or acquiring specialized knowledge from peers. Additionally, edge devices are aware of their Communication Capabilities, including Wi-Fi, Bluetooth, or LoRaWAN, enabling connections with multiple peers across different channels. Each edge device is designed to share its intelligence with others while also utilizing intelligence from other devices. When an edge device receives information from its peers, it is called external intelligence. For example, an edge device connected to a CCTV camera can capture footage of its surroundings and detect signs of a fire. At the same time, it may receive external intelligence from nearby devices, such as a sudden increase in temperature. By cross-verifying this information, the device can confirm the presence of a fire with greater confidence. This verified conclusion can also be referred to as core intelligence.
• Management of Device and Intelligence: The management of devices in DCC directly affects the associated intelligence, as the two are tightly linked. Conversely, in ICC, while intelligence management is crucial, it does not directly impact the management of devices. This makes the system management more independent and flexible. • Clustering Strategy Implementation: In DCC, clustering is implemented at the device level, requiring updates and coordination across all devices involved in the cluster. In ICC, the clustering strategy is independent of the device and implemented at the Edge gateway or controller level. • Discoverability and observability: In DCC, during the cluster formation process, the discoverability and observability mechanisms primarily focus on the devices themselves, regardless of the specific intelligence they are capable of. In contrast, ICC shifts the emphasis in discoverability and observability to the intelligence, rather than the underlying devices. This approach in ICC prioritizes the accessibility and management of intelligence, treating the devices as secondary to the information they produce. • Security and privacy: In DCC, security and privacy rely on the imposed security mechanisms, which may involve sending sensitive data to the cloud [51]. On the other hand, in ICC, the security and privacy can be enhanced in a way that the intelligences are derived on-board and only shared with others. The system is not promoting the sharing of data with other edge devices.
Cloud Environment Intelligence Inventory
Discovery Mechanism Regulator
4. CEI Architecture
Sharing Mechanism Marketplace
Edge Gateway or Controller
Consume
Consume
Core intelligence
Core intelligence
Core intelligence
Knowledge set
Knowledge set
Knowledge set
Data
External intelligence
Data Edge Device
4.1. Edge device layer Figure 7 showcases the internal components of an individual edge device within the CEI architecture. Each edge device is connected with multiple sensors, collecting environmental data such as temperature, humidity, noise levels, and luminosity at varying intervals. These
Discovery Mechanism
Publish
Publish
Consume
Intelligence Inventory
Publish
Figure 6 illustrates the three-tiered architecture of Clustered Edge Intelligence (CEI), consisting of edge devices at the foundational layer, an edge gateway or controller in the intermediary layer, and the cloud at the top. The architecture’s core objective is the effective management of intelligence within the edge infrastructure. Edge devices exhibit two forms of intelligence: Core intelligence, derived from their own data, and External Intelligence, which integrates insights from other devices.
Data
Data
External intelligence
Data Edge Device
Data
Data
External intelligence
Data
Data
Edge Device
Figure 6: Overall Architecture of CEI
4.2. Edge controller layer As mentioned earlier, each edge device is designed to share and collaborate with others. However, the key ques10
tion is: How can they effectively communicate, collaborate, and share their knowledge and intelligence? The answer lies in Figure 5, which illustrates the overall architecture of Clustered Edge Intelligence (CEI). To enable communication between edge devices, the cloud, and their peers, an edge gateway (or controller) is introduced as an intermediary layer. The edge controller is equipped with two key components: the Intelligence Inventory and the Discovery Mechanism. The Intelligence Inventory is responsible for maintaining a record of all available intelligence across connected edge devices. This intelligence is typically shared by multiple devices. For example, an intelligence record may indicate that traffic is 50% congested or that a road is blocked due to an accident. Each record includes essential metadata shared by the edge devices, such as:
Data
Data
Data
sensor-1
sensor-2
sensor-3
Communication capabilities
Knowledge set
Requirements & Capabilities list
Core intelligence
External intelligence
Edge Device
Figure 7: A detailed view of edge device in CEI.
• The intelligence (e.g. “amount of electricity produced by solar panels in all city buses”, “level of CO2 at any given point of time”, “level of noise in a specific city area”, etc.) • Category (e.g., temperature, humidity, CO2 , PM2.5, soil condition, traffic condition, road condition, etc.) • Access URL (to retrieve intelligence) • Input and output format (which defines how another edge device should structure its message and what format to expect in response) • Confidence score (indicating the reliability of intelligence) • Application protocol (e.g., MQTT, CoAP, REST API) • Communication channel (e.g., LoRaWAN, Bluetooth Low Energy (BLE), Wi-Fi) The Discovery Mechanism is designed to allow edge devices to access external intelligence shared by others. Unlike traditional IoT/edge infrastructure, which focuses on discovering physical devices, this mechanism is dedicated to discovering intelligence. If an edge device requires intelligence that is not available locally, the discovery mechanism searches for relevant intelligence across nearby edge devices or intelligence clusters. This enables peer-to-peer intelligence sharing without relying on cloud intervention. It is important to note that the Core Intelligence component within the edge device is responsible for publishing intelligence along with its essential metadata. Meanwhile, the External Intelligence component is responsible for consuming shared intelligence by communicating with the edge controller.
or even millions of edge controllers. The cloud environment at the top, is equipped with a global view and is responsible for managing and coordinating all edge controllers. Similar to an edge controller, the cloud layer also consists of an Intelligence Inventory and a Discovery Mechanism components. The discovery mechanism enables an edge controller to identify and access intelligence from other edge controllers’ intelligence inventories. Additionally, the cloud layer serves as a global, centralized intelligence repository. Furthermore, with the Intelligence Marketplace component, the cloud functions as a marketplace where organizationswhether they own AI models, edge devices, or data-can buy, sell, and exchange AI models, decision rules, and other intelligence. The Regulator component within the cloud enforces access control policies, security restrictions, and compliance rules to prevent unauthorized intelligence sharing. The Sharing Mechanism provides necessary instructions and policies to edge devices, defining what intelligence should be shared, with which peers, and how frequently. 4.4. CEI Characteristics The effectiveness of CEI depends on a set of core characteristics that make the edge agents independent, adaptable/dynamic, capable of exchanging contextual information, and collaborate with other agents. Some of the primary characteristics are summarized below: • Dynamicity of the environment: In such a large-scale distributed edge systems, devices and agents may frequently join or leave the network due to mobility (e.g., drones, autonomous vehicles), power constraints, network availability, or hardware failures. This dynamic nature demands that edge agents be context-aware and capable of adjusting their behavior in real time - for instance, re-routing communication paths when a node goes offline, or switching from visual to audio input when visibility is poor.
4.3. Cloud Environment layer An edge controller has limitations in communicating with and controlling a large number of edge devices. In a large-scale edge infrastructure, there could be thousands
• Autonomy: The fundamental and inherent characteristic of edge agents in such a distributed system is 11
the autonomy, which refers to the ability of each edge agent to operate independently, understand the data and derive all possible intelligence related to that data/information, and execute tasks without requiring constant supervision from a cloud-based central entity. The autonomy also in the context of discovering other edge agents with different capabilities. For example, in a smart agriculture scenario, an edge agent responsible for crop identification may need to discover and coordinate with another edge agentpossibly on a different drone-tasked with pesticide spraying, without relying on central oversight.
communication mechanisms, and others, that can provide the foundation for implementing CEI. 5.1. Abstraction In layman terms, abstraction is simply the process of hiding some complex information and allowing the system to operate with minimal necessary input from the human. For example, high level programming languages hide the complexity of interacting with the hardware and each and every assembly level executions. The primary advantage of abstraction is that it helps manage complexity. With multiple layers of abstraction, developers and system administrator can design and implement complex systems by breaking them down into manageable, interconnected, independent components/modules. Some of the benefits lies in portability, flexibility and simplicity. However, excessive or poorly designed abstraction can leverage the complexity and inefficiency of the whole system. Hence, it is essential to find the right balance that provides clarity and utility without compromising the efficiency [53]. The concept of abstraction can be applied in various contexts, primarily in hardware, software, network, data, and functionality, as summarized below:
• Context Awareness & Sharing: Context awareness allows each agent to interpret its environment, understand its current state, and recognize the relevance of that information to a shared task. For effective collaborative learning, this becomes important when multiple edge agents are capable of performing the same or similar tasks, but in different environmental or situational contexts. For example, in a Smart Agriculture setting, edge devices such as soil moisture sensors, drones, and weather stations operate across different fields and zones. A drone monitoring crop health in one area might detect signs of pest infestation. By sharing this context with nearby drones or ground sensors, those agents can proactively adjust their scanning or sampling behaviors, even if they operate in slightly different microclimates or crop types. This shared situational awareness enhances the overall system’s intelligence and its reliability.
• Hardware abstraction: Hardware abstraction is the process of hiding hardware-specific details, offering a generalized interface for the interaction of software [54]. It ensures software can function across diverse heterogeneous hardware components without the need of specification of each component provided by the corresponding vendor. Operating System (OS) and device drivers are some of the de-facto examples of Hardware Abstraction Layer (HAL). An OS enables software to smoothly interact uniformly with the underlined hardware. On the other hand a device driver bridges a specific hardware component from a specific vendor to the broader software system, ensuring standardized interactions.
• Inter-Agent Communication Protocols: Because edge environments are resource-constrained and often operate under dynamic network conditions, communication must be lightweight, resilient, and low-latency [52]. Common communication paradigms include publish-subscribe models (e.g., MQTT), remote procedure calls (e.g., gRPC), and peer-to-peer messaging using custom APIs or agent frameworks. Edge agents should be capable enough to take advantage of available communication channel to announce their presence, share capabilities, request assistance, or synchronize on tasks. For example, a fall detection AI on a wearable device might send an alert to a nearby surveillance camera to verify the event visually before escalating. 5. Baseline Technologies The implementation of the CEI architecture and its capabilities requires a set of enabling technologies that support the representation, communication, discovery, management, and clustering of intelligence across distributed edge infrastructures. In this section, we discuss the key baseline technologies, including abstraction, knowledge graphs, 12
• Software abstraction: The Modern software system is becoming complex day-by-day. To make is simplified and reduce its complexity for the end user, especially the developer, introducing multiple layers of abstraction, also referred to as software abstraction (SA), is the key. SA allows developers to interact with the system on a higher level, without understanding the underlying architecture [55, 56]. For instance, developers can use high-level generalpurpose programming languages, including Python4 , Java5 , C++6 , etc., to write code for a variety of applications, without any knowledge about the underlying architecture. Some other examples of SA can be an API, libraries like jQuery, graphics engine like OpenGL, etc [55]. 4 https://www.python.org/ 5 https://www.java.com/en/ 6 https://isocpp.org/
• Network abstraction: In a similar fashion, abstraction can also be applied to the network resource management. The network abstraction (NA) can be implemented and achieved by introducing higher-level interface for network-related operations like routing, switching, multicasting, tunneling, etc [57]. The application of NA can be found in a wide variety of domains„ including network protocols, wireless network, cloud network, virtual network, etc [58]. • Data abstraction: In the era of data as next generation gold, its difficult to understand and get insight into the huge amount of unstructured data that comes from large-scale IoT infrastructure and other computing systems. When abstraction is applied to the data, we no longer deal with the raw data and the data container, e.g. stack, queue, linked ilst, etc., rather the knowledge or the information [59, 11, 60]. This is where Machine Learning (ML) and AI related advanced tools come into the picture. DA is especially implemented in the modern database management system. High-level programming languages can be used to implement DA layer for any modern dataintensive application [60]. For example, the concept of “abstract”, “interface” can be used to create custom data type for real-world physical objects. • Functional abstraction: Functional abstract is introduced to hide complexity of an algorithm which is designed to perform a specific task so that the use can simply focus on high-level business logic. A user should only pay attention to the calling convention and assume that the function will return the desired result [61]. As an example, in the sqrt function in Java or Python or any mathematical related library, implementation details are abstracted from the exact steps involved in finding the square root of any integer value.
Figure 8: Different layers of abstraction in an edge device.
provides low-level control over the device’s hardware. Examples include ARM IoT firmware and CISCO IoT firmware7 . Essentially, firmware bridges the gap between hardware and other software tools, ensuring hardware components can execute more complex tasks directed by high-level software layers. • Embedded OS Layer : This layer incorporates the OS specifically tailored for embedded systems. Such operating systems are optimized for performance and resource usage in devices with constraints on memory, processing power, or battery life. Examples include FreeRTOS8 , Zephyr9 , Mbed10 , and Raspberry Pi OS11 . These operating systems manage hardware resources and provide essential services for the software applications above.
5.1.1. Layers of abstraction in an edge device Figure 8 shows different layers of abstraction in a typical edge device capable of preprocessing the input data based on either AI or non-AI-based business logic. Each layer of abstraction hide certain level of complexity making it easier to collect and interact surrounding intelligence. A brief overview of each layer is presented below: • Device Hardware Layer : This is the foundational layer of any edge device, including fitness trackers, medical sensors, and smart fire alarms, consists of the physical components and hardware modules [62]. For example, it includes the processor, memory, storage, attached sensors (e.g. MPU-6050 gyroscope and accelerometer sensor, HC-SR04 ultrasonic distance sensor), and interfaces such as USB and Wi-Fi.
• Device Connectivity Layer : Serving as the communication layer, it focuses on protocols and technologies that facilitate the device’s ability to connect and 7 https://developer.cisco.com/docs/iotod/ 8 https://www.freertos.org/ 9 https://www.zephyrproject.org/
• Firmware Layer : Positioned just above the hardware layer, the firmware layer consists of software that
10 https://os.mbed.com/ 11 https://www.raspberrypi.com/software/operating-systems/
13
communicate with other devices or networks. Protocols such as MQTT12 , CoAP13 , and BLE ensure seamless data transfer and device interoperability. The MQTT protocol, for example, allows IoT devices and the cloud to communicate with one another.
Further, while implementing the edge inventory component, which tracks what intelligences are available (e.g., “Average CO2 in Zone B”), which edge device provides it and the corresponding metadata, one candidate solution to implement is the use of a simple key-value pair. However, this approach lacks semantic understanding (e.g., what is i1 about in Figure 5?). The discovery mechanism would also be entirely based on the intelligence ID and not context-based (i.e., “What’s relevant to air pollution?” is not possible to derive). With such a simple key-value pairbased approach, there is no way to relate different intelligences (e.g., “N oiseLevelHigh” − > “CO2 LikelyHigh”). Google recently introduced Open Knowledge Format (OKF) 16 , which is an open and vendor-neutral specification introduced for representing and exchanging knowledge in a form readable by both humans and AI agents. This recent development may complement and can be used as the enabling technology for intelligence sharing among edge agents. In the context of CEI, an OKF-inspired representation could be used to describe intelligence together with its metadata, such as type, description, provider, timestamp, location, confidence, access endpoint, and others. However, OKF provides a general knowledge representation format rather than a complete intelligencemanagement solution; therefore, CEI-specific extensions would still be required to capture specific properties such as freshness, clustering relationships, provenance, spatial and temporal validity, lifecycle state, etc.
• Device Management Layer : This layer emphasizes the administrative aspect of edge devices. It encompasses functionalities like remote registration, configuration, provisioning, and monitoring of the device. It ensures that devices function efficiently, securely, and can be updated or diagnosed remotely. • API Layer : Standing for Application Programming Interface, this layer offers a set of routines, protocols, and tools that allow for the building and interaction of software applications. Interfaces such as gRPC14 , Apache Thrift15 , and REST ensure that different software components can communicate and work together, providing a consistent and unified approach to accessing device functionalities. • Application Layer : This layer includes end-user applications that run on the edge device. The applications are designed for specific tasks like self-driving functionalities, smart grid management, smart agriculture practices, and more. They utilize services and functions from all the underlying layers to deliver specific user-centric functionalities. • Intelligence Layer : This top layer is envisioned as a core capability of the proposed CEI architecture. It maintains and manages the intelligence available on an edge device, regardless of whether that intelligence is generated through an AI model, a nonAI algorithm, a rule-based mechanism, or custom business logic. By abstracting intelligence from the underlying complex implementation and operations, this layer enables intelligence to be identified, published, discovered, shared, reused, and consumed by other edge agents.
5.2.1. Knowledge-data mapping At the edge, raw sensor data is meaningless unless it is contextualized, as discussed before. To participate in intelligence clustering, sharing, or reasoning, the device and the agent deployed on top of it must know what its data means - and map that data into a knowledge graph node (a semantic description). Without this, devices would just generate and share numbers with each other (e.g., 45, 97), without understanding the meaning of “Temperature” or “PM2.5 pollution”. Thus at the device level, mapping of data to the corresponding knowledge graph would help in sharing a meaningful edge intelligence within the CEI system.
5.2. Knowledge Graph At its core, CEI aims to abstract the intelligence from devices, which will enable discoverability, contextual linking, and collaborative reasoning across edge nodes. However, representing intelligences (e.g., CO2 level trend, temperature anomaly) as nodes, connecting them via semantic relationships (e.g., affects, correlatesWith, dependsOn) and capturing temporal, spatial, and contextual metadata need the use of knowledge graph as one of the baseline technology [63].
5.3. Communication For a smooth collaboration among edge agents in a clustered edge intelligence (CEI) ecosystem, communication is the backbone. The communication protocols used in these systems must be lightweight, resilient, and lowlatency to meet the demands of dynamic, resource-constrained edge environments. Several communication methods that can be employed in a CEI system to facilitate peer-topeer intelligence sharing and overall system coordination includes:
12 https://mqtt.org/ 13 https://coap.space/ 14 https://grpc.io/
16 https://cloud.google.com/blog/products/data-analytics/ how-the-open-knowledge-format-can-improve-data-sharing
15 https://thrift.apache.org/
14
5.3.1. REST REST (Representational State Transfer) is a widelyused architectural style that enables communication between systems over HTTP [64]. In the context of CEI, REST can be used for device-to-cloud communication. While RESTful APIs are not optimized for real-time communication, they are effective for scenarios where simplicity and scalability are crucial, such as periodic updates from edge devices to the cloud or between devices sharing non-timecritical data [65, 66]. 5.3.2. Pub-Sub The Publish-Subscribe (Pub-Sub) model is particularly useful in environments where devices need to communicate asynchronously. In CEI, edge agents can act as either publishers or subscribers. A publisher sends messages (data, intelligence) to a channel, while subscribers receive updates from the channel without knowing the identity of the publisher [67]. Pub-Sub systems, such as MQTT or CoAP, are ideal for CEI systems, where edge agents need to broadcast information (such as sensor readings or alerts) to multiple other devices or services without direct, oneto-one communication on a real-time basis. 5.3.3. 5G/6G and LoRAWAN Advanced communication technologies like 5G and 6G are becoming critical for supporting high-speed, low-latency, and wide-coverage communications in edge computing networks. These technologies are capable of handling the vast amounts of data generated by edge devices, ensuring that the exchange of intelligence by edge agents is fast and reliable [68, 69]. For example, in a smart city application, edge agents distributed across a large geographic area may use 5G or 6G networks to send and receive data about traffic, pollution levels, or emergency responses. Further, LoRaWAN (Long Range Wide Area Network) ensures that intelligence can be shared among devices even in isolated areas, thanks to its design for low-power and long-range communication. 5.4. Discoverability Discoverability is one of fundamental enablers in CEI. In contrast to traditional edge and IoT systems (i.e., where discovery mechanisms primarily focus on devices, services, or network endpoints [70, 71]), CEI requires mechanisms that treat intelligence itself as a first-class, discoverable, and observable (more on observability in next subsection) entity. This shift is essential because CEI operates in highly dynamic, large-scale environments where intelligence is continuously generated, updated, expires, or migrates across heterogeneous edge devices. Discoverability in CEI refers to the ability of edge agents and controllers to identify the intelligence that exists, where it is available, and under what context it is relevant. Unlike conventional service discovery, which relies on static identifiers such as service names or IP addresses, intelligence
Load Balancer
Accessing Service
Service Client A
API Registry client
Service Consumer/Client
Registry client
Service Request
10.1.1.12:4002 Registry client Service D API Registration
Service Registry
10.1.1.11:4001 Service C
API
Query service location & instance
User
10.1.1.10:4000 Service B
10.1.1.13:4003 Registry client Service E API
Figure 9: A general architecture of service discovery
discovery must be semantic- and context-aware. These approaches have been shown to scale well for cloud-native services [72, 73], but they lack native support for semantic reasoning, contextual filtering, and confidence-aware selection, which are required in intelligence-centric systems. To address these limitations, CEI aligns with research in semantic IoT and context-aware discovery, where resources are described using metadata and semantic annotations rather than raw identifiers. Sheth et al. demonstrated that semantically enriched discovery significantly improves interoperability and automated reasoning in sensor-rich environments [4]. Such works support the need for intelligence descriptors that include category, spatial scope, temporal validity, and dependencies on other intelligences. Large-scale distributed systems rely on dedicated discovery mechanisms operating across different architectural layers, including service, device, and network discovery. Such mechanisms can be implemented using either a centralized or a distributed (client–server) model, depending on scalability, latency, and fault-tolerance requirements. The following subsection describes these discovery mechanisms and discusses their relevance in the context of intelligencecentric edge systems. 5.4.1. Service Discovery Service Discovery (SD), commonly used in distributed systems, is the process of automatically discovering and locating a (or a set of) service(s) based on the request made by the user or another service [74], as shown in Figure 9. The request may comprise functional and non-functional requirements, as well as a service description. A Service Registry (SR) is typically implemented as a database of all services, including their network locations and service instances. It is the service provider’s responsibility to register their services with the service registry. The service registry can be implemented on the client or server side. Irrespective of the implementation, SR should be highly available and provide real-time information on the health and availability of all services. Several open-source and commercial service discovery solutions are available in the market. Some of the popular SD solutions are described below: • Docker Hub17 , a server-side discovery service, is an 17 https://hub.docker.com/
15
essential component of the Docker ecosystem, providing a central location for managing and sharing Docker images. Users can create and manage their own repositories for Docker images, as well as search for and download public images created by other users. When a user searches for or pushes an updated Docker image through the Docker Command Line Interface (CLI) or REST API, the request is sent to Docker Hub’s servers. Upon receiving the request, Docker Hub uses its server-side search to find the requested image or stores the newly created image in the appropriate repository. Additional features help streamline Docker container development, including automated builds, webhooks, and access control.
it in its registry in key-value pairs. Such data can be updated in real time without service interruption caused by a restart event. Clients can then use Consul to discover the service location using the service name. Consul can also be used for distributed locking, enabling coordination between services in a distributed system. • Etcd 21 , an open-source tool, is a distributed keyvalue store popular for its reliability, consistency, and scalability features. This is used for implementing a configuration registry and service discovery for distributed systems, such as Kubernetes. Similar to ZooKeeper and Consul, Etcd is a server-side discovery service that provides a centralized mechanism for distributed systems to coordinate and synchronize.
• Apache ZooKeeper18 , a server-side discovery service and an open-source Apache project, offers a centralized and reliable way for distributed systems to coordinate and synchronize with each other. It provides a hierarchical namespace, similar to a file system, where clients can store and access data. Clients can use ZooKeeper to discover service locations, monitor configuration changes, and synchronize data across multiple nodes in a distributed system. Coordination services are provided for client applications by a set of ZooKeeper servers. ZooKeeper’s server-side infrastructure coordinates and synchronizes interactions between client applications and the distributed coordination service.
• Traefik22 is an open-source tool designed to provide automatic service discovery and incoming request routing ability to backend services. Request routing is based on a set of predefined rules configured manually. This can be integrated with other service registries, including the above-mentioned Consul, Etcd. In addition, this can be considered a reverse proxy and load balancer. • SkyDNS23 is another popular server-side discovery tool commonly used with other tools in the Kubernetes ecosystem. In the SkyDNS service discovery and management solution, containerized microservice instances are given a unique Domain Name System (DNS) name and a mapped IP address. Users send the service name, and SkyDNS returns the mapped IP address to use for communication.
• Hysterix19 is an old project created to fulfill Netflix’s needs. Hystrix, a client-side discovery service, is a latency- and fault-tolerance library designed to isolate access to remote services, systems, and libraries. By doing this, Hysterix protects against and controls potential latency and failures caused by external dependencies. In addition, Hysterix facilitates monitoring, alerting, and operational control in near real time. Hysterix relies primarily on manual and preconfigured settings and is not suitable for modern and real-time applications.
5.4.2. Device Discovery Device Discovery (DD) is the first step toward enabling a successful Device-to-Device (D2D) communication. It is the process of finding, locating, and eventually establishing communication among nearby devices or devices connected to LAN, WAN, or the Internet [75]. A Device can be a computer, printer, router, switch, smartphone, or any other IoT device with a unique IP address for communication. Similar to the SD, DD can be of two types based on its implementation: (a) centralized and (b) distributed. In a centralized DD, a central server (e.g., a base station or access point) assists in discovering, monitoring, and managing all connected devices. The centralized entity maintains a database of all the devices, including their unique IP, MAC address, hostname, device status, and other relevant information. The discovery client sends the discovery request to this centralized entity and acquires the required information. On the other hand, in distributed DD, each
• HashiCorp Consul 20 is a service discovery, configuration, network automation, and orchestration tool. It provides a centralized registry of services and their metadata, allowing clients to discover services, securely establish communication among services, automate network routing, and control service access. When a new service is added to a network, it registers itself with the Consul agent running on each node. The agent sends the metadata of the newly added service, including IP address, port number, and health status, to the Consul server, which stores 18 https://zookeeper.apache.org/
21 https://www.cncf.io/projects/etcd/
19 https://github.com/netflix/hystrix
22 https://traefik.io/
20 https://www.consul.io/
23 https://github.com/skynetservices/skydns
16
device is responsible for performing discovery by periodically sending its own information, including IP address, MAC address, and other relevant data, to other devices on the same network. Other devices respond to this information based on the requirement and update the local database if there is any update [76]. Some of the recent use cases where DD is critical are health care services, smart home applications, industry 4.0, Industrial IoT, smart transportation, VANETs, and emergency services [77]. A wide range of protocols and tools are used for device discovery, including Simple Network Management Protocol (SNMP), Ping, Internet Control Message Protocol (ICMP), and network scanner tools such as Nmap 24 and SolarWinds 25 .
the intelligence produced and consumed across heterogeneous edge agents in a distributed computing continuum. It must understood and implemented as intelligence-level observability. This must have the capabilities to continuously observe and collect signals related to the state, quality, reliability, and usability of available intelligence. Since intelligence in CEI is treated as a first-class entity, it is not enough to know whether an edge device is active or whether a communication link is available. The observability system must provides insights on whether a particular intelligence is fresh, valid, reliable, semantically consistent, and suitable for the current context. For example, in a fire-detection intelligence cluster, the intelligence related to recent CO2 increase, temperature rise, humidity change, and air-quality variation must be observable in terms of their generation time, update frequency, confidence score, and current availability. It is to be noted that discoverability (as discussed before) and observability are different and complementary to each other. Discoverability answers the question: “Which intelligence is available and where can it be accessed?”. On the other hand, observability answers a different question: “Can this intelligence be trusted and used at this moment?”. An edge agent may discover multiple sources of intelligence that provide similar information, such as traffic congestion level in a specific road segment. However, the selection of the most suitable source depends on observable properties such as freshness, confidence, latency, and provenance. Similar conclusions have been reached in knowledge-centric networking and fog computing research, where observable and semantically described knowledge artifacts outperform static, data-centric approaches in dynamic environments [88, 89, 85]. There are several tools/implementation available in the research and industrial practice that can be used as the baseline. This includes, OpenTelemetry26 , MLflow27 , Dynatrace28 , AgentTrace29 . However designing and implementing observability within CEI architecture is still an open-ended research question and needs further investigation. Some of the key research challenges include deciding between fully centralized and distributed observability mechanisms, design of lightweight, hierarchical, and adaptive observability mechanism, and design of standard observable attributes for intelligence.
5.4.3. Network Discovery Unlink DD, Network Discovery (ND) focuses on knowing the entire network including the devices and all the pairs of existing connections among the devices. It is also essential to know if there is no connection between a pair of devices [78]. It allows the network administrator to identify and map the topology of the network, understand the flows of data through the network, and identify potential bottlenecks or vulnerabilities [79]. The topology of the network can be obtained in several ways, such as (a) by broadcasting discovery messages to neighboring devices with information about their network interfaces, IP addresses, and other relevant information, (b) by using a scanning tool, such as nmap, to scan all connected devices and accessing device information, and (c) by analyzing the network traffic which provides the source and target device information. 5.5. Observability Complementary to discoverability (as discussed in Section 5.4), observability is another important baseline capability for realizing CEI. Observability, in conventional distributed< systems, refers to the ability to understand and obtain the insights of the internal state and behavior of a system by continuously observing or collecting the available signals, such as logs, metrics, and traces [80]. Observability can be seen in a wide range of domains ranging from microservices [81, 82, 83], to containers and from cluster of computing nodes to edge-cloud distributed computing continuum [84, 85, 86]. In such systems, observability is mainly used to observe the systems behavior under different circumstances and loads, monitor service health, execution latency, resource utilization, failures, and communication bottlenecks [87]. However, this conventional observability system is not sufficient for CEI, because the primary entity being managed is not only a IoT sensor, edge device, service, network endpoint, or deployed AI models and business logic, but
5.6. Clustering mechanisms The heart of CEI is the mechanism on how to efficiently cluster the intelligence components that are available at the edge through large number of edge devices. To the best of our knowledge, no specific clustering algorithm is developed to perform the same for inteligence. However, 26 https://opentelemetry.io/ 27 https://mlflow.org/ 28 https://www.dynatrace.com/monitoring/platform/observabilitysolution 29 https://github.com/tensorstax/agenttrace
24 https://nmap.org/ 25 https://www.solarwinds.com/
17
we believe that the existing clustering mechanisms, such as K-means, Affinity Propagation (AP), Mean-shift, Spectral clustering, Ward hierarchical clustering, Agglomerative clustering, DBSCAN [90], OPTICS, Gaussian mixtures, BIRCH [91], etc. can be used as the baseline for design and development of clustering mechanisms only for intelligence. Further, it is believed that not a single clustering approach can address the problem where, the clustering can be done based on different context, type and other parameters. Centroid-based methods, such as minibatch K-means and distributed K-means variants (e.g., kmeans), are commonly employed when the approximate number of clusters is known. However, for scenarios where intelligences may exhibit irregular relationships, noise, or varying cluster densities, hierarchical density-based methods, such as scalable or approximate variants of HDBSCAN can be adopted over k-nearest-neighbor graphs, as they can discover clusters of arbitrary shape while filtering out weak or noisy intelligences. In highly dynamic environments where intelligences continuously appear, evolve, or expire, streaming clustering frameworks based on microcluster and macro-cluster abstractions (e.g., CluStream-, DenStream-, or ClusTree-style approaches) can be adopted. Since the research is in its very early stage, how the existing clustering mechanisms can be used as baseline is still needs to be investigated.
buyer. On the contrary, the similar term marketplace intelligence refers to the process of collection and making use of the data and knowledge related to different marketplaces, e.g. data marketplace and e-commerce marketplace, The discussion in this paper revolves around the concept of EIM (and not marketplace intelligence), different potential stakeholders and their interests and interactions. For example, a city government for traffic management may buy intelligence (based on AI models or algorithm) that predict congestion based on vehicle counter and optimize traffic light timing on a real-time. Such intelligence can be remotely deployed on the edge device inside traffic light by the intelligence provider (IP) or seller. Another example in the context of smart agriculture is that a farm operator might buy and deploy rule-based or AI-based intelligence for optimal prediction of watering times based on weather data and soil moisture sensors and control basic irrigation. The farm operator may remotely deploy this intelligence on an edge device installed near the farmland.
Smart agriculture operators
Healthcare service providers
City government (smart city authority, traffic management agencies, etc.)
Intelligence Consumer (IC) Requirement specification
5.7. Other baseline technologies Before intelligence can be clustered, it must be developed, deployed, managed in an effective way. This management spans the entire intelligence lifecycle - from development and packaging to deployment, execution, update, and uninstall. All the operations require supporting technologies that are already well matured, including IaC, containerization, orchestration, blockchain for security [92] and several others. Infrastructure-as-Code (IaC) can be used to declare the specification of intelligence, including the underlined dependencies of hosting environments, such as edge and cloud resources, as a code. Containerization technologies can be used to package intelligence, along with the corresponding AI models, algorithms, or rule-based logic and their runtime dependencies. On top of this, orchestration frameworks can be used to coordinate the deployment, scaling, migration, and termination of intelligence instances across edge-cloud computing continuum.
Compatibility validation
Logistics and supplychain operators
Intelligence discovery & Selection Deployment automation
Requirement specification
Utility companies (energy, water)
Transportation department
Marketplace Provider (MP)
IoT analytics companies
Traffic analytics companies
Intelligence catalog
Deployment orchestration
Metadata management
Pricing & licensing
Requirement– capability matching
Recommendation engine
Intelligence Provider (IP) Intelligence development
Intelligence packaging
Metadata annotation
Intelligence publishing
Versioning & updates
AI/ML solution vendors
Industrial automation vendors
Figure 10: Interaction of stakeholders within Intelligence Marketplace
6. Research Dimensions
Potential stakeholders of Edge Intelligence Marketplace, as shown in Figure 10:
6.1. Intelligence marketplace (IM) Intelligence Marketplace, a.k.a Edge Intelligence Marketplace (EIM), can be considered as a platform or a ecosystem where intelligence (derived using AI-based models, algorithms, or even if-else rules) can be exchanged, bought or sold between the Edge Intelligence Producer (EIP) and the
• Intelligence Provider (IP): These can be developers, companies, or individuals who create AI models, new algorithms, or decision rules. They publish these on the marketplace. For examples, IoT and sensor data analytics companies (e.g. PTC (ThingWorx), 18
Siemens MindSphere, or Cisco IoT) can provide/sell intelligence such as predictive maintenance of sensors, edge devices, anomaly detection in edge device, etc. Traffic and transportation analytics firms may provide/sell AI-models based intelligence that give the current traffic congestion and insights.
intelligences, would be both network- and computationintensive. Conversely, a fully centralized observability system introduces inherent and obvious challenges, including single points of failure, scalability bottlenecks, and increased latency.
6.3. Intelligence lifecycle automation Similar to lifecycle automation of software systems in production environment, intelligence lifecycle involves development (using AI models, algorithms, or rule-based logic), packaging, deployment to edge devices or controllers, configuration, execution, monitoring, update or adaptation, and finally uninstalling the intelligence. Automation can be applied on multiple levels of this lifecycle, from basic deployment and start/stop operations to advanced capabilities including automated scaling, migration, versioning, rollback, and context-driven reconfiguration based on environmental changes [86]. While implementing such automation the potential research challenges that we may encounter are how to maintain consistency, how to handle inter-dependencies among multiple intelligences, and • Marketplace provider : The Marketplace Provider (MP) manage heterogeneous device capabilities. The implementation of automation become more cumbersome in coorcan be seen as a trusted platform where intelligence consumer and provider gathered. Intelligence provider dinating lifecycle actions across distributed edge agents while maintaining low latency, scalability, and fault toler(IP) lists the intelligence with a set of metadata ance. and and pricing/licensing terms. In addition, the The existing orchestrators are matured enough to carry IP should also be responsible for updating and mainout the lifecycle operations. However, such solutions are taining intelligences. On the other hand, intelligence not suitable for performing the lifecycle operations on the consumer would specify his/her requirements, search resource devices at the edge of the network. One approach and select intelligences, acquire usage rights, deploy to address this research challenge is to ensure the aumation them onto its edge resources, and utilize their outwhile the orchestrator being deployed on the cloud enviputs for application-level decision-making. The MP ronment. However, this will not be an distributed solution could be seen as a combination of e-commerce platand hence issues such as network stability may become an form (like Amazon) and model hub (like Hugging obstacle in ensuring lifecycle automation. The other apFace Hub) with the additional capabilities of matchproach is to design the orchestrator small enough so that ing consumer’s requirement and providers capabiliit can be installed on the resource constraint edge devices ties, recommendation system for both IP and IC and and responsible to taking care of each edge device sepaan cost-analytics system. rately. • Intelligence Consumer : Intelligence Consumers (IC) typically the business entities or individuals or governments have existing infrastructure and systems in place (e.g., sensors, edge devices and infrastructure). Such consumers may lack deep AI expertise or resources to build models from scratch and hence rely on pre-built intelligence from the marketplace to solve their specific problems. They need specific domain-specific intelligence that is aligned with their operational goals or industry challenges. IC browses the marketplace and select the intelligence (whether it is based on a complex AI model, a simpler algorithm, or an if-else rule set) that suits their needs and compatible with their existing edge infrastructure.
6.2. Intelligence discoverability & observability
6.4. Clustering mechanism This is one of the primary focus area, as an efficient clustering mechanism would determine how individual intelligences are grouped, managed and leveraged for collaborative decision-making. The clustering may follow a manual process, a policy-driven approach defined by administrators, an algorithmic approach based on classical clustering mechanisms, ML methods, or Large Language Model (LLM)-assisted semantic reasoning. The open-ended research questions are where clustering logic resides and who decides the clusters? Clusters may be formed locally by the edge agents, coordinated by the edge controllers or gateways or orchestrated at the cloud level depending on the latency and scalability. Another research aspect is the nature of the cluster itself: in CEI, clusters do not necessarily have a physical manifestation and are primarily logical/virtual constructs that represent relationships
In a system with minimal human intervention, an edge agent must be capable of discovering other relevant intelligences in order to cross-verify and validate its own intelligence before publishing it. Furthermore, when new intelligence becomes available within the edge infrastructure system, an edge agent should be able to utilize it to improve the reliability and accuracy of its own intelligence. A key research challenge, however, lies in how to design, develop, and implement an intelligence discovery mechanism in resource-constrained environments. Existing discovery mechanisms-whether service discovery, device discovery, or network discovery-cannot be directly adopted for intelligence discovery. Another major challenge is the observability of intelligence, such as monitoring its availability and freshness. Designing a fully distributed observability mechanism, where each edge agent monitors other 19
among intelligences rather than underlined co-located devices. Such logical clusters may be instantiated, maintained, or visualized at different layers of the edge-cloud computing continuum. Considering above questions, the research challenges may include designing adaptive clustering mechanisms that can operate across computing layers and ensures that the clusters are meaningful, consistent and dynamic.
7. Use cases 7.1. Industry 4.0 In industry 4.0, such CEI can be applied on a wide range of activities, including predictive maintenance cluster for CNC machines, conveyor belt fault validation cluster, robotic arm collision-risk cluster, industrial fire and gas leak detection cluster, compressor health check, vibrationtemperature-pressure anomaly cluster, industrial wastewater discharge compliance cluster, and many more [93]. For example, one edge device attached to a gas sensor may publish intelligence like “abnormal rise of methane in Zone B”. Another edge device with a thermal camera may publish “heat anomaly near valve section”. A third edge device may publish “real-time increase of smoke”, while a fourth edge device may publish “unstable ventilation detected”. These separate intelligences can be clustered semantically. If gas anomaly, heat anomaly, and smoke-related intelligence appear together within the same time window and location, the cluster can raise a highconfidence “probable fire/gas leak incident in Zone B” alert. Similarly, if vibration anomaly combines with unusual temperature rise and pressure deviation from the same machine section, the cluster may infer “high likelihood of mechanical degradation”, “possible bearing failure”, “seal leakage”, or “pump cavitation risk”, depending on the context and rules used.
6.5. Use of language models Use of language model could be one potential solution to identify and establish relationships among multiple intelligences (e.g., which intelligences participate in remote fire detection or traffic congestion estimation) instead of manually curating clusters. LLM can fill the gap where CEI needs context-aware, multi-constraint clustering and discovery over several artifacts, including AI models, algorithms, rules. LLMs can help in translating naturallanguage intents (e.g., “form a fire-detection cluster for forest zone X with freshness < 2 min and confidence > 0.8”) into structured queries over inventories/knowledge graphs using specific language. Modern language models can also help in figuring out how intelligence from edge agents are related (like what affects what (correlatesWith), what relies on what (dependsOn), and what confirms what (validates)) to suggest candidate clusters and justifications. Language models can be used to complement the formation of clusters. This means LLM can help us find the relationship among multiple intelligence. For example, what are the intelligences that can take participate in detecting the fire in a remote location. Another example is “what are the intelligence that can be considered to estimate the traffic congestion?”
7.2. Smart City & Traffic Management System Similar to Industry 4.0, the potential of CEI can be seen on a wide range of smart city and traffic management related activities, including pedestrian crowding risk cluster, flooded road detection cluster, road accident confirmation cluster, emergency vehicle priority cluster, illegal parking validation cluster, traffic congestion estimation cluster, low-visibility hazard cluster, intersection collisionrisk cluster, and many more [94]. For instance, an edge device with roadside camera may publish intelligence like “high pedestrian density near crossing A”. Another edge device at a smart traffic signal may publish “extended pedestrian waiting time”. A third edge device near a bus stop may publish “unusual crowd buildup”, while a fourth edge device may publish “sharp increase in nearby device presence”. Clustering the above intelligence, a highconfidence “probable pedestrian crowding risk near crossing A” event can be derived reliably. Similarly, if the same cluster also observes repeated road crossing attempts and reduced walking space, it may infer “high likelihood of unsafe pedestrian spillover”.
6.6. Interoperability and Standardization As discussed before, there could be multiple IPs and ICs. In such a situation, intelligences may be generated by different vendors, implemented using different AI models, algorithms, or rule-based logic, and executed on edge devices with varying hardware capabilities, OS, and communication protocols. Intelligence management will become fragmented and difficult to scale across computing continuum if there is no common standard for intelligence representation, description of metadata, communication interface and protocols, and lifecycle operations. For example, the IP publishing the traffic congestion related intelligence may expect the vehicle count data in a JSON format for MQTT. Similarly, the IP publishing pollution monitoringrelated intelligence may expect the CO2 data in a different schema over REST. Without standardized data formats, interface, and metadata description, these two intelligences cannot be combined or clustered.
7.3. Smart Agricultural With recent advancements, drone technology is becoming more affordable, reliable, and efficient. Drones equipped with multiple sensors, along with complementary technologies such as satellite image analysis, are increasingly being used in smart agriculture [95]. In the context of CEI, one 20
edge device connected to a wind sensor may publish intelligence such as “strong gust activity detected in Plot A”. Another edge device attached to a weather station may publish “rapid drop in atmospheric pressure”. A third edge device mounted on a crop-monitoring camera may publish “visible crop bending beyond normal range”, while a fourth edge device may publish “soil surface dryness reducing root grip”. These separate intelligences can be clustered semantically. For instance, if strong gust activity, pressure instability, and abnormal crop bending appear together within the same time window and location, the cluster can generate a high-confidence intelligence alert such as “probable wind-damage risk in Plot A”. CEI, together with drone technology, can support a wide range of agricultural activities, including waterlogging detection, crop disease early warning, rainfall–soil moisture inconsistency detection and prediction, and many more.
learning techniques as an enabler for adaptive and robust decision-making; Zhou et al. [89] strike a similar chord, discussing the importance of deep learning for edge environments to split and distribute intelligence within an architecture. Lastly, Gill et al. [102] widen the scope and describe the architectures to facilitate the interaction and collaborating of distributed software components. Despite this compact body of work, few ideas revolve around the third paradigm discussed in this paper: EI3F—the intelligence available at the edge, where the abstract, derived insights themselves—rather than raw data or hardware models—are treated as independent, firstclass manageable entities. This work addresses this gap. 8.2. Semantic Representation and Abstraction at the Edge To codify, comprehend, and exchange knowledge across heterogeneous edge environments, current research heavily leverages semantic web technologies, ontologies, and knowledge graphs. Early foundational efforts standardized the physical layer: Haller et al. [104] introduced the modular SOSA/SSN ontologies for sensors, actuators, and observations, while Li et al. [105] evaluated Web of Things (WoT) ontologies to drive cross-layer data interoperability and federation. As these standards matured, modeling shifted toward edge infrastructure; for instance, Ghrab et al. [11] proposed EdgeOnto to automate the IoT service lifecycle by mapping node capabilities and Quality of Service (QoS) constraints. Lastly, to elevate raw data into actionable EI, Zhang et al. [108] highlighted context-aware and real-time information fusion as core mechanisms . While standardized ontologies effectively describe data sources and infrastructure, they fail to manage abstract insights as independent assets. To federate intelligence effectively, edge knowledge graphs must model these insights as distinct nodes connected by explicit semantic relationships (e.g., affects, correlatesWith) rather than flat, semantics-free key-value pairs. This shows a gap for ontologies that explicitly encapsulate metadata native to abstract insights, including provenance, confidence levels, geographic scope, temporal freshness, clustering relationships, and operational lifecycle states.
7.4. Other use cases Beyond the above use cases, CEI can also be applied in healthcare monitoring systems, home automation, industrial safety, transportation [96], public infrastructure monitoring, environmental protection, collision avoidance in UAV system [97, 98] and many other domains. The key idea remains the same across all these use cases: instead of relying on a single edge device or a single type of intelligence, multiple related intelligences generated at the edge can be semantically clustered to produce a more reliable and context-aware outcome. 8. Literature Survey This section reviews the academic landscape around EI and CEI across three core dimensions: first, we examine the shifting taxonomies of edge intelligence, highlighting the transition to intelligence as independent assets. Second, we survey the semantic technologies required to represent and codify knowledge across heterogeneous nodes in a system. Third, we analyze the paradigm shift from hardware-bound device networks to insight-driven clusters. Table 4 summarizes the literature considered.
8.3. From Device-Centric to Clustering-Centric The coordination of distributed edge nodes is shifting from physical topologies to intelligence-based grouping. Traditional DCC is grouping edge nodes or IoT devices rigidly by geographic proximity, communication distance, or sensing coverage [47]. For instance, Achkouty et al. [8] and Takale et al. [9] propose spatial clustering to manage device coverage, energy, and storage limits. While these methods optimize data aggregation and network lifetime, they treat nodes strictly as physical containers bounded by hardware and network constraints. Even distributed architectures, like FL, commonly rely on centralized cloud coordinators; this shows a need for multi-edge clustering and direct node communication [107]. For instance, Dehury et
8.1. Paradigms and Taxonomy of Edge Intelligence Existing surveys and taxonomies in EI have targeted mainly the first two paradigms discussed in this paper: EI1F and EI-2F. Closely aligning to these two paradigms, Deng et al. [99] distinguish between “intelligence-enabled edge computing” and “artificial intelligence on the edge”. Xu et al. [100] take an additional step in their survey, where they identify four cross-cutting applications of EI: edge caching, edge training, edge inference, and edge offloading. Stadnicka et al. [103] conducted a large-scale industrial review where they investigated the potentials of EI for various domains. Across these potential applications, Chen et al. [101] review the importance of deep 21
Ref. Zhou et al. [89] Deng et al. [99] Xu et al. [100] Chen et al. [101] Gill et al. [102] Stadnicka et al. [103] Zhang et al. [96] Ghrab et al. [11] Haller et al. [104] Li et al. [105] Dehury et al. [47] Achkouty et al. [8] Takale et al. [9] Zhang et al. [106] Mughal et al. [107]
Year 2019 2020 2020 2019 2025 2022 2022 2023 2019 2019 2022 2024 2025 2025 2024
Focus Area Paradigms and Taxonomy Paradigms and Taxonomy Paradigms and Taxonomy Paradigms and Taxonomy Paradigms and Taxonomy Paradigms and Taxonomy Semantics and Abstractions Semantics and Abstractions Semantics and Abstractions Semantics and Abstractions Intelligence Clustering Intelligence Clustering Intelligence Clustering Intelligence Clustering Intelligence Clustering
Article Type Review Vision Survey Review Survey Review Survey Review Review Review Research Research Research Survey Research
Topic Technologies and frameworks enabling EI Overview of EI paradigms Fundamental components of EI Implications of deep learning for EI Collaborative EI architectures Industrial applications of EI Information fusion as an enabler for EI Ontologies for IoT and Edge computing Ontologies for sensing and processing Ontologies for Internet of Things Emergence of CEI based on topologies Intelligence clusters in IoT networks Usage of CEI for improved FL Usage of agentic AI as enabler for EI Usage of CEI for improved FL
Table 4: Overview of literature revolving around EI and its applications, categorized into three focus areas.
al. [47] form logical clusters based on the specific intelligence devices possess rather than their physical parameters, enabling decentralized peer collaboration with minimal energy expenditure. Furthermore, Zhang et al. [106] show how agentic AI enables edge systems to run continuous perception–reasoning–action loops, allowing nodes to cluster autonomously based on semantic relevance. By treating intelligence as a manageable asset, ICC bypasses hardware vendor lock-in and fosters decentralized collaboration across dynamic edge continuums. In summary, we can report that DCC emerges as a powerful enabler for CEI, allowing distributed entities to act as autonomous and self-organized agents [106], or bypass traditional centralized architectures [107]. Thus, we conclude that CEI is an emerging field with high potential, with most available works focusing on original research, rather than established surveys and reviews (Table 4).
evaluated system. To support the realization of CEI, we discussed a set of baseline technologies, including abstraction, knowledge graphs, communication mechanisms, intelligence discovery and observability, clustering techniques, life cycle automation, containerization, orchestration, and intelligence marketplaces. Furthermore, we investigated several potential research directions in this emerging domain, such as intelligence-centric clustering, distributed intelligence inventories, semantic discoverability, intelligence observability, life cycle management, interoperability, standardization, and the use of LLMs for identifying relationships among heterogeneous intelligence. Despite the wide range of potential research directions, we are currently focusing on developing a CEI prototype and defining a standard representation for intelligence, including its identity, provenance, confidence, freshness, and communication requirements. Our future work also includes designing lightweight and adaptive mechanisms for intelligence discovery, observability, and lifecycle management across the distributed computing continuum.
9. Conclusions and Future Works In this paper, we introduced Clustered Edge Intelligence (CEI) as a vision that redefine Edge Intelligence beyond the conventional use of AI for edge resource management and the deployment of lightweight AI models on edge devices. The existing edge systems provide limited support for managing derived intelligence as an independent entity that can be discovered, shared, reused, and combined across heterogeneous edge devices. The proposed CEI framework filled this gap by first interpreting edge intelligence as the intelligence available at the edge and treating such intelligence as a first-class entity. This intelligence may be derived using AI models, conventional algorithms, rule-based mechanisms, or custom business logic. The proposed CEI framework shifts the focus from device-centric clustering to intelligence-centric clustering. Such clusters allow edge agents to combine and crossvalidate intelligence generated by heterogeneous devices without necessarily sharing raw data. It is important to note that CEI is presented as a vision and reference architecture rather than as an implemented and empirically
Acknowledgement This work is partially funded by IISER Berhampur, institutional seed grant SG/CKD/050225. This work is also partially supported by the EU’s TALENTS (101299722) project. References [1] J. Rowley, The wisdom hierarchy: representations of the dikw hierarchy, Journal of information science 33 (2) (2007) 163–180. [2] S. J. Russell, Artificial intelligence a modern approach, Pearson Education, Inc., 2010. [3] Ieee standard glossary of software engineering terminology, IEEE Std 610.12-1990 (1990) 1– 84doi:10.1109/IEEESTD.1990.101064. [4] A. Sheth, C. Henson, S. S. Sahoo, Semantic sensor web, IEEE Internet computing 12 (4) (2008) 78–83. 22
[5] C. K. Gately, L. R. Hutyra, S. Peterson, I. S. Wing, Urban emissions hotspots: Quantifying vehicle congestion and air pollution using mobile phone gps data, Environmental pollution 229 (2017) 496–504.
[17] D. Vyas, M. d. Prado, T. Verbelen, Towards smart and adaptive agents for active sensing on edge devices (Jan. 2025). doi:10.48550/arXiv.2501.06262. [18] K. S. Velaga, Y. Guo, W. Yu, Edge ai for smart cities: Foundations, challenges, and opportunities, Smart Cities 8 (6) (2025) 211.
[6] J. Espadaler-Clapés, E. Barmpounakis, N. Geroliminis, Traffic congestion and noise emissions with detailed vehicle trajectories from uavs, Transportation Research Part D: Transport and Environment 121 (2023) 103822.
[19] Y. Hu, Q. Li, Y. Chai, D. Wu, L. Lu, N. Shi, Y. Teng, Y. Zhang, Ai service deployment and resource allocation optimization based on human-like networking architecture, IEEE Internet of Things Journal 11 (14) (2024) 24795–24813. doi:10.1109/JIOT.2024.3384546.
[7] J. Du, L. Zhao, J. Feng, X. Chu, Computation offloading and resource allocation in mixed Fog/Cloud computing systems with min-max fairness guarantee, IEEE Transactions on Communications 66 (4) (2018) 1594–1608. doi:10.1109/TCOMM.2017.2787700.
[20] Y. Xie, Q. Fang, An energy-aware generative ai edge inference framework for lowpower iot devices, Electronics 14 (20) (2025). doi:10.3390/electronics14204086. URL https://www.mdpi.com/2079-9292/14/20/ 4086
[8] F. Achkouty, L. Gallon, R. Chbeir, Rdsc: Rangebased device spatial clustering for iot networks, Sensors 24 (17) (2024) 5851. [9] A. K. Takele, B. Villányi, Resource-efficient clustered federated learning framework for industry 4.0 edge devices, Ai 6 (2) (2025) 30.
[21] S. D. A. Shah, M. A. Gregory, F. Bouhafs, F. D. Hartog, Artificial intelligence-defined wireless networking for computational offloading and resource allocation in edge computing networks, IEEE Open Journal of the Communications Society 5 (2024) 2039– 2057. doi:10.1109/OJCOMS.2024.3382265.
[10] B. A. Begum, S. V. Nandury, Data aggregation protocols for wsn and iot applications–a comprehensive survey, Journal of King Saud University-Computer and Information Sciences 35 (2) (2023) 651–681.
[22] C. K. Dehury, S. Poojara, S. N. Srirama, Defdrel: Towards a sustainable serverless functions deployment strategy for fog-cloud environments using deep reinforcement learning, Applied Soft Computing 152 (2024) 111179. doi:https://doi.org/10.1016/j.asoc.2023.111179. URL https://www.sciencedirect.com/science/ article/pii/S1568494623011973
[11] S. Ghrab, I. Lahyani, S. Yangui, M. Jmaiel, A core iot ontology for automation support in edge computing, Service Oriented Computing and Applications 17 (1) (2023) 25–37. [12] R. Singh, S. S. Gill, Edge ai: a survey, Internet of Things and Cyber-Physical Systems 3 (2023) 71–92.
[23] P. Nisha, R. Vinay, K. Laad, P. Sasmal, T. Haraki, C. Juyal, M. A. Gabir Elbakri, A. Acharyya, Run-time prevention of thermal throttling on the edge using reinforcement-learning based predictive thermal aware power and performance management, in: 2024 22nd IEEE Interregional NEWCAS Conference (NEWCAS), 2024, pp. 273–277. doi:10.1109/NewCAS58973.2024.10666109.
[13] A. Choudhury, M. Ghose, A. Islam, et al., Machine learning-based computation offloading in multiaccess edge computing: A survey, Journal of Systems Architecture 148 (2024) 103090. [14] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L.C. Chen, Mobilenetv2: Inverted residuals and linear bottlenecks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4510–4520.
[24] B. Sedlak, P. Raith, A. Morichetta, V. C. Pujol, S. Dustdar, Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices (Oct. 2025). doi:10.48550/arXiv.2510.06882.
[15] F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, K. Keutzer, SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size, arXiv preprint arXiv:1602.07360 (2016).
[25] O. Pandith, M. K. Pandit, A. Gujarati, A. Mehta, R. Sen, Iceedge: Thermal-aware machine learning inference serving for emerging edge applications, ACM Trans. Sen. Netw. (Mar. 2026). doi:10.1145/3801093. URL https://doi.org/10.1145/3801093
[16] X. Zhang, X. Zhou, M. Lin, J. Sun, Shufflenet: An extremely efficient convolutional neural network for mobile devices, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6848–6856. 23
[26] A. Mazloomi, H. Sami, J. Bentahar, H. Otrok, A. Mourad, Reinforcement learning framework for server placement and workload allocation in multiaccess edge computing, IEEE Internet of Things Journal 10 (2) (2023) 1376–1390. doi:10.1109/JIOT.2022.3205051.
Blockchain-Assisted Edge Intelligence Approach, in: GLOBECOM 2020 - 2020 IEEE Global Communications Conference, 2020, pp. 1–6, iSSN: 2576-6813. doi:10.1109/GLOBECOM42002.2020.9348271. [35] T. Rausch, S. Dustdar, Edge Intelligence: The Convergence of Humans, Things, and AI, in: 2019 IEEE International Conference on Cloud Engineering (IC2E), 2019, pp. 86–96. doi:10.1109/IC2E.2019.00022.
[27] C. K. Dehury, S. N. Srirama, An efficient service dispersal mechanism for fog and cloud computing using deep reinforcement learning, in: 2020 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID), 2020, pp. 589–598. doi:10.1109/CCGrid49817.2020.00-34.
[36] G. Plastiras, M. Terzi, C. Kyrkou, T. Theocharidcs, Edge Intelligence: Challenges and Opportunities of Near-Sensor Machine Learning Applications, in: 2018 IEEE 29th International Conference on Application-specific Systems, Architectures and Processors (ASAP), 2018, pp. 1–7, iSSN: 2160-052X. doi:10.1109/ASAP.2018.8445118.
[28] Y. Chang, R. Thompson, C. Hixenbaugh, Tiny longshort term memory model for resource-constrained prediction of battery cycle life, in: G. Nicosia, V. Ojha, S. Giesselbach, M. P. Pardalos, R. Umeton (Eds.), Machine Learning, Optimization, and Data Science, Springer Nature Switzerland, Cham, 2025, pp. 131–144.
[37] E. Peltonen, M. Bennis, M. Capobianco, M. Debbah, A. Ding, F. Gil-Castiñeira, M. Jurmu, T. Karvonen, M. Kelanti, A. Kliks, T. Leppänen, L. Lovén, T. Mikkonen, A. Rao, S. Samarakoon, K. Seppänen, P. Sroka, S. Tarkoma, T. Yang, 6G White Paper on Edge Intelligence, arXiv:2004.14850 [cs] (Apr. 2020). doi:10.48550/arXiv.2004.14850. URL http://arxiv.org/abs/2004.14850
[29] A. Albanese, M. Nardello, D. Brunelli, Automated Pest Detection With DNN on the Edge for Precision Agriculture, IEEE Journal on Emerging and Selected Topics in Circuits and Systems 11 (3) (2021) 458–467, conference Name: IEEE Journal on Emerging and Selected Topics in Circuits and Systems. doi:10.1109/JETCAS.2021.3101740.
[38] D. N. Molokomme, A. J. Onumanyi, A. M. AbuMahfouz, Edge Intelligence in Smart Grids: A Survey on Architectures, Offloading Models, Cyber Security Measures, and Challenges, Journal of Sensor and Actuator Networks 11 (3) (2022) 47, number: 3 Publisher: Multidisciplinary Digital Publishing Institute. doi:10.3390/jsan11030047. URL https://www.mdpi.com/2224-2708/11/3/47
[30] D. Aghi, S. Cerrato, V. Mazzia, M. Chiaberge, Deep Semantic Segmentation at the Edge for Autonomous Navigation in Vineyard Rows, in: 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021, pp. 3421–3428, iSSN: 21530866. doi:10.1109/IROS51168.2021.9635969. [31] P. Radanliev, D. De Roure, K. Page, J. R. C. Nurse, R. Mantilla Montalvo, O. Santos, L. Maddox, P. Burnap, Cyber risk at the edge: current and future trends on cyber risk analytics and artificial intelligence in the industrial internet of things and industry 4.0 supply chains, Cybersecurity 3 (1) (2020) 13. doi:10.1186/s42400-020-00052-8. URL https://doi.org/10.1186/ s42400-020-00052-8
[39] B. Parekh, K. Amin, Edge Intelligence: A Robust Reinforcement of Edge Computing and Artificial Intelligence, in: P. K. Singh, Z. Polkowski, S. Tanwar, S. K. Pandey, G. Matei, D. Pirvu (Eds.), Innovations in Information and Communication Technologies (IICT-2020), Advances in Science, Technology & Innovation, Springer International Publishing, Cham, 2021, pp. 461–468. [40] B. Sedlak, I. Murturi, P. K. Donta, S. Dustdar, A Privacy Enforcing Framework for Transforming Data Streams on the Edge, IEEE Transactions on Emerging Topics in Computing (2023). doi:10.1109/TETC.2023.3315131.
[32] S. N. Srirama, Fog computing and large language models: A vision for mutual beneficiaries, IEEE Computer (2026). doi:10.1109/MC.2026.3708686. [33] B. Sedlak, A. Morichetta, Y. Wang, Y. Fei, L. Wang, S. Dustdar, X. Qu, SLO-Aware Task Offloading Within Collaborative Vehicle Platoons, in: W. Gaaloul, M. Sheng, Q. Yu, S. Yangui (Eds.), Service-Oriented Computing, Springer Nature, Singapore, 2024, pp. 72–86.
[41] P. Thakur, S. Goel, E. Puthooran, Edge ai enabled iot framework for secure smart home infrastructure, Procedia Computer Science 235 (2024) 3369–3378. [42] C. Savaglio, P. Mazzei, G. Fortino, Edge intelligence for industrial iot: Opportunities and limitations, Procedia Computer Science 232 (2024) 397–405.
[34] C. Qiu, X. Wang, H. Yao, Z. Xiong, F. R. Yu, V. C. M. Leung, Bring Intelligence among Edges: A 24
[43] T. Shen, L. Zhang, R. Geng, S. Li, B. Sun, Traffic prediction for diverse edge iot data using graph network, Journal of Cloud Computing 13 (1) (2024) 84.
[54] G. Simmann, V. Veeranna, R. Kriesten, Design of an alternative hardware abstraction layer for embedded systems with time-controlled hardware access, in: 2024 Stuttgart International Symposium, SAE Technical Paper, 2024.
[44] W. Qi, Q. Li, Q. Song, L. Guo, A. Jamalipour, Extensive edge intelligence for future vehicular networks in 6g, IEEE Wireless Communications 28 (4) (2021) 128–135.
[55] A. Lercher, J. Glock, C. Macho, M. Pinzger, Microservice api evolution in practice: A study on strategies and challenges, Journal of Systems and Software 215 (2024) 112110.
[45] B. McMahan, E. Moore, D. Ramage, S. Hampson, B. A. y Arcas, Communication-efficient learning of deep networks from decentralized data, in: Artificial intelligence and statistics, Pmlr, 2017, pp. 1273– 1282.
[56] I. Bartolini, M. Patella, A stream processing abstraction framework, Frontiers in Big Data 6 (2023) 1227156. [57] M. Silva, J. Santos, M. Curado, The path towards virtualized wireless communications: A survey and research challenges, Journal of Network and Systems Management 32 (1) (2024) 12.
[46] Y. Li, X. Wang, R. Zeng, P. K. Donta, I. Murturi, M. Huang, S. Dustdar, Federated domain generalization: A survey, Proceedings of the IEEE (2025). [47] C. K. Dehury, P. K. Donta, S. Dustdar, S. N. Srirama, CCEI-IoT: Clustered and cohesive edge intelligence in Internet of Things, in: 2022 IEEE International Conference on Edge Computing and Communications (EDGE), IEEE, 2022, pp. 33–40.
[58] P. Krishnan, K. Jain, A. Aldweesh, P. Prabu, R. Buyya, Openstackdp: a scalable network security framework for sdn-based openstack cloud infrastructure, Journal of Cloud Computing 12 (1) (2023) 26. [59] W. R. Cook, On understanding data abstraction, revisited, in: Proceedings of the 24th ACM SIGPLAN conference on Object oriented programming systems languages and applications, 2009, pp. 557–572.
[48] M. Silva, T. Gomes, M. Ekpanyapong, A. Tavares, S. Pinto, Chameliot: A tightly-and loosely-coupled hardware-assisted os framework for low-end iot devices, Real-Time Systems 60 (1) (2024) 150–196.
[60] M. Shirali, M. F. Sani, Z. Ahmadi, E. Serral, Llmbased event abstraction and integration for iotsourced logs, in: International Conference on Business Process Management, Springer, 2024, pp. 138– 149.
[49] R. Chan, W. K. Yan, J. M. Ma, K. M. Loh, T. Yu, M. Y. H. Low, K. P. Yar, H. Rehman, T. C. Phua, Iot devices deployment challenges and studies in building management system, Frontiers in the Internet of Things 2 (2023) 1254160.
[61] E. Villagrossi, M. Delledonne, M. Faroni, M. Beschi, N. Pedrocchi, Hiding task-oriented programming complexity: an industrial case study, International Journal of Computer Integrated Manufacturing 36 (11) (2023) 1629–1648.
[50] J. Höglund, S. Bouget, M. Furuhed, J. Preuß Mattsson, G. Selander, S. Raza, Autopki: public key infrastructure for iot with automated trust transfer, International Journal of Information Security 23 (3) (2024) 1859–1875.
[62] L. Kong, J. Tan, J. Huang, G. Chen, S. Wang, X. Jin, P. Zeng, M. Khan, S. K. Das, Edge-computingdriven internet of things: A survey, ACM computing surveys 55 (8) (2022) 1–41.
[51] M. Pathak, K. N. Mishra, S. P. Singh, Securing data and preserving privacy in cloud iot-based technologies an analysis of assessing threats and developing effective safeguard, Artificial intelligence review 57 (10) (2024) 269.
[63] A. Hogan, E. Blomqvist, M. Cochez, C. d’Amato, G. D. Melo, C. Gutierrez, S. Kirrane, J. E. L. Gayo, R. Navigli, S. Neumaier, et al., Knowledge graphs, ACM Computing Surveys (Csur) 54 (4) (2021) 1–37.
[52] O. Dib, Communication-efficient agentic edge intelligence: A survey of federated multiagent information fusion for llm-driven edge ai, Information Fusion 136 (2026) 104577. doi:https://doi.org/10.1016/j.inffus.2026.104577.
[64] C. K. Dehury, S. N. Srirama, Mastering DevOps: A Cloud Engineering and Data Science Perspective, Morgan Kaufmann, 2026. doi:https://doi.org/ 10.1016/C2025-0-00512-4.
[53] S. Wagner, F. Deissenboeck, Abstractness, specificity, and complexity in software design, in: Proceedings of the 2nd International Workshop on the Role of Abstraction in Software Engineering, 2008, pp. 35–42.
[65] A. Dauda, O. Flauzac, F. Nolot, A survey on iot application architectures, Sensors 24 (16) (2024) 5320.
25
[66] T. F. D. B. Sarasola, A. García, J. L. Ferrando, Iiot protocols for edge/fog and cloud computing in industrial ai: A high frequency perspective, International Journal of Cloud Applications and Computing (IJCAC) 14 (1) (2024) 1–30.
on Communication Systems and Networks (COMSNETS 2012), 2012, pp. 1–9, iSSN: 2155-2509. doi:10.1109/COMSNETS.2012.6151335. [77] P. C. Ccori, L. C. C. De Biase, M. K. Zuffo, F. S. C. da Silva, Device discovery strategies for the IoT, in: 2016 IEEE International Symposium on Consumer Electronics (ISCE), 2016, pp. 97–98, iSSN: 21591423. doi:10.1109/ISCE.2016.7797388.
[67] A. Lazidis, K. Tsakos, E. G. Petrakis, Publish– subscribe approaches for the iot and the cloud: Functional and performance evaluation of open-source systems, Internet of Things 19 (2022) 100538.
[78] Z. Beerliova, F. Eberhard, T. Erlebach, A. Hall, M. Hoffmann, M. Mihal’ak, L. S. Ram, Network Discovery and Verification, IEEE Journal on Selected Areas in Communications 24 (12) (2006) 2168–2181, conference Name: IEEE Journal on Selected Areas in Communications. doi:10.1109/JSAC.2006.884015.
[68] R. Gupta, D. Reebadiya, S. Tanwar, 6g-enabled edge intelligence for ultra-reliable low latency applications: Vision and mission, Computer Standards & Interfaces 77 (2021) 103521. [69] K. Huang, Q. Lan, Z. Liu, L. Yang, Semantic data sourcing for 6g edge intelligence, IEEE Communications Magazine 61 (12) (2023) 70–76.
[79] R. I. Ansari, C. Chrysostomou, S. A. Hassan, M. Guizani, S. Mumtaz, J. Rodriguez, J. J. P. C. Rodrigues, 5G D2D Networks: Techniques, Challenges, and Future Prospects, IEEE Systems Journal 12 (4) (2018) 3970–3984, conference Name: IEEE Systems Journal. doi:10.1109/JSYST.2017.2773633.
[70] I. Murturi, C. Avasalcai, C. Tsigkanos, S. Dustdar, Edge-to-edge resource discovery using metadata replication, in: 2019 IEEE 3rd International Conference on Fog and Edge Computing (ICFEC), IEEE, 2019, pp. 1–6.
[80] J. Kosińska, B. Baliś, M. Konieczny, M. Malawski, S. Zieliński, Toward the observability of cloud-native applications: The overview of the state-of-the-art, IEEE Access 11 (2023) 73036–73052.
[71] I. Murturi, S. Dustdar, A Decentralized Approach for Resource Discovery using Metadata Replication in Edge Networks, IEEE Transactions on Services Computing 15 (5) (2022) 2526–2537.
[81] U. Faseeha, H. J. Syed, F. Samad, S. Zehra, H. Ahmed, Observability in microservices: An indepth exploration of frameworks, challenges, and deployment paradigms, IEEE Access 13 (2025) 72011– 72039.
[72] N. Dragoni, S. Giallorenzo, A. L. Lafuente, M. Mazzara, F. Montesi, R. Mustafin, L. Safina, Microservices: yesterday, today, and tomorrow, Present and ulterior software engineering (2017) 195–216.
[82] B. Li, X. Peng, Q. Xiang, H. Wang, T. Xie, J. Sun, X. Liu, Enjoy your observability: an industrial survey of microservice tracing and analysis, Empirical Software Engineering 27 (1) (2022) 25.
[73] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, J. Wilkes, Borg, omega, and kubernetes, Communications of the ACM 59 (5) (2016) 50–57. [74] S. E. Czerwinski, B. Y. Zhao, T. D. Hodes, A. D. Joseph, R. H. Katz, An architecture for a secure service discovery service, in: Proceedings of the 5th annual ACM/IEEE international conference on Mobile computing and networking, ACM, Seattle Washington USA, 1999, pp. 24–35. doi:10.1145/313451.313462. URL https://dl.acm.org/doi/10.1145/313451. 313462
[83] F. Gomes, P. Rego, F. Trinta, A systematic mapping study on observability of microservices-based applications: fundamentals, classifications, and challenges, Computing 107 (9) (2025) 183. [84] B. Costa, A. Banerjee, P. P. Jayaraman, L. R. Carvalho, J. Bachiega Jr, A. Araujo, Achieving observability on fog computing with the use of open-source tools, in: International conference on mobile and ubiquitous systems: computing, networking, and services, Springer, 2023, pp. 319–340.
[75] O. Hayat, R. Ngah, S. Z. Mohd Hashim, M. H. Dahri, R. Firsandaya Malik, Y. Rahayu, Device Discovery in D2D Communication: A Survey, IEEE Access 7 (2019) 131114–131134, conference Name: IEEE Access. doi:10.1109/ACCESS.2019.2941138.
[85] V. C. Pujol, P. K. Donta, A. Morichetta, I. Murturi, S. Dustdar, Edge intelligence—research opportunities for distributed computing continuum systems, IEEE Internet Computing 27 (4) (2023) 53–74.
[76] F. Baccelli, N. Khude, R. Laroia, J. Li, T. Richardson, S. Shakkottai, S. Tavildar, X. Wu, On the design of device-to-device autonomous discovery, in: 2012 Fourth International Conference
[86] B. Sedlak, V. C. Pujol, I. M. de Abril, P. K. Donta, A. N. Toosi, S. Dustdar, Service Orchestration in the Computing Continuum: Structural Challenges 26
and Vision, IEEE Internet Computing 30 (2) (2026) 88–98. doi:10.1109/MIC.2026.3662064.
[97] H. S. Khargharia, A. Ouali, S. Shakya, S. Ahmad, Collision avoidance in uav swarms: A learning-centric perspective on collaborative intelligence, Neurocomputing 663 (2026) 132020. doi:https://doi.org/10.1016/j.neucom.2025.132020. URL https://www.sciencedirect.com/science/ article/pii/S092523122502692X
[87] S. Amgothu, A. K. Vedantham, S. P. M. Gowda, Observability and aiops in cloud-scale devops: Technologies, architectures, challenges, and future trends, in: 2025 38th Conference of Open Innovations Association (FRUCT), IEEE, 2025, pp. 12–20.
[98] J. Zhang, Y. Hu, M. Shao, Y. Tang, L. Wang, X. Li, Learning enhanced scheduling and resource allocation for heterogeneous uav swarms in edge assisted remote sensing, Scientific Reports (2026).
[88] F. Bonomi, R. Milito, J. Zhu, S. Addepalli, Fog computing and its role in the internet of things, in: Proceedings of the first edition of the MCC workshop on Mobile cloud computing, 2012, pp. 13–16.
[99] S. Deng, H. Zhao, W. Fang, J. Yin, S. Dustdar, A. Y. Zomaya, Edge Intelligence: The Confluence of Edge Computing and Artificial Intelligence, IEEE Internet of Things Journal (Aug. 2020). doi:10.1109/JIOT.2020.2984887.
[89] Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, J. Zhang, Edge intelligence: Paving the last mile of artificial intelligence with edge computing, Proceedings of the IEEE 107 (8) (2019) 1738–1762. [90] M. Ester, H.-P. Kriegel, J. Sander, X. Xu, et al., A density-based algorithm for discovering clusters in large spatial databases with noise, in: kdd, Vol. 96, 1996, pp. 226–231.
[100] D. Xu, T. Li, Y. Li, X. Su, S. Tarkoma, T. Jiang, J. Crowcroft, P. Hui, Edge Intelligence: Architectures, Challenges, and Applications, number: arXiv:2003.12172 arXiv:2003.12172 [cs] (Jun. 2020). URL http://arxiv.org/abs/2003.12172
[91] T. Zhang, R. Ramakrishnan, M. Livny, Birch: an efficient data clustering method for very large databases, ACM sigmod record 25 (2) (1996) 103– 114.
[101] J. Chen, X. Ran, Deep Learning With Edge Computing: A Review, Proceedings of the IEEE 107 (8) (2019) 1655–1674. doi:10.1109/JPROC.2019.2921977.
[92] C. K. Dehury, S. N. Srirama, P. K. Donta, S. Dustdar, Securing clustered edge intelligence with blockchain, IEEE Consumer Electronics Magazine 13 (1) (2024) 22–29.
[102] S. S. Gill, M. Golec, J. Hu, M. Xu, J. Du, H. Wu, G. K. Walia, S. S. Murugesan, B. Ali, M. Kumar, et al., Edge ai: A taxonomy, systematic review and future directions, Cluster Computing 28 (1) (2025) 18.
[93] A. Carvalho, N. O. Mahony, L. Krpalkova, S. Campbell, J. Walsh, P. Doody, At the edge of industry 4.0, Procedia Computer Science 155 (2019) 276–281, the 16th International Conference on Mobile Systems and Pervasive Computing (MobiSPC 2019),The 14th International Conference on Future Networks and Communications (FNC-2019),The 9th International Conference on Sustainable Energy Information Technology. doi:https://doi.org/10.1016/j.procs.2019.08.039.
[103] D. Stadnicka, J. Sęp, R. Amadio, D. Mazzei, M. Tyrovolas, C. Stylios, A. Carreras-Coch, J. A. Merino, T. Żabiński, J. Navarro, Industrial Needs in the Fields of Artificial Intelligence, Internet of Things and Edge Computing, Sensors 22 (12) (2022) 4501, number: 12 Publisher: Multidisciplinary Digital Publishing Institute. doi:10.3390/s22124501. URL https://www.mdpi.com/1424-8220/22/12/ 4501
[94] M. Trigka, E. Dritsas, Edge and cloud computing in smart cities, Future Internet 17 (3) (2025). doi:10.3390/fi17030118.
[104] A. Haller, K. Janowicz, S. J. Cox, M. Lefrançois, K. Taylor, D. Le Phuoc, J. Lieberman, R. GarcíaCastro, R. Atkinson, C. Stadler, The modular SSN ontology: A joint W3C and OGC standard specifying the semantics of sensors, observations, sampling, and actuation, Semantic Web 10 (1) (2019) 9–32. doi:10.3233/SW-180320.
[95] M. U. Tariq, S. M. Saqib, T. Mazhar, M. A. Khan, T. Shahzad, H. Hamam, Edge-enabled smart agriculture framework: Integrating iot, lightweight deep learning, and agentic ai for context-aware farming, Results in Engineering 28 (2025) 107342. doi:https://doi.org/10.1016/j.rineng.2025.107342.
[105] W. Li, G. Tropea, A. Abid, A. Detti, F. Le Gall, Review of Standard Ontologies for the Web of Things, in: 2019 Global IoT Summit (GIoTS), 2019, pp. 1–6. doi:10.1109/GIOTS.2019.8766377.
[96] Y. Zhang, C. Jiang, B. Yue, J. Wan, M. Guizani, Information fusion for edge intelligence: A survey, Information Fusion 81 (2022) 171–186. 27
[106] R. Zhang, G. Liu, Y. Liu, C. Zhao, J. Wang, Y. Xu, D. Niyato, J. Kang, Y. Li, S. Mao, S. Sun, X. Shen, D. I. Kim, Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions, arXiv:2508.18725 [cs.NI] (Aug. 2025). doi:10.48550/arXiv.2508.18725. URL http://arxiv.org/abs/2508.18725 [107] F. R. Mughal, J. He, B. Das, F. A. Dharejo, N. Zhu, S. B. Khan, S. Alzahrani, Adaptive federated learning for resource-constrained iot devices through edge intelligence and multi-edge clustering, Scientific Reports 14 (1) (2024) 28746. [108] Y. Zhang, C. Jiang, B. Yue, J. Wan, M. Guizani, Information fusion for edge intelligence: A survey, Information Fusion 81 (2022) 171–186. doi:10.1016/j.inffus.2021.11.018. URL https://www.sciencedirect.com/science/ article/pii/S1566253521002438
28