1 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) <
AISC deployment in dynamic UAV-assisted MEC network: a reinforcement learning method based on heterogeneous graph attention neural network Hanzhi Chang, Jing Bai, Xin Tang, Xiaomei Liu Abstract—Unmanned aerial vehicles-assisted mobile edge computing (UMEC) can execute compute-intensive and latencycritical artificial intelligence (AI) services, which can be provided by multiple UAVs collaborating in the air to perform inference tasks. Completing an AI service requires multiple inferences, each of which is implemented by an AI service chain consisting of multiple virtual network functions (VNFs). The application of AISC relies on an efficient AISC deployment strategy to determine which UAV to deploy VNF on. However, the UMEC network topology is highly dynamic due to the high-speed movement of UAVs or their departure/arrival, which makes the AISC deployment in the UMEC network challenging. In addition, the intricate relationships between UMEC environment and AISC, as well as between individual VNFs in an AISC, can also affect the effectiveness of AISC deployment strategy. Moreover, under the constraints of energy consumption and load balancing, it is also difficult to optimize the AISC strategy to minimize AISC completion time for enhancing the quality of AI service. To address the above challenges, this paper proposes a double deep attention Q-network based on heterogeneous graph neural networks, which incorporates heterogeneous graph to capture diverse relationships in UMEC and utilizes attention mechanisms to adaptively focus on critical nodes and links for intelligent AISC deployment. The experimental results demonstrate that the proposed algorithm performs excellently in AISC completion time, AISC completion rate, load balancing and energy consumption. Index Terms—Artificial intelligence service chain (AISC), unmanned aerial vehicles (UAV), mobile edge computing network, reinforcement learning (RL), graph neural network (GNN).
U
I. INTRODUCTION
NMANNED aerial vehicles (UAVs) have the advantages of flexible deployment, rapid response and wide coverage to make up for the lack of ground serves [1][2], enabling mobile edge computing (MEC) to execute compute-intensive and latency-critical artificial intelligence (AI) services such as face recognition, unmanned driving and augmented reality [3]. In UAV-assisted MEC (UMEC) network, multiple UAVs collaborate in the air to provide AI services for users by performing inference tasks [4]. An AI service consists of multiple inference tasks [9], each of which can be treated as an AI service chain (AISC) provided by several virtual network functions (VNFs) connected in series Corresponding author: Jing Bai Hanzhi Chang, Jing Bai, Xin Tang and Xiaomei Liu are with the School of Cyber Science and Engineering, University of International Relations, Beijing, China. E-mail: {hanzhi.chang, baijing, xtang, liuxiaomei}@uir.edu.cn.
[5]. As a practical example, consider an AI-driven video analytics service chain deployed across multiple low-energy integrated communication units. In such a scenario, raw video frames are first preprocessed and compressed in the initial VNF. Then, object detection or feature extraction is performed in the second VNF. Finally, the last VNF conducts tasks such as activity recognition or anomaly detection. Fig. 1 shows the framework of AISC deployment. In the User Layer, users in different areas initiate AISC requests through their associated base stations. These service requests are transmitted to the UMEC Network Layer, where multiple UAVs form a flexible airborne MEC infrastructure. Each VNF in an AISC is mapped to a specific UAV for execution. As defined in [6], with the assistance of the orchestrator, each VNF is placed on an appropriate UAV and executed on its onboard computing module, which typically consists of an embedded CPU/GPU platform. UMEC networks are increasingly used to execute latencycritical AI services such as video analytics, surveillance, and environmental monitoring, where inference tasks must be processed collaboratively across multiple UAVs. Since each VNF in an AISC relies on the computing and communication resources of its hosting UAV, the placement of VNFs directly determines the end-to-end service delay, energy efficiency, and overall service reliability. An inappropriate deployment may lead to significant performance degradation. For example, placing successive VNFs on distant or overloaded UAVs can cause excessive transmission delay or even service interruption when the UAV becomes unavailable. Therefore, designing an efficient and intelligent AISC deployment strategy is essential for ensuring stable, reliable, and high-quality AI services in UMEC networks. In order to fully utilize UAV resources and optimize the quality of AI services, an efficient AISC deployment strategy is required to decide on which UAV each VNF in an AISC is deployed. Fig. 2 (a) shows the illustration of AISC deployment in the UMEC network. However, there are the following challenges in deploying AISC in the UMEC network. The high-speed movement of UAVs leads to the high dynamic UMEC network topology [7]. Once the network topology changes, the current deployment may not meet user needs. The virtual links of AISC must be remapped
2 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) < to new links, as shown in Fig. 2 (b). In addition, when UAVs leave the network, VNFs deployed on these UAVs must be migrated and the AISCs whose data transmission through the departing UAVs need to be redeployed, as shown in Fig. 2 (c). When new UAVs join the network, the deployment space expands, allowing subsequent AISCs to be deployed on the newly joined UAVs, as shown in Fig. 2 (d). Therefore, the real-time deployment of AISC in the dynamic UMEC network is a problem that needs to be solved. AISC completion time can be used to effectively characterize the quality of AI service. However, minimizing AISC completion time and optimizing load balancing restrict each other. In addition, AISC completion time is also affected by energy consumption. If the UAV leaves the network due to energy exhaustion, VNFs deployed on it must be migrated, which increases AISC completion time [8]. Therefore, how to minimize AISC completion time while optimizing load balancing and energy consumption is a problem that needs to be solved. Different from individual services, the deployment of AISC composed of multiple VNFs in the UMEC network is affected by the correlation relationships between UAVs, between AISCs and VNFs, between VNFs, between VNFs and UAVs, and between virtual links connecting VNFs and U2U links. Therefore, how to incorporate these correlation relationships into the AISC deployment process to improve the performance of the algorithm is a problem that need to be solved. However, existing studies either assumed that the network topology is static [15][18][19][20][25], or focused on the deployment of individual service in the UMEC network [11][23][24], or optimized only one or two of the completion time, energy consumption and load balancing [12][14][16][17][20]. To address the above challenges, this paper proposes a double deep attention Q-network based on heterogeneous graph neural networks (DDAQ-HGNN). To the best of our knowledge, this is the first work to incorporate heterogeneous graph modeling into the AISC deployment process. The main contributions of this paper are summarized as follows:
●
We model AISC deployment problem in the dynamic UMEC network as an Integer Linear Programming (ILP) problem to jointly optimize three critical objectives: energy consumption, AISC completion time and network load balancing. In particular, we propose a migration model for capturing the system behaviors in scenarios where the network topology changes over time and UAVs leave/join the network. ● The proposed DDAQ-HGNN algorithm consists of three components: heterogeneous graph encoder, heterogeneous attention Q-network and heterogeneous nodes decoder. The encoder integrates all correlation relationships in the UMEC network into a unified heterogenous graph representation. The Q-network incorporates a heterogeneous graph attention mechanism to effectively extract features from multi-type nodes and edges. The decoder predicts UAVs selection based on the features learned from the heterogeneous graph, enabling decisionmaking under dynamic environments. ● To enable DDAQ-HGNN to better evaluate long-term deployment performance, we design an asynchronous reward function. That is, the agent does not receive immediate rewards upon each VNF deployment, but calculates and returns rewards after the entire AISC completes its inference process or fails, thereby guiding the algorithm to make more informed decisions. ● We design and conduct extensive experiments to evaluate the adaptability and performance of the proposed algorithm in the dynamic UMEC network. We verify the performance of the algorithm under different network topologies. Additionally, we assess the adaptability of the algorithm under varying UMEC networks topology change probabilities. The results demonstrate that our proposed algorithm significantly outperforms existing baseline algorithms in terms of AISC completion time, network load balancing, AISC completion rate and energy consumption. The rest of this paper is organized as follows. Section II discusses the related work. Section III introduces the system model and problem formulation. Section IV presents the proposed DDAQ-HGNN in detail. Section V shows the results of performance evaluations. Finally, we conclude the paper and discuss future research directions in Section VI.
Fig. 1. The framework of AISC deployment Fig. 2. Illustration of AISC deployment in the dynamic UMEC network
3 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) < II. RELATED WORKS In this section, we introduce the existing studies related to our work from three aspects: RL for service deployment and scheduling, GNN for network modeling and scheduling and multi-objective optimization in UMEC networks. A. RL for Service Deployment and Scheduling Currently, solutions for service placement and scheduling are generally divided into two categories: heuristic solutions and RL solutions. Liu et al. [12] considered cost-oriented and delayconstrained anycasting in space–air–ground networks as a mixed-integer program refined by a greedy search, assuming quasi-static links and fixed number of nodes. Xu et al. [14] proposed a cloud-centric heuristic method for mapping UAVsourced video service chain requests, which minimizes network cost but depends on centrally controlled and stable infrastructures. Wang et al. [16] proposed a two-stage offline scheduler, ToRu, targeting heterogeneous multi-UAV edge computing, which first deploys the chains in parallel when resources are abundant and then serially reorders tasks to maximize revenue under CPU/FPGA constraints. Wang et al. [17] proposed embedding and migration algorithms for UAV IoT scenarios, specifically addressing the challenge of maintaining service continuity under UAV mobility and changing IoT demands, with the main objective of reducing service latency and migration overhead. Because these heuristic solutions lack the ability to dynamically learn from real-time environmental feedback, it is often difficult for them to effectively adapt to the rapid and frequent topology changes in dynamic UMEC networks [12][14][16][17]. RL has become a powerful solution for service deployment and scheduling, owing to its advantages in adaptability, realtime decision-making, and the capability to learn effective policies from environmental interactions without relying on explicit network models. Wu et al. [15] developed a deep reinforcement learning (DRL)-based QoE-aware service chain deployment framework for UAV networks, focusing on improving service quality by dynamically adjusting service chain configurations. Zhang et al. [19] proposed a mobilityaware service chain deployment strategy in network function virtualization (NFV)-based edge-cloud environments to proactively reduce service downtime caused by UAV mobility using predictive migration techniques. Additionally, Wang et al. [20] proposed a dual-agent DRL framework for MEC scenarios, which jointly optimizes partial task offloading and service chain mapping to effectively reduce overall latency and improve resource utilization efficiency. Bao et al. [25] applied multi-agent DRL with action masking to achieve sustainable task offloading in UAV-assisted smart farm networks, focusing on secure and efficient resource management. The methods in [15][20][25] assumed that the network is static and were not suitable for analyzing the dynamic UMEC network, where the topology and the number of UAVs fluctuate frequently. Moreover, the methods in [18][19][20] optimized only one or two of the completion time, energy consumption and load balancing, ignoring the synergistic impact of these metrics on AISC deployment strategy.
B. GNN for Network Modeling and Scheduling GNNs have become a powerful tool for network resource management and service placement, because they are able to learn from graph-structured data to capture the relational and topological characteristics of networks. Ping et al. [22] proposed a GNN-based solution for task scheduling-dependent QoE optimization in edge-cloud computing networks, demonstrating the benefits of graph-structured modeling in task coordination and resource management. Feng et al. [24] utilized graph attention mechanisms combined with RL to design UAV trajectories and allocate communication resources. Gao et al. [23] integrated federated RL with GNNs to address trajectory optimization in UAV networks. However, most of these solutions are designed for relatively static scenarios, where the network topology and the number of nodes remain stable over time [22][23][24]. Such settings simplify the modeling process but limits generalization to more dynamic environments. In contrast, deploying AISC in UMEC networks introduces unique challenges: UAVs may join or leave the network dynamically, inter-node relationships evolve with mobility, and the overall network topology changes over time. Additionally, since AISC deployment involves placing VNFs one by one, it is necessary to consider sequential decision dependencies. These characteristics require adaptive models that can capture dynamic topological changes and heterogeneous relationships. Therefore, non-heterogeneous graph GNNs are not suitable for supporting the adaptive and real-time decision-making required for AISC deployment in UMEC networks. C. AISC Deployment in MEC networks There are fewer studies on the AISC deployment approach in UMEC networks. Lu et al. [18] proposed a RL-based framework for dynamically deploying service chain in UAV networks under security threats, optimizing deployment success rate and security robustness by intelligently avoiding vulnerable network paths. However, this study only considered deployment success rate and cost, without incorporating delay and load balancing, both of which are critical for ensuring timely end-to-end service execution and preventing resource congestion in practical UMEC environments. Wang et al. [21] proposed a distributed generative reinforcement learning proposed a distributed generative RL framework to achieve stable service chain orchestration in highly dynamic UAV swarm networks. Their method leveraged generative models to predict future network states, thereby supporting coordinated decision-making among distributed UAVs, which improves orchestration stability and reduces the impact of network fluctuations. However, this study only considered the deployment problem in static UMEC networks and did not take into account scenarios in which the number of UAVs changes or the UAV network topology evolves. These factors are crucial in real-world UMEC systems, because UAV arrivals, departures, and topology variations can significantly affect link availability, resource allocation, and end-to-end service performance. Qiu et al. [3] proposed an online security-aware and reliability-guaranteed provisioning scheme for AISCs in edge intelligence clouds. Their model integrated security constraints and reliability requirements into a stochastic optimization framework to improve service stability. However,
4 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) < this approach relied on heuristic algorithm, leading to its limited adaptability to dynamic network states. Moreover, in terms of optimization objectives, it only focused on minimizing cost and improving security levels without considering delay performance. Li et al. [13] presented a slicing-based AISC provisioning method that balanced AI service performance and data-management resource consumption. However, this study relied on static edge servers as the execution platforms for VNFs and did not consider UAVs as MEC service nodes, ignoring the flexibility and mobility advantages offered by UAV-assisted MEC systems. In addition, the method only optimized the cost, ignoring other essential performance indicators such as end-to-end delay, load balancing, and service success rate. These factors are critical in practical UMEC environments, because they directly influence user experience, system stability, and the feasibility of sustained AI service deployment. Moreover, none of these studies considered the problem of service migration when nodes fail. In contrast, we address the limitations identified above. Specifically, we employ a heterogeneous graph neural network to support AISC deployment when both the number of UAV and the UMEC network topology change over time. In addition, our model explicitly incorporates VNF migration to maintain service continuity in dynamic UMEC networks. By integrating attention mechanisms with reinforcement learning, the proposed approach achieves joint optimization of energy consumption, AISC completion time, and network load balancing. III. SYSTEM MODEL AND PROBLEM FORMULATION In this section, we propose a system model for AISC deployment in the dynamic UMEC network, which includes the physical model, the migration model, and the computation model. Based on the system model, we propose the problem constraints and formulate the optimization problem as an ILP problem. A. Physical Model At time slot t, the UMEC network is represented as a connected graph 𝐺𝑡 = {𝑈𝑡 , 𝐸𝑡 } , where 𝑈𝑡 denotes the set of UAVs and 𝐸𝑡 denotes the set of U2U links. We use |𝑈𝑡 | and |𝐸𝑡 | to denote the total number of UAVs and U2U links, respectively. Each UAV is denoted as 𝑢𝑡,𝑖 ∈ 𝑈𝑡 (𝑖 ∈ [1, |𝑈𝑡 |]). Each U2U link is denoted as 𝑒𝑡,𝑗 ∈ 𝐸𝑡 , representing an undirected U2U link between two UAVs 𝑢𝑡,𝑖 and 𝑢𝑡,𝑖 ′ (𝑖 ′ ∈ [1, |𝑈𝑡 |], 𝑖 ≠ 𝑖 ′ , 𝑗 ∈ [1, |𝐸𝑡 |]). The attributes of the UAVs and U2U links are listed in Table I. Each AISC is represented as a directed graph ℱ𝑓 = {𝒱𝑓 , ℒ𝑓 }, where 𝑓 denotes the index of the AISC. 𝒱𝑓 is the set of VNFs, and ℒ𝑓 is the set of virtual links. We use |𝒱𝑓 | and |ℒ𝑓 | to denote the number of VNFs and virtual links, respectively. Each VNF is denoted as 𝑣𝑓,𝑚 ∈ 𝒱𝑓 ( 𝑚 ∈ [1, |𝒱𝑓 |] ) . A virtual link is represented as 𝑙𝑓,𝑛 ∈ ℒ𝑓 , indicating a directed virtual link between two connected VNF 𝑣𝑓,𝑚 and 𝑣𝑓,𝑚′ ∈ 𝒱 𝑓 ( 𝑚′ ∈ [1, |𝒱𝑓 |], 𝑚 ≠ 𝑚′, 𝑛 ∈ [1, |ℒ𝑓 |]). The attributes of VNFs and virtual links are listed in Table II.
B. Problem Constraints At any time slot t, the deployment and migration of AISCs must satisfy the following constraints: Mapping Constraints. At each time slot t, each VNF can only be deployed on one UAV, while each UAV can host multiple VNFs. This constraint is denoted as: |𝒱𝑓 |
|𝑈𝑡 |
∑ 𝜆𝑡,𝑓,𝑚,𝑖 = 1, ∑ ∑ 𝜆𝑡,𝑓,𝑚,𝑖 ≥ 0 𝑖=1
(1)
𝑓 𝑚=1
where 𝜆𝑡,𝑓,𝑚,𝑖 ∈ {0,1}, and 𝜆𝑡,𝑓,𝑚,𝑖 = 1 indicates that VNF 𝑣𝑓,𝑚 is deployed on UAV 𝑢𝑡,𝑖 at time slot 𝑡; otherwise, it is not deployed. At each time slot 𝑡, any virtual link can be mapped to multiple U2U links, and each U2U link can carry multiple virtual links. This constraint is denoted as: |ℒ 𝑓 |
|𝐸𝑡 |
∑ 𝜆𝑡,𝑓,𝑛,𝑗 ≥ 1, ∑ ∑ 𝜆𝑡,𝑓,𝑛,𝑗 ≥ 0 𝑗=1
(2)
𝑓 𝑛=1
where 𝜆𝑡,𝑓,𝑛,𝑗 ∈ {0,1}, and 𝜆𝑡,𝑓,𝑛,𝑗 = 1 indicates that virtual link 𝑙𝑓,𝑛 is deployed on U2U link 𝑒𝑡,𝑗 at time slot 𝑡; otherwise, it is not deployed. Capacity Constraints. At each time slot 𝑡, every UAV and U2U link must satisfy capacity constraints. For UAVs, the constraints include computing unit capacity, memory capacity, and energy availability, which are denoted as: |𝒱𝑓 | capnum
(3)
capmem mem ∑ ∑ 𝜆𝑡,𝑓,𝑚,𝑖 ∗ 𝑐𝑓,𝑚 ≤ 𝑐𝑡,𝑖
(4)
num ∑ ∑ 𝜆𝑡,𝑓,𝑚,𝑖 ∗ 𝑐𝑓,𝑚 ≤ 𝑐𝑡,𝑖 𝑓 𝑚=1 |𝒱𝑓 |
𝑓 𝑚=1
𝑐 capeng (𝑢𝑖𝑡 ) > 0 For U2U links, the bandwidth constraint is given by:
(5)
|ℒ 𝑓 | capbw bw ∑ ∑ 𝜆𝑡,𝑓,𝑛,𝑗 ∗ 𝑐𝑓,𝑛 ≤ 𝑐𝑡,𝑗
(6)
𝑓 𝑛=1
AISC Completion Time Constraints. To ensure service quality, the transmission delay of each AISC must not exceed its maximum delay requirement, which is denoted as: delay delay (7) time𝑓 ≤ 𝑐𝑓 Each AISC is allowed to occupy network resources only when its inference step does not meet the requirement, which is denoted as: stp reqstp (8) 𝑐𝑓 < 𝑐𝑓 IV. SOLUTION To address the challenges posed by multi-dimensional states, complex topologies, and multi-objective optimization in the dynamic UMEC network for AISC deployment, this paper proposes a novel RL algorithm—double deep attention Qnetwork based on heterogeneous graph neural networks (DDAQ-HGNN). By constructing the UMEC network as a heterogeneous graph, integrating attention mechanisms with Qlearning strategies, and predicting UAV selection based on the learned heterogeneous representations, the algorithm achieves efficient AISC deployment in dynamic UMEC networks. The
5 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) < entire algorithm can be formalized as an MDP, as illustrated in Fig. 3. In the MDP framework, DDAQ-HGNN acts as the agent that captures observations from the environment, executes actions, and receives feedback in the form of rewards. In each interaction, the environment provides an observation that consists of the current UMEC network topology, node and edge attributes, and pending AISC deployment. DDAQ-HGNN agent consists of three main components. The heterogeneous graph encoder encodes the input observation into a unified graph representation, preserving type-specific information and contextual dependencies within the UMEC network to provide a structured foundation for decision-making. The heterogeneous attention Q-network applies a graph attention mechanism to perform cross-type information fusion and importance modeling for different types of nodes and edges. This process generates high-dimensional feature vectors representing candidate deployment nodes (i.e., UAVs). This design enables the model to capture the joint impact of local attributes and global topology dynamics on deployment decisions, thereby improving the generalization and robustness of the learned policy. The heterogeneous nodes decoder maps each node embedding to a scalar Q-value and selects the optimal UAV node for the current VNF deployment based on the highest Q-value.
Fig. 3. The MDP framework of DDAQ-HGNN for AISC deployment in dynamic UMEC networks
Fig. 4. Detailed architecture of DDAQ-HGNN The environment then provides rewards based on the effectiveness of the AISC deployment action, considering factors such as energy consumption, network load balancing, and resource conflicts to guide policy updates. In addition, we design a custom asynchronous reward function to support multi-objective optimization during the AISC deployment process. This reward function simultaneously accounts for key performance metrics including
energy consumption, average AISC completion time, network load balancing, and AISC completion rate. Through continuous interaction between the RL agent and the dynamic environment, DDAQ-HGNN progressively converges to a stable and efficient deployment policy under variable topologies and uncertain resource conditions. The rest of this section will introduce the three components of DDAQ-HGNN and training algorithm in detail based on Fig. 4. Remark. The UMEC environment features highly dynamic network topologies due to the continuous movement of UAVs. It also contains heterogeneous structures composed of multiple node types, such as UAVs and VNFs, and multiple edge types with attributes like bandwidth and delay that directly affect deployment decisions. To effectively capture these characteristics, we adopt HGNN, which supports type-specific message passing and incorporates edge features, enabling a more expressive modeling of the UMEC network than homogeneous GNNs. Furthermore, DDAQ leverages attention mechanisms to adaptively emphasize the most informative nodes and links under dynamic conditions, allowing the agent to learn robust deployment policies. Therefore, DDAQ-HGNN address the challenges of heterogeneous relationships, topology variability, and multi-objective optimization in AISC deployment. A. Heterogeneous Attention Q-Network The heterogeneous attention Q-network employs the graph attention mechanism to fuse the features of different types of nodes and edges to produce high-dimensional feature vectors representing each candidate UAV. 1) Edges Feature Aggregation For each edge in the heterogeneous graph, the edge attribute and edge type embeddings are concatenated to form the edge eattr ⃑type eattr feature, denoted as 𝑒⃑𝑥𝑦 = [𝑑⃑𝑥𝑦 ||𝑑𝑥𝑦 ] . 𝑑⃑𝑥𝑦 is the type eattr corresponding row vector from 𝐬𝑡 , and 𝑑⃑𝑥𝑦 is the type
corresponding column vector from 𝐬𝑡
. The symbol 𝑥𝑦
indicates that the edge starts from node 𝑥 and ends at node 𝑦 ( 𝑥, 𝑦 ∈ 𝒩𝑡 ). For each destination node 𝑦 , the edge feature + 𝑒⃑𝑥𝑦 = [𝑠⃑𝑦𝑡 ||𝑒⃑𝑥𝑦 ]is combined with the node’s own attribute to form an enriched input representation for the attention mechanism. 𝑠⃑𝑡,𝑦 is the corresponding row vector of node 𝑦 in + 𝐬𝑡nattr . The vector 𝑒⃑𝑥𝑦 is then used to compute the attention coefficient 𝛼𝑥𝑦 by the shared mechanism. The computation is defined as follows:
6 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) < + exp(LeakyReLU((𝑤 ⃑⃑⃑ att )T ∙ [𝑠⃑𝑡,𝑥 ||𝑒⃑𝑥𝑦 ])) (9) att T + ])) ∑𝑧∈𝒩𝑡,𝑥 exp(LeakyReLU((𝑤 ) ⃑⃑⃑ ∙ [𝑠⃑𝑡,𝑥 ||𝑒⃑𝑥𝑧 where 𝒩𝑡,𝑥 denotes the set of neighboring nodes of node 𝑥 in the heterogeneous graph and 𝑠⃑𝑡,𝑥 is the corresponding row vector for node 𝑥 in 𝐬𝑡nattr . A shared single-layer perception network 𝑤 ⃑⃑⃑ att of size 2 + 𝑘 𝒩 + 𝑘 type is used to compute attention scores. The attention coefficient 𝛼𝑥𝑦 reflects the importance of node 𝑦 to node 𝑥, considering both node and edge features. The LeakyReLU activation function is defined as: 𝑝 𝑝>0 (10) LeakyReLU(𝑝) = { 𝛽𝑝 𝑝 ≤ 0 where 𝛽 ∈ ℝ is the leakage factor. 2) Multi-Head Attention The attention coefficients are used to update the feature representation of each node. To incorporate both edge attributes and node attributes, a weighted fusion is performed to compute ′ the updated node feature (𝑠⃑𝑡,𝑥 ) :
𝛼𝑥𝑦 =
′
eattr (𝑠⃑𝑡,𝑥 ) = 𝜎 ( ∑ 𝛼𝑥𝑦 𝑤 ⃑⃑⃑ node [𝑑⃑𝑥𝑦 ||𝑠⃑𝑡,𝑥 ])
(11)
𝑦∈𝒩𝑡,𝑥 (𝑘𝒩 +𝑘ℰ )
where 𝑤 ⃑⃑⃑ node ∈ ℝ is the perception vector applied to connect edge attribute and node feature. The edge type type embedding 𝐬𝑡 is excluded from this fusion because it is a discrete feature and has already been incorporated in the attention coefficient calculation. The activation function 𝜎(∙) is defined as: 1 (12) 𝜎(𝑝) = 1 + 𝜃 −𝑝 where 𝜃 denotes the natural base and is used here to avoid notation conflicts. Finally, in order to stabilize the self-attention mechanism, the heterogeneous attention Q-network employs multiple independent attention heads. The final node feature is obtained by concatenating the outputs of each attention head: K ′
κ 𝜅 eattr (𝑠⃑𝑡,𝑥 ) = [𝜎 ( ∑ 𝛼𝑥𝑦 𝑤 ⃑⃑⃑node [𝑑⃑𝑥𝑦 ||𝑠⃑𝑡,𝑥 ])] 𝑦∈𝒩𝑡,𝑥
(13)
𝜅
where K is the number of attention heads, and 𝜅 indexes each head. B. Heterogeneous Node Decoder The heterogeneous node decoder maps each node feature to a scalar Q-value and selects the optimal UAV for the current VNF deployment based on these values. The decoder first aggregates the updated node features into a out matrix 𝐏𝑡 ∈ ℝ|𝒩𝑡 |×𝑘 , which is calculated as: ′ T
𝐏𝑡 =
((𝑠⃑𝑡,𝑥=1 ) ) ⋯
′ T
(14)
[((𝑠⃑𝑡,𝑥=|𝒩 𝑡| ) ) ] Next, the matrix 𝐏𝑡 is decoded into a matrix 𝐏𝑡val by using a value perception network, as follows: (15) 𝐏𝑡val = ReLU(𝐏𝑡 𝐖 val + 𝐛val )
out
val
where 𝐖 val ∈ ℝ𝑘 ×𝑘 is the value perception matrix, and val 𝐛val ∈ ℝ|𝒩𝑡 |×𝑘 is the bias matrix. The ReLU activation function is defined as: (16) ReLU(𝑝) = max(0, 𝑝) The value matrix is projected into a value vector 𝑞⃑𝑡 , which represents the Q-value for each node: (17) 𝑞⃑𝑡 = 𝐏𝑡val 𝑤 ⃑⃑⃑ val val 𝑘 val where 𝑤 ⃑⃑⃑ ∈ ℝ is a perception vector used for value projection. Therefore, the final output of the agent is defined as: (18) 𝑎𝑡 = argmax 𝑞⃑𝑡 𝑞∈𝑞⃑⃑𝑡UAV
where 𝑞⃑𝑡UAV denotes the subset of 𝑞⃑𝑡 corresponding to UAV nodes. Remark. All trainable parameter matrices and vectors (𝐖 UAV , 𝑏⃑⃑𝑖 , 𝐖 AISC , 𝑏⃑⃑𝑓 , 𝐖 VNF , 𝑏⃑⃑𝑚 , 𝐖 plink , 𝑏⃑⃑𝑗 , 𝐖 vlink , 𝑏⃑⃑𝑛 , 𝐖 type , 𝑤 ⃑⃑⃑ att node val val val 𝑤 ⃑⃑⃑ ,𝐖 ,𝐛 ,𝑤 ⃑⃑⃑ ) depend only on the latent dimensions 𝑘 𝒩 ,𝑘 ℰ ,𝑘 type ,𝑘 out , and 𝑘 val and are independent of the number of nodes and edges in the heterogeneous graph. Therefore, DDAQ-HGNN is able to handle AISC deployment in dynamic UMEC networks. V. SIMULATION AND PERFORMANCE EVALUATIONS In this section, we present extensive experimental results to validate the performance of the proposed DDAQ-HGNN for deploying AISC in dynamic UMEC networks. All experiments are implemented using PyTorch [26] and NetworkX on a laptop equipped with an AMD Ryzen 7745HX CPU, 64 GB RAM, and an RTX 4060 GPU with 8 GB of memory. A. Evaluation Results and Analysis Fig. 5 presents the average AISC completion rate achieved by different algorithms for four network topologies. DDAQHGNN achieves the highest completion rate under all topologies and topology change probabilities, demonstrating its superior deployment stability and task execution capability in dynamic network environments. It is worth noting that even in the extreme condition where the topology change probability reaches 30%, DDAQ-HGNN maintains a significantly higher completion rate than other algorithms, highlighting its robustness and generalization ability. As the topology change probability increases, the AISC completion rate of all algorithms decrease, indicating that dynamic topological changes have a negative impact on the deployment success rate. However, DDAQ-HGNN has the smallest performance degradation, indicating that it has a stronger adaptability to network topology perturbations. PPO and HANConv rank second in AISC completion rate, but their stability declines under high topology change probability, particularly in the Watts-Strogatz and Random Regular topologies. Their completion rates drop rapidly, revealing their limited ability to handle frequent structural changes. Fig. 6 illustrates the average reward of different algorithms for four networks. Fig. 6(a) shows the average reward of the six algorithms under different training episodes. The proposed DDAQ-HGNN achieves the highest average reward in the later stages of training, indicating its superior long-term performance. In the early stage, DDAQ-HGNN also demonstrates a
7 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) < competitive convergence speed second only to PPO, reflecting its strong training efficiency. Moreover, algorithms that leverage heterogeneous graph structures generally outperform those that do not, indicating that heterogeneous graph modeling significantly enhances representational capacity and decisionmaking performance when addressing AISC deployment in the dynamic UMEC network. Fig. 6(b) shows the average rewards of each algorithm under four different network topologies. DDAQ-HGNN consistently achieves the highest average reward under four topologies, demonstrating its strong generalization ability to diverse network structures. In contrast, algorithms such as PPO, HANConv, and GCN show more diverse performance depending on the topology. Traditional strategies such as Greedy and Random perform poorly under four scenarios, indicating that they cannot adapt to the dynamic demands of AISC deployment in the UMEC network.
Fig. 5. Average AISC completion rate of different algorithms under varying topology change probabilities for four network topologies
Fig. 6. Average reward of different algorithms for four network topologies VII. CONCLUSION In this paper, we propose a novel AISC deployment model tailored for the dynamic UMEC network. Based on this model, we introduce the DDAQ-HGNN algorithm, which can learn structure-aware features from heterogeneous graph representation with multiple node and edge types and utilizes attention mechanisms to achieve intelligent AISC deployment in the dynamic UMEC network. Extensive experimental results demonstrate the effectiveness of DDAQ-HGNN in collaboratively optimizing AISC completion time, network
load balancing, AISC completion rate and energy consumption. Compared with the baseline algorithms, DDAQ-HGNN exhibits superior adaptability and generalization capability, consistently achieving optimal performance under different network topologies and different UMEC network topology change probabilities. REFERENCES [1]
Y. Xu, T. Zhang, D. Yang, Y. Liu, and M. Tao, “Joint Resource and Trajectory Optimization for Security in UAV-Assisted MEC Systems,” IEEE Transactions on Communications, vol. 69, no. 1, pp. 573–588, 2021, doi: 10.1109/TCOMM.2020.3025910. [2] Y. Xu, T. Zhang, Y. Liu, D. Yang, L. Xiao, and M. Tao, “UAV-Assisted MEC Networks With Aerial and Ground Cooperation,” IEEE Transactions on Wireless Communications, vol. 20, no. 12, pp. 7712– 7727, 2021, doi: 10.1109/TWC.2021.3086521. [3] Y. Qiu, J. Liang, V. C. M. Leung, and M. Chen, “Online Security-Aware and Reliability-Guaranteed AI Service Chains Provisioning in Edge Intelligence Cloud,” IEEE Transactions on Mobile Computing, vol. 23, no. 5, pp. 5933–5948, 2024, doi: 10.1109/TMC.2023.3314580. [4] C. Deng, X. Fang, and X. Wang, “UAV-Enabled Mobile-Edge Computing for AI Applications: Joint Model Decision, Resource Allocation, and Trajectory Optimization,” IEEE Internet of Things Journal, vol. 10, no. 7, pp. 5662–5675, 2023, doi: 10.1109/JIOT.2022.3151619. [5] J. Liu, X. Wang, K. Ren, Y. Zhou, and M. Li, “Secure Service Function Chain Provisioning for Task Offloading in Device-Edge-Cloud Computing,” IEEE Transactions on Information Forensics and Security, vol. 20, pp. 3717–3730, 2025, doi: 10.1109/TIFS.2025.3553013. [6] B. Li, Z. Fei, and Y. Zhang, “UAV Communications for 5G and Beyond: Recent Advances and Future Trends,” IEEE Internet of Things Journal, vol. 6, no. 2, pp. 2241–2263, 2019, doi: 10.1109/JIOT.2018.2887086. [7] A. H. Wheeb, R. Nordin, A. A. Samah, M. H. Alsharif, and M. A. Khan, “Topology-Based Routing Protocols and Mobility Models for Flying Ad Hoc Networks: A Contemporary Review and Future Research Directions,” Drones, vol. 6, no. 1, 2022, doi: 10.3390/drones6010009. [8] M. Pourghasemian, M. R. Abedi, S. S. Hosseini, N. Mokari, M. R. Javan, and E. A. Jorswieck, “AI-Based Mobility-Aware Energy Efficient Resource Allocation and Trajectory Design for NFV Enabled Aerial Networks,” IEEE Transactions on Green Communications and Networking, vol. 7, no. 1, pp. 281–297, 2023, doi: 10.1109/TGCN.2022.3186911. [9] L. Wu, W. A. Hanafy, A. Souza, T. Abdelzaher, G. Verma, and P. Shenoy, “Enhancing Resilience in Distributed ML Inference Pipelines for Edge Computing,” in MILCOM 2024 - 2024 IEEE Military Communications Conference (MILCOM), 2024, pp. 1–6. doi: 10.1109/MILCOM61039.2024.10773652. [10] M. Abdel-Basset, R. Mohamed, I. M. Hezam, K. M. Sallam, A. Foul, and I. A. Hameed, “Multiobjective trajectory optimization algorithms for solving multi-UAV-assisted mobile edge computing problem,” J Cloud Comp, vol. 13, no. 1, p. 35, Feb. 2024, doi: 10.1186/s13677-024-00594z. [11] S. Tong, Y. Liu, J. Mišić, X. Chang, Z. Zhang, and C. Wang, “Joint Task Offloading and Resource Allocation for Fog-Based Intelligent Transportation Systems: A UAV-Enabled Multi-Hop Collaboration Paradigm,” IEEE Trans. Intell. Transport. Syst., vol. 24, no. 11, pp. 12933–12948, Nov. 2023, doi: 10.1109/TITS.2022.3163804. [12] Y. Liu et al., “Cost-Oriented and Delay-Constrained Anycasting for Service Function Chain Provisioning Leveraging Cloud-Edge Collaboration in Space-Air-Ground Integrated Networks,” IEEE Internet Things J., pp. 1–1, 2024, doi: 10.1109/JIOT.2024.3485640. [13] M. Li, J. Gao, C. Zhou, X. S. Shen, and W. Zhuang, “Slicing-Based Artificial Intelligence Service Provisioning on the Network Edge: Balancing AI Service Performance and Resource Consumption of Data Management,” IEEE Vehicular Technology Magazine, vol. 16, no. 4, pp. 16–26, Dec. 2021, doi: 10.1109/MVT.2021.3114655. [14] D. Xu, X. Tian, K. Pham, E. Blasch, and G. Chen, “Virtual Network Function Placement for Mapping SFC Requests of UAV-Sourced Video Streaming in Cloud Networks,” in 2024 IEEE International Conference on Communications Workshops (ICC Workshops), Denver, CO, USA:
8 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (DOUBLE-CLICK HERE TO EDIT) < IEEE, Jun. 2024, pp. 523–528. doi: 10.1109/ICCWorkshops59551.2024.10615913. [15] Y. Wu, Z. Jia, Q. Wu, and Z. Lu, “Adaptive QoE-Aware SFC Orchestration in UAV Networks: A Deep Reinforcement Learning Approach,” IEEE Trans. Netw. Sci. Eng., vol. 11, no. 6, pp. 6052–6065, Nov. 2024, doi: 10.1109/TNSE.2024.3442857. [16] Y. Wang et al., “Service Function Chain Scheduling in Heterogeneous Multi-UAV Edge Computing,” Drones, vol. 7, no. 2, p. 132, Feb. 2023, doi: 10.3390/drones7020132. [17] X. Wang, S. Shi, and C. Wu, “Research on Service Function Chain Embedding and Migration Algorithm for UAV IoT,” Drones, vol. 8, no. 4, p. 117, Mar. 2024, doi: 10.3390/drones8040117. [18] Y. Lu, C. Jiang, L. Tan, J. Zhang, P. Zhang, and C. Rong, “UAV Dynamic Service Function Chains Deployment Based on Security Considerations: A Reinforcement Learning Method,” IEEE Internet Things J., vol. 11, no. 24, pp. 39731–39743, Dec. 2024, doi: 10.1109/JIOT.2024.3450886. [19] Y. Zhang, R. Wang, Q. Wu, J. Hao, and Z. Xiong, “Mobility-Aware Service Function Chain Deployment with Migration in NFV-Based EdgeCloud,” in 2023 21st International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks (WiOpt), Singapore, Singapore: IEEE, Aug. 2023, pp. 87–94. doi: 10.23919/WiOpt58741.2023.10349842. [20] X. Wang, H. Xing, F. Song, S. Luo, P. Dai, and B. Zhao, “On Jointly Optimizing Partial Offloading and SFC Mapping: A Cooperative DualAgent Deep Reinforcement Learning Approach,” IEEE Trans. Parallel Distrib. Syst., vol. 34, no. 8, pp. 2479–2497, Aug. 2023, doi: 10.1109/TPDS.2023.3287633. [21] Z. Wang, H. Yao, T. Mai, and D. Wu, “Distributed Generative Reinforcement Learning for Stable Service Function Chain Orchestration in Highly Dynamic UAV Swarm Networks,” IEEE Trans. Veh. Technol., pp. 1–15, 2025, doi: 10.1109/TVT.2025.3585912. [22] Y. Ping, K. Xie, X. Huang, C. Li, and Y. Zhang, “GNN-Based QoE Optimization for Dependent Task Scheduling in Edge-Cloud Computing Network,” in 2024 IEEE Wireless Communications and Networking Conference (WCNC), Dubai, United Arab Emirates: IEEE, Apr. 2024, pp. 1–6. doi: 10.1109/WCNC57260.2024.10571289. [23] Y. Gao, M. Liu, X. Yuan, Y. Hu, P. Sun, and A. Schmeink, “Federated deep reinforcement learning based trajectory design for UAV-assisted networks with mobile ground devices,” Sci Rep, vol. 14, no. 1, p. 22753, Oct. 2024, doi: 10.1038/s41598-024-72654-y. [24] Z. Feng, D. Wu, M. Huang, and C. Yuen, “Graph-Attention-Based Reinforcement Learning for Trajectory Design and Resource Assignment in Multi-UAV-Assisted Communication,” IEEE Internet Things J., vol. 11, no. 16, pp. 27421–27434, Aug. 2024, doi: 10.1109/JIOT.2024.3397823. [25] T. Bao, A. Syed, W. S. Kennedy, and M. Erol-Kantarci, “Sustainable Task Offloading in Secure UAV-Assisted Smart Farm Networks: A MultiAgent DRL With Action Mask Approach,” IEEE Trans. Netw. Serv. Manage., pp. 1–1, 2024, doi: 10.1109/TNSM.2024.3486288. [26] A. Paszke, “Pytorch: An imperative style, high-performance deep learning library,” arXiv preprint arXiv:1912.01703, 2019. [27] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” Aug. 28, 2017, arXiv: arXiv:1707.06347. doi: 10.48550/arXiv.1707.06347. [28] T. N. Kipf and M. Welling, “Semi-Supervised Classification with Graph Convolutional Networks,” Feb. 22, 2017, arXiv: arXiv:1609.02907. Accessed: Nov. 11, 2023. [Online]. Available: http://arxiv.org/abs/1609.02907 [29] X. Wang et al., “Heterogeneous Graph Attention Network,” in The World Wide Web Conference, in WWW ’19. New York, NY, USA: Association for Computing Machinery, 2019, pp. 2022–2032. doi: 10.1145/3308558.3313562. [30] P. Erdős and A. Renyi, “On random matrices,” Magyar Tud. Akad. Mat. KutatóInt. Közl, vol. 8, pp. 455–461, 1964. [31] R. Albert and A.-L. Barabási, “Statistical mechanics of complex networks,” Rev. Mod. Phys., vol. 74, no. 1, pp. 47–97, Jan. 2002, doi: 10.1103/RevModPhys.74.47. [32] B. Bollobás, “Random Graphs,” in Modern Graph Theory, New York, NY: Springer New York, 1998, pp. 215–252. doi: 10.1007/978-1-46120619-4_7. [33] D. J. Watts and S. H. Strogatz, “Collective dynamics of ‘small-world’ networks,” Nature, vol. 393, no. 6684, pp. 440–442, Jun. 1998, doi: 10.1038/30918.
Hanzhi Chang (Student Member, IEEE) received his B.S. degree from the Department of Cyber Science and Engineering, University of International Relations, Beijing, China, in 2023. He is currently pursuing for his M.S. degree in the Department of Cyber Science and Engineering, University of International Relations, Beijing, China. His research interests include network function virtualization, network resource orchestration and management, and reinforcement learning algorithms. Jing Bai received the PhD degree in cyberspace security from Beijing Jiaotong University in 2023. She is currently a Lecturer at School of Cyber Science and Engineering, University of International Relations. Her interests include software trustworthiness analysis and cloud security. Xin Tang received the PhD degree in computer science from Beijing University of Posts and Telecommunications, Beijing, China, in 2015. He worked as a post-doctoral fellow in Department of Electronic Engineering at Tsinghua University, Beijing, China, from 2015 to 2017. He is currently an associate professor with the School of Cyber Science and Engineering, University of International Relations, Beijing, China. His current research interests are focused on secure cloud computing and reversible data hiding. Xiaomei Liu received the Master's degree in computer technology from University of Chinese Academy of Sciences in 2018. She is currently an Engineer at School of Cyber Science and Engineering, University of International Relations. Her interests include cloud-edge computing and network security.