On Identifying Adversarial Intent Injection in AI-Native 6G Networks Nilesh Charkraborty1 , Petar Djukic2 , Burak Kantarci1 1
University of Ottawa, Ottawa, ON, Canada Nokia Bell Labs, 600 March Road, Kanata, ON K2K 2E6, Canada 1 {nchakrab, burak.kantarci}@uottawa.ca, 2 [email protected]
arXiv:2609.12144v1 [cs.NI] 10 Sep 2026
2
Abstract—AI-native 6G networks have brought IntentBased Networking (IBN) to the forefront, enabling high-level goals to be translated into network configurations. However, this abstraction opens new attack surfaces, primarily adversarial intent injection, where malicious policies are disguised within benign intent flows. The detection of attack instances might become significantly more difficult if the adversaries adopt a stealthy mode of malicious intent injection. With all these in mind, we first define a fine-grained threat model that facilitates the threat of malicious intent injection in an AI-native network. Alongside, we investigate four malicious intent injection strategies− stealth-mode, random distribution, increasing frequency, and decreasing frequency− and propose a dual-path detection framework: (i) a CNN using TF-IDF features for supervised malicious intent detection, and (ii) an AutoEncoder trained exclusively on benign data for one-class malicious intent detection. Our evaluation demonstrates strong detection performance, with accuracy improving to 0.97 (≈ 9% gain) and F1-score to 0.98 (≈ 36% gain) over the state-of-the-art baseline. Index Terms− AI-Native Networks, Intent, Adversarial Intent Injection, Threat Detection, Network Security
I. Introduction With the rapid evolution of next-generation networks, the network landscape is shifting towards higher levels of automation, adaptability, and intelligence [1]. Among the emerging paradigms, Intent-Based Networking (IBN) has attracted significant attention for its ability to translate high-level declarative intents into complex configurations [2], [3]. Intents are typically represented in machinereadable formats such as JSON structures or service templates like TOSCA. In addition, several industry initiatives have proposed domain-specific intent languages, including the TM Forum Intent Ontology and related frameworks developed within telecom standardization bodies such as GSMA. By leveraging artificial intelligence and natural language processing (NLP), IBN enables network operators to specify desired outcomes without dealing with low-level implementation details. The system interprets, assembles, and translates these intents into concrete network policies, thereby enabling streamlined management, improved agility, and reduced operational overhead [4], [5]. However, this abstraction also introduces new security challenges. As IBN frameworks increasingly rely on automation and semantic parsing, they become exposed to a
new class of risks-particularly adversarial or unauthorized intent injections, where malicious configurations closely mimic legitimate ones, undermining system integrity [6], [7]. For example, in [8] the authors consider a scenario in which a legitimate application installs an intent to establish connectivity between two hosts. An adversary with access to the intent interface may submit a second intent that mimics the original policy but with a different identifier or priority. When the controller processes this new intent, it may overwrite or interfere with the flow rules associated with the legitimate one. By subsequently withdrawing the malicious intent, the attacker can silently remove the corresponding forwarding rules, leaving the original intent logically present in the controller but functionally ineffective in the data plane. Therefore, such threats, if undetected, can lead to undesirable behaviors ranging from misconfigurations to sophisticated cyberattacks. These concerns highlight the importance of building robust IBN systems that are not only intelligent and efficient but also, resilient to manipulation and adversarial exploitation. In this paper, we consider a threat model where attackers inject malicious intents into the benign flow of intent traffic using varying frequency patterns- such as stealth-mode injection or gradual increases in malicious activity. These strategies rely on subtle manipulations of contextual dependencies across sequences of intents, where malicious behaviors emerge from the distribution and timing behaviour [9] of injected intents. To the best of our knowledge, this is the first work to explore the possibility of such threats by considering the distribution patterns of malicious intents as contextual cues, and to detect adversarial attempts from the flow of intent traffic. Based on this primary motivation, the contributions of this paper are as follows. Contribution 1: We propose a comprehensive adversarial threat model targeting the intent acquisition stage of the IBN pipeline. To emulate realistic attack scenarios, we generate a diverse intent dataset capturing four malicious injection strategies: stealth, random, decreasing frequency, and increasing frequency, where frequency is explicitly modeled as a practical threat dimension [10].
Contribution 2: We introduce a dual-path detection framework that leverages TF-IDF [11]-based features with (i) a supervised CNN classifier and (ii) a one-class AutoEncoder trained exclusively on benign intent sequences. Both models employ a sliding window-based temporal representation to segment intent streams into contextaware windows, enabling the detection of localized and temporally correlated adversarial manipulations. Despite being trained exclusively on benign data, the AutoEncoder achieves strong performance, with accuracy and F1 -score exceeding 0.85 in most scenarios, indicating its effectiveness in detecting previously unseen threat patterns [12]. The supervised CNN further enhances performance, achieving accuracy and F1-scores above 0.95 in several cases. The remainder of the paper is organized as follows. Section II establishes the proposed threat model and reviews the related work for gap identification. Section III details the construction of the threat detection models, followed by their performance evaluation in Section IV. Finally, Section V presents concluding remarks. II. Background and Related Work This section first unveils the proposed threat model targeting the intent acquisition stage of the IBN pipeline. Following this, we examine state-of-the-art methods closely related to our work to identify any existing gaps. A. Threat Model In IBN systems, where intents are expressed in formats such as JSON or TOSCA, an adversary can inject malicious intents into benign flows by exploiting the intent ingestion layer or vulnerable APIs [6]. By crafting configurations that closely resemble legitimate requests, these malicious intents evade detection. With the increasing prevalence of API-related threats [13], compromised APIs significantly amplify the risk of unauthorized policy injection. Fig. 1 illustrates this scenario, where an attacker uses a compromised API key to inject intents that may cause denial of service, privilege escalation [14], traffic redirection [15], or persistent backdoors [16], while appearing as routine updates. We consider four injection strategies with distinct temporal characteristics: (i) stealth injection, where malicious intents follow a Poisson arrival process with a fixed rate; (ii) increasing-frequency injection, where the rate of malicious intents grows over time; (iii) decreasingfrequency injection, where the rate gradually declines; and (iv) random injection, where malicious intents are uniformly distributed across the stream without a predefined arrival model. These strategies capture diverse adversarial behaviors and introduce a temporal dimension to intent injection.
Has Access
Intent Assembly & Normalization Layer
Policy Generation & Harmonization Layer
Intent Parsing & Understand -ing Layer
Policy Translation & Deployment Layer
Intent Ingestion Layer
Network Infrastructure Layer
Compromised API Key Entry point to AI-native network Attacker
Malicious JSON code
Feedback & Monitoring Layer
Fig. 1: Threat model−Malicious intent injection through vulnerable API B. Related work IBN has started gaining traction as a promising paradigm for automating and simplifying network management by translating high-level, goal-oriented intents into low-level configurations [17]. Although most early research in IBN has focused on intent formulation, orchestration, and operational efficiency [4], [18], to date, only limited work has explored the associated security risks arising from intent-driven network automation. For example, Bringhenti et al. [19] present a comprehensive survey of automation challenges in network security, emphasizing the need for context-aware validation mechanisms. Kim et al. [20] highlight that the separation between high-level intents and low-level configurations in IBN introduces a semantic gap that can be exploited by adversaries. Trizio et al. [6] propose a rule-based malicious intent detection approach for enterprise networks, where outof-scope intents are classified as malicious. However, this work primarily focuses on static rule-level analysis and does not consider temporal variations or stealthy injection strategies, thereby limiting its ability to capture contextual information. Phantom Link attacks further highlight temporal vulnerabilities, where adversaries exploit flow installation delays without altering or manipulating the intents themselves [21]. To the best of our knowledge, this work is among the first to emphasize context-based malicious intent detection by varying injection strategies across datasets. Although the intent content remains fixed, the positional context−the sequence and spacing of malicious samples−captures adversarial behaviors such as stealthy probing and bursty attacks [22]. Overall, our approach models, simulates, and detects adversarial intent injections with diverse temporal profiles, exposing an underexplored vulnerability and informing potential detection mechanisms [23]. III. Methodology We propose a context-aware detection framework for identifying malicious intent injections in IBN by analyz-
ing sequences of intents and their distribution patterns, capturing adversarial behaviors not observable at the individual intent level. TF-IDF-based feature extraction is used as the foundation for two complementary models: • CNN-based classification (supervised) • AutoEncoder-based reconstruction (one-class) Compared to dense embeddings such as Word2Vec, which require large corpora and may blur discrete policy actions (e.g., treating ALLOW, BYPASS, and ENFORCE as semantically similar despite distinct operational meanings), TF-IDF preserves term distinctiveness, enabling more discriminative modeling of intent features. A 1D CNN is employed due to the sliding-window design with limited context, where detecting short-range dependencies is critical. In contrast, models such as LSTMs target longer temporal dependencies, introduce higher parameter complexity, and are more prone to overfitting under limited data. CNNs, through weight sharing and architectural simplicity, provide an efficient and well-generalizing alternative. To support detection under limited attack knowledge, a Conv1D AutoEncoder is trained on TF-IDF sequences using sliding windows to learn the temporal structure of benign intent flows. During inference, deviations caused by malicious injection patterns (e.g., stealthy or bursty behavior) are identified via reconstruction error. Unlike traditional one-class methods such as Isolation Forest, One-Class SVM, or k-means, which operate on independent samples, the proposed approach captures localized sequential dependencies, improving sensitivity to subtle temporal anomalies while maintaining efficiency and generalizability. A. Supervised Model The intent dataset, consisting of text-based descriptions with each record labeled malicious (1) or safe (0), serves as the input for the supervised algorithm, which follows the key steps outlined below. Step 1: The intents are converted into numerical vectors using TF-IDF vectorization, with a vocabulary size limited to 500 features. This captures term-level importance across the corpus. Step 2: To incorporate contextual information, a sliding window of size six is applied across the TF-IDF matrix to generate input sequences. Each sequence is labeled as malicious if any rule within the window is malicious, following a max-label logic. Step 3: The sequences are split into training (75%) and testing (25%) sets using a fixed random seed (41). Step 4: To mitigate label imbalance, we compute class weights (Wc ) by employing the following equation and apply them during training to give more emphasis to the minority class as Wc = ns /(nc × ntc ) where ns denotes the total number of training samples, nc represents the number of unique classes and ntc denotes number of training samples belonging to the class c.
Step 5: 1D CNN architecture is optimized for the detection of temporal patterns in short intent sequences. Step 6: To further address class imbalance and focus on hard-to-classify samples, we use focal loss with the following formulation: L(ytrue , ypred ) = −α(1 − pt )γ log(pt ) where pt is the predicted probability for the true class, α is the weighting factor to balance classes, and γ is the focusing parameter that down-weights easy examples. Step 7: The CNN is trained using the Adam optimizer and the defined focal loss. Step 8: During inference, an elevated decision threshold is applied to reduce false positives. The model is evaluated using standard performance metrics, including accuracy, precision, recall, and F1-score. B. One-class learning Model This model follows an encoder-decoder architecture and repeats Step 1 to Step 3 from Section III-A. The following three functional features can represent the remaining workflow of this model. Encoder: The encoder employs one-dimensional convolutions to extract localized semantic-temporal patterns from intent sequences, followed by normalization and pooling to stabilize training and reduce dimensionality. A fully connected layer projects the result into a compact latent representation capturing the most salient features. Decoder: The decoder reconstructs the original input from the latent representation using fully connected layers, reshaping the output to the original sequence format with a bounded activation to preserve normalization. Reconstruction loss: During inference, the AutoEncoder computes the mean absolute reconstruction error for each intent window. A detection threshold is set using a highpercentile reconstruction loss from benign windows; windows exceeding this threshold are flagged as anomalous, indicating potential malicious intent injection. IV. Performance Evaluation Alongside a detailed description of the intent dataset, this section presents the results of our experiments conducted on it. We also compare our findings with those reported in [6], which, to the best of our knowledge, is the only existing work closely related to our contribution. We begin by highlighting the key characteristics and structure of the developed dataset. A. Dataset Introduction We construct a dataset of 1,100 network intents (250 malicious, 850 benign), partially assisted by a pre-trained LLM, comparable in scale to the proprietary dataset of 755 intents reported in [6]. Malicious intents are generated under four adversarial conditions. Fig. 2 illustrates the distribution using t-SNE, where intents are encoded via TF-IDF and projected into two dimensions. To systematically generate malicious samples, we curate 20 manually verified base intents covering realistic threat scenarios,