Conceptio › Archive › arXiv CS
arXiv CSopen access

Few-Shot Learning for Network Intrusion Detection: Methods, Datasets, and Performance

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Few-Shot Learning for Network Intrusion Detection: Methods, Datasets, and Performance

arXiv:2609.11275v1 [cs.CR] 10 Sep 2026

Arne Roszeitis, Victor Jüttner, and Erik Buchmann Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI) Dresden/Leipzig, Leipzig University, Leipzig, Germany Email: {arne.roszeitis, victor.juettner, erik.buchmann}@uni-leipzig.de

Abstract. Anomaly-based network intrusion detection systems (NIDS) are an important first line of defense. However, training NIDS for new attack types is challenging, because labeled attack data are rarely available. Few-shot learning (FSL) addresses this problem by learning from few samples. However, the approaches and evaluation settings, that have been investigated so far, vary widely. This work systematically reviews FSL approaches for NIDS published from 2022 to 2026. We conduct a systematic literature review with PRISMA 2020-like reporting to search ACM Digital Library, IEEE Xplore, and Scopus. From a set of 1,358 initial records, we retain 21 studies after screening, deduplication, and quality filtering. We classify the applied FSL approaches, datasets, and experimental parameters and compare reported performance. Meta-learning and convolutional neural networks are the most common approaches, with 8 and 10 studies, respectively. Most studies evaluate five or fewer samples per class, although settings vary. CIC-IDS2017 and CSE-CICIDS2018 are the most frequently used datasets. Missing parameters and source code limit reproducibility and direct comparison between approaches.

Keywords: Network Intrusion Detection · Few-Shot Learning · Review

1

Introduction

Network intrusion detection systems [23] (NIDS) identify attacks that cross network boundaries, complement host-based defenses, and provide control points in heterogeneous environments [3]. As today’s attackers increasingly use generative AI [17] to modify payloads, alter protocol interactions, and tailor attacks to specific systems [2], NIDS must adapt quickly to new attacks. Signature-based detection cannot recognize attacks for which no rule exists, and anomaly-based approaches raise alerts without identifying the attack type [23,3]. Supervised, learning-based NIDS close this gap by generalizing from labeled traffic; however, they require labeled samples of every attack class they are expected to recognize. For novel attacks, such labels are scarce, while benign traffic dominates the available data [15]. Few-shot learning [45] (FSL) addresses

exactly this situation by learning new classes from a small number of labeled samples, making it a natural fit for intrusion detection under label scarcity. Driven by this promise, a growing body of studies has applied FSL to network intrusion detection in recent years. However, the field has not yet converged on common practices. There is no consensus on which FSL techniques or combinations perform best, nor on how they should be evaluated. Moreover, the meaning of “few” remains inconsistent, with recent studies using between one and 20 samples per class [51,53]. Studies differ in the number of samples per class, the number of classes, the datasets used, preprocessing steps, and reported metrics, and reporting practices vary widely. Without such common ground, reported results cannot be compared across studies and reproduction is difficult wherever parameters remain undisclosed. To structure the current landscape, we answer three research questions: RQ1 Which learning techniques are used for few-shot NIDS? RQ2 Which datasets are used to evaluate these approaches? RQ3 How are these approaches evaluated? In this paper, we contribute a systematic literature review following [24] with a PRISMA 2020-like reporting [36] of recent studies on FSL for NIDS. We searched the ACM Digital Library, IEEE Xplore, and Scopus for peer-reviewed studies published between 2022 and 2026. Our search returned 1,358 records, of which we retained 21 studies for detailed analysis after deduplication and filtering according to predefined criteria. For each study, we extract the applied learning technique, datasets, evaluation settings, and reported FSL parameters where available. We organize the studies by their main techniques and compare their reported results where possible. Most reviewed studies combine multiple learning techniques (17 of 21), with CNNs (10 of 21) and meta-learning (8 of 21) being the most common. Among the 17 studies reporting k, k=5 is the most frequent setting. CIC-IDS2017 and CSE-CIC-IDS2018 are the most widely used datasets, appearing in 11 and 6 studies, respectively. Accuracy, precision, recall, and F1-score are the dominant evaluation metrics, with 15 studies reporting F1-scores. However, evaluation settings vary substantially: five studies omit k, four omit the number of classes n, and most provide no accessible source code. These gaps limit reproducibility and direct comparison.

2

Related Work

Intrusion detection has been surveyed extensively. Aldweesh et al. [4] provide a taxonomy of deep-learning approaches for anomaly-based IDS together with open issues. Maseer et al. [31] present a meta review of NIDS, finding deep learning approaches to perform better than traditional machine learning. Also, they criticize the use of out-of-date data sets. More recently, Lansky et al. [25] provide an in-depth review of IDS systems and their deep learning architectures. Ashraf and Masoodi [6] examine over 120 data sets used for intrusion detection

with regard to attack types, temporal features and other categories. Few-shot learning was only once mentioned in [4] and [25]; and not at all in [31]. Few-Shot Learning Techniques Coming from the other angle, few-shot learning itself has been systematized by Wang et al. [45], who categorize FSL techniques by how prior knowledge is used to augment data, constrain the model, or guide the optimization algorithm. Song et al. [40] review over 200 recent FSL studies and demarcate few-shot learning from the related concepts of transfer learning and meta-learning. For the related problem of continuously arriving new classes, Zhou et al. [55] survey class-incremental learning. For the neighboring task of encrypted traffic classification, Yang et al. [50] survey few-shot approaches specifically. These studies are not primarily concerned with network intrusion, however, Yang et al. [50] are very close. Closest Work Even closer to our scope, Duan et al. [13] review few-shot learning for intrusion detection, covering data enrichment, graph embedding, and metalearning approaches published up to 2021. Independent of our systematic search, Winiecki et al. [46] evaluate selected FSL techniques, including a matching network with attention and prototyping, against classical baselines such as kNN and random forests for network intrusion detection. Research Gap In summary, existing surveys either address intrusion detection broadly without a few-shot focus, or, in the case of Duan et al. [13], predate the rapid development of the field after 2021. To the best of our knowledge, no systematic review of few-shot learning for network intrusion detection exists since then. This work closes that gap by systematically reviewing studies published between 2022 and 2026, with a focus on the applied techniques, the datasets, and the evaluation settings, which together determine how comparable the reported results are.

3

Methodology

We conduct a systematic literature review following [24], with a PRISMA 2020like reporting [36]. Our study selection process is illustrated in Figure 1. We queried the ACM Digital Library [1], IEEE Xplore [21], and Scopus [7]. Each query combines intrusion detection and prevention terms (intrusion detection, IDS, intrusion prevention, IDPS ) with few-shot and one-shot learning terms (few shot learning, FSL, one-shot learning, one shot learning, OSL), because OSL is contained in FSL. The broader term learning was included to capture studies not using these terms in the indexed metadata. The full queries are shown in the Appendix.

Identification

Screening

Deduplication

Eligibility

Inclusion

Database search (No=1358) ACM:1034, IEEE:183, Scopus:141

Keyword search & publication criteria (No=162)

After deduplication (No=124)

Full-text eligibility check (No=34)

Full-text articles included (No=21)

Excluded (No=1053)

Excluded (No=90)

Excluded (No=13)

Fig. 1: Our study selection process

The search returned 1034 (ACM), 183 (IEEE), and 141 (Scopus) records. We compared these 1358 studies against the eligibility criteria listed in Table 1. In particular, we automatically filtered titles and abstracts for the terms intrusion, detection, shot, and learning, all of which had to match, leaving 162 studies. After deduplication, 124 studies remained. Of these, 34 were published in a venue meeting our quality criterion, i.e., a journal with an h4 -index of 53 or above as listed in the OOIR database [35] or a conference ranked B or better in the CORE ranking [20]. The h4 -index denotes the maximum number of a journal’s studies with at least that many citations within the last four years.

Table 1: Eligibility criteria for included studies. Criterion

Requirement

Publication period Language & Type Study Keywords Venue Datasets

January 2022 – August 2026 English peer-reviewed research Few-shot or one-shot learning applied to NIDS intrusion, detection, shot, and learning Journal h4 -index ≥ 53 or CORE-ranked conference ≥ B Excluded if ≥ 50% of evaluated datasets predate 2007

After automated filtering, we manually tested the remaining 34 studies for eligibility against the criteria listed in Table 1, resulting in 21 studies being included in the review. Four of the excluded studies evaluated datasets predating 2007, six were not concerned with few-shot learning for network intrusion detection or reported no information about it, one targeted zero-day attacks, which is out of scope, and one did not evaluate its approach conclusively. The remaining exclusion was a review, not a study; while it presents no novel approach itself, we reference it for context [9]. For each included study, we extracted the applied learning techniques, the datasets used for evaluation, the evaluation setting (the number of classes n and samples per class k, where reported), and the reported performance. Most studies

evaluate their approach on more than one dataset and report different subsets of metrics. For studies evaluating multiple settings, we report the performance of the 5-shot setting (k=5), as this is the most commonly evaluated setting. If no 5-shot experiment is reported, we use the study’s best-performing setting. All numbers are truncated after the third decimal place. The full per-dataset results at the selected setting, including accuracy, precision, recall, and F1-score, are given in the Appendix. We structure the review by grouping studies according to their main FSL techniques. The set of techniques was derived from the corpus itself.

Meta-Learning MAML CNN Auto-Encoder FSCIL Twin NN Self-supervised Learning Data generation Attention Graph NN

Table 2: Publication year, number of datasets and learning techniques.

Year

Study

No. Datasets

2026

Martinez-Lopez et al. [30] Wu et al. [47] Yin et al. [52] Jamshidi et al. [22] Zhang et al. [54] Asante and Abass [5] Lu et al. [28]

4 3 2 1 2 1 3

Mao et al. [29] Zhang et al. [53] Xu et al. [48] Xu et al. [49] Qiu et al. [37]

3 3 2 3 5

2024

Tong and Zhang [42] Du et al. [12]

3 2

✓ ✓ ✓✓

✓ ✓

✓ ✓

2023

Ayesha et al. [10] Hang et al. [18] Lu et al. [27] Mirsadeghi et al. [33] Miao et al. [32] Sun et al. [41]

3 5 1 1 3 1

✓ ✓

✓ ✓

✓

2022

Ye et al. [51]

2

Count

21

2025

✓ ✓✓ ✓ ✓

✓

✓✓ ✓ ✓✓

✓ ✓

✓

✓ ✓ ✓

✓

✓ ✓✓ ✓ ✓

✓✓ ✓ ✓ ✓ ✓ ✓ ✓

✓ ✓

✓ ✓ ✓

8 4 10 4 2 3 5 5 5 3

In a first pass over all included studies, we listed every learning technique applied, and then iteratively merged synonymous and subordinate techniques until ten categories remained that together cover every but one included studies. The categories are neither exclusive nor exhaustive. Most studies combine several FSL techniques. Table 2 lists the techniques per study and Figure 2 shows their co-occurrence and usage per year.

4

Review Results

This section provides an overview of our review, lists the datasets and evaluation settings used, and then examines the FSL technique families in more detail. 4.1

Overview

Table 2 is the central artifact of this review: It provides an overview of the 21 included studies, along with the publication date, the learning techniques combined in each and the number of datasets evaluated. Full per-dataset results, including the evaluation parameters n and k, are provided in the Appendix.

8

4

5

4

2

1

10 1

1

4

1

3

1

1

3

4

2

1 3

Meta-learning

3

MAML

1

3

CNN

4

1

3

Auto-encoder

2

2

1

1

FSCIL

1

5

3 5

1

Data generation

5

Attention

2 1

1 1

2

1

1 2

1

2

2

1 2

1

2023

2024 Year

2025

2026

Graph NN 2022

2

M

et

a-

le

ar ni n M g AM Au to L -e CN nc N od Se FS er T CI l f w D -s at up in L a e N ge rv N ne ise r d At ati te o G nt n ra io ph n N N

3

2

Self-supervised

1

4 3

3

1

Twin NN

1

1

Fig. 2: Left: Topic co-occurrences. Diagonal is the total number of occurrences of this topic. Right: Technique usage per year. Figure 2 summarizes the combinations of learning techniques and their temporal distribution. The left panel shows pairwise co-occurrences between techniques; diagonal entries indicate the total number of studies using each technique, while off-diagonal entries show how often two techniques are combined. CNNs (10 studies) and meta-learning (8 studies) are the most common, and their combination occurs most frequently (5 studies). Overall, 17 of the 21 studies combine at least two techniques. The right panel shows the number of studies using each technique per publication year as a heatmap, with darker cells indicating higher

frequencies. CNNs are used throughout the reviewed period, while meta-learning is most frequent in 2026 and graph neural networks appear only in 2025–2026. Figure 2 further shows that approaches for learning from few labeled samples (Meta-learning, MAML, FSCIL, and Twin NN) are more common than datageneration approaches for augmenting limited training data. The temporal distribution also shows a shift in focus. Meta-learning is absent in 2024 and appears only once in 2025 before increasing again in 2026. CNNs are used throughout 2023–2026, while Auto-Encoders appear only in 2023 and 2024 despite the strong performance reported by the respective studies.

4.2

Datasets

To characterize the empirical basis of the reviewed studies, we next examine the datasets used for evaluation. We consider both their frequency of use and their overlap across studies, as common datasets provide the main basis for comparing reported results. Table 3 summarizes how often which dataset was used in a study; the appendix contains all details.

Table 3: Frequency of the datasets used. Dataset

Freq.

CIC-IDS2017

11

CSE-CIC-IDS2018

6

USTC-TFC2016

5

Edge-IIoT, CIC-BoT-IoT

3

FSIDS-IoT, IoT-23, ISCX-IDS2012, NF-CSE-CIC-IDS2018-v2, CIC-ToN-IoT

2

CIC-EVSE2024 Network, CIC-EVSE2024 PowerB, CICIoV2024, Cross-platform, FSIDS-IoT-v2, InSDN, ISCX-Tor2016, ISCX-VPN-2016 (App), ISCX-VPN-2016 (Service), NF-BoT-IoT-v2, NF-ToN-IoT-v2, NF-UNSW-NB15-v2, NSLKDD, UNSW-NB15, WuSTl-EHMS, WuSTl-IIoT

1

CIC-IDS2017 [38] and CSE-CIC-IDS2018 [39] were by far the most popular choices (11 and 6 uses), followed by USTC-TFC2016 [44] and UNSW-NB15 [34] (5 and 1 uses). Additionally, CIC-IDS2017, CSE-CIC-IDS2018 and UNSWNB15 were used in other datasets (FSIDS-IoT(-v2), prefixed NF-[dataset](-v2)), adding between two and five uses depending on the counting method. Each study evaluated a median of three datasets. At the same time, the long tail of single-use datasets means that many results cannot be compared against any other study.

4.3

Evaluation and Performance

We next examine the evaluation settings reported by the included studies. Figure 3 summarizes the n-way and k-shot configurations. Only 16 of the 21 studies explicitly specify k. Of these, 12 evaluate settings with five or fewer samples per class. The most common setting is k=5. The number of classes is typically between five and eight, with values ranging from two to 18. Five studies do not report k and five do not report n, limiting reproducibility and comparability. Fig. 3: Reported evaluation parameters per study. Left: number of classes n; right: samples per class k. Studies not reporting a parameter are omitted. k=5 Ye ’22 [51] Ayesha ’23 [10] Hang ’23 [18] Lu ’23 [27] Mirsadeghi ’23 [33] Miao ’23 [32] Sun ’23 [41] Tong and Zhang ’24 [42] Du ’24 [12] Mao ’25 [29] Zhang ’25 [53] Xu ’25 [48] Xu ’25 [49] Qiu ’25 [37] Martinez-Lopez ’26 [30] Wu ’26 [47] Yin ’26 [52] Jamshidi ’26 [22] Zhang ’26 [54] Asante and Abass ’26 [5] Lu ’26 [28]

no number given

no number given

no number given no number given no number given

no number given

no number given no number given

no number given

0

5

10 15 n (classes)

20

0

5 10 15 20 k (samples per class)

Regarding reported performance, 15 of the 21 studies provide an F1-score. Their best reported F1-scores range from 0.700 to 0.999, and 12 of these 15 studies achieve at least 0.900 on one dataset. Five studies [51,48,37,52,5] report particularly high performance F1 ≈ 0.96, Acc ≈ 0.96, F1 ≈ 0.97, P recision ≈ 0.98, P recision = 0.996). However, these results are difficult to compare directly, even when using the same dataset. Among the six studies reporting an F1-score on CIC-IDS2017, values range from 0.673 to 0.995. Similar variation appears within individual studies: Ayesha et al. [10] report F1-scores between 0.596 and 0.980 across three datasets, while Tong and Zhang [42] report 0.913 on NSL-KDD and 0.997 on

CSE-CIC-IDS2018 without specifying k. These differences indicate that dataset choice and evaluation settings substantially affect the reported performance and limit comparisons between techniques. 4.4

Technique-Level Analysis

This subsection examines the main learning techniques in detail. For each technique family, we introduce its application in few-shot intrusion detection, describe how it is used in the studies, and summarize the reported results. Meta-Learning trains an outer optimization loop over many small tasks so that the resulting model adapts to new classes from only a few labeled samples; the MAML framework [14] is the most common instantiation, used by four of the eight meta-learning studies. Inner learning rates of α = 0.01 and meta-learning rates β = 0.001 where reported. Two MAML studies pair the framework with a CNN as the inner model. Lu et al. [27] convert flows into 3-channel images and optimize the CNN through MAML in a 5-way setting, while Zhang et al. [54] move detection to fog devices and extend MAML with multi-step loss optimization and a learn-to-forget mechanism. Wu et al. [47] pseudo-label formerly unlabeled data via negative learning and use bagging against class imbalance. Easing the few-shot problem itself, Lu et al. [28] generate training samples with a denoising diffusion model [19] enhanced by a drift loss. The remaining meta-learning studies constructed other frameworks: Miao et al. [32] train twin 3D-CNNs to find and handle out-of-distribution classes through prototype distances, Sun et al. [41] combine capsule networks to extract spatial and temporal features, as well as attention-based prototypes, and Xu et al. [49] use bidirectional matching between labels and flows via cosine similarity and a random walk. The results within this family vary widely: where F1-scores are reported, they range from 0.915 [54] to 0.996 [5]. Meanwhile, the MAML-using studies are located in a narrower band between 0.915 and 0.960. Meta-learning is elegant, but appears hard to get right in practice. Convolutional neural networks (CNN) serve as the feature extractor of choice across the included studies: beyond the four studies discussed here, six more use CNNs as FSL technique (see Table 2). Tong and Zhang [42] combine attention, a customized CNN, an Auto-Encoder and an adversarial discriminator into a selfsupervised NIDS. Zhang et al. [53] pre-train and prune a CNN on general traffic. Then, they fine-tune it in a teacher-student twin setup, yielding a lightweight model. Their extensive parameter reporting is a positive example among the included studies. Xu et al. [48] fuse two modalities, namely grayscale images of packets handled by a CNN and statistical flow features handled by a transformer. Self-attention fusion worked best. Learned features are then preserved through transfer learning. In addition to OSL, Mirsadeghi et al. [33] utilize sampling and generation (SMOTE [8], GANs [16]) and weighted random forests in the software defined networking (SDN) domain. Regarding OSL, they use a twin CNN, which was outperformed by the weighted random forest regarding precision on some

classes. CNNs bring in mixed results, but overall better than Meta-Learning approaches. Especially [48] and [52] achieve very good results on popular data sets, making them easily comparable. Graph Neural Networks (Graph NN) represent flows or hosts as nodes and their interactions as edges, allowing models to exploit the topology of attacks. Asante and Abass [5] combine adversarial graph contrastive learning with metalearning on heterogeneous behavior graphs, achieving the best reported precision and F1-score of the included studies (0.996 on IoT-23). Qiu et al. [37] trace the inconsistent outcomes of earlier GNNs such as E-GraphSAGE [26] back to entangled feature distributions in the graph construction and address this with a five-stage model that disentangles graph representations. Invariants are captured by a projection-based few-shot model. Mao et al. [29] distribute contrastive and classification tasks over federated clients sharing only model weights. Graph approaches perform above average, with F1 -score ranging from 0.868 to 0.996. All three evaluate on datasets rarely used by other studies, which limits comparability. Few-Shot Class-Incremental Learning (FSCIL) addresses so-called catastrophic forgetting by adding new classes over time as they appear from few samples. Du et al. [12] reuse the encoder part of an Auto-Encoder trained on flow-based images within a vision-transformer architecture [11] and add session-trained projection layers, fusing them. Yin et al. [52] employ FSL to fine-tune a pre-trained CNN. This FSL is done by constructing a prototypical network using cosine similarity and a CosFace margin [43], mitigating new attack types through an incremental few-shot process with a specialized loss. Although Yin et al. report no F1-score, their precision of 0.978 is among the best in the included studies. Auto-Encoders learn compressed representations of traffic, either for reconstruction"=based pre-training or for generating new samples. Hang et al. [18] pretrain a masked Auto-Encoder on unlabeled traffic and then replace the decoder with a linear classifier for fine-tuning using FSL. Ayesha et al. [10] chain selfsupervised learning with FSL using an Auto-Encoder and contrastive learning and last a kNN classifier. Auto-Encoders also appear as learning technique in two other studies [42,12]. This approach, combined with other techniques seems a promising attempt at few-shot learning. Other Approaches Adapting from the field of natural language processing, Ye et al. [51] propose a text-based workflow based on Dirichlet generative learning, augmenting their few samples with synthetic ones based on semantics. MartinezLopez et al. [30] show that fusing four distance metrics (Chebyshev, Cosine, Euclidean, Wasserstein) outperforms single-metric prototypical few-shot learning. They use these prototypes to generate new datapoints. Jamshidi et al. [22] base their IoT defense on edge devices. They first construct an anomaly-detection, testing various ML techniques. After a detection, an LLM API call is issued, using the LLM to triage the incident. It is the only one among our included studies, that was built around an LLM.

5

Discussion

This review set out to answer three research questions. Regarding RQ1 (learning technique), the field is dominated by meta-learning (8 of 21 studies) and convolutional neural networks (10 of 21), which are rarely used in isolation: 17 of 21 studies combine at least two techniques, most frequently meta-learning with CNNs (5 studies). Regarding RQ2 (datasets), CIC-IDS2017 and CSE-CICIDS2018 have emerged as de facto benchmark datasets (11 and 6 uses), and the reviewed studies evaluate on a median of three datasets each; at the same time, the long tail of single-use datasets means that many studies cannot be compared against any other. Regarding RQ3 (evaluation setting), most studies evaluate five or fewer samples per class, with k = 5 being the most common setting among the studies reporting k, and 15 of 21 studies report an F1 -score. Reporting is incomplete in a relevant share of studies: five omit k, five omit n, and most provide neither accessible source code nor full hyperparameters. 5.1

Interpretation and Implications

Interpretation. Our results support four main observations. First, the most widespread techniques (meta-learning with CNNs or builds on MAML) do not yield the strongest results, whereas graph neural networks and Auto-Encoderbased combinations achieve the highest reported scores. Second, technique adoption appears trend-driven: Auto-Encoders vanish from our set of studies after 2024 despite strong results, and graph neural networks emerge only in 2025–2026. Third, datasets shape reported performance at least as much as techniques on CIC-IDS2017 alone, reported F1 -scores range from 0.673 to 0.995 although the convergence on this benchmark family provides a partial basis for comparison. Implications. For practitioners, CNN-based approaches evaluated on the de facto benchmark datasets currently offer the best comparability and hence the most reliable basis for adoption decisions. For researchers, our findings argue for a shared evaluation protocol: fixed n-way/k-shot settings on common datasets, a defined metric set covering at least accuracy, precision, recall, and F1 -score, and mandatory disclosure of code and hyperparameters. 5.2

Limitations

Our venue-based quality filter (a journal h4 -index of at least 53 according to [35], or a conference ranked B or better in the CORE ranking) was necessary to keep the screening feasible. We acknowledge, however, that citation-based metrics vary considerably between research communities, so relevant work published in newer or lower-ranked venues may have been missed. Screening, eligibility decisions, data extraction, and the design of the search queries were carried out by one researcher. Although all decisions are documented through PRISMA-like reporting, which makes them traceable, a second reviewer would have reduced the subjectivity.

Finally, our synthesis is limited by the primary studies themselves: missing parameters and unavailable source code prevented us from verifying reported numbers, so our performance comparison relies on self-reported results obtained under heterogeneous settings. 5.3

Future Work

Two directions for future work follow directly from our findings. First, a benchmark study that re-implements representative approaches from each technique family and evaluates them under a unified protocol with fixed n and k, common datasets, identical preprocessing, and a full metric set. This would turn the indications observed here into verifiable rankings. Publishing such a benchmark as a community artifact, together with a reporting checklist for few-shot NIDS studies, would directly address the reproducibility gaps identified above. Second, generalization across datasets remains essentially untested: none of our included studies trains and evaluates on disjoint datasets, so the robustness of these approaches to unseen network environments is unknown.

6

Conclusion

Few-shot learning addresses a central obstacle to learning-based intrusion detection: the scarcity of labeled attack data. In this work, we systematically reviewed recent few-shot learning approaches for NIDS using records retrieved from the ACM Digital Library, IEEE Xplore, and Scopus. Starting from 1,358 records published between 2022 and August 2026, we retained 21 peer-reviewed studies and analyzed their learning techniques, datasets, and evaluation settings. Our analysis shows a field dominated by convolutional neural networks and meta-learning, typically combined with other techniques. The datasets CICIDS2017 and CSE-CIC-IDS2018 have emerged as de facto benchmarks and provide a partial basis for comparison. However, heterogeneous evaluation settings, incomplete parameter reporting, and largely unavailable source code still limit direct comparison and independent verification of reported results. For few-shot NIDS to mature into deployable defenses, the community needs shared evaluation protocols and stronger reproducibility practices. We therefore advocate fixed n-way/k-shot settings on common datasets, consistent reporting of evaluation parameters and metrics, and the disclosure of source code and hyperparameters. A unified, openly reproducible benchmark, ideally complemented by cross-dataset evaluation, would be the most valuable next step for assessing both performance and generalization. Data availability statement The data and scripts required to reproduce our results are available at https://speicherwolke.uni-leipzig.de/index.php/s/ AFTL6xKRPaA3BfW.

7

Acknowledgment

The authors acknowledge the financial support by the Federal Ministry of Research, Technology and Space of Germany and by Sächsische Staatsministerium für Wissenschaft, Kultur und Tourismus in the programme Center of Excellence for AI-research „Center for Scalable Data Analytics and Artificial Intelligence Dresden/Leipzig“, project identification number: ScaDS.AI

References 1. ACM: Acm digital library, https://dl.acm.org/, last accessed on Aug 9 2026 2. Ahmed, M., Abdullah, Q.: Network intrusion detection systems: Machine learningbased attack and remedy strategies – a review. Al-Salam Journal for Engineering and Technology (2025), https://api.semanticscholar.org/CorpusId: 280766952 3. Ahmed, M., Mahmood, A., Hu, J.: A survey of network anomaly detection techniques. J. Netw. Comput. Appl. 60, 19–31 (2016), https://doi.org/10.1016/j. jnca.2015.11.016 4. Aldweesh, A., Derhab, A., Emam, A.Z.: Deep learning approaches for anomalybased intrusion detection systems: A survey, taxonomy, and open issues. Knowledge-Based Systems 189, 105124 (2020). https://doi.org/https://doi. org/10.1016/j.knosys.2019.105124 5. Asante, I.O., Abass, F.: Robust few-shot malware detection in iot via adversarial contrastive learning on heterogeneous behavior graphs. IEEE Internet of Things Journal pp. 1–1 (2026). https://doi.org/10.1109/JIOT.2026.3671717 6. Ashraf, W., Masoodi, F.S.: A survey of intrusion detection datasets for communication networks. Discover Computing 29 (2026). https://doi.org/https: //doi.org/10.1007/s10791-026-10486-2 7. B.V., E.: Scopus, https://www.scopus.com, last accessed on Aug 11 2026 8. Chawla, N.V., Bowyer, K.W., Hall, L.O., Kegelmeyer, W.P.: SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16, 321–357 (2002). https://doi.org/10.1613/jair.953 9. Chen, H., Qamar, F., Guo, G., Rehman, M.H.u., Abdullah, S.N.H.S.: Graph neural networks for iomt intrusion detection: Modeling architectures, learning paradigms, and deployment challenges. IEEE Internet of Things Journal 13(16), 38313–38347 (2026). https://doi.org/10.1109/JIOT.2026.3704452 10. Dina, A.S., Siddique, M.A.B., Manivannan, D.: Fs3: Few-shot and self-supervised framework for efficient intrusion detection in internet of things networks. In: Proceedings of the 39th Annual Computer Security Applications Conference. p. 138–149. ACSAC ’23, ACM (2023). https://doi.org/10.1145/3627106.3627193 11. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. In: International Conference on Learning Representations (2021), https://openreview. net/forum?id=YicbFdNTTy 12. Du, L., Gu, Z., Wang, Y., Wang, L., Jia, Y.: A few-shot class-incremental learning method for network intrusion detection. IEEE Transactions on Network and Service Management 21(2), 2389–2401 (2024). https://doi.org/10.1109/TNSM. 2023.3332284

13. Duan, R., Li, D., Tong, Q., Yang, T., Liu, X., Liu, X.: A survey of few-shot learning: An effective method for intrusion detection. Security and Communication Networks 2021(1), 4259629 (2021). https://doi.org/https://doi.org/10. 1155/2021/4259629 14. Finn, C., Abbeel, P., Levine, S.: Model-agnostic meta-learning for fast adaptation of deep networks. In: Proceedings of the 34th International Conference on Machine Learning - Volume 70. p. 1126–1135. ICML’17, PLMR (2017), https: //proceedings.mlr.press/v70/finn17a.html 15. Garg, A.: A review of advancements in deep learning approaches for intrusion detection systems. Journal on Artificial Intelligence (2026), https://api. semanticscholar.org/CorpusId:289371002 16. Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: Advances in Neural Information Processing Systems. vol. 27. Curran Associates, Inc. (2014), https://proceedings.neurips.cc/paper_files/paper/2014/file/ f033ed80deb0234979a61f95710dbe25-Paper.pdf 17. Gupta, M., Akiri, C., Aryal, K., Parker, E., Praharaj, L.: From chatgpt to threatgpt: Impact of generative ai in cybersecurity and privacy. IEEE access 11, 80218– 80245 (2023). https://doi.org/10.1109/ACCESS.2023.3300381 18. Hang, Z., Lu, Y., Wang, Y., Xie, Y.: Flow-mae: Leveraging masked autoencoder for accurate, efficient and robust malicious traffic classification. In: Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses. p. 297–314. RAID ’23, ACM (2023). https://doi.org/10.1145/3607199.3607206 19. Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. In: Advances in Neural Information Processing Systems. vol. 33, pp. 6840–6851. Curran Associates, Inc. (2020), https://proceedings.neurips.cc/paper_files/paper/2020/ file/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf 20. ICORE: Icore conference rankings portal, https://portal.core.edu.au/ conf-ranks/, last accessed on Aug 5, 2026 21. IEEE: Ieee xplore, https://ieeexplore.ieee.org/Xplore/home.jsp, last accessed on Aug 10 2026 22. Jamshidi, S., Wahab, O.A., Herrero, R., Khomh, F., Bellaïche, M., Keivanpour, S., Shahabi, N., Nikanjam, A., Nafi, K.W.: Think fast: Real-time iot intrusion reasoning using ids and llms at the edge gateway. IEEE Internet of Things Journal 13(8), 15485–15513 (2026). https://doi.org/10.1109/JIOT.2026.3656738 23. Khraisat, A., Gondal, I., Vamplew, P., Kamruzzaman, J.: Survey of intrusion detection systems: techniques, datasets and challenges. Cybersecurity 2 (2019), https://doi.org/10.1186/s42400-019-0038-7 24. Kitchenham, B., Charters, S.: Guidelines for performing systematic literature reviews in software engineering. EBSE Technical Report EBSE-2007-01, Keele University and University of Durham (2007), https://legacyfileshare.elsevier. com/promis_misc/525444systematicreviewsguide.pdf, version 2.3 25. Lansky, J., Ali, S., Mohammadi, M., Majeed, M.K., Karim, S.H.T., Rashidi, S., Hosseinzadeh, M., Rahmani, A.M.: Deep learning-based intrusion detection systems: A systematic review. IEEE Access 9, 101574–101599 (2021). https: //doi.org/10.1109/ACCESS.2021.3097247 26. Lo, W.W., Layeghy, S., Sarhan, M., Gallagher, M., Portmann, M.: E-graphsage: A graph neural network based intrusion detection system for iot. In: NOMS 2022-2022 IEEE/IFIP Network Operations and Management Symposium. pp. 1–9 (2022). https://doi.org/10.1109/NOMS54207.2022.9789878

27. Lu, C., Wang, X., Yang, A., Liu, Y., Dong, Z.: A few-shot-based model-agnostic meta-learning for intrusion detection in security of internet of things. IEEE Internet of Things Journal 10(24), 21309–21321 (2023). https://doi.org/10.1109/JIOT. 2023.3283408 28. Lu, Y., Lu, Y., Wang, Z., Wu, W., Ke, C., Liu, S.: D-maml: A few-shot learning algorithm for intrusion detection based on meta-learning and diffusion models. IEEE Internet of Things Journal (2026). https://doi.org/10.1109/JIOT.2026. 3674767 29. Mao, Q., Lin, X., Xu, W., Qi, Y., Su, X., Li, G., Li, J.: Fecograph: Label-aware federated graph contrastive learning for few-shot network intrusion detection. IEEE Transactions on Information Forensics and Security 20, 2266–2280 (2025). https: //doi.org/10.1109/TIFS.2025.3541890 30. Martinez-Lopez, F., Santana, L., Rahouti, M., Chehri, A., Al-Maliki, S., Jeon, G.: Learning in multiple spaces: Prototypical few-shot learning with metric fusion for next-generation network security. IEEE Transactions on Network and Service Management 23, 3156–3165 (2026). https://doi.org/10.1109/TNSM.2026.3665647 31. Maseer, Z.K., Kadhim, Q.K., Al-Bander, B., Yusof, R., Saif, A.: Meta-analysis and systematic review for anomaly network intrusion detection systems: Detection methods, dataset, validation methodology, and challenges. IET Networks 13(5-6), 339–376 (2024). https://doi.org/https://doi.org/10.1049/ntw2.12128 32. Miao, G., Wu, G., Zhang, Z., Tong, Y., Lu, B.: Spn: A method of few-shot traffic classification with out-of-distribution detection based on siamese prototypical network. IEEE Access 11, 114403–114414 (2023). https://doi.org/10.1109/ ACCESS.2023.3325065 33. Mirsadeghi, S.M.H., Bahsi, H., Vaarandi, R., Inoubli, W.: Learning from few cyberattacks: Addressing the class imbalance problem in machine learning-based intrusion detection in software-defined networking. IEEE Access 11, 140428–140442 (2023). https://doi.org/10.1109/ACCESS.2023.3341755 34. Moustafa, N., Slay, J.: UNSW-NB15: a comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set). In: 2015 Military Communications and Information Systems Conference (MilCIS). pp. 1–6 (2015). https://doi.org/10.1109/MilCIS.2015.7348942 35. Pacher, A.: The observatory of international research, https://ooir.org, last accessed on Aug 5, 2026 36. Page, M.J., McKenzie, J.E., Bossuyt, P.M., Boutron, I., Hoffmann, T.C., Mulrow, C.D., Shamseer, L., Tetzlaff, J.M., Akl, E.A., Brennan, S.E., Chou, R., Glanville, J., Grimshaw, J.M., Hróbjartsson, A., Lalu, M.M., Li, T., Loder, E.W., MayoWilson, E., McDonald, S., McGuinness, L.A., Stewart, L.A., Thomas, J., Tricco, A.C., Welch, V.A., Whiting, P., Moher, D.: The prisma 2020 statement: an updated guideline for reporting systematic reviews. BMJ 372 (2021). https://doi.org/10. 1136/bmj.n71 37. Qiu, C., Nan, G., Xia, H., Weng, Z., Wang, X., Shen, M., Tao, X., Liu, J.: Disentangled dynamic intrusion detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 47(11), 10528–10545 (2025). https://doi.org/10.1109/TPAMI. 2025.3595671 38. Sharafaldin, I., Lashkari, A.H., Ghorbani, A.A.: Toward generating a new intrusion detection dataset and intrusion traffic characterization. In: International Conference on Information Systems Security and Privacy (2018), https://www.unb.ca/ cic/datasets/ids-2017.html

39. Sharafaldin, I., Lashkari, A.H., Ghorbani, A.A.: Toward generating a new intrusion detection dataset and intrusion traffic characterization. In: International Conference on Information Systems Security and Privacy (2018), https://registry. opendata.aws/cse-cic-ids2018/ 40. Song, Y., Wang, T., Cai, P., Mondal, S.K., Sahoo, J.P.: A comprehensive survey of few-shot learning: Evolution, applications, challenges, and opportunities. ACM Computing Surveys 55(13s), 1–40 (2023). https://doi.org/10.1145/3582688, https://doi.org/10.1145/3582688 41. Sun, H., Wan, L., Liu, M., Wang, B.: Few-shot network intrusion detection based on prototypical capsule network with attention mechanism. PLoS ONE 18(4 April) (2023). https://doi.org/10.1371/journal.pone.0284632 42. Tong, J., Zhang, Y.: A real-time label-free self-supervised deep learning intrusion detection for handling new type and few-shot attacks in iot networks. IEEE Internet of Things Journal 11(19), 30769–30786 (2024). https://doi.org/10.1109/JIOT. 2024.3414492 43. Wang, H., Wang, Y., Zhou, Z., Ji, X., Gong, D., Zhou, J., Li, Z., Liu, W.: Cosface: Large margin cosine loss for deep face recognition. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 5265–5274 (2018). https://doi. org/10.1109/CVPR.2018.00552 44. Wang, W., Zhu, M., Zeng, X., Ye, X., Sheng, Y.: Malware traffic classification using convolutional neural network for representation learning. In: 2017 International Conference on Information Networking (ICOIN). pp. 712–717 (2017). https:// doi.org/10.1109/ICOIN.2017.7899588 45. Wang, Y., Yao, Q., Kwok, J.T., Ni, L.M.: Generalizing from a few examples: A survey on few-shot learning. ACM computing surveys (csur) 53(3), 1–34 (2020). https://doi.org/10.1145/3386252 46. Winiecki, E., Pawlicki, M., Pawlicka, A., Kozik, R., Choraś, M.: Evaluation of selected few-shot learning methods in network intrusion detection. In: Advanced Information Networking and Applications. pp. 10–20. Springer Nature Switzerland (2025). https://doi.org/10.1007/978-3-031-87778-0_2 47. Wu, M., Zheng, Y., Yang, Y., Luo, J., Wong, D.S.H.: A semi-supervised metanegative-learning approach to few-shot network intrusion detection in internet of things. IEEE Internet of Things Journal 13(7), 12974–12987 (2026). https://doi. org/10.1109/JIOT.2026.3653672 48. Xu, C., Zhan, Y., Wang, Z., Yang, J.: Multimodal fusion based few-shot network intrusion detection system. Scientific Reports 15(1) (2025). https://doi.org/10. 1038/s41598-025-05217-4 49. Xu, C., Zhang, F., Yang, Z., Zhou, Z., Zheng, Y.: A few-shot network intrusion detection method based on mutual centralized learning. Scientific Reports 15(1) (2025). https://doi.org/10.1038/s41598-025-93185-0 50. Yang, C., Gu, Z., Bai, J., Li, Z., Xiong, G., Gou, G., Yao, S., Chen, X.: Fewshot encrypted traffic classification: A survey. In: 2024 Asia-Pacific Conference on Image Processing, Electronics and Computers (IPEC). pp. 646–652 (2024). https://doi.org/10.1109/IPEC61310.2024.00115 51. Ye, T., Li, G., Ahmad, I., Zhang, C., Lin, X., Li, J.: Flag: Few-shot latent dirichlet generative learning for semantic-aware traffic detection. IEEE Transactions on Network and Service Management 19(1), 73–88 (2022). https://doi.org/10.1109/ TNSM.2021.3131266 52. Yin, Z., Wang, W., Cheng, Y., Liu, Y., Lu, X.: Efficient intrusion detection for edge network via multi-stage few-shot class-incremental learning. IEEE Transactions

on Network Science and Engineering 13, 4689–4706 (2026). https://doi.org/10. 1109/TNSE.2025.3644385 53. Zhang, X., Wang, Y., Han, G., Gui, G.: Advanced few-shot network intrusion detection method using lightweight transfer learning. IEEE Internet of Things Journal 12(22), 48678–48688 (2025). https://doi.org/10.1109/JIOT.2025.3606497 54. Zhang, Y., Liu, Y., Wu, Y., Zhao, S., Liu, Z.: Maml-samiot: A cloud-fog computing and maml-based few-shot iot intrusion detection model. Computer Networks 275 (2026). https://doi.org/10.1016/j.comnet.2025.111884 55. Zhou, D.W., Wang, Q.W., Qi, Z.H., Ye, H.J., Zhan, D.C., Liu, Z.: Class-incremental learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 46(12), 9851–9873 (2024). https://doi.org/10.1109/TPAMI.2024.3429383

Appendix Table 4: Database-specific search queries. Source

Query

ACM

[[All: ids] OR [All: intrusion detection] OR [All: intrusion prevention] OR [All: idps]] AND [[All: learning] OR [All: few shot learning] OR [All: fsl] OR [All: one-shot learning] OR [All: one shot learning] OR [All: osl]] AND [E-Publication Date: (01/01/2022 TO 08/31/2026)]

IEEE Xplore

(("All Metadata":intrusion detection) OR ("All Metadata":ids) OR ("All Metadata":intrusion prevention) OR ("All Metadata":idps)) AND (("All Metadata":learning) OR ("All Metadata":few shot learning) OR ("All Metadata":one shot learning) OR ("All Metadata":fsl) OR ("All Metadata":one-shot learning) OR ("All Metadata":osl))

Scopus

(TITLE-ABS-KEY("intrusion detection") OR TITLE-ABS-KEY("ids") OR TITLE-ABS-KEY("intrusion prevention") OR TITLE-ABS-KEY("idps")) AND (TITLE-ABS-KEY("few shot learning") OR TITLE-ABS-KEY("one shot learning") OR TITLE-ABS-KEY("fsl") OR TITLE-ABS-KEY("osl") OR TITLE-ABS-KEY("learning") OR TITLE-ABS-KEY("one-shot-learning")) AND LIMIT-TO(LANGUAGE, "English") AND PUBYEAR > 2022 AND PUBYEAR < 2026

Table 5: Datasets, parameters and results. Results averaged over all datasets of a study are marked with * and balanced accuracy with **. Macro-averaged scores are marked with an indicative (M). Year Author

Dataset

n

k

2022

Ye et al. [51]

USTC-TFC2016 CIC-IDS2017

18

1,10

Ayesha et al. [10]

WuSTl-EHMS WuSTl-IIoT CIC-BoT-IoT

2 5 5

5,10 5,10 5,10

Hang et al. [18]

CSE-CIC-IDS2018 USTC-TFC2016 ISCX-VPN-2016 (App) ISCX-VPN-2016 (Service) ISCX-Tor-2016 Cross-platform

Lu et al. [27]

FSIDS-IoT

5

1,5,10

Mirsadeghi et al. [33]

InSDN

8

1

Miao et al. [32]

ISCX-IDS2012 CIC-IDS2017 USTC-TFC2016

Sun et al. [41]

CSE-CIC-IDS2018

Tong and Zhang [42]

NSL-KDD CIC-IDS2017 CSE-CIC-IDS2018

2023

2024

Accuracy Precision

Recall

F1-Score 0.951(M) 0.962(M)

0.999 0.998 0.998 0.991 0.994 0.992

0.979 0.889 0.619

0.981 0.680 0.603

0.980 0.701 0.596

0.999 0.998 0.999 0.992 0.994 0.992

0.999 0.998 0.998 0.991 0.994 0.992

0.999 0.998 0.999 0.991 0.994 0.992

0.896 0.700(M) 0.983 0.996 0.991

0.924(M) 0.937(M) 0.924(M) 0.965(M) 0.966(M) 0.965(M)

8

0.952

0.988(M)

5 7 7

0.923 0.997 0.996

1,3,5,10

0.940 0.993 0.997

0.887 0.997 0.997

0.913 0.995 0.997

Continued on next page

Year Author

Dataset

n

k

Du et al. [12]

CIC-IDS2017 CSE-CIC-IDS2018

12 14

5

Mao et al. [29]

NF-BoT-IoT-v2 NF-ToN-IoT-v2 NF-CSE-CIC-IDS2018-v2

5 10 7

Zhang et al. [53]

Edge-IIoT CIC-BoT-IoT USTC-TFC2016

5

1,5,10,20

0.922*

Xu et al. [48]

CIC-IDS2017 CSE-CIC-IDS2018

2,3,4

5,10,15

0.934 0.985

Xu et al. [49]

ISCX-IDS2012 CIC-IDS2017 CSE-CIC-IDS2018

2,4,5

5,10

0.882 0.922 0.887

Qiu et al. [37]

CIC-ToN-IoT CIC-BoT-IoT Edge-IIoT NF-UNSW-NB15-v2 NF-CSE-CIC-IDS2018-v2

2025

2026

CIC-IDS2017 Edge-IIoT IoT-23

5

Recall

F1-Score

0.931 0.901

0.931 0.909

0.931 0.901

0.929 0.897

0.984 0.943 0.995

0.980 0.742 0.844

0.799 0.724 0.790

0.868 0.733 0.813

0.883 0.922 0.884 0.961 0.982 0.968 0.954 0.963

5

CIC-IDS2017 CIC-EVSE2024 Network Martinez-Lopez et al. [30] CIC-EVSE2024 PowerB CIC-IoV2024 Wu et al. [47]

Accuracy Precision

2

0.880** 0.820** 0.738** 0.887**

5,10

0.743 0.962 0.840

0.673 0.850 0.552 0.754

Continued on next page

Year Author

Dataset

n

k

Yin et al. [52]

CIC-IDS2017 USTC-TFC2016

5

5

Jamshidi et al. [22]

CIC-IDS2017

15

Zhang et al. [54]

FSIDS-IoT FSIDS-IoT-v2

5

1,5,10

Asante and Abass [5]

IoT-23

5

Lu et al. [28]

CIC-IDS2017 UNSW-NB15 CIC-ToN_IoT

2,5

Accuracy Precision 0.998 0.983

Recall

F1-Score

0.978 0.918

0.981 0.903

0.915 0.905

0.914 0.905

0.915 0.905

0.915 0.905

1,5

0.997

0.996

0.997

0.996

5

0.902 0.967 0.969

0.892 0.960 0.963

0.875 0.954 0.958

0.883 0.957 0.960

0.989

Record · ID 673411 · SHA-256 2232145b1776467d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.