Conceptio › Archive › arXiv CS
arXiv CSopen access

A GAN-Based Framework for Robust DDoS Attack Detection

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

A GAN-Based Framework for Robust DDoS Attack Detection Makram Chehayeb1 , Walid Fahs1,2 , Amina Rizk2 , Rida Khatoun3 , Omran Berjawi3 1

Islamic University of Lebanon, Faculty of Engineering, Lebanon 2 Lebanese University, Faculty of Sciences, Lebanon 3 Polytechnic Institute of Paris, Télécom Paris, LTCI, France

arXiv:2609.18281v1 [cs.LG] 16 Sep 2026

[email protected], [email protected], [email protected] [email protected], [email protected]

Abstract—The availability and consistency of online services remain vulnerable due to Distributed Denial of Service (DDoS) attacks. These attacks are evolving by adopting more complex strategies to evade traditional network security systems. Despite the effectiveness of machine learning models in detecting DDoS traffic, targeted adversarial attacks can degrade their classification accuracy. This work proposes a robust detection framework that integrates generative adversarial modelling with advanced machine learning models. We trained Random Forests, Deep Neural Ensembles, and Transformer-based models using the CICDDoS2019 dataset to establish the framework’s baseline performance. To enhance the models’ defensive capacity, we generated synthetic adversarial flows that simulate potential evasion attempts and adversarial traffic using a Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP). Then, we combined the generated traffic with benign and malicious traffic to construct hybrid datasets to train the models to learn more generalizable decision boundaries. The experimental results indicate that the proposed methodology significantly enhances detection accuracy and resilience, especially against unseen adversarial traffic. We also tested the designed framework using real-world generated traffic, which demonstrates its capability in practical settings. The scalable and efficient solution against adversarial DDoS attacks, introduced in this work, paves the way towards more resilient and adaptive network defense systems that combine generative adversarial augmentation with recent advances in learning models. Index Terms—Distributed Denial of Service (DDoS), Adversarial Machine Learning, WGAN-GP, Robust Detection, Network Security.

cloud-based scrubbing services [3]. This market is currently led by several key players, including Cloudflare, Akamai, NETSCOUT (Arbor), Radware, and AWS Shield. These providers employ a mixture of techniques, including rate limiting, behavioral analytics, and Artificial Intelligence (AI), to create a multi-level defense. Despite these defensive mechanisms, a significant gap persists: adversarial attacks engineered to bypass AI-based detection engines [4]. Figure 1 illustrates a typical AI-assisted DDoS mitigation workflow in which malicious traffic generated by distributed botnets is filtered through intelligent scrubbing infrastructure before reaching the protected server.

I. I NTRODUCTION Nowadays, the increase in services being provided via cloud environments, the number of Internet of Things (IoT) devices connected to networks, and the availability of larger botnets allow attackers to conduct some Distributed Denial of Service (DDoS) campaigns that are larger and more complex, defeating or bypassing many legacy defenses. These attacks are critical threats that have critical economic consequences, and they highlight the vulnerabilities of online services that rely on consistent service uptime, such as remote medical telesurveillance and smart home health monitoring. [1], [2]. Due to the complexity and rising scale of modern cyberattacks, the DDoS protection domain has experienced a considerable expansion recently. These defensive solutions range from local physical appliances (on-premise) to massive

Fig. 1: General architecture of AI-assisted DDoS mitigation using traffic scrubbing centers and intelligent filtering. To overcome these challenges, machine learning-based Intrusion Detection Systems (IDS) have been widely adopted due to their capacity to map complex traffic patterns and detect anomalies with higher accuracy than static, rule-based systems [5]. However, traditional IDS do not generalize well to new attack variants, and modern machine learning models are highly susceptible to adversarial evasion attempts [6]. Moreover, attackers can alter some features of the malicious traffic to keep the original malicious intent while changing the label to be interpreted as benign traffic by classification algorithms [7]. This constitutes a major design flaw because

© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

systems that are created to detect anomalies have become targets for anomalies, which is alarming, as future use would seem unlikely. These challenges motivate the design of detection frameworks that are immune to adversarial attacks [8]. In this context, this study aims to prove that the integration of an adversarial approach as an augmentation strategy of datasets can enhance the robustness of detection systems. In this study, we use Wasserstein Generative Adversarial Networks with Gradient Penalty (WGAN-GP) to generate realistic adversarial DDoS traffic. Unlike standard GANs, WGAN-GP generates high-quality synthetic samples possessing improved training stability. WGAN-GP allows fine-tuning of adversarial traffic that replicates real attack traffic while remaining diverse. These flows are then appended to the training of machine learning models to harden them against evasion. The primary contributions of this work include: • Comprehensive Vulnerability & Hardening Benchmark: We quantify the extreme fragility of clean-trained RF, DNE, and Transformer architectures under featurespace evasion (Adv-5, Adv-9) and evaluate their postaugmentation recovery. • Protocol-Aware Feature Realizability Analysis: We provide a rigorous qualitative and structural discussion on feature dependency constraints, examining how flow-level perturbations map to actual protocol behaviors. • Multi-Phase Live Deployment Simulation: We evaluate model performance across a three-phase operational cycle (Before/Normal, During/Attack Peak, After/Recovery) to highlight practical failure modes like alert fatigue and conservative under-detection. • System Resource & Operational Profiling: We provide concrete latency, throughput, and memory consumption figures, establishing the practical feasibility of deploying Transformer-based detection at the network edge. This paper is structured as follows: Section II reviews the literature on DDoS detection, adversarial learning, and generative approaches. The methodology, in which we present the datasets and the detection models used, is presented in Section III. The experimental setup and the results covering both baseline and enhanced models are discussed in Section IV. Finally, the summary of findings, practical implications, and future research directions are presented in Section V. II. L ITERATURE R EVIEW DDoS detection research has evolved from static rule-based mechanisms toward deep learning models and adversarial defense frameworks. Classical machine learning classifiers— such as Decision Trees, Naı̈ve Bayes, Support Vector Machines (SVM), and Random Forests—achieve high detection accuracy (> 99%) on benchmark datasets like CAIDA 2007 and CICDDoS2019 under clean conditions [9]. However, these classifiers experience severe performance degradation when exposed to subtle feature perturbations. Nevertheless, their effectiveness was limited to clean traffic and failed under adaptive adversarial conditions. To address

this gap, Ali Mustapha et al. [10] introduced an adversarial neural network framework that uses Generative Adversarial Networks (GANs) to create synthetic, malicious traffic flows to train and harden the detection models. Using a Long ShortTerm Memory (LSTM) network as a baseline, this framework achieved up to 100% accuracy and maintained an F1-score of around 0.99 on real traffic. Even when subjected to adversarial evasion attacks, the system maintained high detection rates (between 91.7% and 100%). Chen et al. [11] developed a hybrid architecture that combines a Deep Belief Network (DBN) and a Long Short-Term Memory (LSTM) network. This architecture was specifically designed for Software-Defined Networking (SDN) environments. It achieved high performance metrics: 96.55% accuracy, 98.53% recall, and a 97.47% F1-score when tested on the CICDDoS2019 dataset. This hybrid framework proved robust also against perturbation attacks, maintaining over 91% detection rates under Fast Gradient Sign Method (FGSM) adversarial testing. Shieh et al. [12] developed the Symmetric Defense GAN (SDGAN), combining adversarial attack generation with defensive classification to achieve high recall (0.977 on CIC-IDS2018) while maintaining an AUC above 0.70 under attack. Similarly, Melo et al. [13] proposed “Anomaly-Flow,” a multi-domain federated GAN framework achieving ROCAUC values between 0.74 and 0.96 across heterogeneous IoT datasets. Additional advancements in robust network intrusion detection, ensemble learning, and edge-assisted threat mitigation have further demonstrated the necessity of continuous model adaptation under dynamic attack vectors [14]–[16]. However, it increased network communication overhead. So, existing literature demonstrates that adversarial training, particularly GAN-based data augmentation, can effectively enhance the resilience of detection models against adaptive cyberattacks. Nevertheless, model testing is mainly restricted to offline and static environments, while real-time testing is rarely addressed. Moreover, there is a lack of research addressing the practical deployment and operational integration of these models into large-scale, ISP-level network defenses. Current DDoS mitigation is deployed across multiple infrastructure layers to ensure comprehensive coverage. At the network edge, Internet Service Providers (ISPs) and transit providers implement upstream filtering and BGP blackholing to discard malicious traffic. High-volume attacks are typically redirected to dedicated scrubbing centers for deep inspection. Cloud-based platforms offer either ”always-on” or ”ondemand” scrubbing, providing the elastic scalability needed for massive volumetric floods. For localized control, enterprises utilize on-premise hardware appliances (e.g., Radware DefensePro, Fortinet FortiDDoS) to achieve real-time detection with minimal latency. Many organizations now adopt hybrid models, combining local appliances with cloud scrubbing for peak-load scalability. Furthermore, defense is integrated at the application level via Web Application Firewalls (WAFs) and behavioral analysis. Ultimately, effective mitigation relies on collaborative strategies between Autonomous System (AS) operators and victim

organizations to coordinate filtering and threat intelligence. Despite these developments, a critical comparative synthesis reveals three persistent limitations in state-of-the-art (SOTA) generative defense frameworks: 1) Synthetic Generation vs. Feature Realizability: Most GAN-based augmentation models (e.g., GADoT [17], SDGAN [12], Mustapha et al. [10]) modify numerical feature vectors directly in mathematical latent space without validating whether the resulting vectors satisfy networking logic (e.g., packet rate vs. byte count consistency). 2) Static vs. Multi-Phase Deployment Testing: Prior works overwhelmingly evaluate trained models on static, offline test partitions [4], [10], [12]. They fail to observe model behavior under operational shift across sequential traffic phases (e.g., normal baseline traffic transition into peak attack and post-attack cleanup phases). 3) Omission of Operational Overhead Metrics: While complex deep learning ensembles and GAN pipelines report high detection scores, they routinely omit inference latency, memory footprint, and throughput figures—metrics critical for high-speed edge hardware or ISP scrubbing center deployment [3], [11], [13]. To overcome these challenges, this study adopts an adversarial training framework that has the following characteristics: (i) it uses WGAN-GP to synthesize high-diversity evasive flows; (ii) it evaluates baseline ML/DL and ensemble models under both adversarial and non-adversarial regimes; and (iii) it investigates generalization and feasibility in realistic network conditions. By combining robust data augmentation with realworld testing, we move beyond simple offline accuracy to build a defense system that is truly resilient in live environments. III. M ETHODOLOGY This section presents the proposed framework as illustrated in Fig. 2. The CICDDoS2019 dataset is supplemented with synthetic adversarial flows generated by WGAN-GP. These flows are combined with benign and malicious traffic to create robust hybrid datasets for training three models: The Random Forest (RF), Deep Ensemble (DNE), and Transformer (TF). The models were evaluated under both normal and adversarial conditions, including a simulated real-time deployment. A. DDoS attacks model We model a DDoS attack against a server as a multi-stage queuing network, where traffic originates from a botnet and traverses a series of routers within the ISP infrastructure before reaching the target server. The network is represented as a directed graph G = (N , L), where N denotes the set of nodes and L the set of links with finite bandwidth capacities Cl (in Mbps) and propagation delays dl (in ms). Each router n ∈ N (last ISP) is modeled as a finite-buffer queue with a First-InFirst-Out (FIFO) discipline, characterized by: • Arrival process: legitimate traffic follows a Poisson process with rate λlegit,n , while attack traffic is modeled as a

Markov Modulated Poisson Process (MMPP) to capture burstiness, with state transition matrix Q and arrival rates Λ = [Λ1 , Λ2 , . . . , Λm ] for m attack intensity states. • Service process: the service time at router n is exponenn , where Lpkt is the tially distributed with rate µn = LCpkt average packet size (in bits). The service rate is capped by the link capacity Cn . • Queue dynamics: each router has a buffer of size Kn packets. The queue evolution is governed by the balance equations for an M/M/1/K system, with steady-state probabilities πn,j for j ∈ [0, Kn ] packets in the queue. The probability of packet drops due to buffer overflow is given by: Kn    λtotal,n λtotal,n 1 − µn µn , (1) Pdrop,n = πn,Kn = Kn +1  λ 1 − total,n µn where λtotal,n = λlegit,n + λattack,n is the aggregate arrival rate at router n. • Transmission errors: packets are subject to transmission errors with probability pe , modeled as a Bernoulli process independent of the queue state. The last ISP router (denoted as nlast ) is augmented with an AI-based DDoS detector to classify incoming traffic as either legitimate or malicious. The detector operates as follows: • Feature extraction: for each incoming packet flow, a feature vector xi = [f1 , f2 , . . . , fd ] is extracted, where fj represents features such as: – Packet inter-arrival time. – Packet size and entropy. – Source/destination IP reputation scores. – Flow duration and connection rate. • Classification model: a supervised machine learning model (e.g., Random Forest (RF), Deep Ensemble Neural Network (DNE), and Transformer (TF)) is trained offline on labeled datasets of normal and attack traffic. The model outputs a probability pmalicious ∈ [0, 1] that a flow is part of a DDoS attack. • Mitigation action: if pmalicious > θ (where θ ∈ [0, 1] is a tunable threshold), the flow is classified as malicious and discarded. Otherwise, it is forwarded to the next hop. The detector introduces an additional processing delay Ddetect (in ms) for classification. The detection model is periodically retrained to adapt to evolving attack patterns, with a false positive rate F P R and false negative rate F N R. B. CICDDoS2019 Dataset The primary dataset used in this work is CICDDoS2019 [18]. It contains over 80 flow-based features extracted via CICFlowMeter, and it represents a mix of benign traffic along with several DDoS attacks (UDP Flood, HTTP Flood, SYN Flood, DNS Flood, MSSQL, LDAP, NTP, NetBIOS, SSDP, UDP-Lag, and WebDDoS attacks). The dataset is imbalanced

TABLE I: Comparative Analysis of DDoS Protection Architectures and Deployment Models Solution Edge / ISP Filtering Scrubbing Centers Cloud Protection On-Premise Appliances Hybrid Models Application-Level Defense AI/ML Detection Proposed Framework

Deployment ISP edge Cloud / ISP Cloud Enterprise network Local + Cloud Servers / WAF IDS / Monitoring systems IDS / ISP / Cloud

Strength Early scalable filtering Handles volumetric attacks Highly scalable Low latency control Scalability + visibility Detects Layer-7 attacks Detects unknown attacks Robust against evasive attacks

Limitation Weak against Layer-7 attacks Adds latency Third-party dependency Limited scalability Complex management Resource intensive Vulnerable to adversarial traffic Retraining overhead

Fig. 2: DDoS Attack Detection Framework Integrating Generative Adversarial Modeling with Advanced Machine Learning

with more attack instances than benign ones. First, we preprocessed the data by removing non-informative columns, imputing missing values, and standardizing all features. Second, we performed feature selection using the ANOVA F-test (SelectKBest), and we retained the top 20 most discriminative features for all experiments. C. Adversarial Datasets To generate adversarial attack traffic, we use a Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP). The WGAN-GP framework is composed of two neural networks engaged in a minimax game: a Generator G and a Critic C. The standard WGAN objective uses the Wasserstein distance (Earth-Mover distance) to measure the difference between real and generated distributions: min max Ex∼Pr [C(x)] − Ex̃∼Pg [C(x̃)] G

C∈L

( (i) xadv =

(j)

xbenign x̃(i)

if i ∈ Fmod otherwise

(2)

(3)

where Pr is the real data distribution (attack flows from CICDDoS2019), Pg is the generator’s distribution, and L is

the set of 1-Lipschitz functions. To enforce the Lipschitz constraint, WGAN-GP introduces a gradient penalty term: λEx̂∼Px̂ [(∥∇x̂ C(x̂)∥2 − 1)2 ]

(4)

where x̂ is a random interpolation between real and generated samples: x̂ = ϵx + (1 − ϵ)x̃ with ϵ ∼ U [0, 1], and λ is the penalty coefficient (set to 10 in our experiments). The complete loss functions are: LC = Ex̃∼Pg [C(x̃)]−Ex∼Pr [C(x)]+λEx̂∼Px̂ [(∥∇x̂ C(x̂)∥2 −1)2 ] (5) LG = −Ex̃∼Pg [C(x̃)]

(6)

1) Architecture Specifications: • Generator G: Input: 100-dim noise z ∼ N (0, 1). Architecture: Dense(128) → ReLU → Dense(256) → ReLU → Dense(20) → Tanh. Output: Synthetic attack feature vector x̃. • Critic C: Architecture: Dense(256) → LeakyReLU(0.2) → Dense(128) → LeakyReLU(0.2) → Dense(1) → Linear. Output: Scalar ”realness” score. The model was trained for 100 epochs with batch size 256 using the Adam optimizer (β1 = 0.5, β2 = 0.999, learning

rate = 2 × 10−4 ), with the critic updated 4 times per generator update to ensure proper convergence. 2) Adversarial Dataset Generation: The trained generator produces synthetic attack flows x̃ ∼ Pg . To create adversarial examples that mimic benign traffic while retaining malicious functionality, we apply a targeted feature perturbation: ( (j) xbenign if i ∈ Fmod (i) xadv = (7) x̃(i) otherwise (j)

where Fmod is the set of modified features, and xbenign is a value sampled from the empirical distribution of benign traffic for feature i. We create two categories of datasets: • Hybrid Training Datasets: For adversarial augmentation, the original CICDDoS2019 training set is combined with adversarial samples where |Fmod | = {2, 4, 6} randomly selected features are modified. • Adversarial Test Sets: For robustness evaluation, we generate test sets where |Fmod | = {5, 9} of the most discriminative features (ranked by ANOVA F-score) are perturbed, creating Adv-5 and Adv-9 datasets. D. Detection Models 1) Random Forest (RF): We implemented a Random Forest classifier as a baseline. It was configured with 100 decision trees with no explicit depth limitation. To overcome dataset imbalance, we also used balanced class weighting. 2) Deep Ensemble Neural Network (DNE): We built an ensemble of five independently trained feed-forward neural networks to optimize deep learning and keep the model stable. Each individual network has an identical architecture: Input → Dense(64, ReLU) → Dropout(0.3) → Dense(32, ReLU) → Dropout(0.3) → Dense(1, Sigmoid). We obtained final predictions by averaging the outputs of all five models. 3) Transformer-based Model (TF): To capture global dependencies among traffic features, we designed a model based on the Transformer. To process the traffic data, our architecture embeds the 20-dimensional input into a 32-dimensional space. It then utilizes four attention heads to identify key feature patterns, incorporating layer normalization and residual links for better flow, and concludes with a standard feed-forward block. The output is flattened and passed through a final classification head (Dense(64, ReLU) → Dropout(0.3) → Sigmoid). We trained neural models (DNE, TF) using the Adam optimizer with binary cross-entropy loss for 20 epochs (batch size=256), with a validation size of 10% of the training data. The proposed framework functions as an intelligent DDoS detection component designed for integration within diverse security environments, including enterprise gateways, cloudbased IDS, and ISP edge monitoring systems. By using a standardized feature-extraction pipeline, the system can process real-time flows captured directly from network infrastructure, such as routers and switches. The core of the architecture—a Transformer-based model hardened by WGAN-GP adversarial augmentation—is specifically engineered to maintain balanced

performance and robustness against stealthy, evasive traffic patterns. This versatility allows the framework to operate either as a standalone module or as a resilient layer within a broader defense stack (alongside firewalls and scrubbing services), supporting both localized enterprise security and large-scale ISP operations. IV. E XPERIMENTAL S ETUP AND R ESULTS A. Phase 1: Baseline Model Performance In the first phase, we focused on establishing a baseline for later comparison. We chose three models: Random Forest (RF), Deep Ensemble Neural Network (DNE), and Transformer (TF), and trained them on the clean CICDDoS2019 dataset. These models were evaluated not only using the clean dataset, but also against two adversarial sets: Adv-5 and Adv9; these represent traffic where the top-5 and top-9 most discriminative features were modified, respectively. Since the main priority was to avoid missing any DDoS attacks, we chose Recall as the main metric. We also used the F1-score to balance Precision. The results of the evaluation on the clean dataset show excellent performance. All three models achieved Recall scores above 0.99 as shown in Table II. These results confirm the models’ ability to identify standard DDoS patterns. TABLE II: Baseline model performance on the clean CICDDoS2019 test set. Model Random Forest (RF) Deep Ensemble (DNE)

Recall 0.9991 0.9960

F1-Score 0.9995 0.9980

However, the models’ performance dropped significantly when evaluated on adversarially modified traffic (Table III). This reveals a generalization gap as training models only on a clean dataset makes them unable to adapt to modified features. The RF and DNE models were the most affected by this degradation; Recall on Adv-9 dropped to 0.1036 and 0.0465, respectively. The TF model achieved a recall score of 0.2045 and demonstrated greater relative resilience compared to the other architectures. TABLE III: Baseline model vulnerability to adversarial test sets. Model RF DNE TF

Recall (Adv-5) 0.5451 0.2910 0.6504

F1 (Adv-5) 0.7056 0.4509 0.7882

Recall (Adv-9) 0.1036 0.0465 0.2045

F1 (Adv-9) 0.1877 0.0888 0.3396

This significant decline in performance illustrates that traditional models are not sufficient for modern security as they cannot adapt to evolving attacks. Consequently, they need to be strengthened through adversarial training. B. Phase 2: Robustness via Adversarial Augmentation To fix the weaknesses identified in the first phase, we retrained the three models using hybrid datasets which combine the original clean dataset (CICDDoS2019) with adversarial samples generated by the WGAN-GP, where 5 randomly

selected features were modified. The retrained models were then re-evaluated on the same test sets (Adv-5 and Adv-9) to quantify the detection improvement. The results confirm that adversarial augmentation of the training dataset yielded a significant improvement in model performance, as shown in Table IV. The RF model achieved high Recall scores on both adversarial test sets,s representing a substantial enhancement over its Phase 1 results. The TF model showed remarkable performance, while maintaining a high Recall score of 0.8002 on the hardest evaluation (Adv-9). The DNE showed an improvement on Adv-5 but failed to sustain it on Adv-9, remaining the weakest model among the three. TABLE IV: Model performance on adversarial test sets after adversarial augmentation training. Model RF DNE TF

Recall (Adv-5) 1.0000 0.5848 0.9994

F1 (Adv-5) 1.0000 0.7380 0.9997

Recall (Adv-9) 0.9689 0.1631 0.8002

F1 (Adv-9) 0.9842 0.2805 0.8890

The results demonstrate that WGAN-GP-based adversarial augmentation effectively strengthens DDoS detectors against feature-space evasion. The Transformer’s consistent high performance across both phases highlights its inherent capacity to learn robust, generalized feature representations. C. Real-Time Traffic Evaluation To assess the viability of the models in a real-world scenario, we tested them in a simulation that mimics a real-time pipeline using three unlabeled traffic segments that represent realistic network conditions: The first one represents normal network operation and was used as a baseline for false positives. The second segment captures the peak of a DDoS attack, and the third one is the recovery phase with residual malicious traffic. To mimic a realistic scenario, we performed inference flowby-flow using only existing preprocessing tools (scaler, feature mask) without any additional training or adjustment during the simulation. The models exhibited distinct performance patterns, which reveal the necessity of balance between sensitivity and specificity. TABLE V: Model performance on unlabeled real-time traffic segments. Segment Before (Normal) During (Attack) After (Recovery)

Metric Accuracy Recall Recall

RF 1.00 0.01 0.10

DNE 0.00 1.00 1.00

TF 1.00 0.69 0.80

The RF model achieved an accuracy score of 1.00 during normal operations, indicating it did not produce false alarms; however, it almost entirely failed to detect attacks during the attack peak (recall score during the attack phase was 0.01). This conservative approach makes it useless for realworld applications as a primary detector. The DNE labeled all traffic as malicious across the three periods (before, during, and after). This means the false positive rate is 100% during

normal operation, rendering it impractical due to alert fatigue. The TF maintained perfect accuracy during normal operation, resulting in zero false alarms; furthermore, it was able to detect the majority of attacks (recall score during the attack was 0.69). Moreover, it was able to detect malicious traffic during the recovery phase. This balance makes it the most suitable model for real-world deployment. D. Discussion By critically evaluating the results of this study, we can draw the following conclusions: On one hand, the results of phase 1 (Table II) demonstrate that relying exclusively on clean datasets can be misleading. While the three models achieved near-perfect accuracy, their performance decreased significantly when exposed to adversarial attacks. On the other hand, the recovery of the RF and TF models demonstrates that adversarial augmentation using WGAN-GP is an efficient strategy to enhance model generalization and resilience. By analyzing each architecture’s performance, we observe that design directly influences operational behavior: • The RF can be described as an extreme conservative; it failed to flag attacks in real-time tests. • The DNE was hypersensitive; it triggered excessive false alarms. • The Transformer model achieved significant improvements from adversarial training and is the most robust model for real-world deployment. Moreover, it provides a balance between sensitivity (recall) and precision, avoiding overreactions to normal traffic variations. V. C ONCLUSION AND F UTURE W ORK In this study, we evaluated the effectiveness of machine learning and deep learning models in detecting DDoS attacks in both controlled and real-world settings. The main requirement for a practical DDoS defense system is to achieve a balance between recall (to avoid missing attacks) and precision (to avoid triggering false alarms). Failing to achieve this balance makes the system ineffective for real-world deployment. The results show that Random Forest (RF) performs well on normal traffic but fails against active attacks. The Deep Ensemble (DE) models performed well at detecting attacks (high recall), but triggering too many false positives undermined their practical utility. These results demonstrate that the high performance of models on standard datasets degrades in real-world scenarios due to adversarial attacks; consequently, these models lack the adaptability to detect dynamic and changing attacks. The Transformer-based model achieved high recall and sensitivity, but it requires more tuning to achieve better precision. To move this model from laboratory environments to the real-world, we must harden it against evasion attacks by employing adversarial augmentation techniques using WGAN-GP. This strategy was successful, as evidenced by the evaluation metrics. Future work will expand this framework along two primary vectors: Problem-Space Adversarial Testing: Implementing

packet-level evasion tools (e.g., ptf-agent, craft-based packet injection) to evaluate closed-loop, protocol-constrained attack realizability in physical testbeds. Cross-Dataset Generalization: Validating the WGAN-GP pipeline across additional benchmark suites (e.g., Bot-IoT, ToN-IoT) and exploring online learning techniques for real-time model updating under dynamic zero-day attack patterns. R EFERENCES [1] A. Choukeir, B. Fneish, N. Zaarour, W. Fahs, and M. Ayache, “Health smart home,” International Journal of Computer Science Issues (IJCSI), vol. 7, no. 6, p. 126, 2010. [2] A. Singh and B. B. Gupta, “Distributed denial-of-service (ddos) attacks and defense mechanisms in various web-enabled computing platforms: Issues, challenges, and future research directions,” International Journal on Semantic Web and Information Systems (IJSWIS), vol. 18, no. 1, pp. 1–43, 2022. [3] J. Wu, Z. He, W. Wang, H. Luo, S. Hu, Y. Miao, and J. Deng, “CADA: A flexible and elastic DDoS mitigation architecture,” IEEE Access, 2025. [4] E. SABRINE, F. A. B. I. O. DE GASPARI, H. DORJAN, A. K. BIDI, and L. V. MANCINI, “Adversarial challenges in network intrusion detection systems: Research insights and future prospects,” arXiv preprint arXiv:2409.18736, 2024. [5] O. Berjawi, A. El Attar, F. Chbib, R. Khatoun, and W. Fahs, “Cyberattacks detection through behavior analysis of internet traffic,” Procedia Computer Science, vol. 224, pp. 52–59, 2023. [6] S. Dyrmishi, S. Ghamizi, T. Simonetto, Y. Le Traon, and M. Cordy, “On the empirical effectiveness of unrealistic adversarial hardening against realistic adversarial attacks,” in 2023 IEEE symposium on security and privacy (SP). IEEE, 2023, pp. 1384–1400. [7] M. Macas, C. Wu, and W. Fuertes, “Adversarial examples: A survey of attacks and defenses in deep learning-enabled cybersecurity systems,” Expert Systems with Applications, vol. 238, p. 122223, 2024. [8] M. Abdelaty, S. Scott-Hayward, R. Doriguzzi-Corin, and D. Siracusa, “Gadot: Gan-based adversarial training for robust ddos attack detection,” in 2021 IEEE Conference on Communications and Network Security (CNS). IEEE, 2021, pp. 119–127. [9] K. Kumari and M. Mrunalini, “Detecting denial of service attacks using machine learning algorithms,” Journal of Big Data, vol. 9, no. 1, p. 56, 2022. [10] A. Mustapha, R. Khatoun, S. Zeadally, F. Chbib, A. Fadlallah, W. Fahs, and A. El Attar, “Detecting DDoS attacks using adversarial neural network,” Computers & Security, vol. 127, p. 103117, 2023. [11] L. Chen, Z. Wang, R. Huo, and T. Huang, “An adversarial DBNLSTM method for detecting and defending against DDoS attacks in SDN environments,” Algorithms, vol. 16, no. 4, p. 197, 2023. [12] C. S. Shieh, T. T. Nguyen, W. W. Lin, W. K. Lai, M. F. Horng, and D. Miu, “Detection of adversarial DDoS attacks using symmetric defense generative adversarial networks,” Electronics, vol. 11, no. 13, p. 1977, 2022. [13] L. H. De Melo, G. de Carvalho Bertoli, M. Nogueira, A. L. Dos Santos, and L. A. Pereira, “Anomaly-flow: A multi-domain federated generative adversarial network for distributed denial-of-service detection,” IEEE Network, 2025. [14] R. V. Kulkarni et al., “Robust iot threat detection via generative augmentation,” IEEE Internet of Things Journal, 2025. [15] M. A. Al-Garadi et al., “Generative adversarial networks for cyber security: A comprehensive review,” Applied Soft Computing, vol. 112, p. 112455, 2024. [16] S. H. C. Chilukuri et al., “Adversarial machine learning in network intrusion detection: Vulnerabilities and defense mechanisms,” IEEE Transactions on Network and Service Management, 2023. [17] M. Abdelaty, S. Scott-Hayward, R. Doriguzzi-Corin, and D. Siracusa, “Gadot: GAN-based adversarial training for robust DDoS attack detection,” in 2021 IEEE Conference on Communications and Network Security (CNS). IEEE, 2021, pp. 119–127. [18] I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein gans,” in Advances in neural information processing systems, vol. 30, 2017.

Record · ID 965346 · SHA-256 ecc9a12f24a39218
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.