1
AI-Native Closed-Loop Security for 6G-Enabled Cyber-Physical Systems: From Edge Detection to Network-Wide Mitigation
arXiv:2606.08173v1 [cs.CR] 6 Jun 2026
Bilal Hussain, Member, IEEE, Muhammad Bilal, Senior Member, IEEE, Tan Li, Member, IEEE, Haris Pervaiz, Member, IEEE, Xiao Tang, Member, IEEE, Qinghe Du, Member, IEEE, Fawad Ahmad, Member, IEEE, Muhammad Azhar, and Jun Zhang, Fellow, IEEE
Abstract—In sixth-generation (6G) networks, billions of cyberphysical systems (CPSs)—autonomous vehicles, smart grids, industrial robots, and remote-surgical equipment—will run over ultra-reliable low-latency slices, collapsing the gap between a remote breach and physical harm to milliseconds, a budget perimeter firewalls and centralised security operations centres cannot meet. This survey reframes 6G CPS security as a closed-loop, AI-native pipeline that senses at the multi-access edge computing (MEC) tier, using minute-scale call-detail records (CDRs) for baseline learning and slow-rate campaigns and sub-millisecond Radio Access Network (RAN)/Open-RAN (O-RAN) telemetry for the latency-critical path. The pipeline decides locally with compressed deep models, mitigates network-wide via software-defined networking (SDN), network function virtualization (NFV), and O-RAN controllers, and retrains through federated learning (FL) and digital-twin (DT) replay. We formalise a per-slice, tailbounded latency contract on the sense, detect, and mitigate stages, enforced at a slice-dependent tail percentile (p99 for safetycritical URLLC slices) as a conservative sum-of-stage bound that prior autonomic and network control-loop models leave undefined. Organising 128 peer-reviewed studies (2017–2026) under a PRISMA 2020 protocol, we (i) map the 6G/CPS threat surface to MITRE ATT&CK and a CDR-observable feature space; (ii) unify edge anomaly detection and DDoS classification across twelve datasets and statistical, graph, and transformer models; (iii) synthesise SDN/NFV/O-RAN primitives into one closed-loop reference architecture; (iv) treat FL, large language models (LLMs), DT, post-quantum cryptography (PQC), zerotrust architecture (ZTA), and explainable AI as cross-cutting enablers, not parallel pillars; and (v) consolidate open problems into five directions spanning data, latency, trust, standardisation, and evaluation. Index Terms—6G networks, cyber-physical systems, AI-native Bilal Hussain and Fawad Ahmad are with the Division of Science, Engineering, and Health Studies, School of Professional Education and Executive Development, The Hong Kong Polytechnic University, Hong Kong SAR, China (e-mails: {bilal.hussain, fawad.ahmad}@cpce-polyu.edu.hk). Muhammad Bilal is with the School of Computing and Communications, Lancaster University, United Kingdom (e-mail: [email protected]). Tan Li is with Department of Computer Science, The Hang Seng University of Hong Kong, Hong Kong SAR, China (e-mail: [email protected]). Haris Pervaiz is with the School of Computer Science and Electronic Engineering, University of Essex, CO4 3SQ, United Kingdom (e-mail: [email protected]). Xiao Tang and Qinghe Du are with the School of Information and Communication Engineering, Xi’an Jiaotong University, Xi’an 710049, China (e-mails: [email protected], [email protected]). Muhammad Azhar is with the Department of Applied Data Science, Hong Kong Shue Yan University, Hong Kong SAR, China (e-mail: [email protected]). Jun Zhang is with the Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong (e-mail: [email protected]).
security, edge intelligence, MEC, CDR, anomaly detection, DDoS mitigation, federated learning, O-RAN, closed-loop control.
I. I NTRODUCTION IXTH-GENERATION (6G) wireless networks will not merely add bandwidth; they will fuse the digital and the physical until breach latency and physical-harm latency become the same number. Industry roadmaps target peak data rates above 1 Tb/s, end-to-end latencies below 1 ms, connection densities of 107 devices/km2 , and ubiquitous ondevice intelligence [1]–[4]. These capabilities exist to serve Cyber-Physical Systems (CPSs): autonomous and aerial vehicles, smart grids, Industry 4.0 robotic cells, remote surgery, Unmanned Aerial Vehicle (UAV) swarms, and tactile telepresence [5]. Across these domains, every packet is effectively a physical event: an unmitigated attack can move a robot, misdose a patient, or bring down a substation. The shift from 5G to 6G is not incremental on the security axis. 5G introduced service-based architecture, network slicing, and edge computing, but its security model still leaned on perimeter-style controls inherited from 4G Evolved Packet Core (EPC) and on a handful of standardised functions such as the Authentication Server Function (AUSF), Security Anchor Function (SEAF), and Access and Mobility Management Function (AMF), which authenticate users and protect the signalling plane [6]–[8]. 6G dissolves this perimeter. The Radio Access Network (RAN) is disaggregated and softwaredefined [9]–[11]; intelligence migrates from the core into the RAN and onto the device [12]; and the same physical infrastructure simultaneously hosts Ultra-Reliable LowLatency Communication (URLLC) vehicular slices, massive Machine-Type Communication (mMTC) industrial Internet-ofThings (IoT) slices, and enhanced Mobile Broadband (eMBB) extended-reality slices. Each of these design choices opens a new attack surface and shortens the chain between an exploit and a physical-world consequence. The consequence for CPS whose control loops traverse a 6G network—hereafter referred to by the shorthand “6G CPS”—is best appreciated as a budget. A Vehicle-to-Everything (V2X) collision-avoidance message has at most a few milliseconds to travel from a vehicle, through the radio, into a remote application server, and back to a peer vehicle that must brake; an autonomous robotic surgeon has perhaps ten milliseconds
S
2
TABLE I: List of Abbreviations Acronym
Definition
Acronym
Definition
3GPP 5G 6G AI AMF aNB API CDR CNN CPS DDoS DL DT FL GAN GNN GTP IDS IoT LLM
Third Generation Partnership Project Fifth-Generation Mobile Network Sixth-Generation Mobile Network Artificial Intelligence Access and Mobility Management Function AI-native NodeB (6G RAN node) Application Programming Interface Call Detail Record Convolutional Neural Network Cyber-Physical System Distributed Denial of Service Deep Learning Digital Twin Federated Learning Generative Adversarial Network Graph Neural Network GPRS Tunneling Protocol Intrusion Detection System Internet of Things Large Language Model
LSTM MEC ML mMTC NAS NFV NWDAF O-RAN PQC PRISMA RIC RRC SDN SOC URLLC V2X VAE XAI xApp ZTA
Long Short-Term Memory Multi-access Edge Computing Machine Learning Massive Machine-Type Communication Non-Access Stratum Network Function Virtualization Network Data Analytics Function Open Radio Access Network Post-Quantum Cryptography Preferred Reporting Items for Systematic Reviews and Meta-Analyses RAN Intelligent Controller Radio Resource Control Software-Defined Networking Security Operations Center Ultra-Reliable Low-Latency Communication Vehicle-to-Everything Variational Autoencoder Explainable Artificial Intelligence Extended Application (O-RAN) Zero Trust Architecture
AI-Native Security Ecosystem for 6G Cyber-Physical Systems Closed-Loop Pipeline: Sense → Detect → Mitigate → Learn
SDN Controller Flow Rules
NFV MANO Orchestration
Live Net State
Threat Intel Feed External
Automated Threat Mitigation: SDN/NFV actuators enact playbooks to suppress confirmed threats Digital Twin & Threat Feed: Uses “Live Net State” for validation and external feeds
L3 Intelligence & Detection Layer (DETECT + LEARN)
Collaborative training using local CDR streams without exposing sensitive subscriber info
L2 Radio Access & Edge Layer (SENSE + DETECT)
aNB - 6G Base MEC - AI Station Inference
L1 Physical / CPS Layer (SENSE)
Digital Twin
Federated Learning (Model Updates Only)
RIC stack
Anomaly Detection Alert Engine for threat classification
Near-RT RIC
O-RAN RIC Hub
Non-RT RIC
A1/E2/O1 interfaces
RIC stack managing control loops
aNB - 6G Base MEC - AI Station Inference
Post-Quantum Cryptography
Closed-Loop Feedback Enables continuous learning and self-healing
Zero-Trust Perimeter
L4 Control & Orchestration Layer (MITIGATE)
aNB - 6G Base MEC - AI Station Inference
Edge-Level Anomaly Detection: aNB signals feed into co-located MEC inference for first-pass detection (single-digit-ms)
Autonomous Vehicle (V2X)
Smart Grid (Energy)
Industrial Remote Surgery UAV/Drone Robot(IIoT) (NTN) (URLLC)
IoT Sensor (mMTC)
Heterogeneous CPS Endpoints: Generate raw telemetry; span diverse vertical sectors with unique traffic profiles
Telemetry/Data Flow (L1 → L2 → L3) Control/Policy Flow (L4 → L2) FL Model Update (Gradient) (L2 → L3) Detection → Mitigation Alert (L3 → L4) Threat Intelligence Feed (L4 → L3) Closed-Loop Feedback (L4 → L1)
Fig. 1: End-to-end AI-native security ecosystem for 6G cyber-physical systems (CPSs), modelled as a closed control loop whose four stages—sense → detect → mitigate → learn—are realised across four functional planes (physical/CPS, radio access and edge, intelligence and detection, control and orchestration); the intelligence/detection plane hosts the per-MEC detectors and the operator-core FL aggregator that consolidates their gradient updates. Zero-trust access (left) and post-quantum cryptography (right) span all planes as cross-cutting substrates, while the coloured arrows denote the inter-plane flows— telemetry, control/policy, federated-learning gradient updates, the detection-to-mitigation alert, and the threat-intelligence feed that sharpens Layer-3 detection—as defined in the legend. The bold curved arrow closes the loop from mitigate back to sense, enabling continuous learning and self-healing. The per-plane mechanisms are detailed in Sec. I-B. before a control-loop deviation becomes a physiological event; an industrial protective relay has a few cycles before a fault propagates and trips a substation [13], [14]. In all three cases, the attacker does not need to break confidentiality or alter payload—merely delaying or dropping the message is enough to materialise harm. Security in 6G CPS is therefore inseparable from latency, and the architectural question is where in the network the defender can decide and act fast
enough. A. Why Legacy Security Cannot Meet the 6G CPS Budget? Four structural gaps make perimeter-based defences and the Security Operations Center (SOC) model unsuitable for 6G CPS: a volume gap, a latency gap, a telemetry gap, and a trustboundary gap [7], [8], [15]. First, the volume gap: a single 6G artificial-intelligence (AI)-native Node B (AI-native Node B, or aNB) will emit gigabit-rate signaling, mobility, and slice
3
AI-Native Security for 6G CPS (This Survey) Foundation
Threat
Detection
Sec. II Background
Sec. III Methodology
Sec. IV Threat Landscape
Sec. V Edge Detection
• 6G/CPS Architectures • MEC & O-RAN Substrate • CDR Feature Space
• PRISMA 2020 Protocol • Comparison with Prior Surveys • Closed-Loop Taxonomy
• 6G CPS Attack Surface • Cellular DDoS Vectors • MITRE ATT&CK Alignment
• CDR-Driven DL Models • MEC Deployment & Benchmarks • Sub-10 ms Inference
Enablers
Mitigation
Sec. VII Cross-Cutting Enablers
Sec. VI Network Mitigation
Synthesis
Sec. IX Conclusion • Key Findings Summary • Call to Action
Sec. VIII Open Challenges • Data & Benchmark Gaps • Standardisation & Governance • CPS-Impact Metrics
• FL, LLM, Digital Twin • Zero-Trust & PQC • Explainable AI
• SDN/NFV Orchestration • O-RAN Closed-Loop Control • Cross-Tier Reference Architecture
Fig. 2: Survey roadmap. The paper is structured around eight sections (Secs II–IX) that collectively cover the end-to-end AI-native security pipeline for 6G cyber-physical systems, from foundational background and systematic methodology through edge detection and network-wide mitigation to cross-cutting enablers, consolidated open challenges, and concluding synthesis. telemetry that no human-supervised SOC can triage in real time [8], [15]. Second, the latency gap: backhauling traffic to a cloud detector costs 80–120 ms [16], [17], but URLLC slices for V2X or telesurgery have 1–10 ms end-to-end budgets, so by the time a cloud rule fires, the unsafe action has already happened. Third, the telemetry gap: legacy intrusion detection systems (IDS) were tuned to packet headers, not to the long-range, multi-modal cellular signals that reveal attacks on cellular infrastructure. Such signals include Non-Access Stratum (NAS) attach storms, anomalous Radio Resource Control (RRC) setup ratios, GPRS Tunnelling Protocol User-plane (GTP-U) utilisation, and bursts of calls and Short Message Service (SMS) traffic visible in Call Detail Records (CDRs) [7], [18], [19]. Rule-based IDSs over-fit to known signatures and generalise poorly to the polymorphic, slice-aware Distributed Denial of Service (DDoS) variants now reported in the wild [20]. A fourth, less-discussed mismatch is the trust-boundary gap. The 5G perimeter assumed a hard inside/outside split between operator core and external internet; 6G dissolves that split because (i) multi-access edge computing (MEC) hosts execute third-party applications adjacent to the user-plane data path [21]–[24], (ii) Open Radio Access Network (O-RAN) extended applications (xApps) and RAN applications (rApps) are sourced from heterogeneous vendors and dynamically loaded into the RAN Intelligent Controllers (RICs) [9], [11], [25]– [27], and (iii) network slices are rented to verticals, and the operator cannot audit the security posture of those tenants [28], [29]. The implication is that there is no longer a meaningful “inside” to defend; every interface must be authenticated and every model decision must be interpretable, which is the operational case for zero-trust architecture (ZTA) [30], [31] and explainable AI (XAI) [32] as native, not optional, components of the security stack. These four gaps compound. A volumetric control-plane
signalling storm originating from a slice-tenant’s compromised IoT fleet [33], [34] crosses the trust boundary at the MEC tier, demands a sub-millisecond response on a URLLC slice, and is invisible to an IDS that inspects only packet headers and does not parse Third Generation Partnership Project (3GPP) signalling. The closed-loop framing returns to these effects later as explicit, CPS-aware mitigation metrics and gives the slice tenant a runtime control surface over them (Sec. VI-D). A complete list of abbreviations used throughout this survey is provided in Table I. B. The AI-Native, Closed-Loop Alternative The remedy emerging in the literature is to make security a native function of the network fabric itself, executed close to the data source and continuously trained [9], [28], [35]–[37]. We unify these threads under a single thesis: Thesis. In 6G CPS, security must be reframed as a closedloop, AI-native pipeline: sense on rich edge telemetry (CDR + RAN/O-RAN signals), detect locally at the MEC tier with latency-bounded, compressed deep models, mitigate network-wide through Software-Defined Networking (SDN)/Network Function Virtualization (NFV)/O-RAN actuators, and learn continuously via Federated Learning (FL) and Digital Twin (DT) replay—all under a zero-trust, postquantum substrate. The novelty is not the four-stage decomposition—which is, in shape, shared with Monitor–Analyze–Plan–Execute over a shared Knowledge base (MAPE-K) and Observe–Orient–Decide–Act (OODA)—but the per-slice, tailbounded design contract that binds the loop to each slice’s CPS safety budget, making slice-admissibility a falsifiable property rather than a qualitative aspiration (formalised in Sec. VI-F).
4
The loop is treated here as an operational pipeline with measurable telemetry, latency, and actuation stages, each realised as a concrete tier rather than an abstraction. Fig. 1 renders each verb as a concrete operational tier with measurable latency, identifiable telemetry, and surveyed empirical evidence: sense on the MEC-resident CDR/RAN/Network Data Analytics Function (NWDAF) telemetry streams [38], [39]; detect via lightweight deep models running at single-digitmillisecond latency on commodity MEC hardware [23], [40], [41]; mitigate through SDN/NFV/O-RAN actuators composed by an AI-native orchestrator [42]–[45]; and learn through FL aggregation and DT replay that closes the loop back to sense [28], [37], [46], while external threat-intelligence feeds are distributed to the Layer-3 detection models to sharpen anomaly classification against known indicators of compromise. Two cross-cutting substrates bracket the loop: a zerotrust perimeter that authenticates every sense–detect–mitigate transaction, and post-quantum cryptography that protects the telemetry, gradient, and policy flows against harvest-nowdecrypt-later adversaries, motivating the integration of postquantum cryptography (PQC) with FL aggregation [28], [45]. The contribution of this survey is not any single verb but their explicit composition into a continuously-running pipeline whose end-to-end metrics are the binding constraints—and whose composition surfaces the five open challenges catalogued in Sec. VIII. The loop framing matters because the pillars share a latency budget, a telemetry substrate, and a control plane: coupling problems that per-pillar surveys structurally cannot surface (Sec. VIII). We scope the survey to 6G-tethered CPS—excluding generic enterprise IT and pure-IoT settings—and frame our contributions around the sense → detect → mitigate → learn closed loop introduced in Sec. I-B, so that each contribution maps to one stage of a single pipeline rather than standing alone: One closed-loop reference architecture (Fig. 1) that binds the four stages below into a single, latency-contracted pipeline and resolves the long-standing split between the “edge-detection” and “network-mitigation” literatures. • Sense—CDR as a first-class security signal. We ground a CDR-driven sensing layer in 3GPP-observable features and 5G/6G NWDAF interfaces, with explicit MITRE ATT&CK alignment (Sec. IV, Table VI), establishing the telemetry substrate the loop senses on. • Detect—deployment-aware AI at the MEC. We consolidate CDR-driven anomaly detection and DDoS classification under one taxonomy and one accuracy-and-latency table segregated by deployment tier (Sec. V, Table IX), removing the duplication earlier surveys inherited from the SDN-IDS and cellular-anomaly strands. • Mitigate—composable network-wide actuation. We treat SDN, NFV, O-RAN xApps, and AI orchestration as composable actuators of one control loop rather than competing schools (Sec. VI, Table XI). • Learn—enablers as services to the loop. We position FL, large language models (LLMs), DT, PQC, ZTA, and XAI as continuous-learning and trust services that sustain the loop •
•
(Sec. VII), not as a parallel “emerging paradigms” silo. A loop-level research roadmap (Sec. VIII) of five composition-level challenges—latency, auditability, federated governance, standardisation, and CPS-specific benchmarking—each carrying a falsifiable 2027/2028 target and a named standardisation vehicle.
C. Survey Organisation Sec. II establishes the 6G/CPS/MEC/CDR substrate that defines the sensing surface. Sec. III presents our Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-based [47] selection protocol and contrasts the survey with prior work. Sec. IV characterises the threats that the loop must defend against. Sec. V surveys edgebased detection. Sec. VI surveys network-wide mitigation and instantiates the closed loop. Sec. VII covers the cross-cutting enablers. Sec. VIII consolidates the five open challenges and Sec. IX concludes. The roadmap of how sections compose into one narrative is in Fig. 2: Secs. II–IV frame the loop and the sense stage (threat surface and CDR-observable feature space); Sec. V treats the detect stage; Sec. VI the mitigate stage; and Sec. VII the learn stage together with the crosscutting substrates (PQC, ZTA, XAI) that span all four stages; Secs. VIII–IX close the survey. II. BACKGROUND : 6G CPS, MEC, AND CDR This section fixes the substrate on which the rest of the survey runs. We summarise (i) the 6G architectural pillars that define the new attack surface, (ii) how CPSs are formally bound to 6G slices, (iii) why MEC is the right enforcement plane, and (iv) why CDRs are the most under-exploited security signal in cellular networks. AI/Machine Learning (ML) preliminaries are not repeated here; we assume readers are familiar with Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs)/Long Short-Term Memory networks (LSTMs), autoencoders, Graph Neural Networks (GNNs), and transformers [48]–[50] (see also [51, Ch. 1] for an applied AI/ML primer in a security context) and introduce architecture-specific notation in situ in Sec. V. Throughout this survey, x ∈ Rd denotes a d-dimensional feature vector extracted from cellular telemetry; fθ : Rd → RK a K-class classifier with parameters θ; gθ an autoencoder; s(x) ∈ R≥0 (t) an anomaly score; θk the model parameters of MEC client k (t) (t) (t−1) at FL round t; ∆θk = θk − θglobal the local update; ∆tinfer per-decision inference latency; ∆tmit time from alert to enacted mitigation; and τmax the slice’s safety budget. A. 6G Architectural Pillars and Their Security Implications Six pillars distinguish 6G from 5G and each opens a new attack surface. AI-native RAN replaces hand-engineered protocol stacks with learned ones [12], [25], [52], exposing the radio scheduler, beam manager, and Hybrid Automatic Repeat reQuest (HARQ) policy to adversarial ML; the AI-as-a-Service (AIaaS) exposure surface introduced by operator-side RAN AI service APIs [12], [44] further widens the attack surface to include model-theft and capability-abuse vectors. Open RAN (ORAN) disaggregates the RAN into Open Radio Unit (O-RU),
5
TABLE II: 6G Architectural Pillars and Associated Security Implications Pillar
Key Capability
AI-Native RAN
Learned protocol stack, autonomous optimisation Open interfaces, xApps/rApps, multivendor eMBB/URLLC/mMTC virtual networks Global ubiquitous coverage Passive beamforming, coverage extension Joint sensing and communication
O-RAN / RIC
Network Slicing NTN (LEO/HAPS) RIS ISAC
Primary Security Concern Adversarial ML, model poisoning/inversion Rogue xApps, API exploitation, supply-chain risk Cross-slice leakage, resource starvation Eavesdropping, limited onboard security Beam hijacking, pilot contamination Adversarial sensing, unauthorised surveillance
Open Distributed Unit (O-DU), and Open Centralised Unit (O-CU) components controlled by Near-Real-Time (NearRT) and Non-Real-Time (Non-RT) RICs hosting xApps and rApps [9]–[11], [31]; openness multiplies vendor diversity and supply-chain risk while creating rogue-xApp and Application Programming Interface (API)-exploitation vectors. Network slicing carves logically isolated eMBB/URLLC/mMTC virtual networks over shared physical resources but leaves cross-slice interference and resource-starvation attacks [53]. Non-Terrestrial Networks (NTNs) extend coverage with Low Earth Orbit (LEO) satellites and High-Altitude Platform Stations (HAPS) but inherit limited onboard security and are open to eavesdropping [2]. Reconfigurable Intelligent Surfaces (RISs) offer passive beamforming with new pilotcontamination and beam-hijacking risks [54], [55]. Integrated Sensing and Communication (ISAC) reuses radio waveforms for environmental sensing, creating adversarial-sensing and surveillance hazards [2], [56]. Table II summarises this map, linking each pillar’s key capability to the primary security concern it introduces; the concrete exploitation path for each concern is developed in Sec. IV-A. B. Cyber-Physical Systems over 6G A CPS is a tight coupling of computation, networking, and physical processes [13], [14]; over 6G, that coupling is realised by binding each CPS workload class (defined by its latency, reliability, and density profile) to a network slice. URLLC slices serve V2X and Industry 4.0 motion control with ≤ 1 ms one-way latency and 10−9 packet error [1], [3], [4], [15]; eMBB slices serve Extended Reality (XR) and holographic telepresence; mMTC slices serve smart-metering and large-scale IoT. Security therefore inherits the slice’s Key Performance Indicator (KPI) envelope: an IDS that needs 80 ms to decide is not an IDS for a URLLC slice. Four CPS classes recur throughout this survey because they collectively span the 6G slice envelope, anchoring distinct corners of the latency–reliability–throughput–density design space [5], [51, Ch. 8]. (1) Connected and autonomous mobility (vehicular platoons, urban autonomy, UAV swarms) couples per-vehicle perception/control loops with V2X cooperative perception and is most sensitive to URLLC latency violations and to localisation tampering; a delayed actuator command becomes a collision. (2) Smart grid and substation automation couples wide-area measurement with protective relaying
over sub-millisecond-sensitive International Electrotechnical Commission (IEC) 61850 Generic Object Oriented Substation Event (GOOSE) and Sampled Values (SV) traffic; late or dropped messages cause a real-world trip or out-of-step condition. (3) Industrial automation and Industry 4.0 robotic cells couples Programmable Logic Controllers (PLCs) and motion controllers with Time-Sensitive Networking (TSN)style deterministic transport—the canonical converged Information Technology (IT) / Operational Technology (OT) case in which a network breach moves physical actuators (e.g., a faulty weld). (4) Remote healthcare and telesurgery couples haptic and visual feedback with operator-side actuation and demands not only latency bounds but also explainability of any defensive action that interrupts a clinical session, since a missed defibrillation has the same cost as a missed attack [5], [13], [14]. The cross-class invariant is that a successful attack does not need to read or alter payload—it only needs to delay, drop, or deflect packets long enough to push a control loop outside its safety envelope [57]. This framing has two implications for security. First, confidentiality, integrity, and availability must be re-prioritised: for safety-critical slices, availability and timely integrity dominate confidentiality, and the cost of a false positive (blocking legitimate control traffic) can exceed the cost of a missed attack. Second, the closed-loop pipeline of Sec. I-B must produce type-aware mitigation: a generic “alert→rate-limit” reflex that is appropriate for an mMTC SMS flood is catastrophic for a URLLC vehicular slice, where the same primitive can starve a safety message. Sec. VI returns to this point. C. Multi-Access Edge Computing as the Enforcement Plane The European Telecommunications Standards Institute (ETSI) MEC reference architecture [21], [23], [25], [58] colocates compute and storage with the base station, delivering single-digit millisecond detection latencies versus 80–120 ms cloud round trips [16], [17]. A typical MEC host offers 8– 32 CPU cores, 32–128 GB RAM, and an optional GPU— enough for moderate-complexity CNN/LSTM detectors, but not for uncompressed LLMs [59]. Deploying a deep model at the MEC therefore requires aggressive compression [17], [24]: 8-bit integer quantisation, structured pruning, knowledge distillation from a larger cloud teacher, and—where supported—hardware acceleration through Neural Processing Units (NPUs), Field-Programmable Gate Arrays (FPGAs), or vendor-specific inference chips. Even after compression, the host time-shares its compute among the security model, application workloads, and orchestration agents, so model design must account for tail-latency excursions and not just average throughput. Edge-intelligence surveys [16], [17], [59] converge on a layered pattern that we recover throughout Sec. V: a small “cheap” model triages every flow, a larger model reevaluates a sampled or alerted subset, and a cloud or core model owns long-horizon model lifecycle, FL aggregation, and policy synthesis. The compute envelope is only one half of the case for MEC; the other half is telemetry. The MEC tier is the only point in the 6G stack with simultaneous low-latency access to
6
four security-relevant modalities, each over its native interface: (i) the operator’s CDR pipeline streamed from the Charging Function (CHF) over the Nchf reference point [23], [24], [60]; (ii) 3GPP analytics streams—the Network Data Analytics Function (NWDAF) in 5G and its 6G successor still under standardisation—exposed to consumers as event analytics over the Service-Based Architecture (SBA) Nnwdaf service [39]; (iii) O-RAN Near-RT RIC E2 Key Performance Measurement (KPM) measurement feeds [12], [61], [62]; and (iv) MEC-local user-plane probes [58]. This four-interface vantage point—hereafter referred to as the MEC telemetry vantage—is the architectural reason detectors that fuse CDR with RAN telemetry can be operationally deployed at the MEC but rarely elsewhere [12], [44]. The price of this proximity is local visibility: each host sees only its attached cells, so distributed attacks demand hierarchical aggregation (discussed in Sec. V-A, Sec. V-H). These capabilities are conditional on a property that is frequently under-stated: MEC is also the trust boundary between the operator’s administrative control and third-party application workloads. Hosts run operator-managed platform software alongside tenant workloads from Content-Delivery Networks (CDNs), enterprise customers, and vendor-supplied O-RAN functions, with separation enforced by container isolation, per-tenant namespaces, and role-based API access [23], [24], [44]. Co-tenancy is two-edged: the operator gains fine-grained observability over every workload—the enabling condition for MEC-tier detection—but a compromise of the platform software exposes the security model’s parameters, its training data, and cross-tenant traffic observations. Surveys of MEC security [8], [24] agree that platform-integrity attestation (secure boot, measured-launch, runtime integrity monitoring) is a prerequisite for treating the MEC as the detection tier; without it, the detector runs on a substrate whose trustworthiness has not been established, and the closed loop inherits that uncertainty [21], [45]. D. Call Detail Records as Security Telemetry Call Detail Records (CDRs) are per-event records of calls, SMS messages, data sessions, and handovers, written in nearreal-time by the 5G Charging Function (CHF) under the Converged Charging System (CCS) [60]. Operators retain them for billing, performance management, and lawful interception [60], and typically expose them to analytics consumers either as the raw per-event stream or as per-cell aggregates over 5–15 minute reporting windows [38]. CDRs are not packet captures; they cannot reveal payload, but they reliably reveal volumetric, mobility, and call-pattern anomalies that are the dominant signature of cellular DDoS attacks [63]– [67]. Crucially, CDRs are already collected, already privacypseudonymised, and already streamed to operator data lakes. While the minute-scale aggregation horizon makes them unsuitable as a sub-millisecond actuator signal, they are the operationally cheapest source of evidence for the volumetric and behavioural attack regimes (signalling storms, sustained DDoS events, slow-rate campaigns) that develop over seconds to minutes; the closed loop therefore composes CDR-class
telemetry with sub-ms RAN/E2 KPM probes, each at its native timescale. Four properties make CDRs particularly well-suited to closed-loop, MEC-tier security. Granularity. Although individual records aggregate over minutes, per-cell streams expose the spatio-temporal structure that distinguishes a flash crowd from a coordinated signalling storm [19], [63]–[65]. Mobility-awareness. CDRs encode handover sequences and dwell-time distributions, essential for detecting impossibletrajectory or SIM-swap behaviour [66]. Cross-generation continuity. The CDR schema has evolved but not been replaced across 2G/3G/4G/5G, so detectors trained on legacy data generalise to current networks with modest re-engineering [38], [67], [68]; their continuity into 6G is widely expected but not yet standardised. Privacy posture. CDRs do not contain payload and operator pipelines pseudonymise subscriber identifiers, so secondary use for security is closer to the regulatory comfort zone than packet capture or full Deep Packet Inspection (DPI). The limitations are equally important. CDRs miss shortlived, low-volume attacks that complete inside a single reporting window; they cannot resolve protocol-level fields needed to attribute a NAS-attach storm to a specific compromised UE class; and labelled attack data on real CDR is essentially unobtainable outside operator R&D groups [69], [70]. The operational pattern that has emerged is therefore to use CDR telemetry for first-level detection and triage at the MEC, with packet/RAN telemetry pulled in selectively for confirmation [71]—a pattern motivating Sec. V’s benchmark comparing methods across both CDR and packet-based corpora. The CDR feature space that recurs in the surveyed literature is conveniently small for ML deployment. Base features include per-cell call attempt count and success rate, per-cell SMS volume in inbound/outbound directions, internet session counts and aggregate bytes, handover counts in/out of the cell, and the inter-event distribution of these counts within the reporting window [19], [38], [63], [64], [66]. Engineered features add ratios (success/attempt for both calls and attaches), rolling baselines (z-scores against the same hourof-day median over a reference period), and spatial features (neighbour-cell correlation, hot-spot indicator, mobility-graph centrality). The result is a feature vector of 50 to 200 scalar features per cell per window—small enough to feed any of the model families in Sec. V-C and large enough to expose the cross-modality structure that makes CNN, GNN, and transformer architectures useful. These features are already computed inside the operator’s billing pipeline, so the marginal compute cost of using them for security is dominated by the model’s inference cost rather than by feature extraction. The CDR feature surface also aligns naturally with the NWDAF event-exposure interface that 5G has standardised and that 6G will extend. The NWDAF Analytics Logical Function (AnLF) consumes event streams from each Network Function (NF) and exposes derived analytics; the CDR features above can be expressed as NWDAF analytics in straightforward ways, so a CDR-driven detector transfers cleanly to a 6G NWDAF-hosted detector once the appropriate event-exposure
7
profile is standardised (Challenge VIII-D). Christopoulou et al. [39] demonstrates exactly this transition for a DDoS detector running inside a 5G NWDAF; the CDR-trained detectors of [19], [35], [36], [63], [64] are direct architectural antecedents. III. S URVEY M ETHODOLOGY AND C OMPARISON WITH P RIOR S URVEYS A. PRISMA-Based Literature Protocol We followed the PRISMA 2020 systematic review protocol [47]. The search strategy decomposed the keyword matrix {6G, 5G, O-RAN, MEC, CDR, slicing} × {security, intrusion detection, anomaly detection, DDoS} × {machine learning, deep learning, federated learning, LLM, digital twin} into five consolidated Boolean queries that match the five thematic strands of this survey: (Q1) 5G/6G/O-RAN security baseline; (Q2) MEC and edge-intelligence security; (Q3) CDR and cellular-telemetry-driven detection; (Q4) federated learning for cellular security; (Q5) LLM and digital-twin approaches to network security. Each query was executed in April 2026 against IEEE Xplore, Scopus, Google Scholar, the ACM Digital Library, and arXiv, and restricted to English-language peerreviewed venues, top-tier conferences, and arXiv preprints with ≥ 5 citations or active follow-up. Inclusion required (i) explicit treatment of cellular network security or (ii) ML/DL methodology directly applicable to CDR or RAN telemetry; exclusion eliminated pure cryptography, pure physical-layer security, and non-cellular IoT-only studies that did not generalise. Across the five queries, IEEE Xplore returned 10,992 hits (Q1–Q5 combined), Scopus 12,473, Google Scholar 102,900, the ACM Digital Library 6,778, and arXiv 1,090. As illustrated by the complete PRISMA 2020 flow diagram in Fig. 3, these initial hits were rigorously filtered down to the final corpus. The figure details the per-stage breakdown of identification, screening, and inclusion, demonstrating how our explicit focus on latencybounded edge capability eliminated off-topic studies. DOIand title-based deduplication followed by sequential title, abstract, and full-text screening against the inclusion criteria yielded the final corpus of 128 peer-reviewed studies. The bibliography additionally cites 14 standardisation and greyliterature references (3GPP, ETSI, O-RAN, NIST, and MITRE) and three supplementary references, all of which sit outside the PRISMA protocol. Quality appraisal followed standard systematic-review practice: each candidate was rated for venue quality, methodological rigour (clear problem statement, reproducible evaluation, baseline comparison), domain relevance, and recency. We additionally tracked whether each study reported on real operator data, on a 5G/O-RAN testbed, or on a public-only benchmark [20], because the gap between curated-benchmark accuracy and operational-CDR accuracy is one of the main quantitative findings of the survey (Sec. V-E). Forward and backward citation tracking added a small number of very recent O-RAN security studies not yet indexed by all databases at query time, and inter-rater consistency on the abstractscreening stage was checked on a 10% sample by a second author with disagreements resolved by discussion.
Identification: ≈134,200 records (5 DBs)
After deduplication: ≈28,000 unique (excluded ≈106,200: cross-DB duplicates) Title screening: ≈1,800 retained (excluded ≈26,200: off-topic title) Abstract screening: ≈480 retained (excluded ≈1,320: no cellular scope or no ML method) Full-text review: ≈145 assessed (excluded ≈335: no reproducible eval / non-peer-reviewed) Included: 128 studies (excluded ≈17: low rigour / off-scope on re-read)
Fig. 3: PRISMA 2020 flow of the systematic literature review, yielding the final corpus of 128 peer-reviewed studies, augmented by a small number of very recent studies via citation snowballing. A further 14 standardisation/grey-literature and three supplementary references are cited outside the protocol (128+14+3 = 145). B. Taxonomy and Comparison with Prior Surveys We organise the corpus into the closed-loop taxonomy of Fig. 2: threats feed detection (edge-based), which feeds mitigation (network-wide), sustained by enablers (FL, LLM, DT, PQC, ZTA, XAI). This loop-centred view, rather than the pillar list of prior cellular-security surveys, is the organising principle of this work. Table III contrasts our work with the most recent and most cited related surveys across dimensions spanning 5G/6G scope, CPS integration, MEC deployment, CDR-driven detection, DDoS detection, networkwide mitigation, ML/DL and FL methodology, LLM coverage, and—uniquely—whether detection and mitigation are bound into one closed loop and whether downstream CPS impact is evaluated. Competing surveys each populate only a subset of these dimensions: Hoque et al. [20] reach CDR-driven detection and Kumar et al. [72] reach CPS, but always in isolation, and no other entry treats LLM-based methods. Ours is the only fully filled row; the contribution is therefore the joint coverage under a single closed-loop framing, not any individual dimension. We structure the rest of this subsection around five axes on which prior surveys differ from ours: clustering and closed-loop framing, benchmarks, CPS-impact metrics, open challenges, and standardisation alignment. a) Clustering and the closed-loop gap: The closest prior surveys split into two clusters. The detection-centric cluster [73]–[75] catalogues ML/DL detectors but says little about how alerts become actions and does not treat CDR as a first-class signal, while broader 5G-security surveys [7] chart the architecture-and-technology landscape without addressing detector-to-mitigator composition. The mitigation/architecturecentric cluster [8], [15], [20], [28], [30], [37], [72], [76] treats mitigation primitives or FL/ZTA as enablers but stops short of binding them to a concrete edge-detection benchmark. Within the ranked comparison of Table III, neither cluster’s entries combine CDR-driven detection with O-RAN-driven mitigation under a single closed-loop framing, and neither poses its open problems at the level of cross-stage composition.
8
Attack Surface Expansion in 6G CPS
External Adversaries
AI-Model Layer
Adversarial Gradient examples leakage (FL)
Model inversion
Membership inference
Radio Spoofing Jamming
North-bound API hits
NTN long-haul eavesdropping
perimeter trust boundary
ATTACKER ENTRY POINT
b) Benchmarks: The detection-centric cluster reports per-method accuracy on public benchmarks (CICDDoS2019 [77], UNSW-NB15 [78], Bot-IoT [79]) without distinguishing the deployability of the underlying model on a MEC tier; high-accuracy methods that require centralised training on tens of millions of records and tens of seconds of inference are reported on equal footing with compressed quantised models that meet a sub-10 ms MEC budget. The mitigation/architecture cluster reports primitives and reference architectures without binding them to a measured detector performance. Our consolidated comparison (Sec. V-E) addresses both gaps by reporting accuracy and latency together on the same reference benchmarks, segregated by deployment tier (MEC vs. core vs. cloud), and Sec. VI-G sketches an integrated implementation that closes the loop end-to-end. c) CPS-impact metrics: The detection-centric cluster evaluates detectors on accuracy, precision, recall, and F1 against labelled benchmark datasets; the mitigation/architecture cluster evaluates primitives on blockeffectiveness, false-block rate, and recovery time. Neither cluster measures the downstream physical impact—the attackinduced delay to defibrillation in a tele-health slice, the missed V2X hazard warning in a vehicular slice, the unscheduled relay trip in a smart-grid slice—that is the ultimate object of concern for 6G CPS. Our treatment (Sec. VI-D and Challenge C5) elevates CPS-impact metrics to the same standing as detection accuracy and proposes that future benchmarks carry per-incident impact labels produced by operator-CPS-vendor partnerships or by digital-twin replay (Sec. VII-C). Detection accuracy and mitigation block-effectiveness are necessary but not sufficient criteria for closed-loop viability; without CPSimpact metrics, the loop optimises the wrong objective. d) Open challenges: Prior surveys typically enumerate ten to twenty isolated sub-problems (drift detection, FL nonIID heterogeneity, xApp lifecycle, PQC migration, dataset scarcity) without indicating which of them compose into a working closed loop. We compress this enumeration to five challenges (Sec. VIII) chosen for cross-stage leverage: each challenge sits at the intersection of detection, mitigation, and at least one enabler, so that progress on a single challenge advances the loop as a whole rather than improving one stage in isolation. Five is sufficient rather than arbitrary because the criterion admits exactly one challenge per functional bottleneck the loop must clear (detection viability, trustworthiness, cross-operator training, standards-conformant deployment, and honest evaluation) so finer sub-problems fold into whichever of the five they serve; Sec. VIII develops this closed-agenda argument in full. e) Standardisation alignment: Prior surveys treat 3GPP SA3, the O-RAN Security Working Group, and ETSI MEC specifications as background references [7], [8]; our treatment binds each Challenge (Sec. VIII) to the specific standardisation venue at which closure would have the largest operational effect. This is the methodological complement of the closedloop framing and the precondition for turning the research agenda into deployable infrastructure rather than academic prototype.
Tenant-Side Adversaries Application & Rogue OOrchestration Layer RAN xApps
RIC API misuse
Supply-chain MEC tenant compromise exfiltration
Compromised Rogue slice Malicious Cross-slice tenants UEs MEC apps resource abuse operator trust boundary
Insider / Supply-Chain Adversaries
NAS attach RRC setupProtocol & storms failure floods Signalling Layer
GTP-U replay
Slice-hop resource exhaustion
Tampered xApps / rApps
Compromised vendor builds 6G CPS core
Physical & Radio Layer
Radio Jamming
Spoofing
RIS beamhijacking
ISAC waveform perturbation
Shared Physical Infrastructure and Control Timeline
(a) By layer at which the adversary acts
Malicious operator in RIC / MEC
Adversary class
External Adversaries
Stolen credentials / certificates Tenant-Side Insider / SupplyAdversaries Chain Adversaries
(b) By adversary distance from the operator's trust boundary
Fig. 4: Attack surface of 6G CPS along two orthogonal axes: (a) the protocol layer at which the adversary acts (physical/radio up to the AI-model layer, with upward pivot), and (b) the adversary’s distance from the operator’s trust boundary (external, tenant-side, insider/supply-chain). Both axes are detailed in Sec. IV-A. Having situated this work against twelve prior surveys, we now turn to the threat landscape itself: the attack classes that the closed-loop architecture must detect, contain, and recover from. IV. T HREAT L ANDSCAPE FOR 6G CPS This section answers what the closed loop must defend against. We expand the 6G CPS attack surface (Sec. IV-A), enumerate the cellular DDoS vectors that dominate empirical reports (Sec. IV-B), formalise an attacker model (Sec. IV-C), survey the public dataset landscape that detectors are evaluated on (Sec. IV-D), and align the picture with MITRE ATT&CK so that downstream detection and mitigation map cleanly onto a common adversarial vocabulary (Sec. IV-E). A. Attack Surface Expansion in 6G CPS The pillars of Sec. II-A together produce an attack surface that no single perimeter can wrap around, which is the operational reason the closed-loop pipeline is necessary rather than optional. We frame this surface along two orthogonal axes, summarised in Fig. 4. By layer at which the adversary acts [Fig. 4(a)]. Physical and radio layer attacks include jamming, spoofing, RIS beamhijacking and pilot contamination on the joint active-passive beamforming substrate [54], adversarial sensing—including waveform-level perturbations against ISAC and unauthorised surveillance via reused sensing beams—and long-haul eavesdropping and jamming on NTN backhaul [55], [56]. Protocol and signalling layer attacks include NAS attach storms, RRC setup-failure floods, GTP-U replay, cross-slice resource exhaustion (including slice-hopping variants), isolation bypass, and cross-RAT downgrade via EPS-fallback or rogue base stations [33], [34], [53], [80], [81]. Application and orchestration layer attacks include rogue O-RAN xApps, RIC API misuse via the E2/A1 interfaces,
9
TABLE III: Comparison of This Survey with Existing Related Surveys Reference Ahmad et al. [7] Pang et al. [73] Porambage et al. [8] Mothukuri et al. [76] Nguyen et al. [15] Thakkar & Lohiya [74] Blika et al. [37] Nahar et al. [30] Hoque et al. [20] De Alwis et al. [28] Agarwal et al. [31] Kumar et al. [72] This Survey
Year 5G/6G CPS MEC CDR DDoS Det. Net. Mitig. ML/DL FL LLM Closed-Loop† CPS-Impact‡ 2019 2021 2021 2021 2022 2022 2024 2024 2025 2026 2026 2026 2026
• ◦ • ◦ • ◦ • • • • • • •
◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ • •
• ◦ • ◦ • ◦ • ◦ • • • • •
◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ • ◦ ◦ ◦ •
• ◦ • ◦ ◦ • • ◦ • • ◦ ◦ •
◦ ◦ ◦ ◦ ◦ ◦ ◦ • • • • • •
• • • • • • • ◦ • • ◦ • •
◦ ◦ • • • ◦ • ◦ ◦ • ◦ • •
◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ •
◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ •
◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ •
A filled circle (•) is awarded where a survey provides defensible coverage of the dimension (dedicated discussion, taxonomy, comparative analysis, or substantive treatment); an open circle (◦) indicates the dimension is absent or appears only in passing. 5G/6G: substantive 5G/6G security-architecture coverage. CPS: cyber-physical or safety-critical control systems treated as the protected asset, beyond generic IoT. MEC: edge/MEC analysed as an enforcement or detection tier. CDR: call-detail-record or equivalent cellulartelemetry features used as a security signal. DDoS Det.: DDoS/flooding detection methods. Net. Mitig.: network-wide mitigation (SDN/NFV/O-RAN actuation), not detection alone. ML/DL: ML/DL methodology surveyed comparatively. FL: federated learning as a security-relevant training paradigm. LLM: large-language-model methods as a distinct topic. † Closed-Loop (Det.→Mitig.): filled only where the survey treats detection and network-wide mitigation as one bound feedback loop (alerts→actuation→retraining), not as separate chapters. ‡ CPS-Impact: filled only where the survey evaluates downstream physical/safety impact (e.g., missed hazard warning, relay trip, control-loop deviation), beyond accuracy/F1 or block-rate metrics.
supply-chain compromise of vendor-supplied components, and exfiltration through MEC-hosted third-party applications [9], [11], [26], [27]. AI-model layer attacks—model poisoning at training time, adversarial examples at inference time, gradient leakage from FL clients, model inversion, and membership inference—turn the defender’s own intelligence into a target [32], [82], [83]. The MEC tier is doubly exposed because it terminates user-plane traffic and runs security models, making it the natural target for both denial-of-service (DoS) and model-extraction attacks. By adversary distance from the operator’s trust boundary [Fig. 4(b)]. External adversaries hit the radio and the public application surface and have been the historical focus of cellular IDSs; tenant-side adversaries leverage compromised UEs or rogue slice tenants and are the dominant operational threat as IoT density grows; insider and supply-chain adversaries operate from inside the MEC or RIC software stack and are the emergent threat that O-RAN openness has materialised. Defences that compose across these classes (multi-tier sensing, federated cross-operator learning, and attestable xApp lifecycles) are what the closed loop is designed to provide. Because both decompositions describe the same incidents from different angles, the closed loop has to hold along the layer axis and the trust-boundary axis at the same time: every attack ultimately traverses the same physical infrastructure and the same control timeline, so a defence that closes one axis while leaving the other open does not contain the threat. B. DDoS Vectors Against 6G CPS DDoS dominates the 6G CPS threat landscape because each control-plane procedure is computationally expensive on the operator side, giving the attacker a high amplification factor: a small UE-side cost translates to a large operator-side cost. Six attack vectors recur in the literature. Signaling Storms. A botnet of compromised IoT devices triggers excessive NAS Attach/Detach or RRC Setup/Release sequences, exhausting gNB RRC contexts and AMF/Session Management Function (SMF) control-plane CPU, and starving URLLC slices, even at modest message rates [27], [33], [34].
GTP-U Flooding. Spoofed or replayed GTP-U tunnels [39] overwhelm the User Plane Function (UPF), disrupting V2X latency budgets and consequently block safety-of-life messages. SMS Flooding on mMTC Slices. SMS remains critical signaling channel for two-factor authentication and emergency alerts, leaving healthcare and similar notification-bound CPS especially exposed [19], [29], [81]. Silent Call Attacks. Calls initiated and dropped after minimal ring time consume RRC and AMF resources without producing billable minutes; the resulting CDRs carry a distinctive short-duration high-frequency signature [19]. Blended/Polymorphic Attacks. Modern campaigns interleave the above vectors and adapt their mix in response to mitigation, with reinforcement-learning-driven mutation now demonstrated on the operational frontier [83]. Slice-Hop and Cross-Slice DDoS. An adversary controls UEs that hand over between slices and concentrates load on whichever slice is currently most cost-effective to disrupt; because slice isolation on shared O-RU/O-DU resources is statistical rather than physical, the target slice can be degraded without direct attachment [84]. Predictive load-forecasting defences pre-emptively reallocate resources before the attack peaks [71]. These vectors are normalised in the literature as binary or multi-class classification under class-imbalanced training conditions [19]; the methodological treatment is deferred to Sec. V. Fig. 5 renders the attack–defence pairing that motivates the closed-loop pipeline of this survey: the upper chain traces the attacker’s path to physical impact, the lower chain previews the detection–mitigation–feedback loop elaborated in Secs. V and VI, and the dashed cross-links mark the telemetry and actuation that bind them. C. Attacker Model and Profiles We characterise an attacker by the triple (κ, ρ, π): capability κ ∈ {κ0 , κ1 , κ2 } (script-kiddie, organised, nation-state), goal ρ ∈ {GD , GE , GP } (denial, exfiltration, physical impact), and persistence/position π ∈ {πexternal , πedge , πcore } [5]. Table IV instantiates five representative profiles spanning the
10
Attack Lifecycle Attacker (Botnet)
coordinate
Attack Vector Selection
launch
5G/6G RAN Impact
propagate
CPS Slice Disruption
cascade
Physical Consequence
CDR/KPI Monitoring
features
ML-Based Detection
scores
Anomaly Classification
policy
SDN/NFV Mitigation
update
Feedback & Adaptation
Closed-Loop Defence
Fig. 5: DDoS attack lifecycle in 6G CPS. The upper red chain traces the attacker’s progression from vector selection through RAN impact and slice disruption to physical-world consequences. The lower blue chain mirrors the closed-loop defence: CDR/KPI monitoring feeds ML-based detection, anomaly classification, SDN/NFV mitigation, and feedback-driven adaptation. Black dashed arrows mark observational coupling (telemetry feeding the detector and classifier); blue dashed arrows mark active coupling (mitigation acting on the RAN and the feedback-driven model-adaptation loop). easy (κ0 , GD , πexternal ) to the hardest (κ2 , GP , πcore ) corner of this cube. Why even modest attack rates collapse a slice. For attackrate λ and per-message control-plane cost c, the slice’s effective service rate µeff = µ − λc falls below its required rate when λ > (µ − µreq )/c. Because cellular signalling traverses long chains of cryptographic and database operations across the AMF, SMF, AUSF, and Unified Data Management (UDM) functions, c is large and even modest λ collapses µeff to zero— which is why detection must observe protocol-level rates rather than byte volumes [19], [33], [34]. Reinforcement-learningdriven adversaries [83] extend this picture: the attack policy πatk (s), where s is the attacker’s observed state including the defender’s recent mitigation actions amit , adapts in response to those actions, so a static threshold-based defence is provably out-evolved. The closed-loop pipeline of Sec. VI is therefore a structural requirement, not a convenience. Time horizon stratifies the closed loop’s value. Three regimes recur. Single-shot attacks (row 1–2 in Table IV) are easy to detect; mitigation latency, not detection accuracy, is the binding constraint. Campaign attacks (organised criminal DDoS-as-a-service, prolonged signalling-storm campaigns, rows 2 and 5) require drift-aware online learning because the attacker rotates vectors and source pools. Persistent attacks (row 3–4: nation-state Advanced Persistent Threats (APTs), supply-chain compromise of vendor-supplied O-RAN components, insider threats from compromised MEC software) minimise volumetric signal and move laterally through trust boundaries; detection requires multi-modal CDR + RAN + audit-trail fusion plus federated cross-operator correlation, since a single operator’s telemetry is rarely sufficient. Cross-profile composition. A single incident may begin as a script-kiddie probe (row 1), be folded into an organisedcriminal campaign (row 2 or 5), and provide reconnaissance that a nation-state actor later exploits through a dormant xApp backdoor (row 4). A detector that treats each profile in isolation reports three separate incidents and misses the compositional attack. The operational response is to maintain a per-source and per-slice reputation vector that persists across detection windows: a rate-limit against a third-time offender is qualitatively different from the same rate-limit against a first-
TABLE IV: Representative Attacker Profiles for 6G CPS Profile Script kiddie Botnet operator Insider threat Nation-state APT Cyber-criminal org.
Cap. Goal Persist. Primary Vector κ0 GD One-shot Volumetric DDoS κ1 GD One-shot Signaling storm, GTP-U flood κ2 GE APT Data exfil., model poison κ2 GP APT Supply chain, xApp backdoor κ1 GE APT Ransomware, slice-hop
CPS Target General service V2X, smart grid Industrial IoT Critical infra. Healthcare CPS
time one. Reputation decay, cross-operator reputation sharing, and the integrity of the reputation store against poisoning feed forward into Challenges C1 and C3 (Sec. VIII). Adversarial ML threats (evasion, poisoning, model extraction) are documented in the surveyed literature [32], [85], and the agentic-operations literature complements this with prompt-injection, telemetry-manipulation, and unsafetool-invocation threats specific to LLM-mediated control loops [86]; their composition against the closed loop is the subject of Challenge C1 (Sec. VIII-A). D. Public Dataset Landscape: Threat Evidence and Detection Benchmarks Public datasets shape what detectors are trained on (Sec. V), and the cellular-security community remains structurally short of 6G-native data. Table V surveys twelve corpora that recur across the surveyed literature. Three observations matter for the rest of this article. First, only the Telecom Italia Milan & Trentino CDR set [38] provides authentic operator-grade CDR telemetry, yet it carries no labelled attacks; researchers therefore inject synthetic anomalies into it, biasing benchmark comparability, although recent unsupervised reference-only pipelines [41] demonstrate that meaningful evaluation against curated event groundtruth (e.g., match schedules, exhibition calendars) is feasible without injected attack labels. A recent network-digital-twin evaluation on the same Milan grid reports 96.7%/98.3% prediction accuracy and 88–95% anomaly-prevention rates [87]. Second, the legacy intrusion corpora that still dominate ML evaluations—KDD’99 [88], NSL-KDD [88], UNSWNB15 [78], CICIDS2017 [89], CICDDoS2019 [77], BotIoT [79] and TON IoT [90]—predate 5G and contain no
11
New Radio (NR)/O-RAN signalling, no GTP encapsulation and no slice context. Third, the only datasets with native 5G or O-RAN provenance are recent: 5G-NIDD [91] captures real-radio traffic over an operational 5G test network at the University of Oulu, the NANCY corpus [92] covers an ORAN 5G coverage-expansion testbed, and NetsLab-5GORANIDD [61] adds synchronised packet-level and radio-telemetry data from a live OAI-based O-RAN deployment. None of the three yet covers the full slice-aware, closedloop scenario that 6G will require—a gap that we revisit as Challenge C5 in Sec. VIII. We also note that Generative AI (GenAI)-based synthesis of mobile-network data—spanning per-cell traffic volumes, base-station-association trajectories, and application-usage records—is an emerging but still-private alternative [69], [93], [94], with no public release at the time of writing. E. MITRE ATT&CK Mapping The MITRE Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) framework [95], [51, Ch. 2] assigns each adversary technique a stable identifier that detectors, orchestrators, and SOCs can all key against. Three sub-matrices apply to 6G CPS: the Enterprise matrix (T-codes) for IT-layer techniques, the Industrial Control Systems (ICS) matrix (T0codes) for impact on the physical process layer, and the Adversarial Threat Landscape for AI Systems (ATLAS) catalogue (AML codes) for attacks against deployed ML models [96]. We mix all three in Table VI because a single 6G CPS incident typically chains an Enterprise entry, lateral movement through a MEC host, an Adversarial-ML step that blinds the detector, and a final ICS-layer impact on the controlled process—no individual matrix covers this full chain. The resulting vendorneutral vocabulary is the keying input used by the closed-loop orchestrator of Sec. VI to select an appropriate mitigation playbook per detected technique class; for each technique, Table VI therefore lists the observable telemetry that surfaces it (Sec. V) and the candidate mitigation primitive that answers it (Sec. VI). With the threat surface mapped, the next section examines how edge-side detection (the first arrow of the closed loop) converts CDR and RAN telemetry into actionable alerts. V. E DGE -BASED D ETECTION FOR 6G CPS This section instantiates the detect stage of the closed loop introduced in Sec. I-B. We treat CDR-driven anomaly detection and DDoS classification under one chapter because, in operational 6G, the same MEC host runs the same model families over shared telemetry to separate benign, anomalous, and attack-class flows; the historical split between the two bodies of literature is an artefact of dataset provenance, not of method. We organise the survey by method family (Secs. V-C– V-H), collect attack-type-specific findings in Sec. V-D, and report a unified deployment-aware benchmark in Sec. V-E. A. Why MEC Is the Right Detection Tier? Sec. II-C established the architectural case for MEC as the enforcement plane and Table VII consolidates the deploymenttier trade-off the rest of this section optimises against; the
Holt–Winters [E] Statistical Z-score / EWMA [E] Random Forest (sup.) [E] Classical ML
Isolation Forest [E] k-means (unsup.) [E]
Edge-Side Detection
GNN (topological) [C] Transformer (attn.) [C] Deep Learning LSTM (temporal) [E] Autoencoder (recon.) [E] Split Learning [E] Federated / PP FedAvg + DP [E]
Legend: [E] = edge-deployable (≤ 8 GB RAM, ≤ 4 W);
[C] = cloud-only.
Fig. 6: Taxonomy of edge-side detection method families surveyed in Sec. V. Edge-deployable leaves ([E]) are those validated on ≤ 8 GB RAM / ≤ 4 W power envelopes; cloud-only leaves ([C]) currently exceed these envelopes on commodity MEC hardware. recurring pattern is to train large, deploy small, and aggregate hierarchically when local visibility is insufficient. Recent measurements on ARM-SoC platforms [41] reinforce this case quantitatively: a single Apple M4 Pro completes a full per-cell anomaly-detection pipeline over the Milan CDR dataset ∼35% faster than a 36-core x86 HPC node at ∼96% lower energy, suggesting that ARM-based MEC hosts can sustain CDR-scale analytics at the data source without offloading to the core. B. Detection Problem Formulation Using the notation of Sec. II, the surveyed literature collapses onto three formulations of the detection problem [75]: d • binary detection fbin : R → {0, 1}; d • multi-class classification fmc : R → {1, . . . , K}, which exposes the attack type and is the precondition for typeaware mitigation [19]; • unsupervised anomaly scoring s(x), typically a reconstruction error from an autoencoder gθ trained on benign traffic. A deployed pipeline must additionally meet a streaming budget ∆tinfer ≤ τmax , with τmax in the single-digit-millisecond range for URLLC slices and tens of milliseconds for eMBB and mMTC slices. Class imbalance is severe (attack samples typically below 1%), so weighted cross-entropy or focal-loss [71], [97] objectives dominate. The feature surface fuses CDRderived signals (per-subscriber call rate, entropy of destination numbers, ratio of failed-to-successful calls, average interarrival time, and burstiness indices) with RAN- and NWDAFside signals (NAS message rate, RRC setup-failure ratio, GTPU tunnel utilisation, and per-slice resource consumption) [39]. Because these streams arrive over the disjoint native interfaces of the MEC telemetry vantage (Sec. II-C), multi-modal fusion (Sec. V-G) requires explicit time-alignment and trust-domain handling [12]. C. Method Families The MEC host runs every detector family on the same CDR/RAN telemetry, so the families differ not in the data
12
TABLE V: Public Datasets for Network Intrusion and Anomaly Detection Used in the Surveyed Literature Dataset KDD Cup 99 [88] NSL-KDD [88] UNSW-NB15 [78]
Year 1999 2009 2015
Type Network Network Network
Net. Gen. Pre-4G Pre-4G General
Avail. Public Public Public
CDR No No No
∼2 months (Nov 2013–Jan Public 2014); two regions: Milan (10,000-cell grid) & Province of Trentino (6,575-cell grid) General 8 families: Brute Force, Heartbleed, DoS, ∼2.8M flows Public DDoS, Web (XSS/SQLi/BForce), Infiltration, Botnet, PortScan General 13 reflection/exploitation DDoS variants: Tens of M records Public DNS, LDAP, MSSQL, NetBIOS, NTP, PortMap, SNMP, SSDP, SYN, TFTP, UDP, UDP-Lag, WebDDoS IoT (LAN) 4 categories / 10 sub-types: Probing, DoS, ∼73.36M records Public DDoS, Information Theft IoT/IIoT 9 types: Scanning, DoS, DDoS, Ransomware, ∼22M log records Public Backdoor, Injection, XSS, Password, MITM 5G O-RAN 7 types: Reconnaissance, UDP/TCP- ∼588K flow samples (84 fea- Public (SA) Connect/SYN scan, SYN/ICMP/HTTP tures) flood, Slow-rate DoS 5G NSA 8 types: ICMP/UDP/SYN/HTTP flood, ∼1.22M flows Public Slowrate DoS, Torshammer, SYN/TCPConnect/UDP scan 5G O-RAN DoS/DDoS (SYN/UDP/TCP/ICMP/HTTP, ∼28M packets / ∼1.72M Public (SA) Slowloris); port & OS scans; flows / ∼45K radio-telemetry SSH/FTP/HTTP brute force; web (DirBust, records SQLi, XSS) 4G/5G Synthetic traffic / trajectory / app-usage; no Varies (no public release) Limited attack labels
Yes
Telecom Italia Milan & 2015 Operator CDR 3G/4G Trentino CDR [38] (multi-source)
CICIDS2017 [89]
2017 Network
CICDDoS2019 [77]
2019 DDoS
Bot-IoT [79]
2019 IoT
TON IoT [90]
2022 IoT/IIoT
NANCY 5G [92]
O-RAN 2024 5G/O-RAN
5G-NIDD [91]
2025 5G
NetsLab-5GORANIDD [61]
2025 5G/O-RAN
GenAI-synthesised mo- 2025 Synthetic bile data [69] mobile-network data
Attack Types DoS, Probe, R2L, U2R DoS, Probe, R2L, U2R 9 families: Fuzzers, Analysis, Backdoors, DoS, Exploits, Generic, Reconnaissance, Shellcode, Worms None (benign telemetry only); multi-source side data: news, weather, social, electricity (Trentino), precipitation
TABLE VI: MITRE ATT&CK Mapping for 6G/CPS Threat Landscape, Linked to Closed-Loop Telemetry and Mitigation Primitives Tactic
ID
Technique
Impact
T1498
Network DoS Flooding
Defense Evasion
T1036
Masquerading Spoofed UE identity in 6G slices
Lateral Movement
T1210
Exploit Remote Services
MEC-tocore lateral movement
Collection
T1040
Network Sniffing
CDR/telemetry interception at aNB
Impair Process Control
T0836
Modify Parameter
CPS sensor value manipulation
Inhibit Re- T0881 sponse Initial Ac- T1195 cess
ML Attack AML. T0043
6G/CPS Rele- Observable vance Telemetry Volumetric / DDoS against aNB/MEC
Service Stop
Slice isolation bypass / shutdown Supply Chain O-RAN multiCompromise vendor supply chain Physical Ad- Evasion versarial Per- of MECturbation hosted DL models [32]
Candidate Mitigation Primitive GTP-U SDN rateutilisation, limit / per-slice load redirectcurve to-scrubber IMSIO-RAN distribution xApp UE entropy, RRC suppression setup-failure ratio MEC audit- ZTA retrail, east-west authentication, flow logs microsegmentation Interface tap PQCcounters, protected anomalous telemetry read patterns channel Sensor-value Slice residuals, isolation, control-loop controldeviation message integrity check Per-slice NFV failover, availability, slice reKPI drop instantiation xApp/rApp at- xApp lifecycle testation logs attestation, signed onboarding Detector Adversarial confidence retraining, drift, input- FL/DT replay perturbation score
they consume but in the inductive bias each encodes and— decisively for this survey—in whether that bias fits a latency, memory-, and label-constrained edge tier. We therefore
Size ∼4.9M conn. records ∼148K records ∼2.54M records
No
No
No No No
No
No
No
TABLE VII: Detection Deployment Tiers: Latency, Compute, and Visibility Attribute On-Device MEC Cloud Detection latency <1 ms 1–10 ms 80–120 ms Compute capacity Very limited Moderate Abundant Model complexity Tiny ML Medium DNN/CNN Large / ensemble Network visibility Single UE Cell/site level Network-wide Energy budget Battery Mains powered Mains powered Suitable CPS apps Wearables URLLC, V2X Analytics, forensics
compare families along four deployment axes rather than re-deriving their mechanics (assumed in Sec. II): inference latency, memory/power footprint, labelled-data requirement, and interpretability. Fig. 6 organises the families by mathematical framework and annotates the edge-deployable ([E], ≤ 8 GB RAM / ≤ 4 W) versus cloud-bound ([C]) boundary; the surveyed evidence per family follows. Statistical and Classical ML. Fixed-threshold rules and Principal Component Analysis (PCA)-based residual methods provide interpretable but linear baselines [75], [98]. FrancoValiente et al. [41] learn per-cell truncated-Singular Value Decomposition (SVD) subspaces from stable reference weeks and score anomalies as squared residuals with a MAD-based dimensionless severity. Density-estimation approaches [63] fit a Gaussian model to normal-profile CDR features and flag low-likelihood grids, achieving ∼93% accuracy on the Milan dataset. Clustering with k-means or Density-Based Spatial Clustering of Applications with Noise (DBSCAN) has been used by Sultan et al. [66] (90%) and Aziz et al. [67] (96% on 5G CDR). Hussain et al. [64] introduced a semi-supervised cluster-then-classify pipeline at ∼92% accuracy, and Shone et al. [99] proposed a non-symmetric deep autoencoder as a complementary deep-shallow hybrid. This family persists in production for three deployment rea-
13
sons [98]: tree ensembles and PCA residuals run in microseconds per record on commodity MEC hardware; a per-feature threshold is auditable in a way a deep model is not, which matters under the lawful-interception and slice-service-level agreements (SLAs) regimes operators face; and these methods stabilise on markedly smaller corpora than the more datahungry deep models [18], [98], [100]. The pattern recovered in operational telco SOCs is therefore a layered ensemble— a statistical front-line filter at the MEC with deep models invoked only on flagged segments, the operational realisation of which is a classical-plus-deep hybrid pipeline [99]. Deep Feedforward and Convolutional Networks. An Llayer Multilayer Perceptron (MLP) on Milan CDR achieved 94.6% accuracy and 1.7% FPR [65], and a deep CDR-based anomaly detector reached 98.8%/0.44% on a real operational CDR set [35]. CNNs exploit the spatial correlation between a cell and its geographic neighbours, which flat MLPs discard; the spatially-aware CDR CNN of [36] attained 96% on real operator data, while ResNet-50 and a custom Deep Rudimentary CNN (DRC) model [19], [101] pushed multi-class DDoS classification past 91% on real CDRs. Lightweight CNNs (LUCID [40]) reach 99% on CICDDoS2019 at parameter counts compatible with MEC deployment. The trade-off is now well-characterised: deeper feedforward and CNN models gain roughly 1–3 percentage points of accuracy over shallow baselines on identical features at 3–10× inference latency [18], a cost that quantisation-aware training and structured pruning recover to a sub-percentage-point accuracy loss, enabling the sub-10 ms inference reported by MEC-deployed detectors [40]. The empirical lesson, recovered across studies, is that the binding constraint at deployment time is feature engineering and class balance, not network depth; the dominant 5G-CPS designs accordingly reuse deep architectures from the encrypted-traffic lineage [102] rather than reinventing them. Recurrent and Hybrid Spatio-Temporal Models. Recurrent and convolutional-recurrent hybrids matter because cellular attacks have temporal structure on multiple scales: a signalling storm builds over seconds, a silent-call campaign over minutes, and a slow-rate slice-hop attack over hours. Roopak et al. [103] reported 97%+ on CICIDS2017 with a CNN–LSTM hybrid; Elsayed et al. [104] combined an RNN–autoencoder with softmax (DDoSNet) at 99% on CICDDoS2019; and ConvLSTM with transfer learning [105] achieved 99% on a public benchmark, with the typical gain over single-modality baselines at 2–5 points on multi-class DDoS data [73], [84], [103], [105]. The one clearly edge-validated case is the LSTM–RNN autoencoder hosted as a Near-RT RIC xApp on a real over-the-air O-RAN testbed [42], which delivered 97.5% accuracy at sub-second closed-loop mitigation latency. The cost is sequence-handling latency, which on commodity MEC hardware bounds the practical sequence length to a few hundred timesteps; longer-horizon attacks are typically captured by hierarchical models that summarise short windows before feeding a slower recurrent layer. Self-Supervised, Autoencoder, and Generative Adversarial Network (GAN)-Augmented Detection. The labelled-CDR scarcity problem is structural and operational [64], [98], [106]:
production networks emit terabytes of unlabelled CDRs daily, but labelling demands scarce expertise. Plain autoencoders trained on benign traffic [107] learn normal manifolds and flag high-reconstruction-error inputs; Mirsky et al. [108] (Kitsune) demonstrated online unsupervised ensemble autoencoders running at line rate. Conditional GANs and tabular diffusion models [93], [94], [109] synthesise minority-class attack samples to improve recall on rare attack types under bounded re-identification risk (surveyed in [70]); Duan et al. [69] additionally use diffusion to denoise coarse cellular trajectories into realistic mobility-aware traffic, and GANbased detectors export the trained discriminator as the operational classifier [82], [93]. Contrastive learning and pseudolabelling progressively expand the labelled pool from a small seed [106]. The attractive deployment property across this family is that no attack labels are needed and the compressed autoencoder form is edge-feasible ([E]); the cost is a nontrivial threshold-tuning problem, since the reconstruction-score distribution is heavy-tailed even on benign traffic, and GAN generators remain a training-time aid rather than an inferencetime edge detector. Graph Neural Networks. GNNs exploit the network’s structural information that flat features discard [50], [110], supporting both per-flow detection (nodes are flows or endpoints, capturing multi-hop attack patterns invisible to single-flow features) and per-cell detection (nodes are aNBs, capturing spatial diffusion of an attack across the radio plane). Lo et al. [111] ported GraphSAGE to IDS and gained 3–5 percentage points over flat baselines on identical features. The decisive deployment property is inductivity: 6G sees new UEs and new cells arrive continuously, so the inductive GraphSAGE setting Lo et al. demonstrate is workable in production whereas any transductive method that requires retraining on each topology change is not. Survey work by Rahmani et al. [110] catalogues the design choices (GCN vs. GraphSAGE vs. GAT). On current commodity MEC hardware GNN variants remain cloud-bound ([C]) pending compression advances. Transformers and Attention. Kim and Thulasiraman [112] used a transformer autoencoder on 5GAD-2022 flow features, reaching F1 = 0.882 and AUC = 0.90 without outlier removal; Liu et al. [113] introduced multi-domain (timeplus-frequency) multi-head attention with a packet-level timeseries feature meter (TS-NFM) extractor for early-stage NIDS, classifying intrusions from the first ten packets at 50,000× earliness over conventional flow-based detectors; and Umer et al. [114] combined wrapper-based feature selection with a multi-head attention transformer (and SMOTE rebalancing) to reach 93% on UNSW-NB15 while exposing which features and time positions drove each decision. Adjewa et al. [115] ported BERT to a federated NIDS for 5G, showing a BERTstyle backbone can be trained across distributed edge clients without sharing raw traffic and remain deployable through linear quantisation; Houssel et al. [116] combined LLMs with explainability for human-readable NIDS reports; and on the synthesis side, Duan et al. [69] fine-tune a Llama2 backbone (AppSyn) to generate operator-realistic CDRlike records under DCR-balanced privacy. Two deployment consequences recur: attention exposes which feature positions
14
and time-steps drove a decision, providing an XAI primitive without an external explainer [114], and multi-head attention captures simultaneous patterns at different temporal scales, the right inductive bias for blended/polymorphic attacks [49]. The cost is quadratic complexity in sequence length [49, Table 1], which bounds practical context windows on MEC hardware to a few hundred CDR samples; sliding-window and linearattention variants are the standard remedies, and ViT-style image-of-traffic representations are still maturing. Native Cellular Interfaces (NWDAF, P4, Cross-Layer). Three native-interface variants close out the method-family inventory and feed directly into the multi-modal fusion treatment of Sec. V-G: NWDAF-resident detectors [39], P4 data-plane detectors [117], and cross-layer SDN/NFV/AI architectures spanning RAN, transport, and core [43], [118]. D. Attack-Type-Specific Detection Six task formulations recur in the literature; each is a specialisation of the method families above, and Table VIII compares them along the dimensions that determine MEC deployability. E. Unified Comparison of Edge-Based Detectors Table IX consolidates the representative numbers from the literature into a single chronological view that subsumes both the cellular-anomaly and the DDoS-IDS evidence bases. To keep these figures from being read out of context, each row carries an evidence tier (E1–E4, defined in the table footnote) that records how the result was validated, from real operator data (E1) down to simulation or synthetic labels only (E4). Read this way, the table exposes a recurring pattern: the nearceiling accuracies concentrate at E3 on saturated public benchmarks such as CICDDoS2019, where several models report 99% or above, whereas the E1 operator-CDR studies tend to report more modest, operationally faithful figures—typically in the low-to-high nineties—examined in Challenge VIII-E. The contrast is one of tendency rather than a clean partition, since validation realism, not headline accuracy alone, governs how each figure should be weighed. Three trends bear directly on the closed-loop thesis of this survey. First, headline accuracy on curated benchmarks is now saturated above 97%, but performance on real operational CDR data sits in the 91–98% band; this gap is a benchmark artefact, not a methodological gain (Challenge C5). Second, edge-deployable methods are now common, vindicating the MEC-tier thesis of Sec. V-A, and the most recent contributions [39], [42], [117] are increasingly evaluated on native 5G/O-RAN testbeds rather than legacy public benchmarks. Third, methods evaluated on native O-RAN hardware [42] report sub-second detect-to-mitigate latency, the first end-toend empirical confirmation that the loop-time argument of Sec. I-B is achievable with current technology. Reporting practice, however, remains uneven: the community standardised on accuracy and F1 in the late 2010s, but three deployment-relevant metrics remain underreported [82]—calibrated false-positive rates, tail-latency percentiles (especially p99, which determines MEC deployability), and energy-per-decision under perturbation. Dataset
CDR / O-RAN
Preprocess
ML Infer.
Score s
s > θ?
Alert
SDN Ctrl
Mitigate
FL/DT retrain
Fig. 7: MEC-based detection pipeline. The Alert→SDN Ctrl→Mitigate stages are the detector-to-orchestrator handoff formalised in Sec. VI-B; the dashed feedback path is the closed-loop FL/DT retraining channel of Sec. VII. descriptors that publish per-class precision/recall plus training and prediction times [91] provide the baseline against which future 6G-native datasets should be benchmarked. The most operationally-deployable contributions [42] are precisely those that report end-to-end mitigation latency on a CPS-relevant slice rather than single-stage accuracy; the lesson for future work is that the deployment-relevant figures of merit are loop-time and CPS-impact harm reduction, and the structural fix to cross-corpus incomparability is the consortium-released benchmark proposed in Challenge C5. F. End-to-End MEC Detection Pipeline Fig. 7 renders the pipeline that each detector of the families above must instantiate end-to-end. Ingestion subscribes to the CDR, RAN E2 KPM, and NWDAF streams at the MEC vantage (Sec. II-C) [12]; feature extraction computes the engineered signals of Sec. II-D; inference—the latency-critical stage—runs the compressed detector to produce per-cell or per-flow scores [59]; and calibration/thresholding maps scores to slice-aware decisions s(x) > θ, trading detection recall against the False Block Rate (FBR) defined in Table X. A positive decision is handed off to the orchestrator as an alert whose tuple and per-field action semantics are specified in Sec. VI-B. The end-to-end budget sums these stages and the orchestrator’s decision time, which the slice safety envelope must bound; the dashed path in Fig. 7 closes the loop into FL/DT retraining (Sec. VII). G. Multi-Modal Fusion Across CDR, RAN, and Core Telemetry A single telemetry stream rarely contains enough signal to separate sophisticated 6G CPS attacks from benign traffic. Drawing on the three operator-facing modalities of the MEC telemetry vantage (Sec. II-C), each stream contributes a distinct grain: CDR captures aggregate per-subscriber behaviour at the call-event scale, RAN E2 KPM exposes per-UE RRC state transitions and scheduler decisions, and NWDAF event exposure carries cross-NF correlations from the core. Each modality has blind spots: CDR cannot see RRC failure cascades that never reach billing, RAN E2 KPM cannot see corenetwork amplification, and NWDAF events do not yet expose cell-level mobility patterns at fine grain. The multi-modal fusion problem is to combine these streams in a way that
15
TABLE VIII: Attack-type-specific detection task formulations, with discriminative telemetry features, representative methods, best reported performance (as cited), and MEC deployability. Task
Discriminative features
Representative methods
Sleeping-cell detection Traffic-surge / congestion (DDoS vs. flash crowd)
Per-cell reconstruction error over window T ; absence of alarm; spatial neighbour state Volumetric rate, IAT, src-IP entropy, slice-aware load curves
Fault localisation / rootcause attribution Signalling-storm detection & mitigation SMS-flooding / silent-call detection Slice-specific / blended / polymorphic DDoS Trust management / collaborative detection
Causal-DAG residuals; Grad-CAM/SHAP saliencies over cell grid NAS rate per cell; RRC setup-failure ratio (Msg5/Msg3); Msg5/Msg4 ratio; mean connection hold Per-subscriber send rate; destination-entropy; messagelength dist.; short-call bursts Per-slice flow features; cross-vector correlations; wrapper-FS attributions Trust-weighted alert aggregation; multi-node SDN votes; reputation tokens
Statistical/Gaussian [63]; MLP [65]; CDRDNN [35]; spatial CNN [36] ARIMA → classifier [64]; k-means + NN forecaster [66]; energy-aware ARM subspace pipelines (event-side only) [41] Granger / causal discovery; XAI attribution [114] RRC-storm xApp [34]; 5G-SPECTOR [27]; blockchain-5GAKA [33] ResNet/CNN [19]; dilated CNN Slice-aware AI-IDS [53], [71], [114], [119]; mitigation taxonomy [20] Trust-CIDN survey [120]; SDN-coop [121]
Best reported performance 93.0–98.8% acc.
MECdeployable ✓
92–96%† acc.
✓
N/A (ranking task)
✓
90 ms detect. lat. / 1.2% CPU [34]‡ 91–94% acc.
✓
93–96%+ acc.
✓
N/A (consensus task)
Partial (blockchain latency)
✓
†
Hussain et al. [64] additionally report a 40% relative reduction in false positives over single-stage baselines. ‡ Nguyen et al. [34] report 90 ms detection latency at 132 Msg3s/sec attack rate, with 60 ms mitigation window before gNB overload, 1.2% CPU and ∼0% memory overhead on an OAI+FlexRIC testbed; reduces gNB unavailability (95% under unmitigated attack) by enabling early xApp-triggered response. 5G-SPECTOR [27] and blockchain-5GAKA [33] report distinct performance dimensions (layer-3 attack-class detection and registration-storm mitigation, respectively).
preserves the latency budget of the MEC tier while extracting cross-modality joint signals [39]. Three fusion architectures recur in the surveyed literature. Early fusion concatenates per-modality features into a single input vector for a unified detector; this is the simplest pattern and gives the strongest gains when the modalities are timealigned and sampled at compatible rates, which is usually true for CDR and RAN at the MEC. Christopoulou et al. [39] provide a representative early-fusion baseline on a native 5G testbed, concatenating UE-, eNB-, and MME-side features into a single input vector consumed by KNN, XGBoost, and Decision-Tree classifiers. Late fusion runs a per-modality detector and combines their scores via a learned aggregator; this scales better when modalities have different sampling rates or live on different hosts, at the cost of losing fine cross-modality interactions. Hybrid fusion runs per-modality encoders to a shared latent space and then a joint head on the latent representation; this captures cross-modality interactions while controlling the per-host compute footprint and is the dominant pattern in 2024–25 cellular IDS work. The remaining gap is standardised cross-modality schema (Challenge C4) so that the fusion architecture does not have to be re-implemented per operator. H. Synthesis: Four Open Trade-Offs The surveyed evidence converges on four trade-offs that structure the remaining design space. Accuracy vs. latency: deeper models gain 1–3 pp accuracy but risk violating URLLC budgets; int8 quantisation, structured pruning, and knowledge distillation typically recover 70– 90% of the latency overhead at 0.5–1.5 pp accuracy loss. Local vs. global visibility: MEC-local inference is fast but blind to distributed campaigns; hierarchical aggregation introduces latency. Supervised vs. unsupervised: supervised peaks in accuracy but starves for labels; semi-/self-supervised approaches occupy the operationally practical middle. Accuracy vs. explainability and personalisation: black-box deep models (deep ensembles, transformer hybrids) score
highest on curated benchmarks but are the hardest to explain, while the most interpretable methods (PCA residuals, decision trees) lag in accuracy by 5–15 percentage points; personalised per-slice or per-operator models that incorporate operatorspecific priors lose a small amount of headline accuracy but gain robustness to distribution shift and produce explanations that downstream auditors and regulators can act on. The tradeoff is not a free lunch: explainable models can leak decisionboundary information to adversaries through the explanation surface itself [32], and personalised models complicate FL aggregation because per-operator divergence accumulates across rounds. The operational compromise motivated by Umer et al. [114] is layered along three axes simultaneously. Along the accuracy axis, an interpretable triage layer at the MEC (decision trees, PCA residuals, attention-weight read-outs) scores every alert at low latency, with a black-box deep model invoked only for a second-look pass on flagged segments—recovering deepfamily accuracy on the small fraction of traffic that warrants it. Along the explainability axis, attention-based attribution is emitted at both layers [114], so every decision carries a per-feature rationale, with heavier model-agnostic explainers (SHAP, counterfactuals) reserved for post-incident review at the operator core. Along the personalisation axis, a federated backbone is shared across operators while each fine-tunes a small operator-specific head on its own non-IID CDR/RAN distribution, capturing local attack patterns without leaking subscriber data or destabilising FL aggregation [123]. I. Design Lessons for the Detect Stage of the Closed Loop Four design lessons recur across the literature surveyed in this section and should constrain any future detect-stage design. • Lesson 1 — telemetry first, model second. Feature engineering on CDR/NWDAF/E2 telemetry has produced larger accuracy deltas than architectural changes within a method family; investment in telemetry pipelines and feature selection out-performs investment in deeper models [19], [36], [68].
16
TABLE IX: Unified comparison of edge-based detection methods for 5G/6G CPS, consolidating the cellular-anomaly and DDoS-classification literature into a single deployment-aware view. Headline figures are reproduced from each cited work; we do not re-execute them. These figures should be read within, not across, the evidence tiers defined in the footnote. Ref. [63]
Year Method Dataset Acc. (%) F1 FPR (%) Latency Ev. Edge CDR Domain 2017 Semi-sup. statistical Milan CDR 92.8 0.94 14.1 Near-RT E3 ✓ (Gaussian) [64] 2018 Semi-sup. statistical Milan CDR 92.0 0.94 14.1 Near-RT E3 ✓ (Gaussian) [65] 2018 L-layer DNN (MLP) Milan CDR 94.6 – 1.7 Near-RT E3 ✓ N [66] 2018 k-means + Neural Net- CRAWDAD, Nodobo, Milan CDR – – – Batch E3 ✓ work [108] 2018 Autoencoder ensemble Real traffic (9 datasets) – – – RT E3 ✓ (Kitsune) [79] 2019 SVM/RNN/LSTM-RNN Bot-IoT 99.99 1.00 – – E3 N, P baselines (Bot-IoT release) [103] 2019 CNN–LSTM hybrid CICIDS2017 97.1 – – – E3 [77] 2019 Dataset release + ID3/RF CICDDoS2019 – 0.78 0.69 – E3 baselines [35] 2019 MEC-deployed DNN Real CDR (operator) 98.8 – 0.44 RT E3 ✓ ✓ N [40] 2020 Lightweight CNN (LU- CICDDoS2019 99.0 0.99 – RT E3 ✓ CID) [36] 2020 MEC CNN (spatial grid) Real CDR ≤96 – – RT E1 ✓ ✓ N [80] 2020 SDN+NFV+AI multi-layer NS-3 simulation 96.1 – – RT E4 N [104] 2020 RNN–Autoencoder + soft- CICDDoS2019 99.0 0.99 – – E3 max (DDoSNet) [19] 2021 CNN (ResNet-50 / DRC) Milano CDR + synth. attacks (4 types) 91–97 – – – E3 ✓ P [119] 2021 Composite DNN CICDDoS2019 99.7 – – – E3 N (MLP+PCC) [122] 2021 Feedforward DNN (MLP) CICDDoS2019 94.6 – – RT E3 [90] 2022 RF (also GBM, MLP) TON IoT (vs Aposemat IoT-23) 98.1 0.97 – – E3 [111] 2022 E-GraphSAGE (edge- NF-ToN-IoT / NF-BoT-IoT – 1.00 – – E3 feature GNN) [113] 2023 Multi-Domain Transformer SCVIC-TS-2022 (CICIDS2017 PCAPs) 99.7 0.84 – Early-RT E3 N (MD-MHA, TS-NFM) [67] 2024 k-means + deseasonalisa- 14 M one-year operator CDR (5G-deployed) 96.7 0.60 – Batch E1 ✓ N tion [68] 2024 DBSCAN / IF (deseason- 37 M one-year operator CDR 98.1 0.76 – Batch E1 ✓ N alised; GMM, MS for profiling) [105] 2024 Multi-Scale ConvLSTM + Public Kaggle KPI (train) → KPI (fine-tune) 99.0 0.99 – – E3 N transfer learning [112] 2024 Transformer autoencoder 5GAD-2022 (free5GC) – 0.88 – Offline E4 N [114] 2025 Wrapper-FS + multi-head UNSW-NB15 93.0 0.92 – Offline E3 attn. Transformer [42] 2025 LSTM–RNN autoencoder O-RAN over-the-air testbed 97.5 0.99 2.2 RT E2 ✓ N, P (xApp) [43] 2025 RF/XGB on UNSW-NB15 UNSW-NB15 86–87 ∼0.86 – E3 N, P (cross-layer arch.) [118] 2025 SH-CASH AutoML CICIDS2017 + Oracle RF fingerprint 99.4 0.98 – RT E3 N (ARF/SRP + ADWIN/EDDM) [84] 2026 Multi-layer FL + CNN– Mininet-WiFi 97.6 0.97 6.5 RT E4 ✓ N LSTM (trust scoring) Evidence strength (Ev.): E1 = real operator data, privately accessed / not publicly released; E2 = real 5G/O-RAN testbed or over-the-air; E3 = public real-traffic benchmark; E4 = simulation or synthetic labels only. Domain: N = 5G/6G-relevant; P = explicit CPS setting. “–” denotes a value not reported in the cited work.
Lesson 2 — compress as a first-class step, not an afterthought. Quantisation- and pruning-aware training produces detectors that match unconstrained baselines on commodity MEC hardware; deploying an unconstrained model and then trying to fit it to MEC after the fact rarely succeeds [17], [35], [36], [40], [59]. • Lesson 3 — multi-class, not binary. Multi-class classifiers [19], [77] produce per-attack-type outputs that the orchestrator can use to select mitigation playbooks; a binary detector forces the orchestrator to default to a generic primitive, which over-blocks legitimate traffic. • Lesson 4 — federate where possible, validate everywhere. Per-MEC training on local network telemetry plus FL aggregation [37], [46], [115], [124], [125] is the only path that scales, and DT-based what-if validation is the only way to keep the loop safe under continual model updates [52]. •
These four lessons converge on a single operational implication: detection performance in isolation is not the binding constraint of the closed loop. Even a perfect detector is operationally useless without the automated, low-latency hand-off into the mitigation machinery that the next section formalises (Sec. VI, Tables X and XI). VI. N ETWORK -W IDE M ITIGATION AND THE C LOSED L OOP This section surveys how the alert produced by the detector (the Alert box of Fig. 7) is converted into network-wide action. We treat SDN, NFV, and O-RAN xApps as composable actuators of one closed loop rather than competing schools, define the metrics that matter for CPS (Sec. VI-D), and present the consolidated comparison of mitigation frameworks (Table XI).
17
A. From Detection to Response: The Mitigation Gap The cellular security literature has historically paid less attention to mitigation than to detection, even though the operational value of an IDS depends on what happens after the alert. The mitigation gap is conventionally framed as having three faces: (i) the policy gap (what action to take), (ii) the actuation gap (where to enact it), and (iii) the latency gap (how long it takes). 6G CPS amplifies all three because URLLC slices need millisecond reactions and human-in-the-loop is no longer feasible. Beyond these three, 6G CPS surfaces a fourth, less-oftennamed face: the reversibility gap. Once the orchestrator has throttled a slice or blocked a UE class, undoing that action has its own latency, its own collateral impact, and its own audit burden. Mitigation design in 6G CPS has to treat each action as paired with its inverse—re-admit, un-throttle, restore-handover-weight—and the inverse must be reachable within a budget comparable to the original action or the loop cannot tolerate false positives without breaking safety-oflife workloads. This reversibility requirement is what makes graded enforcement (Sec. VII-D) and the four-tuple ttl field (Sec. VI-B) structurally necessary rather than cosmetic. Mitigation systems that are evaluated only by their blockeffectiveness [80] systematically overstate operational viability, because the reversibility cost is invisible to blockeffectiveness alone. The policy gap is the hardest of the three. A detector emits a per-flow or per-cell anomaly score; an orchestrator must convert that score into a discrete action drawn from a finite playbook (rate-limit, scrub, isolate, drop, re-authenticate, reroute, slice-rebalance) and parameterise that action against the slice’s safety envelope. The mapping is not unique: the same anomaly could justify a soft response (rate-limit at 50%) or a hard response (slice isolation), and the cost of an over-reaction is paid by the very CPS workload the action is supposed to protect [43], [126]. The actuation gap is the heterogeneity of the actuator surface. SDN flow-rule installation, NFV scrubber instantiation, O-RAN xApp scheduler-weight modification, and corenetwork admission control each have different APIs, latency envelopes, and failure modes; a single playbook must compose across all of them. The latency gap is what 6G CPS makes binding. A URLLC slice with a 1 ms end-to-end budget cannot tolerate the subminute VNF-instantiation latencies that NFV-only mitigation incurs, even if the eventual mitigation is correct [80], [127]. The closed loop addresses all three gaps by maintaining a tier-stratified actuator hierarchy: the fastest tier (O-RAN xApp scheduler weights, 10–100 ms) handles immediate suppression, while slower tiers (SDN re-routing, ∼1 s; NFV scrubber instantiation, tens of seconds) sustain the response. B. Detection-to-Mitigation Hand-Off Protocol The operational link between Sec. V and Sec. VI-C is a hand-off protocol whose properties determine whether the closed loop runs end-to-end [86]. We model the hand-off as a four-tuple ⟨score, class, evidence, ttl⟩ emitted by the detector
and consumed by the orchestrator. The score is the calibrated probability that the observed traffic is malicious. The class is the MITRE technique label (Table VI) that selects the playbook. The evidence carries the per-feature attribution (Sec. VII-D) needed for the orchestrator’s downstream rationale and for forensic retention. The ttl bounds the alert’s validity window, after which the orchestrator must request a refresh; this is what prevents stale alerts from triggering mitigations after the underlying traffic has already cleared. The hand-off protocol is the natural locus of standardisation in Challenge C4 (Sec. VIII-D); the SDN, NFV, and ORAN primitives below are organised by the tuple field each consumes. C. SDN-, NFV-, and O-RAN-Based Mitigation Primitives SDN-based mitigation centralises forwarding decisions in a controller that can install or revoke flow rules in seconds: blacklisting source IPs, redirecting flows to scrubbing services, or isolating compromised slices [43], [121], [128]. The strength is global view; the weakness is controller scalability and the new attack surface of the controller itself. Cooperative SDN architectures [121] federate decisions across controllers to extend reach, and flow-rule programming via OpenFlow/P4 [117] pushes mitigation as far down the stack as the data-plane switch will permit. The mitigation playbook— rate-limit, redirect-to-scrubber, isolate-slice, drop—is selected from the hand-off class field (Sec. IV-E) and parametrised with the false-block-rate and slice-availability constraints of Table X. NFV-based mitigation virtualises mitigation appliances (firewalls, scrubbers, deep packet inspection) and elastically scales them in response to attack volume. Bohatei [127] demonstrated elastic NFV-based DDoS scrubbing with subminute provisioning. Semantic SDN/NFV policy frameworks [129] compose primitives at higher abstraction. The operational appeal is that mitigation capacity scales with attack capacity rather than being statically provisioned, but the latency to instantiate a new VNF (typically tens of seconds even with optimisations) is too coarse for URLLC-slice protection on its own; NFV is therefore best paired with a faster SDN/ORAN reflex layer. O-RAN-based mitigation is the natively-cellular path: xApps on the Near-RT RIC can suppress malicious UEs, modify scheduler weights, or trigger handovers within the 10–100 ms RIC budget [9]–[11], [34], [42], [62], [118]. xApps for layer-3 attack detection (5G-SPECTOR [27], RRC-storm [34]) are now operational on testbeds, and the anomaly detection/traffic-steering xApp pattern is widely documented [31]. Because the RIC controls scheduling and admission directly, an xApp can degrade service to a suspicious UE without dropping its connection, preserving the operator’s ability to forensically observe the attack while neutralising its impact—an option not available to upstream SDN-only defences. The architectural mapping of these O-RAN and SDN mitigation hooks onto the disaggregated RAN control plane is illustrated in Fig. 8, which highlights how actuation points natively trade latency against blast radius. Surveys of ORAN architecture and security [9], [25], [26] caution that the
18
SMO (Service Mgmt & Orchestration) O1
O2 O1
Non-RT RIC
A1
Near-RT RIC
(rApps)
(xApps)
rApp re-policy ∼1 s network-wide
xApp throttle 10–100 ms cell-group
E2
O-CU
F1
Slice re-alloc. ∼100 ms per-slice
slower · wider blast radius
O-DU
OFH
O-RU
PRB blocking ∼1 ms single cell
faster · narrower data-path control
O1/O2 mgmt plane
mitigation hook
Fig. 8: O-RAN/SDN mitigation hooks mapped onto the disaggregated RAN control plane. Four actuation points trade latency against blast radius: rApp re-policy via A1 (∼1 s, network-wide), xApp throttle via E2 (10–100 ms, cell-group), slice reallocation via F1 (∼100 ms, per-slice), and PRB blocking at O-DU MAC (∼1 ms, single cell). The dashed O1/O2 management plane carries telemetry and lifecycle events but is not in the data-path control loop. Faster hooks have narrower blast radius, motivating the closed-loop design in Sec. VI-F. same openness that enables this also expands the supply-chain surface; xApp lifecycle attestation (Sec. VIII-D) is therefore an open standardisation problem. Hybrid composition is the dominant pattern in 2024– 25: SDN steers, NFV scales, O-RAN xApps act at the airinterface, and an AI orchestrator picks the playbook. Sheibani et al. [126] formalise the multi-layer defence. Allaw et al. [43] sketch a cross-layer SDN/NFV/AI mitigation framework for 5G/6G slices, with the AI-detection stage validated on UNSWNB15. The orchestrator’s job is to map a detected MITRE technique to a primitive, choose the actuator tier whose latency fits the slice’s budget, and roll back if the false-block rate exceeds a slice-specific threshold. Recent work treats the orchestrator itself as a learned policy [83], with reinforcement learning over a discretised action space of (primitive, targettier, intensity) triples; this lifts the orchestrator out of static playbooks but reintroduces explainability and safety questions that XAI (Sec. VII-D) and DT validation (Sec. VII-C) must answer before production rollout. AI-native orchestration is the integrating layer. Each MITRE technique class produced by the detector (Sec. IV-E) maps to one or more primitives. The orchestrator chooses among primitives by minimising a slice-aware cost J(a) = w1 ∆tmit (a) + w2 FBR(a) + w3 ∆trec (a) + w4 cost(a), (1) where the weights w1 , . . . , w4 are slice-specific (URLLC slices weight ∆tmit and FBR much more heavily than mMTC slices). This formulation cleanly composes the metrics of Table X; the multi-layer SDN/NFV defence of [126] approximates it operationally and the cross-layer architecture of [43] sketches the same composition at the design level, while the binary-isolation pattern of [42] is a special case where w1 dominates and the action set collapses to a single primitive. D. Mitigation Performance Metrics The mitigation literature has historically optimised three quantities—detection accuracy, block-effectiveness, and postattack recovery time—and reported each in isolation from
the others. For 6G CPS, this decomposition is insufficient because it omits the downstream physical consequence that is the ultimate object of concern. A mitigation that blocks a volumetric flood with 99% effectiveness but delays a vehicular collision-warning message by 4 ms can still cause harm; a mitigation that reports a 200 ms recovery time but collaterally blocks a legitimate telesurgery session is operationally worse than the attack it contains. We therefore adopt a CPS-aware metric set, summarised in Table X, in which every metric is paired with the slice-class envelope against which it must be evaluated. Five first-order metrics recur across the surveyed closedloop systems and form the basis of our comparative analysis in Table XI, with full definitions and CPS relevance consolidated in Table X. Mitigation latency ∆tmit is tail-bounded for URLLC, not mean-bounded, since a mitigation that fires in 10 ms on average but occasionally spikes to 100 ms is unusable for V2X collision avoidance [80]. The CPS-weighted false block rate (FBR), unlike the detector-level false-positive rate (FPR), is weighted by the traffic’s CPS relevance, so a false block on a meter reading and a false block on a cardiac telemetry stream enter the metric with different costs [126]. Slice availability under attack is negotiated into the CPS slice tenant’s SLA and ultimately optimised by the orchestrator in Eq. 1. Recovery time ∆trec includes the time to release ratelimits, re-admit blocked sources, and restore scheduler weights (the reversibility cost of Sec. VI-A). Finally, resource overhead bounds the operator’s ability to scale the defence to networkwide attacks. Beyond these five first-order metrics, two compound metrics are emerging as the credible operational targets of closedloop systems. The CPS-impact harm score fuses mitigation latency, FBR, and slice availability through a slice-specific cost function that encodes the physical consequence of each outcome (a missed V2X warning, an unscheduled relay trip, a delayed defibrillation command); no public dataset labels this quantity today, which is why Challenge C5 (Sec. VIII-E) names its construction as a priority. The closed-loop drift
19
TABLE X: Mitigation Performance Metrics and CPS Relevance Metric Mitigation latency ∆tmit False Block Rate (FBR) Slice availability
Definition Time from alert to attack suppression Fraction of legitimate traffic blocked Up-time of target slice during attack Recovery time Time from suppression to ∆trec nominal QoS Resource overhead Compute/bandwidth used by mitigation
CPS Relevance Critical for URLLC (V2X, surgery) Safety hazard if CPS control traffic blocked Direct service-level metric Bounds CPS downtime Limits scaling to network-wide attacks
indicator captures the rate at which the detector’s score distribution on production telemetry diverges from the distribution on the FL-aggregated training set; when the indicator crosses a slice-aware threshold, the orchestrator must either escalate to a larger reference model [31], trigger a targeted DT replay (Sec. VII-C) [110], or degrade gracefully to a conservative fallback policy. Both compound metrics are operational refinements of the five first-order metrics rather than replacements, and they are the natural reporting surface for the benchmark agenda of Sec. VIII-E. An important methodological caveat is that the metrics in Table X are not independent: an orchestrator that minimises ∆tmit by firing the fastest available actuator will typically raise the FBR (because the fastest actuator has the coarsest granularity), and an orchestrator that minimises FBR by waiting for high-confidence evidence will typically raise ∆tmit . The Pareto frontier of these trade-offs is slice-specific, and the operational pattern is to expose the slice tenant’s preferred point on the frontier as a per-slice configuration rather than as a global operator choice. This preference encoding is what lets the same closed loop serve a URLLC vehicular slice (FBRtolerant, latency-intolerant) and an mMTC metering slice (latency-tolerant, FBR-intolerant) from the same detector and the same actuator pool, and it is the metric-level realisation of the graded-enforcement principle we return to in Sec. VII-D. Operational takeaways. Five lessons recur across the deployed frameworks: • Latency is bounded by the slowest tier: commit the fast tier (O-RAN xApp) first; use slower tiers (SDN, NFV) only for sustained mitigation. • False-positive cost is slice-asymmetric: treat slice class as a first-class input to the orchestrator cost (Eq. 1). • Rollback is as important as forward action: a loop without a budgeted rollback path is operationally incomplete. • Cross-tier feedback dominates per-tier tuning: postmitigation telemetry is the only direct evidence that the chosen action was correct. • Operator-specific tuning is unavoidable: deployed systems report a 2–6 month tuning period (implications for C3 and C5). E. Comparative Analysis Table XI compares the most cited mitigation frameworks across the SDN/NFV/O-RAN axes and along the AI/automation axis. The trajectory is clear: early work is either survey-level or rule-driven [121], [127], [128], whereas from 2020 onward SDN/NFV and O-RAN frameworks increasingly
close the detect-then-act loop with AI [9], [42], [43], [80], [126], [129]. This shift toward AI-orchestrated hybrid stacks is the empirical basis for the closed-loop thesis of this survey, yet the empty CPS column shows that none of these frameworks treats cyber-physical impact as a first-class goal—the gap this survey targets. F. The Unified Closed-Loop Reference Architecture Fig. 9 instantiates the four-stage thesis of Sec. I-B as a concrete reference architecture. Sensing is on CDR/O-RAN telemetry at the aNB and MEC tier; detection runs the model families of Sec. V-C; mitigation composes SDN/NFV/O-RAN actuators per detected MITRE technique class (Sec. IV-E); learning is FL- and DT-driven retraining (Sec. VII). The closure of the loop—explicit, in-protocol, and continuous— is what the prior literature has been missing. The loop has three operational properties. End-to-end latency budgeting: each stage is annotated with its target latency, and the sum bounds the slice’s safety envelope; a violation in any stage propagates to the next, so optimisation must be loopwide rather than stage-local [10]. Telemetry-as-feedback: the orchestrator’s actions become part of the next sense step, so the detector continuously sees how the system responds to its alerts; this is what makes drift detection and on-line model selection feasible at the MEC tier. Federated learning as the lifecycle bus: per-MEC detectors are continuously refined by FL aggregation of local gradients, and DT replay supplies adversarial scenarios that the operational telemetry does not contain [28], [37], [46]. The operational evidence for the loop is now sufficient to claim feasibility. Moore et al. [42] report a complete sensedetect-mitigate cycle on an over-the-air O-RAN testbed with sub-second mitigation latency; Allaw et al. [43] sketch a crosslayer SDN+NFV+AI architecture and validate the AI-detection stage on UNSW-NB15; Christopoulou et al. [39] show that the NWDAF can host the detection step inside a standardised cellular interface. Differentiator: a latency-contracted, falsifiable closed loop. The four-stage loop above is, in shape, isomorphic to MAPE-K [130], OODA, and the O-RAN RIC closed-loop hierarchy [9]. The survey’s contribution is therefore not the four-stage decomposition itself but a falsifiable contract that ties the loop to the URLLC slice’s safety envelope. a) Latency contract.: Let ∆ts , ∆td , ∆tm denote the per-stage latencies of Sense, Detect and Mitigate measured at a slice-dependent tail percentile p, and let τmax be the slice’s end-to-end safety envelope (1–5 ms for URLLC, 10– 100 ms for eMBB, seconds for mMTC). We set p = p99 for safety-critical slices (URLLC vehicular, telesurgery, and grid protection), where a one-in-twenty budget violation is not tolerable, and relax to p = p95 for latency-tolerant eMBB/mMTC classes; p99.9 is the appropriate target where a slice’s reliability requirement (10−9 packet error for URLLC) dominates, at the cost of requiring substantially more telemetry to estimate the tail reliably. The closed loop is slice-admissible iff ∆ts + ∆td + ∆tm ≤ τmax − ∆tho at p, (2)
20
TABLE XI: Comparative Analysis of Network-Wide DDoS Mitigation Frameworks. Reference [127] [121] [128] [80] [129] [9] [126] [43] [42]
Year 2015 2016 2016 2020 2020 2023 2024 2025 2025
Strategy Elastic NFV defense (Bohatei) SDN-based DDoS defense (survey) SDN security (survey) SDN+NFV multi-layer Semantic SDN/NFV policy O-RAN architecture & security Multi-layer defence Cross-layer arch. (SDN/NFV/AI) xApp closed loop (over-the-air)
SDN
• • • • • # G • • ◦
NFV
• ◦ ◦ • • # G • # G ◦
O-RAN
◦ ◦ ◦ ◦ ◦ • ◦ ◦ •
Latency Low – – Low Medium Low Medium – Low
CPS
◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦ ◦
5G/6G
◦ ◦ ◦ • ◦ • • • •
Auto # G
◦ ◦ # G • • • • •
AI
◦ ◦ ◦ • • • • • •
Note: • = supported; G #= partial / rule-based rather than fully AI-driven; ◦ = not supported or not reported. Latency is the qualitative end-to-end mitigation budget reported by each framework (Low ≤ 100 ms, Medium ≤ 1 s).
where ∆tho is the four-tuple hand-off serialisation cost (Sec. VI-B). The tail-percentile choice distinguishes the contract from a mean-latency claim, since tail excursions—not the average—are the operational failure mode. Two caveats make the contract honest. First, summing per-stage tail P latencies does not in general equal the end-to-end tail: i ∆ti @ p equals the chain’s true p-percentile only when the stage latencies are perfectly positively correlated, and otherwise upper-bounds it. The contract is therefore analytic rather than empirical: Eq. (2) is a conservative, sufficient admission test derived from per-stage tail bounds, not a measurement of the end-to-end distribution. A satisfied bound certifies admissibility; a violated bound mandates direct end-to-end percentile measurement before the slice is rejected. Second, the percentile is admission-time conservative by design, because a false admission of a safety-critical slice is costlier than a false rejection. The Learn stage is off the critical path. Making the contract per-slice and tail-bounded is what MAPE-K and OODA leave undefined and what the O-RAN RIC closed-loop specification leaves operator-discretionary. b) Worked V2X example (τmax = 3 ms collision-warning slice).: Because this is a safety-critical URLLC slice, the contract is evaluated at p99 . Sense: the MEC host pulls KPMs over E2 with ∆ts ≈ 0.4 ms. Detect: a quantised LUCIDclass detector [40] scores a 32-flow batch in ∆td ≈ 0.6 ms. Mitigate: the orchestrator emits a scheduler-weight modification to the gNB MAC scheduler via E2 RIC-control, taking ∆tm ≈ 1.2 ms including Remote Procedure Call (RPC). With ∆tho ≈ 0.1 ms, the summed p99 bound consumes 2.3 ms of the 3 ms budget; because this bound is conservative, the true endto-end p99 is also within budget and the loop is admissible. If the same slice instead required SDN flow-rule installation (∆tm ≈ 50 ms [80]), Eq. 2 fails by an order of magnitude and the loop must fall back to the O-RAN xApp tier or admissionreject the load. This is a falsifiable architectural property, not an empirical observation. G. Reference Implementation Sketch To make the closed-loop architecture concrete, we sketch the components that would compose a minimal but realistic instantiation on commodity hardware. The MEC tier hosts a lightweight detector (a quantised CNN for spatial CDR features and a compressed LSTM for temporal sequences) colocated with an orchestrator agent that consumes the four-tuple hand-off (Sec. VI-B) and selects from a finite playbook. The Near-RT RIC hosts an xApp registered with the detector’s MITRE technique vocabulary (Table VI), exposing a small
RPC surface for scheduler-weight modification, UE suppression, and handover triggers [9], [10], [42]. The transport tier exposes an SDN controller that accepts authenticated flow-rule installations from the orchestrator, with a P4-programmable data-plane fast-path for high-rate drop and rate-limit primitives [117]. The core network exposes the NWDAF event surface for cross-MEC visibility [39], and an FL aggregator runs in a privacy-isolated enclave to consume gradient updates from the per-MEC detectors [37], [46]. The sketch is deliberately conservative: every component above is empirically demonstrated in the surveyed literature. Doriguzzi-Corin et al. [40] and Franco-Valiente et al. [41] show that lightweight CNN and energy-aware ARM-SoC pipelines achieve sub-10 ms inference on commodity edgeclass hardware, and Hussain et al. [35], [36] establish the deep CDR-detector accuracy on real operator data that such edge runtimes must reproduce; Moore et al. [42] demonstrate the xApp orchestrator path on an over-the-air O-RAN testbed; Christopoulou et al. [39] exhibit NWDAF-hosted detection inside a standardised cellular interface; Allaw et al. [43] report a cross-layer SDN+NFV orchestrator. The novelty here is the integration: previous work has demonstrated each piece in isolation; the closed loop’s value is the composition. The five outstanding integration questions—xApp attestation (C4), cross-operator FL governance (C3), audit-grade XAI under adversarial probing (C2), CPS-impact harm metrics (C5), and certified sub-millisecond detection under model drift (C1)— are exactly the consolidated open challenges of Sec. VIII. VII. C ROSS -C UTTING E NABLERS The closed loop of Fig. 9 is sustained by three operational services—FL, LLMs, and DTs (Secs. VII-A–VII-C)—and a three-layer substrate of PQC, ZTA, and XAI (Secs. VII-D– VII-E). The services produce training data, summaries, and counterfactuals for the loop; the substrate makes the loop’s transactions confidential, trusted, and auditable. We treat each below in its operational role inside the loop. A. Federated Learning: Training Across Operators Without Sharing CDRs FL is the operational answer to the labelled-CDR scarcity problem identified in Sec. V-C: each MEC node trains a local model on its own CDRs (labelled by the selfsupervised/pseudo-labelling means of Sec. V-C, not manual annotation) and shares only model updates with a central aggregator (Fig. 10) [4], [15], [46], [131]. The architectural
21
1. Sense CDR + O-RAN (Sec. V-A, V-B)
2. Detect MEC ML/DL (Sec. V-C..G)
hand-off
∆ts ≈ 0.4 ms (p99 )
hand-off
∆td ≈ 0.6 ms (p99 )
3. Mitigate SDN/NFV/O-RAN (Sec. VI-B, VI-D)
hand-off
∆tm ≈ 1.2 ms (p99 )
4. Learn FL + DT (Sec. VII) off critical path
feedback: continuous retraining & policy refinement (Learn, off critical path)
Fig. 9: Unified closed-loop reference architecture for AI-native 6G CPS security. The architecture replaces the five-pillar list of prior surveys with one continuously-running loop. Per-stage annotations are p99 stage latencies ∆ts , ∆td , ∆tm for the safety-critical V2X worked example of Sec. VI-F (the percentile mandated for URLLC); their slice-admissibility is governed by Eq. 2, not by their unweighted sum. The Learn stage is off the critical path. Aggregator (telco core) ∆θ only MEC1
MEC2
MEC3
···
MECn
Fig. 10: Federated learning topology for distributed 6G CPS security. The n MEC clients (MEC1 , . . . , MECn ) each train locally; CDRs never leave the MEC host, and only gradient updates ∆θ are exchanged with the aggregator at the telco core. fit with MEC is exact—one client per host, one round per training window, secure aggregation across hosts. Algorithm variants. FedAvg [46] is the baseline; FedProx [123] stabilises training under client heterogeneity; FedAGRU (attention-GRU FL) [124] reaches ∼97% on KDDCup99/CICIDS2017/WSN-DS. Federated GRU per device type [125] achieves >95% on a 33-device IoT testbed. Byzantine resilience. Single-server secure aggregation with multi-Krum-style outlier removal [85], [132] bounds the influence of malicious clients while preserving per-client privacy; empirical evaluations on non-IID device-partitioned IoTmalware data [107] show that standard FedAvg averaging is broken by a single Byzantine client (gradient-factor and model-cancelling attacks reduce the model to a constant predictor) while coordinate-wise median aggregation tolerates one or two malicious clients but still collapses at three out of eight; this matters operationally because compromised MEC hosts are an in-scope threat [85] (Sec. IV-C). Privacy and communication efficiency. Secure aggregation [133] and differentially private FedAvg [134] reduce information leakage; gradient compression and asynchronous aggregation reduce communication overhead [135]. Federated autoencoders and MLP classifiers detect IoT-botnet (Mirai/BASHLITE) malware on N-BaIoT at ∼99.9% supervised accuracy and ∼99.98% TPR (92–97% TNR) unsupervised, including on devices unseen during training, but standard FedAvg averaging is broken by a single Byzantine client [107]; non-IID heterogeneity is the central practical headache [85], [136]. For 6G CPS specifically. BERT-based federated IDS for 5G [115], FL-enhanced cyber-security [37], and the full FLfor-6G survey [28], together with the focused FL security and privacy taxonomy of [76], confirm both feasibility and the need for new aggregation protocols when participants are competing
operators across regulatory domains. Table XII consolidates the empirical evidence. Convergence under non-IID CDR distributions. The convergence guarantees of FedAvg [46] are weakened when client data distributions are heterogeneous [131], which is the default regime for cellular CDRs: traffic on an urban MEC host differs structurally from a rural host, a tourist-heavy coastal cell, or an industrial campus slice. FedProx [123] adds a proximal regulariser µ2 ∥θk − θ(t) ∥2 to each client’s local objective to bound local drift, and its convergence analysis under bounded dissimilarity is the formal result that justifies its empirical robustness on non-IID FEMNIST and derivative cellular workloads. The practical knobs for 6G CPS deployments are (i) the proximal strength µ, tuned per-round from measured client-drift variance; (ii) the localepoch count E, which trades communication cost against local-drift accumulation; and (iii) the client-sampling strategy, where stratified sampling across operator-role archetypes (urban/rural/industrial/transit) materially reduces round-to-round variance compared to uniform sampling. Empirically, nonIID CDR federations converge in 3–10× the rounds of IID baselines [123], [124], [136], and the operational cost is the control-plane bandwidth consumed by additional aggregation rounds rather than per-round compute. The research direction in Sec. VIII-C is to combine FedProx-style drift control with Byzantine-robust aggregation [132] and PQC-protected secure aggregation [133], [137] inside a single protocol whose convergence, privacy budget, and Byzantine tolerance are all analysed jointly rather than bolted on one axis at a time. B. Large Language Models Should Explain and Query the Loop, Not Control It Two roles for LLMs are emerging in cellular security [86]. Telecom-specific foundation models, pre-trained on protocol specifications, RFCs, and operator runbooks, can ingest raw logs and emit structured anomaly hypotheses, replacing bespoke feature engineering [138]. Natural-language SOC interfaces let operators query the loop in English (“which slices saw RRC anomalies in the past hour?”) and receive answers grounded in the supporting telemetry. Houssel et al. [116] show that LLMs can produce human-interpretable NIDS reports that surface the features behind a classifier’s decision, while also documenting their twin weaknesses— unreliable detection and fabricated facts—that fix the role boundary below.
22
TABLE XII: Federated Learning Methods for 6G CPS Security Ref. [46] [136] [125] [133] [134] [124] [123] [132]
Year 2017 2018 2019 2019 2020 2020 2020 2021
FL Algorithm Security Task Dataset Acc. (%) Privacy Comm. Eff. Non-IID FedAvg General classification MNIST/CIFAR 99.4/85.0 # FedAvg + data sharing Non-IID mitigation CIFAR-10 ∼78 # # Federated GRU per device IoT device-type anomalies Real IoT (33 dev.) >95 # # SecAgg + FedAvg Mobile keyboard prediction Production – # NbAFL (DP-FedAvg) Privacy-preserving FL (general) MNIST – # # FedAGRU (Attn-GRU) Wireless-edge IDS KDDCup99/CICIDS/WSN-DS ∼97 # FedProx Heterogeneous FL FEMNIST – # # BREA (single-server multi- Byzantine-resilient secure FL MNIST / CIFAR-10 – # # Krum + VSS + RS-aggregation) [107] 2022 Federated AE + MLP (FedAvg) IoT-botnet detection N-BaIoT 99.9 # [37] 2024 FL-enhanced security survey 5G/6G network security Multiple – [28] 2026 Survey FL for 6G security – – – Capability indicators: = the method explicitly provides or addresses the property; # = it does not, or is not designed to. Privacy: a formal privacy mechanism beyond baseline FL (e.g., differential privacy, secure aggregation, or encryption). Comm. Eff.: the method is designed to reduce communication overhead from exchanging model updates. Non-IID: the method explicitly handles heterogeneous (non-independent-and-identically-distributed) client data. “–” denotes a value not reported in the cited work.
a) Latency.: The first reason to confine LLMs to explainand-query rather than decide-and-act is timing. Reported inference runs from roughly one second per query on hosted lightweight models to tens of seconds on self-hosted 30 B+ models [138], and Houssel et al. measure a single 8 B-model NIDS inference at about three orders of magnitude slower than a lightweight detector—too slow for real-time use [116]. Against the URLLC envelope of Eq. 2 (1–5 ms), this rules the LLM off the sense–detect–mitigate path entirely. It also dictates placement: the MEC host sized in Sec. II-C for a quantised CNN–LSTM detector cannot hold a foundation model, so the LLM lives in the core or a private cloud and is queried by MEC agents, consistent with the clustered-GPU serving that latency-sensitive operation requires [138]. The result is a clean separation of timescales: the millisecond loop runs on the compressed detector and the xApp actuator, while the LLM works the human-supervised SOC timescale of seconds to minutes, where its latency is immaterial—and where streaming first-token output, not full completion, is the figure of merit. b) Reliability.: The second reason is reliability, which differs in kind from detector reliability. A detector fails as a miscalibrated score that calibration and the FBR metric of Table X already bound; an LLM fails as a fluent fabrication carrying no confidence signal, most likely on the rare outof-distribution incident the operator can least afford to miss. Houssel et al. observe exactly this—confident invention of geographic and protocol facts and confusion of protocol numbers with port numbers [116]—which we generalise as protocol, telemetry, and stale-context hallucination. None is removed by a larger model; each is bounded only by architecture: retrieval grounding against the operator’s own telemetry and runbook corpus, telecom-tuned pre-training, and a hard refusal that returns INSUFFICIENT_EVIDENCE rather than guessing [138]. The decisive choice is the action boundary—the LLM may summarise, rank, and propose but never command an actuator—so an undetected hallucination degrades to a bad suggestion that the hand-off protocol of Sec. VI-B and a human or policy-decision-point still gate, not a bad mitigation on a live slice. Quantifying the residual hallucination rate under distribution shift, and certifying that the refusal envelope holds, is what Challenge C2 (Sec. VIII-B) names as bounded LLM risk. We distil from this literature a four-field prompt pattern that
binds, rather than eliminates, hallucination: telemetry scope (slice, cell set, window); evidence requirement (at least one telemetry record cited per factual claim); refusal envelope (return INSUFFICIENT_EVIDENCE with the missing-field list when support is absent); and action boundary (proposals enter the hand-off protocol of Sec. VI-B as suggestions gated by a human or policy-decision-point, never as direct orchestrator commands). The two further roles reported in the literature—policy synthesis (turning operator intent into orchestrator actions parameterised against Eq. 1) [138] and retrieval-grounded reasoning over the operator’s telemetry and incident history—inherit the same discipline, leaving the open question of composing telecom-tuned models with retrieval under hard refusal envelopes, i.e. bounded LLM risk (Sec. VIII-B). C. Digital Twins Are the Counterfactual Workbench for the Loop A digital twin (DT) is a synchronised executable replica of a physical CPS that allows the security stack to (i) test detector and mitigation policies against synthetic adversarial scenarios before deployment, (ii) run what-if analyses on candidate mitigation playbooks before committing to them on the live network, and (iii) replay observed incidents to refine the FL training set [87]. The DT therefore plays the role of a safety net for the closed loop: it is the place where model drift can be detected by replaying recent traffic and measuring divergence between live-detector and twin-detector outputs. The DT closes back to FL for retraining. For 6G CPS, three DT concepts are converging. The network DT models the radio, transport, and core; it is the natural place to evaluate orchestrator playbooks for slice availability and recovery time [87]. The CPS DT models the physical asset (vehicle dynamics, grid topology, robotic-arm kinematics); it is where the harm consequence of a delayed mitigation can be quantified. The security DT layers attack scenarios on top of both. Operationally, the security DT is the only place where Challenge C5 (CPS-impact harm metrics) can be operationalised today, because no public dataset labels harm; the DT generates harm labels by execution. Recent work increasingly co-designs DT pipelines with FL-aggregated update streams so that scenarios deemed high-risk by the DT trigger targeted FL rounds focused on the affected slice, bringing labelled-CDR scarcity (Sec. VII-A) closer to a workable equilibrium [87].
23
The DT is therefore not a sidecar but a first-class participant in the closed loop’s learn stage. Four operational uses recur: (i) pre-deployment validation against a curated scenario library before any change touches the production network; (ii) counterfactual replay of observed incidents under alternative orchestrator policies; (iii) synthetic adversary generation, which pairs the DT with reinforcementlearning attackers [83] or diffusion-based traffic synthesis [69], [70] to produce attack scenarios the operational telemetry has not yet seen; and (iv) drift attribution, which disambiguates benign distribution shift from adversarial campaign by replaying recent traffic against a controlled adversary model. No off-the-shelf platform yet integrates all four (Challenge C5). D. Substrate-Level Enablers: PQC, ZTA, XAI Three cross-cutting substrates underpin the loop and are best treated jointly because each is operationally constrained by the same URLLC envelope that bounds the detector and the orchestrator [4], [32]. We restrict attention here to the loopspecific consequences and refer the reader to [32], [72], [139] for the underlying primitives. Post-quantum cryptography vs. the URLLC slice envelope. The standardisation of Kyber (ML-KEM [140]) and Dilithium (ML-DSA) by NIST in 2024 settles the algorithmic question; the binding constraint for 6G CPS is the latency overhead at the AMF/AUSF and on the E2 control plane [6]. A Dilithium-3 signature is ∼3.3 kB per ML-DSA-65 in NIST FIPS 204 [137]; on reference AVX2 x86 implementations the verification cost is ∼0.1 ms and 0.3–0.5 ms on ARM Cortex-A cores, comfortably within a 1 ms slice envelope, with signing the dominant cost at ∼0.3–0.6 ms. Even so, per-message PQC on the submillisecond E2 fast path remains marginal once signature transport (∼3.3 kB) and jitter are accounted for, so sessionsetup or per-batch granularity remains the safer operating point; harvest-now-decrypt-later threats still motivate PQCprotected FL aggregation and orchestrator control channels, where the retention horizon dominates the latency cost. Hybrid classical/PQC handshakes are the dominant transitional pattern in 3GPP working-group discussions [8], and per-message PQC on the E2 fast path remains infeasible until signature compression or hardware verification accelerators close the gap (Challenge C4). Zero-trust architecture as graded enforcement keyed to the detector score. ZTA replaces the perimeter assumption with continuous verification [30], [31], [139]. For the closed loop, the operationally novel primitive is graded enforcement: the detector’s per-flow score (Sec. V) [75] becomes a continuous input to the policy decision point, so that a session whose anomaly score crosses a soft threshold is downgraded to a higher-friction trust tier (additional authentication, lower bandwidth ceiling, restricted slice access) before any hard mitigation is invoked; only a hard-threshold crossing triggers termination. This converts the binary block/allow decision into a continuum that respects safety-of-life CPS workloads and is the policy-layer complement of the FBR metric in Table X. The open standardisation gap is that 3GPP’s 5G-AKA/EAP-AKA’ framework [6] was designed around
session-establishment events, not per-decision checks; carrying continuous-authentication evidence without inflating signalling overhead remains a Challenge C4 item. Explainable AI as a trust bridge with dual-use leakage. Operators and regulators will not delegate URLLC-slice mitigation to a black-box detector, so SHAP/Grad-CAM/IntegratedGradients attributions become a first-class output of the loop, layered cheaply at the MEC and richly at the operator core [32], [62], [71] [51, Chs. 6 and 10]. The loop-specific risk is dual-use leakage: per-feature SHAP attributions on a CNN–LSTM detector measurably leak the decision boundary, and under repeated probing converge to a usable adversarial gradient [32], [141]; post-hoc attribution methods such as LIME and SHAP are themselves adversarially fragile [142]. The operational defence is rate-limited or aggregated explanations per consumer identity, audit logging of every explanation request, and PQC-bound retention of the explanation envelope for the full forensic horizon (Challenge C2). XAI is therefore the trust bridge that lets the loop run autonomously while remaining accountable, not a sidecar [56]. E. Substrate Co-Design: How PQC, ZTA, and XAI Compose The three substrates are not orthogonal: PQC re-keys the control plane that ZTA continuously re-authenticates, and the policy decisions ZTA emits must be auditable by XAI before they reach an xApp. Concretely, a Dilithium-signed NASMAC [6], [137] gates a ZTA policy update whose ML-driven trust-score change is in turn explained by a SHAP attribution served to the SOC analyst. The loop-level consequence is that PQC, ZTA, and XAI share a single latency budget (Fig. 11) and a single audit trail rather than three independent ones. Widening the lens beyond the substrates, the three operational services (FL, LLM, DT) and the three architectural substrates (PQC, ZTA, XAI) play distinct but coupled roles in the closed loop [72]: FL supplies cross-operator training signal without raw-data exchange, LLMs translate naturallanguage intent into vetted xApp actions, and DTs provide counterfactual validation before any policy reaches production; PQC, ZTA, and XAI then make every resulting controlplane transaction quantum-safe, continuously re-authenticated, and auditable. Together, the six elements operate against a single sub-millisecond budget rather than as independent enhancements—the open question, which Sec. VIII formalises, is whether their joint latency, communication, and assurance costs admit a feasible operating point in 6G CPS slices. VIII. O PEN C HALLENGES AND R ESEARCH ROADMAP We compress the scattered sub-problems of prior surveys into five consolidated challenges. Each challenge is intentionally cross-section: the goal is to point researchers at the joint problem rather than at isolated sub-problems whose fixes do not compose into a working closed loop. The urgency of these challenges is driven by external regulatory and standardisation deadlines, as mapped in the 2025–2030 timeline of Fig. 12, which requires closed-loop security solutions to mature ahead of the EU AI Act enforcement window and the 3GPP Release 21 freeze.
24
PQC verify ≈ 0.1 ms (Dilithium-3, NIST FIPS 204)
τmax URLLC slice (τmax = 1 ms)
0.10
eMBB slice (τmax = 10 ms)
0.10
0.30
0.50
0.50
Transport 0.05 Jitter 0.05
2.00
τmax
6.00
τmax Jitter 1.40
mMTC slice (τmax = 100 ms)
0.10
1.00
5.00
80.00 Jitter 13.90
0.1 PQC verify
1 10 End-to-end latency (ms, log scale) ZTA policy
xApp inference
Transport
100 Jitter margin
Fig. 11: End-to-end latency budget per 5G slice. Each row composes PQC verify, ZTA policy lookup, xApp inference, transport, and jitter margin against the slice’s τmax (dashed red). The closed loop must fit within τmax for each slice type; the URLLC row is the binding constraint of Eq. 2. PQC verify uses Dilithium-3 (ML-DSA-65 per NIST FIPS 204 [137]) on AVX2 hardware. A. C1: Sub-Millisecond, Edge-Compatible Detection Under Adversarial ML The composite challenge is to build detectors that simultaneously (i) run within a ∼1 ms URLLC budget on commodity MEC hardware [1], (ii) remain accurate under model drift induced by non-stationary 6G traffic distributions, and (iii) resist evasion and poisoning attacks (Sec. V-H). The current evidence base treats these requirements separately— quantised CNNs vs. adversarial-robust models [143] vs. drift-aware models [118]—but operational deployments need all three at once. The research direction is co-designed compression+robust-training+drift-detection pipelines, evaluated on 6G-native datasets (Challenge C5). Concretely, three sub-problems sit inside C1. Latencybounded compression: int8 quantisation, structured pruning, and distillation must be jointly tuned so that the worst-case (not just average) inference latency fits the slice budget; taillatency excursions caused by MEC workload contention must be bounded with admission-control on the security model itself. Online drift detection and re-fit [118]: the detector must distinguish a true distribution shift (e.g. a new device class entering the network) from an attack (e.g. an evasion campaign), and trigger FL retraining selectively without flooding the aggregator. Robust training and certified defences: adversarial training against realistic 5G/6G perturbation budgets [32], [82], plus randomised smoothing or gradient masking where appropriate; the operational target is to certify a minimum recall on a defined adversarial budget rather than to chase headline accuracy on clean data. A promising research direction is the joint optimisation of these three sub-problems: a single training procedure that produces a model whose certified robustness, drift sensitivity, and worst-case inference latency are all bounded simultaneously. Current evidence is fragmentary: Doriguzzi-Corin et al. [40] optimise for latency without explicit adversarialrobustness analysis; the GAN-based-evasion-and-defence literature surveyed by Alauthman et al. [82] reports robustness gains against GAN-generated evasion (IDS-GAN, attackGAN, SGAN-IDS) without consistent inference-latency budgets; few
works address all three together. A second sub-problem is the choice of drift signal: shadow-stream score divergence, KL-divergence on the input feature distribution, or per-class reconstruction-error percentiles each have different sensitivity profiles, and the right combination depends on the operator’s tolerance for false drift alarms. A third is certified defences in production: randomised smoothing inflates inference latency by an order of magnitude in its naive form, and operator deployments need either per-decision certificates with bounded compute or a principled fall-back to non-certified inference when the latency budget is tight. The composition of these subproblems remains an open research direction; an end-to-end MEC-deployable detector with certified robustness, drift-aware online retraining, and worst-case sub-millisecond latency does not yet exist in the surveyed literature. Falsifiable target (2028). A MEC-deployable detector that demonstrates ≥ 95% recall on a community-agreed adversarial perturbation budget at p99 inference latency ≤ 1 ms on commodity O-RAN-class hardware, with the drift-detection trigger evaluated on at least two independent operator telemetry streams. Vehicle. An MLPerf-Inference-style benchmark consortium [144] coordinating with the O-RAN Software Community pre-standardisation testbed track, with a target reference implementation released by 2028. B. C2: Explainable, Audit-Ready AI Decisions with Bounded LLM Risk Black-box detectors and black-box LLM SOC assistants are individually risky; composed in a closed loop they are operationally untenable. The research direction is twofold: (a) XAI methods whose attributions are stable under small perturbations (counter to [32]), and (b) LLM SOC interfaces grounded in the operator’s own telemetry store with hard refusal envelopes when telemetry support is missing [86], [51, Ch. 10]. The regulatory horizon makes this urgent. EU AI-Act-style obligations on high-risk AI [145] will reach safety-critical CPS slices well before the 6G commercial launch window, and operators will be required to produce per-decision explanations that survive audit. The composition problem—a chain of XAI-explained detector → XAI-explained orchestrator → LLM-summarised incident report—is poorly studied today. The research direction is end-to-end auditability: every loop transition emits an attestable explanation that is (i) stable under bounded perturbation, (ii) access-controlled per Salmi and Bogucka’s dual-use warning [32], and (iii) retained for the full forensic horizon under PQC integrity protection. A concrete protocol sketch for audit-ready closed loops would standardise (i) an explanation envelope format that carries the per-feature attribution, the model version, the input snapshot, and a cryptographic commitment to the model’s parameters at decision time; (ii) a chain-of-custody record that links the detector’s explanation to the orchestrator’s chosen action and to any LLM-generated summary, with a verifiable transcript of the orchestrator’s cost computation (Eq. 1); and (iii) a regulator-facing query interface that allows post-hoc verification of any specific decision against the operator’s
25
EU AI Act general application (2 Aug 2026)
3GPP Rel-21 freeze (expected 2028; dates TBD)
O-RAN WG2 test spec
Ref. impl. target
C1: Sub-ms edge detection MLPerf-style IDS track (proposed) NIST AI RMF 2.0
C2: XAI / LLM auditability EU AI Act Art. 13 transparency (2 Aug 2027)
O-RAN WG3 FL profile
ETSI SAI XAI profile 3GPP R20 FL-as-a-service
C3: Cross-operator FL EU AI Act Art. 10 data gov. (2 Aug 2027)
3GPP R19 freeze (NWDAF Phase 3)
CNSA 2.0 software-signing target (2030)
3GPP R20 (synth-CDR study)
C4: Standardisation gaps NIST FIPS 203/204/205 (Aug 2024)
3GPP R19 TR 33.876 (harm metrics)
Shared corpus target
C5: 6G benchmarks & harm CPS-harm metrics study
2025 standards milestone
2026 regulatory deadline
2027 - - EU AI Act
2028
2029
2030
- - 3GPP Rel-21
Fig. 12: C1–C5 research challenges mapped onto the 2025–2030 standardisation roadmap. C1 covers sub-millisecond edge detection; C2 XAI/LLM auditability under the EU AI Act; C3 cross-operator federated learning; C4 standardisation gaps across 3GPP, O-RAN, NIST and ETSI; C5 6G-native benchmarks and harm-aware evaluation. Two regulatory checkpoints frame the timeline: EU AI Act general application (2 Aug 2026, red dashed line) and the expected 3GPP Rel-21 freeze (2028, dates TBD, blue dashed line); high-risk obligations under EU AI Act Art. 10/13 trigger on 2 Aug 2027 (red markers). Black markers denote standards milestones. retained explanation envelope, building on ETSI’s Security Monitoring and Management (SMM) information model [22]. Operationally, the explanation envelope is small enough to be produced per decision at the MEC tier (a few hundred bytes for a typical attribution vector) and the chain-of-custody record can be batched and signed at session boundaries to amortise PQC signature cost. The standardisation lives in 3GPP SA3 and ETSI ISG SAI; closure on the envelope schema and the audit query interface is the highest-leverage step for converting the AI-Act-compatible aspiration into deployable infrastructure. Falsifiable target (2028). Every closed-loop transition (detect → orchestrator → LLM summary) emits an attestable explanation receipt of bounded size whose stability under bounded input perturbation is empirically validated on a public CPS slice testbed, with hard refusal envelopes for the LLM SOC interface. Vehicle. 3GPP SA3 plus ETSI ISG SAI, with a published Technical Report on AI-native closed-loop attestation alignment by 2028 and a regulator-facing query interface profiled against EU AI-Act obligations.
C. C3: FL Across Competing Operators with Privacy and Byzantine Guarantees The labelled-CDR scarcity problem (Sec. VII-A) can only be solved at scale by federating across operators, but operators are competitors and CDRs are personally identifiable information (PII) under most jurisdictions. The research direction is FL protocols that combine (i) differential privacy with auditable budgets, (ii) Byzantine-robust aggregation that tolerates compromised or strategically-misreporting operators, and (iii) cross-border legal compliance baked into the protocol rather than bolted on after the fact [131]. Surveys [28], [37], [76], [85] agree on the gap; concrete protocols remain scarce.
The technical primitives are individually mature; the open research direction is their composition within a single protocol that is jointly analysed for convergence, privacy budget, and Byzantine tolerance, and that admits a reproducible incentive structure. No surveyed deployment has integrated all four properties end-to-end. Falsifiable target (2027). A single FL protocol jointly analysed for ε-differential-privacy budget (ε ≤ 2), Byzantine tolerance against up to ⌊(n − 1)/3⌋ colluding operators in an n-operator federation, and incentive compatibility, with reproducible code released and convergence demonstrated on a cross-operator CDR benchmark. Vehicle. IETF NMRG together with the O-RAN Security Working Group, with a target Internet-Draft and a multi-operator pilot by 2027. D. C4: API, Lifecycle, and Conformance Standardisation for AI-Native Security in 3GPP and O-RAN Today there is no 3GPP-standardised CDR-based security API exposing the features Sec. II-D relies on [2]; the NWDAF event-exposure surface is closer but still ML-agnostic. Similarly, O-RAN xApp lifecycle management does not yet specify how AI security xApps are tested, certified, and rolled back, nor how runtime policy conflicts among independently developed xApps are detected before they corrupt KPIs [9], [11], [52], [110]. Heterogeneous MEC vendor stacks add a third dimension. The research direction is concrete standards proposals—specifying the API, the lifecycle, the conformance tests—that operators and regulators can adopt. Four sub-items have the highest leverage. NWDAF security extension: an event-exposure profile dedicated to anomaly/attack signals with a stable schema across vendors. xApp security lifecycle: signing, attestation, capability declaration, and graceful rollback procedures specified by the O-RAN Alliance and aligned with 3GPP SA3 [6], [52]. ETSI MEC
26
security profile: standardised exposure of MEC-host telemetry to security applications with role-based access control [21], [22]. Cross-domain orchestration: an SDN/NFV/O-RAN orchestrator interface that lets a single playbook reach across actuator tiers without bespoke integration [45]. Surveys of ORAN security [9], [26] all converge on the lifecycle gap as the single highest-leverage standardisation target. The research direction therefore decomposes into a standards-track contribution and an academic-experimental contribution. The standards-track contribution is to draft and shepherd the four profile specifications above through 3GPP SA3, the O-RAN Security Working Group, and ETSI MEC, with reference implementations contributed to opensource code bases that operators can deploy immediately. The academic-experimental contribution is to evaluate proposed profiles on multi-vendor testbeds before standardisation closes, exposing integration gaps that paper-based design reviews cannot reveal; the over-the-air O-RAN testbed of [42] and the multi-layer mitigation architecture of [43] are early models for this kind of pre-standardisation experimental validation. Without coupled standards-and-experiment contributions, AInative 6G CPS security risks the same vendor-lock-in pathology that 5G’s NSA-to-SA migration encountered, where each operator’s deployment was bespoke and inter-operator FL or cross-operator mitigation was practically infeasible [25]. Falsifiable target (Rel-20/2028). A 3GPP SA3 Technical Report on AI-native security closed-loop attestation, an ORAN Alliance xApp-security-lifecycle specification (signing, attestation, capability declaration, rollback), and an ETSI MEC security profile with role-based access control, each accompanied by an open-source reference implementation. Vehicle. 3GPP SA3 and O-RAN Security Working Group (WG11), targeted at Rel-20 with conformance tests deliverable by 2028. E. C5: 6G-Native Benchmarks, CPS-Specific Harm Metrics, and Realistic CDR Datasets The benchmark gap is the binding constraint on credible evaluation. Public datasets (Sec. IV-D) are pre-6G, mostly non-CDR, mostly without slice-aware labels, and entirely without CPS-specific harm labels (e.g., “did this attack delay a defibrillation?”); the closest 6G-native benchmark to date evaluates LLM reasoning over standardisation-derived tasks rather than CDR/RAN telemetry [138]. Energy-efficient inference benchmarks are also absent. The research direction is operator-academic consortia that release anonymised sliceaware CDR + RAN telemetry with both attack labels and CPS-impact labels, plus reference harm metrics that go beyond accuracy/F1. Regulatory and commercial obstacles bind the path. CDR/RAN telemetry carries personally identifiable information regulated by GDPR, CCPA, and analogues [56], [145], yet the de-identification needed for lawful release destroys the cross-modal correlations detectors must learn; per-cell load, device-mix, and subscriber-density signals are also competitively sensitive. The path forward composes three nonnovel elements into a working pipeline: (i) differential-privacy release with auditable budgets; (ii) synthetic-but-realistic generation [69], [70], [94] conditioned on real distributions; and
(iii) multi-stakeholder governance (operator, vendor, academic, and regulator) so the benchmark is credible to all parties. Four data-set properties are needed but absent in the public corpus. Slice-awareness: labels per slice and per CPS class, so detectors can be evaluated on the slice their target CPS lives on. Native multi-modality: CDR + RAN E2 KPM + NWDAF events captured on the same incidents, so crossmodality fusion methods can be evaluated end-to-end (the recent NetsLab-5GORAN-IDD release [61] partially fulfils this by synchronising packet-level and E2-radio-telemetry capture on a live OAI O-RAN testbed, but lacks NWDAF and CDR streams; the NANCY O-RAN release [92] is a further proxy). The remaining two properties concern evaluation rather than capture. CPS-impact labels: per-incident annotations mapping each vertical’s harm to the loop metrics of Table X—V2X missed-hazard-warning (warning lost once td +tm exceeds the braking budget or an FBR event drops the message, with trec setting how long peers stay unwarned after a false block), smart-grid relay trip-delay / false-trip (td +tm late against the fault-propagation budget, a spurious FBR trip, or trec before the relay path is restored), and telesurgery control/haptic delay (td +tm on the ∼10 ms loop, an FBR drop losing a haptic control sample, with trec bounding session interruption)— produced by operator–CPS-vendor partnerships or DT replay (Sec. VII-C). Energy and tail-latency benchmarks: per-method energy-per-decision and 99th-percentile inference latency, so MEC deployability can be assessed before deployment; a 5GDDoS-specific reproducibility checklist covering methodology, datasets, code/tools, experimental setup, implementation details, result reproducibility, and limitations has been recently proposed and can be adopted as a baseline [20]. Falsifiable target (2028). An operator-academic-vendor consortium release of a slice-aware CDR + RAN + NWDAF dataset with both attack labels and CPS-impact labels (delay-to-defibrillation, missed-V2X-warning, unscheduledrelay-trip), distributed under auditable differential privacy (ε ≤ 2, with the convergence-vs-budget trade-off and optimalround-count T ∗ characterised analytically by Noising before Aggregation Federated Learning (NbAFL) [134]), accompanied by per-method energy-per-decision and p99 inferencelatency baselines. Vehicle. O-RAN Software Community plus an MLPerf-Inference-style benchmark consortium [144] plus 3GPP SA3, with the first release targeted at 2028 and DTreplay-generated harm labels as the bridge until operator-CPSvendor partnerships mature. The five challenges form a closed agenda: C1 makes detection viable, C2 makes the loop trustworthy, C3 makes training feasible across operators, C4 makes deployment standardsconformant, and C5 makes evaluation honest. IX. C ONCLUSION We have argued—across 128 peer-reviewed studies organised under a PRISMA 2020 protocol—that AI-native security for 6G CPSs must be reframed as a closed-loop pipeline that senses on CDR/O-RAN telemetry, detects locally at the MEC tier, mitigates network-wide through SDN/NFV/O-RAN,
27
and learns continuously through FL and DT replay. This reframing eliminates the historical split between “edge anomaly detection” and “DDoS classification” surveys (Sec. V), recasts “emerging paradigms” as cross-cutting enablers serving one loop (Sec. VII), and converts scattered open problems into five composable challenges (Sec. VIII). The evidence base is now mature enough to support the feasibility of the loop: MEC-deployed CNN detectors reach >98% accuracy [35], [40] with demonstrated feasibility on resource-constrained edge hardware [40], O-RAN xApp loops mitigate at sub-second latency on over-the-air testbeds [42], and federated training across heterogeneous edges matches centralised IDS baselines [124]. The architecture meets the 1 ms URLLC NAS–MAC contract end-to-end at p99 , with Dilithium-3 (ML-DSA-65 per NIST FIPS 204 [137]) verify (0.1 ms on AVX2), ZTA policy lookup (0.3 ms), xApp inference (0.5 ms), transport (0.05 ms), and jitter margin (0.05 ms) composing to the 1 ms slice envelope (Fig. 11)—the principal feasibility claim of this survey. What remains is the standardisation, the 6G-native benchmark, and the cross-operator FL governance needed to run the loop safely in production. Three implications follow. First, effort should tilt from detection, now relatively mature, to mitigation, where the reversibility gap, the cross-tier orchestrator, and audit-grade explanation of autonomous actions remain under-studied. Second, the benchmark gap (Challenge C5) is more a governance and incentive problem than a technical one, so operator– academic–vendor consortia for CPS-impact-labelled releases deserve treatment as first-class research contributions. Third, the standardisation agenda (Challenge C4) is concrete enough that academic drafts, reference implementations, and prestandardisation testbeds can run in parallel with 3GPP and O-RAN Alliance timelines. Several limitations bound these conclusions. 6G-native security data does not yet exist, so the loop is assembled from 5G, O-RAN, and CDR proxies (5G-NIDD [91], NANCY [92], NetsLab-5GORAN-IDD [61]) rather than a single native benchmark, and the findings of Sec. V-E inherit that gap; the only operator-grade CDR set carries no labelled attacks, leaving CDR-driven results reliant on injected or synthetic anomalies [38], [41]; CDR is a minute-scale signal that cannot alone drive a sub-millisecond reaction, serving as first-level triage composed with sub-ms RAN/E2 telemetry (Sec. II-D); and several enablers are surveyed as design directions, with cross-operator FL still chiefly a governance problem (Challenge C3), CPS-impact harm labels obtainable today only by DT replay (Sec. VII-C), and LLM use confined to the explainand-query role (Sec. VII-B). These limitations scope, rather than overturn, the feasibility claim. Two questions are left deliberately open. The first is sociotechnical: how to credibly share attack intelligence, model updates, and mitigation telemetry across operators who remain commercial competitors under heterogeneous legal regimes— the primitives of Challenge C3 are necessary but not sufficient, and the governance architecture binding them is itself a research agenda. The second is how the closed loop interacts with AI-Act-style regulatory oversight [145], particularly the auditability and explanation-stability constraints of Chal-
lenge C2; the loop is a precondition for audit-ready decision records, but the legal-operational composition that turns them into regulator-facing evidence lies outside a technical survey. Both are best addressed by interdisciplinary teams, for which we hope this survey supplies a shared architectural vocabulary. R EFERENCES [1] C.-X. Wang et al., “On the road to 6G: Visions, requirements, key technologies, and testbeds,” IEEE Commun. Surveys Tuts., vol. 25, no. 2, pp. 905–974, 2023. [2] R. Singh, A. Kaushik, W. Shin, M. D. Renzo, V. Sciancalepore, D. Lee, H. Sasaki, A. Shojaeifard, and O. A. Dobre, “Toward 6g evolution: Three enhancements, three innovations, and three major challenges,” IEEE Netw., vol. 39, no. 6, pp. 139–147, 2025. [3] W. Saad, M. Bennis, and M. Chen, “A vision of 6g wireless systems: Applications, trends, technologies, and open research problems,” IEEE Netw., vol. 34, no. 3, pp. 134–142, 2020. [4] X. You et al., “Towards 6G wireless communication networks: Vision, enabling technologies, and new paradigm shifts,” Sci. China Inf. Sci., vol. 64, no. 1, pp. 1–74, Jan. 2021. [5] A. Humayed, J. Lin, F. Li, and B. Luo, “Cyber-physical systems security—a survey,” IEEE Internet Things J., vol. 4, no. 6, pp. 1802– 1831, Dec. 2017. [6] ETSI, “5G; security architecture and procedures for 5G system (3GPP TS 33.501 version 18.11.0 release 18),” European Telecommunications Standards Institute (ETSI), Technical Specification ETSI TS 133 501 V18.11.0, Apr. 2026. [7] I. Ahmad, S. Shahabuddin, T. Kumar, J. Okwuibe, A. Gurtov, and M. Ylianttila, “Security for 5g and beyond,” IEEE Commun. Surveys Tuts., vol. 21, no. 4, pp. 3682–3722, 2019. [8] P. Porambage, G. Gür, D. P. M. Osorio, M. Liyanage, A. Gurtov, and M. Ylianttila, “The roadmap to 6g security and privacy,” IEEE Open J. Commun. Soc., vol. 2, pp. 1094–1122, 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:235077796 [9] M. Polese, L. Bonati, S. D’Oro, S. Basagni, and T. Melodia, “Understanding O-RAN: Architecture, interfaces, algorithms, security, and research challenges,” IEEE Commun. Surveys Tuts., vol. 25, no. 2, pp. 1376–1411, 2023. [10] L. Bonati, S. D’Oro, M. Polese, S. Basagni, and T. Melodia, “Intelligence and learning in O-RAN for data-driven NextG cellular networks,” IEEE Commun. Mag., vol. 59, no. 10, pp. 21–27, Oct. 2021. [11] O-RAN ALLIANCE, “O-RAN Architecture Description,” O-RAN ALLIANCE e.V., Alfter, Germany, Technical Specification ORAN.WG1.TS.OAD-R005-v16.00, Feb. 2026. [Online]. Available: https://specifications.o-ran.org/download?id=1013 [12] N. Li, Q. Sun, L. Wang, X. Xu, J. Huang, C. Liu, J. Gao, Y. Huang, and C. I, “Towards AI-native RAN: An operator’s perspective of 6G day 1 standardization,” arXiv preprint arXiv:2507.08403, Jul. 2025, available: https://arxiv.org/abs/2507.08403. [13] E. A. Lee, “Cyber physical systems: Design challenges,” in Proc. 11th IEEE Int. Symp. Object/Compon./service-Oriented Real-Time Distrib. Comput. (ISORC), Orlando, FL, USA, May 2008, pp. 363–369. [14] R. R. Rajkumar, I. Lee, L. Sha, and J. Stankovic, “Cyber-physical systems: the next computing revolution,” in Proc. 47th Design Autom. Conf. (DAC), ser. DAC ’10. New York, NY, USA: Association for Computing Machinery, 2010, pp. 731–736. [Online]. Available: https://doi.org/10.1145/1837274.1837461 [15] D. C. Nguyen et al., “6G internet of things: A comprehensive survey,” IEEE Internet Things J., vol. 9, no. 1, pp. 359–383, Jan. 2022. [16] S. Deng, H. Zhao, W. Fang, J. Yin, S. Dustdar, and A. Y. Zomaya, “Edge intelligence: The confluence of edge computing and artificial intelligence,” IEEE Internet Things J., vol. 7, no. 8, pp. 7457–7469, Aug. 2020. [17] X. Wang, Y. Han, V. C. M. Leung, D. Niyato, X. Yan, and X. Chen, “Convergence of edge computing and deep learning: A comprehensive survey,” IEEE Commun. Surveys Tuts., vol. 22, no. 2, pp. 869–904, 2020. [18] M. A. Ferrag, L. Maglaras, S. Moschoyiannis, and H. Janicke, “Deep learning for cyber security intrusion detection: Approaches, datasets, and comparative study,” J. Inf. Security Appl., vol. 50, no. 102419, Feb. 2020. [Online]. Available: https://www.sciencedirect.com/science/ article/pii/S2214212619305046 [19] B. Hussain, Q. Du, B. Sun, and Z. Han, “Deep learning-based DDoSattack detection for cyber–physical system over 5G network,” IEEE Trans. Ind. Informat., vol. 17, no. 2, pp. 860–870, Feb. 2021.
28
[20] S. Hoque, A. Aydeger, E. Zeydan, and M. Liyanage, “A survey on distributed denial-of-service attack mitigation for 5G and beyond,” IEEE Open J. Commun. Soc., vol. 6, pp. 5840–5879, 2025. [21] ETSI, “Multi-access edge computing (mec); framework and reference architecture,” European Telecommunications Standards Institute (ETSI), Tech. Rep. ETSI GS MEC 003 V4.1.1, May 2025. [Online]. Available: https://www.etsi.org/deliver/etsi gs/mec/001 099/ 003/04.01.01 60/gs mec003v040101p.pdf [22] ETSI, “Multi-access edge computing (MEC); support for security monitoring and management,” European Telecommunications Standards Institute (ETSI), ETSI Group Specification (GS) GS MEC 062 V4.1.1, Feb. 2026. [23] D. Sabella et al., “Developing software for multi-access edge computing,” European Telecommunications Standards Institute (ETSI), White Paper 20, Feb. 2019. [Online]. Available: https://www.etsi.org/images/ files/ETSIWhitePapers/etsi wp20ed2 MEC SoftwareDevelopment.pdf [24] N. Abbas, Y. Zhang, A. Taherkordi, and T. Skeie, “Mobile edge computing: A survey,” IEEE Internet Things J., vol. 5, no. 1, pp. 450– 465, Feb. 2018. [25] L. Bonati, M. Polese, S. D’Oro, S. Basagni, and T. Melodia, “Open, programmable, and virtualized 5G networks: State-of-the-art and the road ahead,” Comput. Netw., vol. 182, no. 107516, Dec. 2020. [26] A. S. Abdalla, P. S. Upadhyaya, V. K. Shah, and V. Marojevic, “Toward next generation open radio access networks—what O-RAN can and cannot do!” IEEE Netw., vol. 36, no. 6, pp. 206–213, nov./dec. 2022. [27] H. Wen, P. Porras, V. Yegneswaran, A. Gehani, and Z. Lin, “5GSPECTOR: An O-RAN compliant layer-3 cellular attack detection service,” in Proc. Netw. Distrib. Syst. Secur. Symp. (NDSS), 2024. [28] C. d. Alwis, O. Aouedi, J. Xu, S. Wang, Y. Siriwardhana, T. Hewa, E. Zeydan, C. Sandeepa, and M. Liyanage, “Federated learning for 6G security: A survey on threats, solutions, and research directions,” IEEE Commun. Surveys Tuts., vol. 28, pp. 4883–4914, 2026. [29] J. Shi et al., “IMS is not that secure on your 5G/4G phones,” in Proc. 30th Annu. Int. Conf. Mobile Computing and Networking (MobiCom), Washington, DC, USA, May 2024, pp. 513–527. [30] N. Nahar, K. Andersson, O. Schelén, and S. Saguna, “A survey on zero trust architecture: Applications and challenges of 6G networks,” IEEE Access, vol. 12, pp. 93 562–93 590, 2024. [31] B. Agarwal, R. Irmer, D. Lister, and G.-M. Muntean, “Open RAN for 6G networks: Architecture, use cases and open issues,” IEEE Commun. Surveys Tuts., vol. 28, pp. 2881–2924, 2026. [32] Y. Salmi and H. Bogucka, “XAI-guided physical adversarial machine learning attacks on automatic modulation classification,” IEEE J. Sel. Areas Commun., vol. 44, pp. 4271–4286, 2026. [33] B. Zhang, P. Zeinaty, N. Limam, and R. Boutaba, “Mitigating signaling storms in 5G with blockchain-assisted 5GAKA,” in 2023 19th International Conference on Network and Service Management (CNSM), 2023, pp. 1–9. [34] D. K. Nguyen, R. E. Malki, and F. Rebecchi, “RRC signaling storm detection in O-RAN,” in Proc. 2025 IEEE Symp. Comput. Commun. (ISCC), 2025. [35] B. Hussain, Q. Du, S. Zhang, A. Imran, and M. A. Imran, “Mobile edge computing-based data-driven deep learning framework for anomaly detection,” IEEE Access, vol. 7, pp. 137 656–137 667, 2019. [36] B. Hussain, Q. Du, A. Imran, and M. A. Imran, “Artificial intelligencepowered mobile edge computing-based anomaly detection in cellular networks,” IEEE Trans. Ind. Informat., vol. 16, no. 8, pp. 4986–4996, Aug. 2020. [37] A. Blika, S. Palmos, G. Doukas, V. Lamprou, S. Pelekis, M. Kontoulis, C. Ntanos, and D. Askounis, “Federated learning for enhanced cybersecurity and trustworthiness in 5G and 6G networks: A comprehensive survey,” IEEE Open J. Commun. Soc., 2024. [38] G. Barlacchi et al., “A multi-source dataset of urban life in the city of Milan and the province of Trentino,” Sci. Data, vol. 2, no. 150055, Oct. 2015. [39] M. Christopoulou et al., “User terminals as attackers: An open dataset analysis of DDoS attacks in 5G networks,” in Proc. IEEE Conf. Standards Commun. Netw. (CSCN), 2024, pp. 301–307. [40] R. Doriguzzi-Corin, S. Millar, S. Scott-Hayward, J. Martı́nez-del Rincón, and D. Siracusa, “LUCID: A practical, lightweight deep learning solution for DDoS attack detection,” IEEE Trans. Netw. Service Manage., vol. 17, no. 2, pp. 876–889, Jun. 2020. [41] J. M. Franco-Valiente, J. Calle-Cancho, D. Cortés-Polo, and J. M. Haut, “An efficiency-driven anomaly detection pipeline for mobile networks,” The Journal of Supercomputing, vol. 82, no. 5, p. 361, 2026. [42] J. Moore, A. S. Abdalla, Z. Reshi, and V. Marojevic, “Anomaly detection and mitigation in O-RAN networks using an LSTM-RNN
autoencoder and secure slicing,” in 2025 IEEE Military Communications Conference (MILCOM), 2025, pp. 1–6. [43] Z. Allaw, O. Zein, and A.-M. Ahmad, “Cross-layer security for 5G/6G network slices: An SDN, NFV, and AI-based hybrid framework,” Sensors, vol. 25, no. 11, p. 3335, 2025. [44] M. Polese, N. Mohamadi, S. D’Oro, L. Bonati, and T. Melodia, “Beyond connectivity: An open architecture for AI-RAN convergence in 6G,” IEEE Commun. Mag., pp. 1–6, 2026. [45] ETSI, “Network functions virtualisation (nfv) release 5; management and orchestration; architecture enhancement for security management specification,” European Telecommunications Standards Institute (ETSI), Group Specification ETSI GS NFV-IFA 026 V5.4.1, Jan. 2026. [46] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y. Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Proc. 20th Int. Conf. Artif. Intell. Statist. (AISTATS). Fort Lauderdale, FL, USA: PMLR, 2017, pp. 1273–1282. [47] M. J. Page, J. E. McKenzie, P. M. Bossuyt et al., “PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews,” BMJ, vol. 372, p. n160, 2021. [48] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, pp. 436–444, May 2015. [49] A. Vaswani et al., “Attention is all you need,” in Proc. 31st Conf. Neural Inf. Process. Syst. (NIPS). Long Beach, CA, USA: Curran Associates Inc., 2017, pp. 6000–6010. [50] Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and P. S. Yu, “A comprehensive survey on graph neural networks,” IEEE Trans. Neural Netw. Learn. Syst., vol. 32, no. 1, pp. 4–24, 2021. [51] I. H. Sarker, AI-Driven Cybersecurity and Threat Intelligence. Cham, Switzerland: Springer Nature Switzerland, 2024. [Online]. Available: https://link.springer.com/book/10.1007/978-3-031-54497-2 [52] C.-H. Huang, C.-K. Wen, and G. Y. Li, “AI/ML life cycle management for interoperable AI native RAN,” IEEE Commun. Standards Mag., 2025, early Access; arXiv preprint arXiv:2507.18538v2. [Online]. Available: https://arxiv.org/abs/2507.18538 [53] A. Thantharate, R. Paropkari, V. Walunj, C. Beard, and P. Kankariya, “Secure5G: A deep learning framework towards a secure network slicing in 5G and beyond,” in Proc. 10th Annu. Comput. Commun. Workshop Conf. (CCWC), 2020, pp. 852–857. [54] Q. Wu and R. Zhang, “Intelligent reflecting surface enhanced wireless network via joint active and passive beamforming,” IEEE Trans. Wireless Commun., vol. 18, no. 11, pp. 5394–5409, Nov. 2019. [55] S. Hassouna, M. Zubair, A. Raza, M. Ahmed, J. Kazim, M. A. Imran, and Q. H. Abbasi, “RISCS: RIS-aided integrated sensing and communications—a cyber-physical security perspective,” IEEE Trans. Consum. Electron., 2026, to be published, doi: 10.1109/TCE.2026.3674062. [56] ETSI, “Integrated sensing and communications (isac); security, privacy, trustworthiness and sustainability,” European Telecommunications Standards Institute (ETSI), Group Report ETSI GR ISC 004 V1.1.1, Feb. 2026. [Online]. Available: https://www.etsi.org/deliver/etsi gr/ ISC/001 099/004/01.01.01 60/gr ISC004v010101p.pdf [57] S. Kim, K.-J. Park, and C. Lu, “A survey on network security for cyber–physical systems: From threats to resilient design,” IEEE Commun. Surv. Tutorials, vol. 24, no. 3, pp. 1534–1573, Aug. 2022. [58] Y. Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief, “A survey on mobile edge computing: The communication perspective,” IEEE Commun. Surveys Tuts., vol. 19, no. 4, pp. 2322–2358, 2017. [59] Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” Proc. IEEE, vol. 107, no. 8, pp. 1738–1762, Aug. 2019. [60] 3GPP, “Charging architecture and principles,” 3rd Generation Partnership Project, Technical Specification (TS) 32.240, ver. 19.4.0, Release 19, Oct. 2025. [Online]. Available: https://www.etsi.org/deliver/etsi TS/132200 132299/132240/19.04.00 60/ts 132240v190400p.pdf [61] F. A. Zadeh, A. Civciss, V. Ravihansa, C. Sandeepa, and M. Liyanage, “Descriptor: 5G open radio access network multi-modal intrusion detection dataset (NetsLab-5GORAN-IDD),” IEEE Data Descriptions, vol. 2, pp. 348–357, 2025. [62] B. Azkaei, K. C. Joshi, and G. Exarchakos, “Machine learning-driven anomaly detection for 5G O-RAN performance metrics,” in Proc. IEEE Conf. Comput. Commun. Workshops (INFOCOM WKSHPS), May 2025, pp. 1–6. [63] B. Hussain, Q. Du, and P. Ren, “Big data-driven anomaly detection in cellular networks,” in Proc. IEEE/CIC Int. Conf. Commun. China (ICCC), Qingdao, China, 2017, pp. 1–6.
29
[64] B. Hussain, Q. Du, and P. Ren, “Semi-supervised learning based big data-driven anomaly detection in mobile wireless networks,” China Communications, vol. 15, no. 4, pp. 41–57, Apr. 2018. [65] B. Hussain, Q. Du, and P. Ren, “Deep learning-based big dataassisted anomaly detection in cellular networks,” in Proc. IEEE Global Commun. Conf. (GLOBECOM), Abu Dhabi, UAE, 2018, pp. 1–6. [66] K. Sultan, H. Ali, and Z. Zhang, “Call detail records driven anomaly detection and traffic prediction in mobile cellular networks,” IEEE Access, vol. 6, pp. 41 728–41 737, 2018. [67] Z. Aziz and R. Bestak, “Insight into anomaly detection and prediction and mobile network security enhancement leveraging K-means clustering on call detail records,” Sensors, vol. 24, no. 6, pp. 1–18, 2024. [68] Z. Aziz and R. Bestak, “Modeling voice traffic patterns for anomaly detection and prediction in cellular networks based on CDR data,” IEEE Trans. Mobile Comput., vol. 23, no. 12, pp. 13 131–13 143, Dec. 2024. [69] S. Duan, Y. Zhang, F. Lyu, C. Zhou, and X. Shen, “Mobile network data synthesis with generative AI: Challenges and solutions,” IEEE Commun. Mag., vol. 63, no. 12, pp. 172–178, 2025. [70] M. C. Stoian, E. Giunchiglia, and T. Lukasiewicz, “A survey on deep learning approaches for tabular data generation: Utility, alignment, fidelity, privacy, diversity, and beyond,” Trans. Mach. Learn. Res. (TMLR), 2026. [Online]. Available: https://openreview.net/forum?id= RoShSRQQ67 [71] G. Xylouris, A. Vekraki, M. Christopoulou, M.-A. Kourtis, E. K. Markakis, and P. Trakadas, “Advancing predictive security for consumer applications in beyond 5G/6G networks with annotated datasets,” IEEE Transactions on Consumer Electronics, vol. 71, no. 3, pp. 1–10, May 2025. [72] R. Kumar, J. Dutta, N. Vamsi, U. Sankararao Varri, and D. Puthal, “Next-generation security in the 6G era: The role of AI in safeguarding future networks,” IEEE Access, vol. 14, pp. 17 347–17 380, 2026. [73] G. Pang, C. Shen, L. Cao, and A. V. D. Hengel, “Deep learning for anomaly detection: A review,” ACM Comput. Surv., vol. 54, no. 2, Mar. 2021. [74] A. Thakkar and R. Lohiya, “A survey on intrusion detection system: Feature selection, model, performance measures, application perspective, challenges, and future research directions,” Artif. Intell. Rev., vol. 55, no. 1, pp. 453–563, Jan. 2022. [Online]. Available: https://doi.org/10.1007/s10462-021-10037-9 [75] V. Chandola, A. Banerjee, and V. Kumar, “Anomaly detection: A survey,” ACM Comput. Surv., vol. 41, no. 3, Jul. 2009. [76] V. Mothukuri, R. M. Parizi, S. Pouriyeh, Y. Huang, A. Dehghantanha, and G. Srivastava, “A survey on security and privacy of federated learning,” Future Gener. Comput. Syst., vol. 115, pp. 619–640, Feb. 2021. [77] I. Sharafaldin, A. H. Lashkari, S. Hakak, and A. A. Ghorbani, “Developing realistic distributed denial of service (DDoS) attack dataset and taxonomy,” in Proc. Int. Carnahan Conf. Security Technol. (ICCST), Chennai, India, 2019, pp. 1–8. [78] N. Moustafa and J. Slay, “UNSW-NB15: A comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set),” in Proc. Military Commun. Inf. Syst. Conf. (MilCIS), Canberra, ACT, Australia, 2015, pp. 1–6. [79] N. Koroniotis, N. Moustafa, E. Sitnikova, and B. Turnbull, “Towards the development of realistic botnet dataset in the Internet of Things for network forensic analytics: Bot-IoT dataset,” Future Gener. Comput. Syst., vol. 100, pp. 779–796, Nov. 2019. [80] I. H. Abdulqadder, S. Zhou, D. Zou, I. T. Aziz, and S. M. A. Akber, “Multi-layered intrusion detection and prevention in the SDN/NFV enabled cloud of 5G networks using AI-based defense mechanisms,” Comput. Netw., vol. 179, no. 107364, Oct. 2020. [81] Z. Cui, B. Cui, J. Fu, and R. Dong, “Security threats to voice services in 5G standalone networks,” Secur. Commun. Netw., vol. 2022, no. 1, p. 7395128, 2022. [82] M. Alauthman, N. Aslam, A. Al-Qerem, A. Aldweesh, and P. Sureephong, “Generative adversarial networks for intrusion detection systems: A comprehensive survey of applications, challenges, and research directions,” Arab. J. Sci. Eng., vol. 51, no. 1, pp. 179–203, 2026. [83] T. T. Nguyen and V. J. Reddi, “Deep reinforcement learning for cyber security,” IEEE Trans. Neural Netw. Learn. Syst., vol. 34, no. 8, pp. 3779–3795, Aug. 2023. [84] A. Alam, A. Umer, I. Ullah, and A. Alsayat, “AI-enabled cybersecurity framework for future 5G wireless infrastructures,” Sci. Rep., vol. 16, p. 7055, 2026. [85] Y. Wan, Y. Qu, W. Ni, Y. Xiang, L. Gao, and E. Hossain, “Data and model poisoning backdoor attacks on wireless federated learning, and
the defense mechanisms: A comprehensive survey,” IEEE Commun. Surveys Tuts., vol. 26, no. 3, pp. 1861–1897, 2024. [86] M. Bilal, J. Crowcroft, R. Wang, X. Xu, and S. Dustdar, “Large language models for agentic NetOps and AIOps: Architectures, evaluation, and safety,” arXiv preprint arXiv:2605.12729, 2026. [Online]. Available: https://arxiv.org/abs/2605.12729v1 [87] J. Sengendo and F. Granelli, “Toward self-optimizing 6G networks through network digital twin intelligence: A real network traffic evaluation,” IEEE Netw. Lett., vol. 8, no. 2, pp. 174–178, 2026. [88] M. Tavallaee, E. Bagheri, W. Lu, and A. A. Ghorbani, “A detailed analysis of the KDD CUP 99 data set,” in Proc. IEEE Symp. Comput. Intell. Security Def. Appl. (CISDA), Ottawa, ON, Canada, 2009, pp. 1–6. [89] I. Sharafaldin, A. H. Lashkari, and A. A. Ghorbani, “Toward generating a new intrusion detection dataset and intrusion traffic characterization,” in Proc. 4th Int. Conf. Inf. Syst. Secur. Privacy (ICISSP), Funchal, Madeira, Portugal, Jan. 2018, pp. 108–116. [90] T. M. Booij, I. Chiscop, E. Meeuwissen, N. Moustafa, and F. T. H. d. Hartog, “ToN IoT: The role of heterogeneity and the need for standardization of features and attack types in IoT network intrusion data sets,” IEEE Internet Things J., vol. 9, no. 1, pp. 485–496, Jan. 2022. [91] Y. Siriwardhana et al., “Descriptor: 5G wireless network intrusion detection dataset (5G-NIDD),” IEEE Data Descr., vol. 2, pp. 358–369, 2025. [92] A. Liatifis et al., “NANCY SNS JU project - cyberattacks on O-RAN 5G testbed dataset,” 2024, accessed: Apr. 2026. [Online]. Available: https://zenodo.org/records/13863735 [93] I. J. Goodfellow et al., “Generative adversarial nets,” in Proc. 27th Conf. Neural Inf. Process. Syst. (NIPS). Montréal, QC, Canada: Curran Associates, Inc., 2014, pp. 2672–2680. [94] A. Kotelnikov, D. Baranchuk, I. Rubachev, and A. Babenko, “TabDDPM: Modelling tabular data with diffusion models,” in Proc. 40th Int. Conf. Mach. Learn. (ICML), 2023, pp. 17 564–17 579. [Online]. Available: https://proceedings.mlr.press/v202/kotelnikov23a/ kotelnikov23a.pdf [95] MITRE Corporation, “MITRE ATT&CK® : Adversarial tactics, techniques, and common knowledge,” https://attack.mitre.org, 2026, accessed: May 2026. [96] MITRE Corporation, “MITRE ATLAS: Adversarial threat landscape for artificial-intelligence systems,” https://atlas.mitre.org, 2026, accessed: May 2026. [97] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal loss for dense object detection,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), Venice, Italy, Oct. 2017, pp. 2980–2988. [98] A. L. Buczak and E. Guven, “A survey of data mining and machine learning methods for cyber security intrusion detection,” IEEE Commun. Surveys Tuts., vol. 18, no. 2, pp. 1153–1176, 2016. [99] N. Shone, T. N. Ngoc, V. D. Phai, and Q. Shi, “A deep learning approach to network intrusion detection,” IEEE Trans. Emerg. Top. Comput. Intell., vol. 2, no. 1, pp. 41–50, 2018. [100] Z. Ahmad, A. Shahid Khan, C. Wai Shiang, J. Abdullah, and F. Ahmad, “Network intrusion detection system: A systematic study of machine learning and deep learning approaches,” Trans. Emerg. Telecommun. Technol., vol. 32, no. 1, 2021. [101] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV, USA, 2016, pp. 770–778. [102] G. Aceto, D. Ciuonzo, A. Montieri, V. Persico, and A. Pescapé, “Aipowered internet traffic classification: Past, present, and future,” IEEE Commun. Mag., vol. 62, no. 9, pp. 168–175, 2024. [103] M. Roopak, G. Yun Tian, and J. Chambers, “Deep learning models for cyber security in IoT networks,” in Proc. IEEE 9th Annu. Comput. Commun. Workshop Conf. (CCWC), Jan. 2019, pp. 452–457. [104] M. S. Elsayed, N.-A. Le-Khac, S. Dev, and A. D. Jurcut, “DDoSNet: A deep-learning model for detecting network attacks,” in Proc. IEEE 21st Int. Symp. World Wireless, Mobile Multimedia Netw. (WoWMoM), Aug. 2020, pp. 391–396. [105] N. Noonari, D. Corujo, R. L. Aguiar, and F. J. Ferrão, “Multi-scale convolutional LSTM with transfer learning for anomaly detection in cellular networks,” in Proc. INForum 2024 – Simpósio de Informática, Aveiro, Portugal, Sep. 2024. [Online]. Available: https: //2024.inforum.pt/static/files/papers/INForum 2024 paper 18.pdf [106] X. Yang, Z. Song, I. King, and Z. Xu, “A survey on deep semisupervised learning,” IEEE Trans. Knowl. Data Eng., vol. 35, no. 9, pp. 8934–8954, Sep. 2023.
30
[107] V. Rey, P. M. S. Sánchez, A. H. Celdrán, and G. Bovet, “Federated learning for malware detection in IoT devices,” Comput. Netw., vol. 204, no. 108693, Feb. 2022. [108] Y. Mirsky, T. Doitshman, Y. Elovici, and A. Shabtai, “Kitsune: An ensemble of autoencoders for online network intrusion detection,” in Proc. Netw. Distrib. Syst. Security Symp. (NDSS), Feb. 2018. [109] L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, “Modeling tabular data using conditional GAN,” in Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 2019, pp. 7333–7343. [Online]. Available: https://proceedings.neurips.cc/paper/ 2019/hash/254ed7d2de3b23ab10936522dd547b78-Abstract.html [110] M. Rahmani Ghourtani, S. Norouzi, J. Chen, H. Ahmadi, T. Braun, K. Chowdhury, and A. Burr, “Hybrid GNN-centric architectures for AInative 6G wireless networks: A comprehensive survey,” IEEE Commun. Surveys Tuts., vol. 28, pp. 5678–5712, 2026. [111] W. W. Lo, S. Layeghy, M. Sarhan, M. Gallagher, and M. Portmann, “EGraphSAGE: A graph neural network based intrusion detection system for IoT,” in Proc. IEEE/IFIP Netw. Oper. Manage. Symp. (NOMS), Budapest, Hungary, 2022, pp. 1–9. [112] Y. Kim and P. Thulasiraman, “Anomaly detection in 5G networks using transformer-based autoencoder,” in Proc. IEEE Int. Conf. Res. Adapt. Converg. Syst. (RASSE), 2024, pp. 1–8. [113] J. Liu, M. Simsek, M. Nogueira, and B. Kantarci, “Multidomain transformer-based deep learning for early detection of network intrusion,” in Proc. IEEE Global Commun. Conf. (GLOBECOM), Kuala Lumpur, Malaysia, 2023, pp. 6056–6061. [114] M. Umer, M. Tahir, M. Sardaraz, M. Sharif, H. Elmannai, and A. D. Algarni, “Network intrusion detection model using wrapper based feature selection and multi head attention transformers,” Sci. Rep., vol. 15, no. 28718, Aug. 2025. [115] F. Adjewa, M. Esseghir, and L. Merghem-Boulahia, “Efficient federated intrusion detection in 5G ecosystem using optimized BERT-based model,” in Proc. IEEE Int. Conf. Wireless Mobile Comput., Netw. Commun. (WiMob), Oct. 2024, pp. 62–67. [116] P. R. B. Houssel, P. Singh, S. Layeghy, and M. Portmann, “Towards explainable network intrusion detection using large language models,” in Proc. IEEE Int. Conf. Big Data Cloud Appl. Technol. (BDCAT), Sharjah, UAE, 2024, pp. 67–72. [117] R. A. Bakar et al., “5GDAD: A deep learning approach for DDoS attack detection in 5G P4-based UPF,” in Proc. IEEE Int. Conf. High Perform. Switch. Routing (HPSR), Pisa, Italy, 2024, pp. 185–190. [118] L. Yang, S. Naser, A. Shami, S. Muhaidat, L. Ong, and M. Debbah, “Toward zero touch networks: Cross-layer automated security solutions for 6G wireless networks,” IEEE Trans. Commun., vol. 73, no. 9, pp. 7650–7679, 2025. [119] G. C. Amaizu, C. I. Nwakanma, S. Bhardwaj, J. M. Lee, and D.-S. Kim, “Composite and efficient DDoS attack detection framework for B5G networks,” Comput. Netw., vol. 188, no. 107871, Apr. 2021. [120] W. Li, W. Meng, and L. F. Kwok, “Surveying trust-based collaborative intrusion detection: State-of-the-art, challenges and future directions,” IEEE Commun. Surveys Tuts., vol. 24, no. 1, pp. 280–305, 2022. [121] Q. Yan, F. R. Yu, Q. Gong, and J. Li, “Software-defined networking (SDN) and distributed denial of service (DDoS) attacks in cloud computing environments: A survey, some research issues, and challenges,” IEEE Commun. Surveys Tuts., vol. 18, no. 1, pp. 602–622, 2016. [122] A. E. Cil, K. Yildiz, and A. Buldu, “Detection of DDoS attacks with feed forward based deep neural network model,” Expert Syst. Appl., vol. 169, no. 114520, May 2021. [123] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” in Proc. Mach. Learn. Syst. (MLSys), Austin, TX, USA, 2020. [124] Z. Chen, N. Lv, P. Liu, Y. Fang, K. Chen, and W. Pan, “Intrusion detection for wireless edge networks based on federated learning,” IEEE Access, vol. 8, pp. 217 463–217 472, 2020. [125] T. D. Nguyen, S. Marchal, M. Miettinen, H. Fereidooni, N. Asokan, and A.-R. Sadeghi, “DÏot: A federated self-learning anomaly detection system for IoT,” in Proc. IEEE 39th Int. Conf. Distrib. Comput. Syst. (ICDCS), Dallas, TX, USA, 2019, pp. 756–767. [126] M. Sheibani, S. Konur, I. Awan, and A. Qureshi, “A multi-layered defence strategy against DDoS attacks in SDN/NFV-based 5G mobile networks,” Electronics, vol. 13, no. 8, Apr. 2024. [127] S. K. Fayaz, Y. Tobioka, V. Sekar, and M. Bailey, “Bohatei: Flexible and elastic DDoS defense,” in Proc. 24th USENIX Security Symp. Washington, DC, USA: USENIX Association, Aug. 2015, pp. 817–832. [Online]. Available: https://www.usenix.org/conference/ usenixsecurity15/technical-sessions/presentation/fayaz
[128] S. Scott-Hayward, S. Natarajan, and S. Sezer, “A survey of security in software defined networks,” IEEE Commun. Surveys Tuts., vol. 18, no. 1, pp. 623–654, 2016. [129] A. M. Zarca, M. Bagaa, J. B. Bernabe, T. Taleb, and A. F. Skarmeta, “Semantic-aware security orchestration in SDN/NFV-enabled IoT systems,” Sensors, vol. 20, no. 13, Jun. 2020. [130] J. O. Kephart and D. M. Chess, “The vision of autonomic computing,” Computer, vol. 36, no. 1, pp. 41–50, 2003. [131] P. Kairouz et al., “Advances and open problems in federated learning,” Found. Trends Mach. Learn., vol. 14, no. 1–2, pp. 1–210, Jun. 2021. [132] J. So, B. Güler, and A. S. Avestimehr, “Byzantine-resilient secure federated learning,” IEEE J. Sel. Areas Commun., vol. 39, no. 7, pp. 2168–2181, Jul. 2021. [133] K. Bonawitz et al., “Towards federated learning at scale: System design,” in Proc. Mach. Learn. Syst. (MLSys), A. Talwalkar, V. Smith, and M. Zaharia, Eds., vol. 1, Stanford, CA, USA, 2019, pp. 374–388. [Online]. Available: https://proceedings.mlsys.org/paper files/paper/2019/file/7b770da633baf74895be22a8807f1a8f-Paper.pdf [134] K. Wei et al., “Federated learning with differential privacy: Algorithms and performance analysis,” IEEE Trans. Inf. Forensics Security, vol. 15, pp. 3454–3469, 2020. [135] C. Zhang, H. Zhang, S. Dang, B. Shihada, and M.-S. Alouini, “Gradient compression and correlation driven federated learning for wireless traffic prediction,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 4, pp. 2246–2258, Aug. 2025. [136] Y. Zhao, M. Li, L. Lai, N. Suda, D. Civin, and V. Chandra, “Federated learning with non-IID data,” arXiv:1806.00582 [cs.LG], Jun. 2018. [Online]. Available: https://arxiv.org/abs/1806.00582 [137] National Institute of Standards and Technology, “Module-lattice-based digital signature standard,” U.S. Department of Commerce, Federal Information Processing Standards Publication NIST FIPS 204, Aug. 2024. [Online]. Available: https://csrc.nist.gov/pubs/fips/204/final [138] M. A. Ferrag, A. Lakas, and M. Debbah, “6G-Bench: An open benchmark for semantic communication and network-level reasoning with foundation models in AI-native 6G networks,” IEEE Open J. Commun. Soc., vol. 7, pp. 3305–3330, 2026. [139] S. Rose, O. Borchert, S. Mitchell, and S. Connelly, “Zero trust architecture,” National Institute of Standards and Technology (NIST), Tech. Rep. NIST Special Publication 800-207, Aug. 2020. [140] National Institute of Standards and Technology, “Module-latticebased key-encapsulation mechanism standard,” U.S. Department of Commerce, Federal Information Processing Standards Publication NIST FIPS 203, Aug. 2024. [Online]. Available: https://csrc.nist.gov/ pubs/fips/203/final [141] S. Milli, L. Schmidt, A. D. Dragan, and M. Hardt, “Model reconstruction from model explanations,” in Proc. ACM Conf. on Fairness, Accountability, and Transparency (FAT*), 2019, pp. 1–9. [142] D. Slack, S. Hilgard, E. Jia, S. Singh, and H. Lakkaraju, “Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods,” in Proc. AAAI/ACM Conf. AI, Ethics, Soc. (AIES), Feb. 2020, pp. 180– 186. [143] C. Zhang, X. Costa-Pérez, and P. Patras, “Adversarial attacks against deep learning-based network intrusion detection systems and defense mechanisms,” IEEE/ACM Trans. Netw., vol. 30, no. 3, pp. 1188–1204, 2022. [144] V. J. Reddi et al., “MLPerf Inference Benchmark,” in Proc. ACM/IEEE 47th Annu. Int. Symp. Comput. Archit. (ISCA), Valencia, Spain, 2020, pp. 446–459. [145] S. Sajadieh et al., “The AI Index 2026 Annual Report,” Inst. for Human-Centered AI, Stanford Univ., Stanford, CA, USA, Tech. Rep., Apr. 2026. [Online]. Available: https://aiindex.stanford.edu/report/