Conceptio › Archive › arXiv CS
arXiv CSopen access

EdgeDetect: Importance-Aware Gradient Compression with Homomorphic Aggregation for Federated Intrusion Detection

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

1

EdgeDetect: Importance-Aware Gradient Compression with Homomorphic Aggregation for Federated Intrusion Detection

arXiv:2604.14663v1 [cs.CR] 16 Apr 2026

Noor Islam S. Mohammad

Abstract—Federated learning (FL) enables collaborative intrusion detection without raw data exchange, but conventional FL incurs high communication overhead from full-precision gradient transmission and remains vulnerable to gradient inference attacks. This paper presents EdgeDetect, a communication-efficient and privacy-aware federated IDS for bandwidth-constrained 6GIoT environments. EdgeDetect introduces gradient smartification, a median-based statistical binarization that compresses local updates to {+1, −1} representations, reducing uplink payload by 32× while preserving convergence. We further integrate Paillier homomorphic encryption over binarized gradients, protecting against honest-but-curious servers without exposing individual updates. Experiments on CIC-IDS2017 (2.8M flows, 7 attack classes) demonstrate 98.0% multi-class accuracy and 97.9% macro F1-score, matching centralized baselines, while reducing per-round communication from 450 MB to 14 MB (96.9% reduction). Raspberry Pi 4 deployment confirms edge feasibility: 4.2 MB memory, 0.8 ms latency, and 12 mJ per inference with < 0.5% accuracy loss. Under 5% poisoning attacks and severe imbalance, EdgeDetect maintains 87% accuracy and 0.95 minority class F1 (p < 0.001), establishing a practical accuracy–communication–privacy tradeoff for next-generation edge intrusion detection. Index Terms—PPFL, IDS, Edge Computing, 6G Security, IoT Networks, Communication Efficiency, and Machine Learning.

leakage, shared model updates may be reverse-engineered to reconstruct sensitive training samples [7]. To address these challenges, we propose EdgeDetect, a scalable and privacy-aware federated IDS tailored for resourceconstrained 6G-IoT environments [8]. EdgeDetect introduces a novel gradient smartification mechanism that transforms continuous gradient updates into lightweight binarized representations ({+1, −1}) using median-based statistical thresholding. This adaptive, distribution-aware compression reduces uplink payload size by up to 32× while preserving empirical convergence behavior [9], [10]. Unlike fixed-threshold methods (e.g., signSGD), our approach suppresses low-magnitude gradient components below the per-client median, reducing stochastic noise and improving stability under heterogeneous data distributions. We further integrate Paillier homomorphic encryption over the binarized gradients, ensuring that only aggregated model updates are visible to the central server, providing strong cryptographic protection against gradient inversion and honest-but-curious adversaries [10]. The joint optimization of compression and privacy enables EdgeDetect to achieve both communication efficiency and end-to-end confidentiality without compromising detection accuracy.

I. I NTRODUCTION

N

EXT-GENERATION Wireless technologies 5G, 6G, and IoT enable massive machine-type communications and ultra-reliable low-latency services [1], [2] while simultaneously expanding the attack surface for sophisticated cyber threats. As billions of heterogeneous edge devices generate high-volume traffic in smart cities, autonomous vehicles, and Industry 4.0, traditional centralized IDS face fundamental limitations, including scalability bottlenecks, communication latency, single points of failure, and difficulty handling high dimensionality and severe class imbalance in modern network traffic. Machine learning has become central to automated threat identification [3], [4]. However, centralized deployment of deep learning architectures requires aggregating raw sensor readings, enterprise logs, and user data at cloud servers, exposing systems to potential data breaches and regulatory violations [5]. Federated Learning (FL) addresses this limitation by enabling collaborative model training while preserving data locality [6]. Despite its advantages, practical FL implementations face two critical challenges: (1) communication overhead, transmitting high-dimensional gradient vectors from thousands of edge clients consumes excessive bandwidth; and (2) gradient Department of Computer Science, Istanbul Technical University, Maslak, TR (Corresponding author: [email protected]). This research received no external funding.

A. Contribution This work makes the following key contributions: • Alignment-Aware Federated IDS Architecture: We present EdgeDetect, a privacy-preserving federated intrusion detection framework designed for 6G-IoT environments [11]. The architecture integrates PCA-based dimensionality reduction, imbalance-aware sampling, and secure aggregation within a unified decentralized pipeline, enabling collaborative learning without sharing raw network traffic while maintaining scalability and robustness. • Adaptive Median-Based Gradient Smartification with Encrypted Aggregation: We introduce a statistically adaptive median-threshold binarization strategy that compresses gradients into {+1, −1} while preserving directional alignment under heterogeneous and heavy-tailed client distributions. In contrast to fixed zero-threshold signSGD [12], the proposed per-client adaptive rule improves convergence stability. Combined with Paillier homomorphic encryption applied directly to binarized gradients, the method achieves up to 32× communication reduction while mitigating gradient inversion risks [13], [14]. • Quantified Privacy–Utility–Efficiency Trade-off: Extensive ablation and adversarial analyses demonstrate 98.0% multi-class accuracy with 96.9% communication reduction

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

2

on CIC-IDS2017 (2.8M flows), achieving performance D. Distinction from signSGD and Quantized FL comparable to centralized baselines while providing crypUnlike fixed-threshold quantizers such as QSGD [25] or tographic privacy guarantees. The framework maintains TernGrad [32], which apply uniform quantization levels, our > 85% accuracy with 20% malicious clients and reduces median-threshold binarization adapts to the per-client gradient inversion PSNR from 31.7 dB to 15.1 dB [1]. distribution. This property is especially valuable for IDS data, • Edge-Validated Deployment: Real-world deployment on where gradients exhibit heavy tails due to rare attack events. Raspberry Pi 4 devices confirms practical feasibility, requiring only 4.2 MB memory, 0.8 ms latency, and 12 mJ per inference, with less than 0.5% accuracy degradation. τt = median(gt ) (1) These results validate suitability for resource-constrained 6G-IoT edge environments. Thus, smartification preserves relative ordering information within each gradient vector and adapts to heavy-tailed feature distributions typical in IDS models. Distribution-adaptive II. R ELATED W ORK quantization with provable descent guarantees and entropyThe evolution of IDS from signature-based systems to aware privacy strengthening. ML and deep learning paradigms has significantly advanced network security [15], [16]. This section reviews anomaly detection in wireless networks, FL for decentralized security, III. S YSTEM A RCHITECTURE (IDS) and privacy–communication efficiency challenges. A. Deep Learning-Based Anomaly Detection Deep learning has become the standard for detecting complex attack patterns in high-dimensional network traffic. Classical algorithms such as SVMs and random forests remain competitive for structured features [17], [18], while CNN–RNN and LSTM architectures capture temporal dependencies for DDoS and zero-day detection [11]. Image-based encodings of time-series traffic further enhance spatial feature extraction [19]. However, these centralized approaches require large-scale data aggregation, introducing privacy risks and system-level vulnerabilities.

We propose EdgeDetect, a privacy-preserving federated learning architecture for 6G-enabled IoT environments [31], [33], [34]. The system comprises K resource-constrained edge clients and a central aggregation server collaboratively training a global anomaly detection model Mglobal without exposing private local datasets Di . Each communication round consists of four phases: 1) Phase 1: Client-Side Local Training: Let’s S = {1, 2, ..., K} denote the participating clients. At round r, the server broadcasts global parameters W (r) [35], [36]. Each client performs E local epochs minimizing L(Wi , Di ): (r+1)

Wi

(r)

= Wi

(r)

− η∇L(Wi , Di )

(2)

B. Federated Learning in IoT Networks The model update is FL enables decentralized training without sharing raw data [20]. Applications include IoT security, industrial sensor net(r) (r+1) works, and cross-domain intrusion detection [21]. Edge–cloud ∆i = Wi − W (r) . (3) collaborative architectures reduce response latency while preserving data locality [22], [23]. However, standard FL Phase 2: Gradient Smartification. To reduce uplink algorithms such as FedAvg rely on full-precision gradient communication cost, we apply a statistical binarization operator exchange, creating communication bottlenecks in bandwidth- Φ(·): ( limited 6G IoT systems [24], [25]. (r) +1, if ∆i,j ≥ θi bin ∆i,j = −1, otherwise C. Privacy Preservation and Gradient Compression While FL mitigates raw data exposure, it remains vulnerable where θ = median(|∆(r) |) is the median of the absolute i i to gradient inference attacks. Differential Privacy (DP) and values of the local gradient vector. The resulting vector Homomorphic Encryption (HE) improve confidentiality but ∆bin ∈ {+1, −1}d compresses the representation by 32× while i may introduce accuracy or computational overhead [26], preserving directional information. [27]. Communication-efficient methods such as signSGD and 2) Phase 3: Privacy-Preserving Encryption: Each client gradient sparsification reduce bandwidth requirements [12], encrypts ∆bin using a homomorphic encryption scheme E(·) i [28]. [37], [38]: However, few approaches jointly optimize gradient compression and encrypted aggregation in resource-constrained intru(r) Ci = E(∆bin (4) sion detection settings. Our PoL-based gradient smartification i ) mechanism integrates statistical binarization with encrypted aggregation to address both communication efficiency and ensuring that individual updates remain confidential during privacy preservation [29]–[31]. transmission.

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

3) Phase 4: Secure Aggregation and Global Update: Upon receiving ciphertexts from active clients Sr ⊆ S, the server performs encrypted aggregation [9], [39]: ∆bin agg =

1 X (r) D(Ci ) |Sr |

(5)

i∈Sr

The global model is updated as W (r+1) = W (r) + α · ∆bin agg .

3

C. Dimensionality Reduction via Incremental PCA To mitigate multicollinearity (23% of feature pairs with |ρ| > 0.8) and reduce computational overhead, incremental PCA was applied to the standardized feature matrix Z ∈ Rn×d (d = 78) [44], [45]: Cov(Z) =

1 Z ⊤ Z = V ΛV ⊤ . n−1

(10)

(6) The reduced representation was obtained as

IV. M ETHODOLOGY ZPCA = ZVk ,

A. Data Exploration and Preprocessing The CIC-IDS2017 dataset contains 2,830,743 records with 79 features. Exploratory data analysis revealed: (i) 308,381 duplicate rows, removed to mitigate potential overfitting bias; (ii) missing and infinite values in Flow Bytes/s and Flow Packets/s (0.06%), imputed using median statistics to preserve distributional robustness; (iii) high memory consumption (≈ 1.5 GB), mitigated via numerical downcasting (float64 to → float32, int64 to → int32), achieving 47.5% memory reduction; and (iv) severe class imbalance with benign traffic dominating attack categories. To ensure computational feasibility, a 20% stratified sample was extracted. Statistical validation confirmed representativeness, with feature mean deviations below 5% relative to the full dataset.

Temporal features (e.g., flow inter-arrival statistics) capture the bursty nature of volumetric attacks, while entropy-based features quantify the randomness in packet sizes, which often deviates during scanning or exfiltration attempts [40], [41]. Temporal Features: Flow inter-arrival time statistics were computed as

Pk

i=1 λi

Pd

i=1 λi

≥ 0.993.

(12)

This preserves 99.3% of explained variance while reducing feature dimensionality by 55%.

D. Class Balancing Strategies Binary Classification: Random under-sampling balances benign and attack samples [46]: (13)

yielding 15,000 balanced instances. Multi-Class Classification: SMOTE generates synthetic minority samples [47], [48]: xnew = xi + λ(xij − xi ),

λ ∼ U (0, 1).

(14)

Adaptive SMOTE: Density-aware interpolation [49]:

∆tmean =

(7)

Entropy-Based Features: Packet size entropy captures distributional randomness [32], [41]: X H(S) = − p(s) log2 p(s), (8) s∈S

where S denotes unique packet sizes and p(s) their empirical probabilities. Feature Selection: Recursive Feature Elimination (RFE) was applied using Random Forest permutation importance [42], [43]: T  1X  (−j) Ij = I ft (D) ̸= ft (D) , (9) T t=1 (−j)

retaining k = 35 principal components satisfying

Dbal = Dmin ∪ Sample(Dmax , |Dmin |),

B. Feature Engineering and Selection

n 1 X (ti − ti−1 ), n − 1 i=2 v u n u 1 X ∆tstd = t (∆ti − ∆tmean )2 . n − 1 i=2

(11)

where ft denotes a tree t with a feature j permuted. Features were ranked according to Ij and selected prior to dimensionality reduction.

λ ∼ Beta(α, β),

α = 1 + ρi , β = 1 + (1 − ρi ),

(15)

where ρi reflects local minority sparsity.

E. Protocol Flow and Algorithm The protocol follows a privacy-preserving federated optimization pipeline 1. First, the server generates a Paillier homomorphic encryption keypair and broadcasts the public key to all clients. Each client performs local training on its private dataset Di , computes the model update ∆i , and applies median-based binarization to obtain ∆bin i . The binarized gradients are encrypted element-wise and transmitted to the server without revealing raw updates. Using Paillier’s additive homomorphism, the server aggregates encrypted gradients via ciphertext multiplication, decrypts only the summed result, and normalizes it to form the global update. Finally, the updated model W (r+1) is broadcast back to clients, enabling secure and communication-efficient collaborative learning across rounds.

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

Algorithm 1 Secure Binarized Gradient Aggregation 1: INITIALIZATION (One-time setup) 2: Server generates Paillier keypair (pk, sk): 3: Public key pk = (n, g) where n = p·q (2048-bit RSA modulus) 4: Private key sk = (λ, µ) where λ = lcm(p − 1, q − 1) 5: Server broadcasts pk to all K clients 6: Server keeps sk secret 7: 8: LOCAL TRAINING (Each client i ∈ {1, . . . , K}) 9: Train local model on private dataset Di for E epochs 10: Compute gradient update: ∆i = Wnew − Wold 11: Binarize gradient (Equation 3): 12: θi = median(|∆i |) 13: ∆bin i [j] = +1 if ∆i [j] ≥ θi , else −1 14: Encrypt binarized gradient element-wise: ∆bin i [j] · r n mod n2 15: Ci [j] = Encpk (∆bin i [j]) = g

(22)

i∈S

Model Update with Momentum: (r)

M (r+1) = M (r) + η∆global + µ(M (r) − M (r−1) ),

where r ← − Z∗n is random nonce

16: 17: Send ciphertext Ci = {Ci [1], . . . , Ci [d]} to server 18: 19: SECURE AGGREGATION (Server) 20: Receive ciphertexts {C1 , . . . , CK } from active clients 21: Perform homomorphic addition in the encrypted domain:  

PK

bin 22: Cagg [j] = i=1 Ci [j] mod n2 = Encpk i=1 ∆i [j] 23: Decrypt aggregated gradient: λ 24: ∆bin mod n2 )·µ mod n agg [j] = Decsk (Cagg [j]) = L(Cagg [j]

25: where L(x) = (x − 1)/n bin 26: Normalize by client count: ∆bin agg [j] = ∆agg [j]/K 27: 28: GLOBAL UPDATE (Server) 29: Apply aggregated update: W (r+1) = W (r) + α · ∆bin agg 30: Broadcast W (r+1) to all clients for next round

F. Machine Learning Models Logistic Regression (Elastic Net) [50]: P (y = 1|x) = σ(β0 + β ⊤ x),

G. Privacy-Preserving Federated Learning A federated learning framework enables collaborative intrusion detection without raw data exchange [54]. At the communication round r, the client i computes a local update (r) ∆i . Gradient Smartification:   (r) (r) (r) ∆i,bin = sign ∆i − θ , θ = median(∆i ). (21) Secure Aggregation (Paillier): 1 X (r) (r) ∆global = ∆i,bin . |S|

$

QK

4

(16)

(23)

where η = 0.01 and µ = 0.9. Differential Privacy [55]: ∆i ˜i = + N (0, σ 2 C 2 I), ∆ (24) max(1, ∥∆i ∥2 /C) with a clipping threshold C = 0.1 and noise scale σ = 0.01, yielding (ϵ, δ) = (1.0, 10−5 ). The framework achieves 98.7% of centralized accuracy while reducing communication overhead by more than 30×, enabling privacy-aware cross-domain intrusion detection. H. Evaluation Metrics Model performance was evaluated using standard confusionmatrix-based metrics: true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). Discrimination ability was further assessed using the area under the receiver operating characteristic curve (ROC-AUC), defined as AUC = R1 TPR(FPR) d(FPR), which equivalently represents the 0 probability that a randomly chosen positive sample receives a higher score than a randomly chosen negative sample, P (ŷ+ > ŷ− | y+ = 1, y− = 0).

with objective I. Matthews Correlation Coefficient (MCC) 

 1−ρ 2 L = LCE + α ∥β∥2 + ρ∥β∥1 , 2

(17)

where α = 0.01 and ρ = 0.5. SVM (RBF Kernel) [51]: K(xi , xj ) = exp(−γ∥xi − xj ∥2 ),

γ = 0.001,

(18)

optimized via SMO with C = 1.0. Random Forest [52]: An ensemble of T = 100 trees with √ maximum depth 20 and m = ⌊ d⌋ features per split: ŷ(x) = arg max k

TP · TN − FP · FN

MCC = p

T X

I(ht (x) = k).

(19)

t=1

. (TP + FP)(TP + FN)(TN + FP)(TN + FN) (25) Cohen’s Kappa: po − pe κ= . (26) 1 − pe

J. Cross-Validation and Hyperparameter Optimization Stratified K-Fold Cross-Validation k 1X CVscore = Metric(Dival ), k = 5. k i=1 Nested Grid Search:

Gradient Boosting [53]:

∗

Fm (x) = Fm−1 (x) + νhm (x),

(27)

ν = 0.1.

(20)

Neural Network: A multilayer perceptron (MLP) with ReLU activations and dropout (p = 0.5), optimized using Adam (α = 10−3 ). Architecture: 35 → 128 → 64 → K.

θ = arg max

1

θ∈Θ kinner

kX inner

Score(Mθ , Djval ).

(28)

j=1

Early Stopping: Neural and boosting models employed early stopping with patience p = 10 epochs based on validation loss monitoring.

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

V. E XPERIMENTAL S ETUP A. Dataset Construction and Sampling Validation The CIC-IDS2017 dataset contains N = 2,830,540 flow records with 78 features. A stratified 20% subset (n = 504,472) was sampled to reduce computational cost while preserving distributional properties; Kolmogorov–Smirnov tests showed no significant deviations (p > 0.05), with 92% of features exhibiting < 5% mean deviation. After standardization and incremental PCA, dimensionality was reduced to k = 35 (99.3% variance retained). An 80:20 stratified split (seed 42) was applied. Binary classification used a balanced 15,000sample subset (7,500 benign, 7,500 attack), while multi-class detection employed SMOTE-balanced 35,000 samples across seven attack categories (Table I).

5

TABLE II: Feature Analysis: Top-10 Discriminative (PC) Rk

1 2 3 4 5 6 7 8 9 10

Config 1 (C = 0.1, saga)

Config 2 (C = 100, sag)

PC

Coeff.

∆ (%)

PC

Coeff.

Var. Exp.

PC04 PC26 PC24 PC31 PC06 PC23 PC27 PC35 PC29 PC21

−1.607 −1.372 +1.335 +1.102 +0.958 +0.864 −0.830 −0.752 −0.740 −0.539

+27.5 +17.7 +24.5 +35.0 +26.2 +5.4 +36.8 +42.6 +29.5 —

PC04 PC26 PC24 PC31 PC06 PC23 PC27 PC35 PC29 PC14

−2.050 −1.614 +1.662 +1.488 +1.209 +0.911 −1.135 −1.072 −0.958 +0.723

8.2% 1.4% 1.8% 0.9% 5.3% 2.1% 1.2% 0.4% 0.7% 3.6%

Notes: Components ranked by |β|. Positive coefficients indicate attack correlation; negative values indicate benign traffic. ∆ (%) denotes relative coefficient amplification from Config 1 to Config 2. Var. Exp. is the PCA variance contribution of each component.

TABLE I: Distribution and Sampling: CIC-IDS2017 Dataset Sampled

∆ (%)

KS p

F LOW T EMPORAL DYNAMICS (n = 6) Flow Duration (µs) 1.66 × 107 Flow IAT Mean (µs) 1.45 × 106 Flow IAT Std (µs) 3.28 × 106 Fwd IAT Total (µs) 1.62 × 107

1.65 × 107 1.43 × 106 3.26 × 106 1.62 × 107

0.47 0.72 0.56 0.46

0.82 0.76 0.79 0.84

F LOW R ATE M ETRICS (n = 2) Flow Bytes/s Flow Packets/s

1.38 × 106 4.70 × 104

2.42 0.64

0.68 0.81

PACKET VOLUME S TATISTICS (n = 8) Total Fwd Packets 10.28 12.08 Total Bwd Packets 11.57 13.91 Total Fwd Length (B) 611.58 601.84 Total Bwd Length (B) 1.81 × 104 2.44 × 104

17.53 20.30 1.59 34.70∗

0.42 0.38 0.73 0.21

PACKET S IZE D ISTRIBUTIONS (n = 8) Fwd Pkt Max (B) 231.09 Fwd Pkt Mean (B) 63.47 Bwd Pkt Max (B) 974.37 Bwd Pkt Mean (B) 340.41

0.47 0.50 0.23 0.16

0.88 0.85 0.91 0.93

C ONNECTION & I DLE F EATURES (n = 5) Destination Port 8704.76 8686.62 Idle Max (µs) 9.76 × 106 9.71 × 106 5 Idle Mean (µs) 5.65 × 10 5.64 × 105

0.21 0.51 0.20

0.94 0.77 0.89

C LASS BALANCE Attack Prevalence

0.63

0.99

Feature Group

Original

1.41 × 106 4.73 × 104

230.01 63.15 972.14 339.88

0.73

Aggregate Statistics (All 78 Features) Mean Absolute Deviation Median Deviation Features with ∆ < 5% KS Test Rejections (α = 0.05)

0.73 — — — —

3.42 — 0.64 — 72/78 (92.3%) 0/78 (0%)

Notes: Original dataset: N = 2,830,540; sampled: n = 504,472 (stratified 20%). ∆ denotes absolute percentage deviation in feature means. KS p is the Kolmogorov-Smirnov test p-value (null hypothesis: identical distributions). ∗ Higher deviation in backward packet totals reflects temporal clustering of DDoS events; distributional shape remains preserved (KS p = 0.21 > 0.05). Units: µs = microseconds, B = bytes.

B. Hyperparameter and Model Complexity Feature heterogeneity (e.g., ∼ 103 unique ports vs. > 106 flow-metric values) necessitated normalization prior to PCA. Although certain packet totals exhibited larger mean deviations (18–35%), Kolmogorov–Smirnov tests indicated stochastic variation rather than sampling bias, supporting dataset validity. Two configurations guided model selection: Config. 1 prioritized computational efficiency for constrained deployment, whereas Config. 2 maximized accuracy via 3-fold grid search under constraints of algorithmic parity, deployment feasibility, and statistical robustness (Table II). The per-round encryption complexity scales as O(d log n) under a 2048-bit modulus. For d = 35, the resulting overhead remains below one second.

For logistic regression, reduced regularization (C : 1.0 → 0.5) increased the ℓ2 -norm (4.127 → 5.243), strengthening discriminative PCA weights and shifting the intercept (−2.351 → −2.962). Replacing linear SVM (83.0%) with RBF (γ = 0.1) enabled nonlinear separation in the 35-dimensional space. In Random Forest, controlled depth (max depth=20) limited overfitting, while n = 200 trees reduced variance via bagging. Decision tree regularization (min split=5) improved generalization, and reducing KNN neighborhood size (5 → 3) enhanced locality-based discrimination (Table III). TABLE III: Model Hyperparameters and Learned Parameters Alg.

Config.

Hyperparameters

Learned Parameters

LR

Model 1 C = 1.0, ℓ2 , lbfgs; w0 = −2.351, ∥w∥2 = max iter=100 4.127 Model 2 C = 0.5, ℓ2 , lbfgs; w0 = −2.962, ∥w∥2 = 5.243 max iter=100

Model 1 linear, C = 1.0; tol=10−3 b = −0.870, nsupport = 4,237 SVM Model 2 RBF, C = 10.0, γ = 0.1; b = −0.420, nsupport = tol=10−3 2,891

RF

DT

Model 1 n = 100, depth=None; Depth: 28.4, Nodes: 284,320 bootstrap Model 2 n = 200, depth=20; boot- Depth: 20.0, Nodes: 523,600 strap Model 1 depth=None; split=2; gini Depth: 32, Leaves: 1,847 Model 2 depth=15; split=5; gini Depth: 15, Leaves: 892

Model 1 k = 5, Euclidean, uniform — (non-parametric) Model 2 k = 3, Euclidean, uniform — (non-parametric) KNN Notes: w0 = intercept; ∥w∥2 = weight norm; b = SVM bias; nsupport = support vectors. Hyperparameters via 3-fold grid search. RF node count = mean nodes × estimators. KNN stores all training instances.

C. Learned Model Complexity Table IV summarizes the evaluated configurations and model complexity. Logistic regression uses stronger regularization in Config. 1 (C = 0.1) and relaxed regularization in Config. 2 (C = 100); saga supports ℓ1 penalties, while sag accelerates convergence for large C. SVM transitions from a polynomial to an RBF kernel to capture nonlinear boundaries, with fewer support vectors indicating tighter margins. Random Forest Config. 2 increases ensemble size

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

and depth, with max_features=20 improving decorrelation and generalization. Decision Tree Config. 2 deepens partitions while min_impurity_decrease regularizes splits. KNN Config. 2 applies distance weighting and k = 7 to reduce variance while maintaining locality in PCA space. TABLE IV: Learned Model Complexity Algorithm

Config. 1 (Efficiency)

Config. 2 (Expressiveness)

Linear Models Logistic Reg. C = 0.1, saga, ℓ2 , iter=100 C = 100, sag, ℓ2 , iter=100 ∥w∥2 = 2.14 ∥w∥2 = 5.24, w0 = −2.96 Kernel-Based Methods poly, deg=3, C = 1, tol=10−3rbf, C = 1, γ = 0.1, tol=10−3 SVM nSV = 5,124 nSV = 2,891, b = −0.42 Tree-Based Ensemble Random Forest n = 10, depth=6, bootstrap n = 15, depth=8, feat=20 ≈640 nodes ≈3,840 nodes, OOB=0.978 Single Decision Tree Decision Tree depth=6, gini depth=10, gini, imp=10−4 63 leaves 247 leaves Instance-Based Learning KNN k = 5, uniform, Euclid k = 7, distance, Euclid 12k stored 12k stored Notes: C controls regularization; γ is RBF bandwidth; nSV denotes support vectors; OOB = out-of-bag estimate. Configuration 2 was selected via a 3-fold grid search on training data. Random Forest node count approximated as mean nodes per tree ×n. KNN is non-parametric and stores all training samples.

6

Stability and efficiency: Computational efficiency (training and inference time) was measured on a standardized platform (Intel i7-9700K, 32GB RAM, single-threaded execution). Total training time includes hyperparameter search, cross-validation, and final model fitting. VI. E XPERIMENTAL R ESULTS A. Linear and Kernel-Based Models Logistic regression provided a stable linear baseline, achieving 92.21% accuracy (std = 5.81 × 10−3 ). Reducing regularization (C = 0.5) yielded a marginal improvement to 92.51% (+0.30%) with similarly low variance, indicating stable convergence despite partial linear inseparability in the PCA space. SVM exhibited the largest configuration sensitivity. The linear kernel underperformed (83.00%, std = 37.27 × 10−3 ), confirming inadequate linear separation. Replacing it with an RBF kernel (C = 10, γ = 0.1) increased accuracy to 96.14% (+13.14%) while reducing variance to 3.89 × 10−3 , validating the presence of non-linear decision boundaries in intrusion patterns. B. Tree-Based Ensemble Methods

Random Forest achieved the highest overall performance. The baseline model (100 trees, unlimited depth) reached 95.98%, while structured tuning (200 trees, max depth=20) D. Evaluation Protocol and Statistical Validation improved accuracy to 98.09% (+2.11%) and halved variance (3.45 × 10−3 → 1.72 × 10−3 ). Depth restriction mitigated overE. Cross-Validation (Stage 1) Model assessment followed a two-stage protocol to ensure fitting, and ensemble expansion reduced prediction variance via generalizability and statistical reliability. Performance was first bagging. Single Decision Trees showed moderate performance evaluated using 5-fold stratified cross-validation on the training (94.89%) with higher variance due to unconstrained growth. partition (n = 12,000, i.e., 80% of the balanced binary dataset). Imposing max depth=15 and min split=5 increased accuracy Stratification preserved the 50:50 benign-to-attack ratio in to 97.24% (+2.35%), demonstrating the necessity of structural each fold. Fold-to-fold variability was quantified via standard regularization in non-ensemble trees. deviation: v C. Instance-Based Learning u K u 1 X t ¯ 2 , K = 5. K-Nearest Neighbors exhibited strong performance with exσCV = (Acci − Acc) (29) K − 1 i=1 ceptional stability. Model 1 (k = 5) achieved 97.40% accuracy −3 Low variance (σ < 0.01) indicates stable performance across with the lowest variance among all models (std = 0.89×10 ), indicating consistent neighborhood-based predictions across dipartitions, which is essential for production deployment. verse fold partitions. Reducing the neighborhood size to k = 3 Model 2 yielded a marginal +0.53% improvement to 97.93%, F. Hold-Out Testing (Stage 2) The best configuration for each algorithm was retrained suggesting that tighter locality constraints better capture attackspace. The on the full training set and evaluated on a held-out test set specific patterns in the 35-dimensional embedding −3 negligible increase in variance (0.89 → 1.27×10 ) confirms (n = 3,000, 20%). We report accuracy, precision, recall, F1KNN’s robustness to hyperparameter perturbations. score, ROC-AUC, and confusion matrices to capture both global correctness and error structure. For binary classification, ROC VII. C OMPARISON WITH SIGN SGD M ETHODS and Precision-Recall (PR) curves were additionally analyzed to Unlike zero-threshold in V signSGD [12] or stochastic quansupport operating point selection under deployment constraints tization methods (QSGD [25], TernGrad [32]), EdgeDetect’s such as controlling false positives. gradient smartification integrates two key innovations: medianG. Statistical Reliability based adaptive thresholding that adjusts to per-client gradient To mitigate random initialization effects, experiments were distributions and homomorphically encrypted aggregation that repeated with three independent random seeds (42, 123, 456) provides end-to-end confidentiality. As summarized in Table V, existing methods lack adaptivity and privacy integration, and reported with 95% confidence intervals: limitations that critically undermine performance under the σ CI95% = x̄ ± 1.96 · √ , n = 3. (30) heavy-tailed gradient distributions characteristic of IDS data. n

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

A. Convergence and Compression Trade-off Empirically, EdgeDetect achieves convergence parity with full-precision FedAvg at 32× compression across 2.8M CICIDS2017 samples, with no measurable accuracy degradation (∆ < 0.2 pp). This near-lossless compression stems from median thresholding, which preserves directional alignment (cosine similarity 0.87±0.04) while suppressing low-magnitude noise—a property absent in fixed-threshold methods.

B. Privacy Enhancement Through Smartification

7

F. Effect of Regularization Reducing regularization (C : 0.1 → 100) increases mean absolute coefficient magnitude by 27.2% with minimal accuracy gain (92.2% → 92.5%), suggesting near-saturation performance. Rank ordering remains stable, with 8 of the Top-10 components preserved, confirming the robustness of the discriminative subspace. The balanced polarity (18 positive, 17 negative) indicates unbiased evidence representation. Importantly, several highly discriminative components (PC24 , PC31 , PC35 ) are not among the top variance-ranked PCs, demonstrating that explained variance does not imply classification relevance; lower-variance components encode subtle but attack-specific signals.

We quantify gradient inversion resistance across methods: (i). FedAvg (undefended): High-fidelity reconstruction (PSNR 31.7 dB) exposes structured attack signatures. (ii). signSGD: G. Regularization-Induced Weight Scaling Binarization reduces fidelity to 16.8 dB, but zero-thresholding Reducing regularization (λ : 10 → 0.01, equivalently preserves sufficient structure for partial recovery. (iii). EdgeDe- C : 0.1 → 100) produces a near-uniform amplification of tect (median-threshold): Further degrades reconstruction to logistic regression coefficients, with mean magnitude increasing 15.1 dB, rendering feature structure minimally discernible and by 27.2% across all 35 principal components. Similar growth in reducing label recovery to near random guessing (14.3%). the mean, median, maximum, and ℓ2 -norm (27–32%) confirms The addition of Paillier homomorphic encryption provides global scaling rather than selective feature inflation, indicating semantic security under the Decisional Composite Residuosity that weaker regularization relaxes shrinkage without altering the Assumption (DCRA), ensuring IND-CPA guarantees even if learned discriminative structure. Coefficient polarity remains ciphertexts are intercepted. While differential privacy param- unchanged (18 positive, 17 negative), demonstrating stable eters (ε = 1.0, δ = 10−5 ) are applied per round, cumulative class attribution. Both positive (attack-indicative) and negative privacy loss under composition remains future work. (benign-indicative) weights scale proportionally, preserving symmetry and avoiding decision-boundary bias. The modest increase in mean positive weight under weaker regularization reflects slightly stronger attack signatures but does not materiC. Binary Classification Performance Analysis ally affect the precision–recall balance. Table VI reports complete cross-validation results for binary classification, showing that Random Forest Config. 2 achieves H. Binary Classification Visualizations the highest mean accuracy (98.09%, σ = 0.0017), while KNN Figures 1 through 8 present confusion matrices and classifiConfig. 2 provides the strongest stability-efficiency tradeoff cation metrics across configurations and models. (97.93%, σ = 0.0013) with minimal training overhead.

D. Statistical Analysis of Coefficient Distributions Table VII presents a comprehensive statistical analysis of logistic regression coefficients across regularization configurations, quantifying the impact of hyperparameter tuning on learned feature weights.

(a)

(b)

Fig. 1: Comparative performance under two hyperparameter configurations. Model 2 improves detection with higher F1scores, particularly for rare attack classes.

E. Feature and Model Interpretability To enhance transparency, we analyze logistic regression coefficients over the PCA-transformed space. Coefficient magnitude reflects discriminative contribution. Across configurations, PC04 , PC26 , and PC24 consistently rank highest, jointly accounting for 38.2% of the total ℓ2 -norm (Config 2), indicating regularization-invariant importance. Positive weights (e.g., PC24 : +1.662) correspond to attack-indicative patterns—likely volumetric anomalies—while negative weights (e.g., PC04 , PC26 ) capture benign flow regularities. Additional components (PC31 , PC06 ) encode abnormal temporal and packet-rate behaviors.

(a)

(b)

Fig. 2: Model analysis and classification performance. (a) Precision, recall, and F1-score for logistic regression and SVM. (b) Multi-class results across seven traffic categories.

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

8

TABLE V: Comparison of Gradient Compression and Privacy Mechanisms Method

Quantization Rule

Adaptive Threshold

Theoretical Alignment

Privacy Integration

No No No Yes (per-client)

Implicit (unbiased sign) Variance bounded Gradient clipping bound Explicit cosine alignment (γ-bound)

None None None Paillier HE + DP

signSGD [12] sign(gi ) (zero threshold) QSGD [25] Stochastic quantization TernGrad [32] {−1, 0, +1} ternary levels EdgeDetect (Ours) sign(gi − median(g))

Notes: Among all evaluated models, Random Forest (Config. 2) achieved the highest multi-class accuracy (98.0%) with low cross-validation variance (σ = 0.0017), while KNN (Config. 2) exhibited the lowest variance overall (σ = 0.0013), confirming its robustness to data partitioning.

TABLE VI: Binary Classification Performance: 5-Fold Cross-Validation Results with Statistical Analysis Algorithm L INEAR M ODELS Logistic Regression

Config.

F1

F2

F3

F4

F5

Mean

Std

95% CI

CV Range

Config 1 0.9209 0.9253 0.9244 0.9124 0.9276 0.9221 0.0058 ±0.0051 [0.912–0.928] Config 2 0.9249 0.9324 0.9244 0.9138 0.9302 0.9251 0.0072 ±0.0063 [0.914–0.932]

K ERNEL -BASED M ETHODS Config 1 SVM (RBF) Config 2 T REE -BASED E NSEMBLE Config 1 Random Forest Config 2 S INGLE D ECISION T REE Config 1 Decision Tree Config 2 I NSTANCE -BASED L EARNING K-Nearest Neighbors Config 1 Config 2

∆ (%) Rank +0.30

5

0.8160 0.8987 0.8080 0.8058 0.8213 0.8300 0.0373 ±0.0327 [0.806–0.899] +13.14 0.9609 0.9556 0.9627 0.9609 0.9671 0.9614 0.0039 ±0.0034 [0.957–0.967]

4

0.9619 0.9615 0.9638 0.9545 0.9575 0.9598 0.0035 ±0.0031 [0.955–0.964] 0.9808 0.9832 0.9815 0.9781 0.9811 0.9809 0.0017 ±0.0015 [0.978–0.983]

+2.11

1

0.9476 0.9474 0.9486 0.9528 0.9480 0.9489 0.0022 ±0.0019 [0.947–0.953] 0.9678 0.9718 0.9703 0.9760 0.9762 0.9724 0.0035 ±0.0031 [0.970–0.976]

+2.35

3

0.9743 0.9750 0.9726 0.9735 0.9747 0.9740 0.0009 ±0.0008 [0.973–0.975] 0.9783 0.9802 0.9787 0.9781 0.9813 0.9793 0.0013 ±0.0011 [0.978–0.981]

+0.53

2

Notes: Balanced binary dataset (n = 15,000; 7,500 benign, 7,500 attack) with 35 PCA features (99.3% variance retained). All experiments used 5-fold stratified cross-validation (seed 42). Columns: F1–F5 denote fold-wise F1-scores; Mean and Std represent average and standard deviation (σ); 95% CI is computed as √ ±1.96 σ/ 5; CV Range indicates [Min, Max] across folds; ∆ denotes relative improvement. Non-overlapping confidence intervals imply p < 0.05. Key Findings: Random Forest (Config 2) achieves the highest accuracy (98.09%) with low variance (0.0017), indicating stable generalization. KNN (Config 2) exhibits the lowest variance (0.0013) with marginally lower accuracy (97.93%). SVM shows the largest hyperparameter sensitivity (+13.14%), highlighting the importance of kernel selection (linear → RBF). Logistic regression demonstrates marginal gains (+0.30%), suggesting limited linear separability in PCA space. All models exhibit low within-fold variability (range < 2%), confirming reproducibility.

(a)

(b)

Fig. 3: Per-class metrics across attack categories. Model 2 improves recall and F1-score for minority classes.

(a)

(b)

Fig. 5: Binary confusion matrices: Model 2 reduces false negatives while preserving high true-positive rates.

(a)

(b)

Fig. 4: Per-class metrics for classical ML baselines: Random Forest (most consistent), Decision Tree, and K-Nearest Neighbors.

Fig. 6: Logistic regression vs. SVM: SVM exhibits lower false positives and improved class separation.

I. Implications for Intrusion Detection

space, motivating non-linear or ensemble approaches for further improvement. Table VIII reports binary classification results using 5-fold stratified cross-validation on a balanced dataset (n = 15,000) with 35 PCA components preserving 99.3% variance.

Controlled relaxation of regularization enables finer-grained discrimination without compromising numerical stability or interpretability. The modest performance gains suggest proximity to the representational limit of linear classifiers in PCA

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

9

TABLE VII: Logistic Regression Coefficient Statistics: Regularization Impact on Feature Weights Config 1 Config 2 ∆ (%)

Statistical Metric

Interpretation

Central Tendency Measures Mean |βi | Median |βi | Std Dev |βi |

0.543 0.412 0.389

0.691 0.543 0.512

+27.2 Average discriminative strength +31.8 Typical feature importance +31.6 Weight distribution spread

Magnitude Characteristics Max |βi | Min |βi | ℓ2 -norm ∥w∥2 ℓ1 -norm ∥w∥1

1.607 0.004 4.127 19.012

2.050 0.007 5.243 24.185

+27.5 +75.0 +27.0 +27.2

Strongest discriminative PC Weakest discriminative PC Total model complexity Manhattan weight magnitude

Fig. 8: Confusion matrices: Random Forest (diagonal dominance), Decision Tree, K-Nearest Neighbors.

Class Association Distribution Positive coefficients 18 / 35 18 / 35 0.0 Negative coefficients 17 / 35 17 / 35 0.0 Mean β+ +0.621 +0.798 +28.5 Mean β− −0.571 −0.723 +26.6

Attack-indicative features Benign-indicative features Avg. attack feature weight Avg. benign feature weight

TABLE VIII: Binary Classification Performance: CrossValidation and Test Set Evaluation

Weight Concentration Analysis Top-3 PCs (% ℓ2 ) 32.8% Top-10 PCs (% ℓ2 ) 70.3% Bottom-10 PCs (% ℓ2 ) 4.2% Gini coefficient 0.412

34.1% 72.4% 3.8% 0.428

+4.0 +3.0 −9.5 +3.9

Dominance of key features Cumulative importance Low-importance features Weight inequality measure

Sparsity and Regularization Effects Near-zero (|βi | < 0.1) 3 / 35 Low (0.1 ≤ |βi | < 0.5) 15 / 35 Medium (0.5 ≤ |βi | < 1.0) 13 / 35 High (|βi | ≥ 1.0) 4 / 35

2 / 35 −33.3 12 / 35 −20.0 15 / 35 +15.4 6 / 35 +50.0

Weakly discriminative PCs Moderate importance High importance Critical features

Model Characteristics Intercept w0 −2.351 −2.962 +26.0 Effective degrees of freedom 32 33 +3.1 Regularization strength λ 10.0 0.01 −99.9 Condition number κ 18.4 24.3 +32.1

Decision boundary offset Active parameters Penalty magnitude Numerical stability

Performance Correlation CV Accuracy Test Accuracy AUC-ROC

0.922 0.920 0.978

0.925 0.930 0.980

+0.3 +1.1 +0.2

Cross-validation performance Held-out set performance Discriminative ability

Notes: Config 1: C = 0.1 (strong regularization, saga); Config 2: C = , and |βi | 100 (weak regularization, sag). βi denotes the coefficient of PC qiP 35 2 its magnitude. ∆ (%) = 100(Cfg2 − Cfg1)/Cfg1. ∥w∥2 = i=1 βi ; P35 ∥w∥1 = i=1 |βi |. Gini measures the coefficient of inequality (0 = uniform, 1 = concentrated). κ (condition number) = ratio of largest to smallest singular value. Effective degrees of freedom = number of coefficients with |βi | ≥ 0.01. Top-k (% ℓ2 ) = share of total ℓ2 -norm from the k largest coefficients.

(a)

Model

Config.

Linear Models Logistic Reg. C = 0.1, saga Logistic Reg. C = 100, sag

CV Acc.

Test Acc. Prec. Rec. F1 AUC-ROC

0.922±0.006 0.925±0.007

0.920 0.930

0.918 0.923 0.924 0.928 0.932 0.929

0.978 0.980

Kernel-Based Methods SVM poly, C = 1 0.830±0.037 SVM rbf, C = 1, γ = 0.1 0.961±0.004

0.830 0.960

0.826 0.835 0.830 0.958 0.962 0.960

0.892 0.987

Notes: CV Acc. = mean accuracy ± standard deviation across 5 folds. Test metrics computed on held-out 20% partition (n = 3,000). Prec. = precision (macro-averaged); Rec. = recall (macro-averaged); F1 = F1-score; AUC-ROC = area under receiver operating characteristic curve. Green highlight indicates the best overall performance. Random seed 42 for reproducibility.

K. Concentration and Distribution of Discriminative Power Under weaker regularization, the top-10 principal components account for 72.4% of the total ℓ2 -norm (vs. 70.3%), indicating modest concentration of discriminative mass. The Gini coefficient increases slightly (0.412 → 0.428), suggesting mild inequality while importance remains broadly distributed. Lowerranked components contribute less, reflecting suppression of non-informative variance. Both settings remain weakly sparse, with over 94% of components active; weaker regularization activates one additional feature, slightly increasing effective degrees of freedom. Although the condition number rises (18.4 → 24.3), it remains well within stable bounds, and intercept adjustment preserves calibration.

(b)

Fig. 7: Classical classifiers on CIC-IDS2017: KNN shows improved minority-class detection; DT exhibits higher DoSDDoS misclassification.

TABLE IX: Multi-Class Classification Performance (5-Fold Cross-Validation) Model

J. Multi-Class Classification Performance Table IX presents attack categorization performance across seven classes: BENIGN, DoS, DDoS, Port Scan, Brute Force, Web Attack, and Bot. The balanced multi-class dataset (n = 35,000; 5,000 samples per class) was constructed via SMOTE oversampling for minority classes and random undersampling for the majority class, following the removal of classes with fewer than 1,950 instances.

Config.

CV Acc.

Test Acc. Prec.

Rec.

F1

Tree-Based Ensemble Random Forest T =10, d=6 0.960±0.009 Random Forest T =15, d=8, m=20 0.980±0.007

0.971 0.980

0.969 0.970 0.969 0.979 0.980 0.979

Single Decision Trees Decision Tree d=6 Decision Tree d=10

0.948±0.012 0.960±0.012

0.887 0.903

0.882 0.885 0.883 0.901 0.902 0.901

Instance-Based Learning KNN k=5, uniform KNN k=7, distance-wt

0.935±0.015 0.940±0.014

0.945 0.952

0.943 0.946 0.944 0.950 0.953 0.951

Notes: T = number of trees; d = maximum depth; m = max features per split; k = number of neighbors. CV Acc. = mean ± std over 5 folds. Precision, recall, and F1 are macro-averaged across 7 classes. A green highlight indicates the best overall performance.

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

Fig. 9: ROC curve analysis: Model comparison across configurations.

Fig. 10: ROC curve comparison: Algorithm performance across metrics.

10

Fig. 11: ROC analysis: SVM performance across configurations.

Fig. 12: Recall curve analysis: Detection performance across models.

indicates strong directional alignment with the true gradient. This preservation of gradient direction is sufficient to maintain Table X positions EdgeDetect within modern IDS research. convergence parity in convex and near-convex regimes. The proposed federated Random Forest achieves 98.0% accu  racy on CIC-IDS2017 while reducing per-round communication ⟨∆bin , ∇L⟩ by 96.9% (450 MB → 14 MB) and enabling CPU-only edge E[∆bin ] ̸= ∇L, E = 0.87 ± 0.04. ∥∆bin ∥2 ∥∇L∥2 deployment (Raspberry Pi 4: 4.2 MB memory, 0.8 ms latency). (31) Unlike GPU-dependent centralized deep models, EdgeDetect operates in fully federated settings with cryptographic privacy guarantees. The framework integrates four synergistic A. Theoretical Convergence Analysis components: (1) an end-to-end privacy-preserving federated pipeline; (2) hybrid SMOTE–undersampling with PCA yielding Lemma 1 (Descent under Median-Threshold Smartification). 95.0% minority-class F1; (3) gradient smartification providing Let L(W ) be L-smooth and bounded below. Let g̃t denote the ⟨g ,g̃ ⟩ 32× communication compressions without accuracy loss; and smartified gradient with cosine similarity cos(θt ) = ∥gtt∥∥g̃tt ∥ ≥ (4) robustness to heterogeneity, imbalance, and poisoning γ > 0. Then for sufficiently small step size η, (p < 0.001). Random Forest achieves 98.09% F1 in binary detection (0.17% variance) and 98.0% multi-class accuracy Lη 2 (97.9% macro F1), outperforming single trees (90.3%). EnE[L(Wt+1 )] ≤ L(Wt ) − ηγ∥gt ∥2 + ∥g̃t ∥2 . 2 semble scaling (T : 10 → 15, d : 6 → 8) improves accuracy by +2.0 pp and reduces variance by 22%. Under non-IID Proposition 1 (Bias–Variance Tradeoff). Let g be the true t conditions (α = 0.1), FedProx maintains 95.1% accuracy, with gradient and g̃ its smartified version. Then the expected t sub-linear convergence scaling (98 rounds at K = 10 vs. 234 deviation satisfies: at K = 500). Gradient encryption reduces inversion quality (PSNR 15.1 dB vs. 31.7 dB) with only 156.4 ms overhead E[∥gt − g̃t ∥2 ] = Bias2 + Varquant , per round. The system tolerates 20% malicious clients while maintaining >85% accuracy and limiting backdoor success to where median-thresholding reduces Varquant for heavy-tailed <7%. Compared to KNN (95.2%, 3.21 ms), Random Forest gradient distributions. delivers superior throughput (0.87 ms), confirming suitability for high-rate edge deployment. For heavy-tailed IDS gradients, variance reduction dominates Three additional ROC analysis figures are provided (Fig- bias increase, yielding stable convergence. ures 9 through 11) demonstrating AUC-ROC analysis across Theorem 1 (Convergence under Bounded Variance). Assume models. bounded stochastic gradient variance σ 2 and cosine similarity γ > 0. Then after T rounds, VIII. F EDERATED L EARNING C ONVERGENCE A NALYSIS   While binarization introduces coordinate-wise bias (i.e., 1 2 √ min E[∥∇L(Wt )∥ ] = O . E[∆bin ] ̸= ∇L), empirical cosine similarity of 0.87 ± 0.04 t≤T γ T L. Comparative Analysis with State-of-the-Art

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

11

TABLE X: Comparative Analysis with State-of-the-Art Intrusion Detection Systems Study

Year

Centralized Approaches Alam et al. [56] 2023 Ghani et al. [57] 2023 Savic et al. [58] 2021 Cerar et al. [59] 2020 Federated Learning Approaches Liu et al. [2] 2023 Wang et al. [60] 2022 Zhang et al. [61] 2022 Chen et al. [62] 2021 This Work (EdgeDetect) EdgeDetect 2026(Ours) (Binary)

Model

Acc. (%)

F1 (%)

Dataset

Classes

Privacy

Comm. (MB)

Key Innovation

CNN XGBoost LSTM-AE Iso. Forest

97.2 96.1 95.5 93.8

96.8 95.4 94.2 91.6

CIC-IDS2017 CIC-IDS2017 NSL-KDD CIC-IDS2017

Binary 7-class Binary Binary

✗ ✗ ✗ ✗

N/A N/A N/A N/A

Image-encoded traffic Feature visualization Anomaly scoring Unsupervised learning

Fed-DNN Fed-CNN FedAvg-LSTM Fed-XGB

96.3 94.7 93.5 95.8

95.1 93.8 92.4 94.9

UNSW-NB15 CIC-IDS2017 KDD-CUP99 IoT-23

5-class Binary 4-class Binary

DP ✗ DP SecAgg

380 520 410 290

Differential privacy Model aggregation Temporal modeling Gradient encryption

Fed-RF

98.0 96.0

97.9 96.0

CIC-IDS2017

7-class Binary

HE

14 14

Gradient smartification + Paillier encryption

Notes: Acc. = test accuracy; F1 = macro-averaged F1-score. Privacy mechanisms: ✗= none, DP = differential privacy, SecAgg = secure aggregation, HE = homomorphic encryption (Paillier). Comm. = per-round communication cost per client; N/A indicates centralized training with no federated communication. Dataset sizes: CIC-IDS2017 (2.8M samples), UNSW-NB15 (2.5M), NSL-KDD (148K), KDD-CUP99 (4.9M), IoT-23 (325K). EdgeDetect achieves 96.9% communication reduction versus federated baselines (14 MB vs. 290-520 MB) while providing stronger cryptographic guarantees (Paillier HE vs. DP or SecAgg). Green highlighting indicates the best performance. We emphasize that differential privacy (DP) and secure aggregation (SecAgg) address distinct threat models: DP provides formal statistical guarantees against inference attacks on individual data samples, whereas SecAgg cryptographically prevents the server from accessing individual client updates, revealing only their aggregate.

variance from skewed coordinates while preserving directional consistency. For symmetric distributions, E[gi sign(gi − τ )] ≥ c E[gi2 ] for some c > 0 depending on distribution kurtosis. Summing across coordinates yields the global alignment constant γ. 2) Experimental Setup for Federated Scenarios: We evaluate Fig. 13: Recall-precision trade-off analysis across configura- EdgeDetect under realistic federated settings by partitioning CIC-IDS2017 across K ∈ {10, 25, 50, 100, 500} clients using tions. (i) IID balanced sampling; (ii) Non-IID quantity skew via Dirichlet allocation with α ∈ {0.1, 0.5, 1.0, 10.0} (smaller α implies stronger heterogeneity); and (iii) Non-IID label skew, where each client predominantly observes 2–3 attack types (e.g., web servers dominated by Web/Bot traffic). To model intermittent availability, we vary the per-round participation rate C ∈ {0.25, 0.50, 0.75, 1.00}. Table XI summarizes convergence and bandwidth for EdgeDetect and baselines. Unless stated otherwise, we use a local batch size B = 32, E = 5 Fig. 14: Logistic regression recall analysis: Threshold- local epochs, and a global learning rate η = 0.01. dependent performance. B. Interpretation and Edge Detection 1) Proposition: Alignment of Median-Threshold SmartifiVolumetric attack classes exhibit near-linear separability in cation: Proposition 1 (Expected Descent Alignment). Let PCA space. Classes (DoS/DDoS) remain robust (> 0.97 F1 at g ∈ Rd denote the true gradient and g̃ the median-threshold α = 0.1) due to distinctive flow statistics. In contrast, Bot and binarized update defined as Web Attack degrade the most (0.927→0.854 and 0.939→0.881), consistent with rarer and semantically overlapping behaviors g̃i = sign(gi − τ ), τ = median(g). that are fragmented under skewed client partitions. EdgeDetect matches full-precision convergence under IID (R98 =289 vs. 287 Assume each coordinate of g follows a symmetric heavy-tailed for FedAvg) while reducing total bandwidth by 96.9% (4.05 GB distribution with finite second moment and zero median shift. vs. 129.15 GB). Under heterogeneity, the gap to full precision Then there exists a constant γ ∈ (0, 1) such that increases, but EdgeDetect remains competitive: at α = 0.1,   2 EdgeDetect improves over signSGD in both accuracy (94.2% E ⟨g, g̃⟩ ≥ γ∥g∥2 . vs. 92.1%) and rounds (612 vs. 721), and the combination Sketch of Justification. Under symmetric heavy-tailed dis- EdgeDetect+FedProx yields the best heterogeneous result tributions, the median satisfies P(gi ≥ τ ) ≈ 0.5. Unlike zero- (95.1%, 7.88 GB). Scalability is favorable: increasing the threshold signSGD, the adaptive median threshold reduces number of clients from K = 10 to K = 500 raises R98 from

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

TABLE XI: Federated Learning Convergence Analysis K Distribution R95 R98 Acc. (%) Comm./R Total Clients (MB) (GB) IID Distribution (α = ∞) FedAvg 50 IID 142 287 98.2 450.0 129.15 FedProx (µ = 0.01) 50 IID 138 276 98.3 450.0 124.20 signSGD 50 IID 156 312 97.8 14.1 4.40 EdgeDetect 50 IID 145 289 98.0 14.0 4.05 Non-IID (Moderate Heterogeneity, α = 1.0) FedAvg 50 Dir. α = 1.0 201 423 96.4 450.0 190.35 FedProx (µ = 0.01) 50 Dir. α = 1.0 187 389 97.1 450.0 175.05 signSGD 50 Dir. α = 1.0 218 445 95.7 14.1 6.27 EdgeDetect 50 Dir. α = 1.0 192 398 96.8 14.0 5.57 Non-IID (High Heterogeneity, α = 0.1) FedAvg 50 Dir. α = 0.1 312 687 93.8 450.0 309.15 FedProx (µ = 0.01) 50 Dir. α = 0.1 276 591 94.9 450.0 265.95 signSGD 50 Dir. α = 0.1 334 721 92.1 14.1 10.16 signSGD + Momentum 50 Dir. α = 0.1 298 652 93.4 14.1 9.19 EdgeDetect 50 Dir. α = 0.1 287 612 94.2 14.0 8.57 EdgeDetect + FedProx 50 Dir. α = 0.1 264 563 95.1 14.0 7.88 Scalability Analysis (IID) EdgeDetect 10 IID 98 201 98.1 14.0 2.81 EdgeDetect 25 IID 126 254 98.0 14.0 3.56 EdgeDetect 100 IID 178 356 97.9 14.0 4.98 EdgeDetect 500 IID 234 467 97.7 14.0 6.54 Algorithm

Notes: R95 and R98 are rounds to reach 95% and 98% accuracy. Comm./R is per-client per-round communication. Total is the total bandwidth to reach 98% accuracy. FedProx uses µ = 0.01. Results are averaged over 5 runs (different seeds).

201 to 467 (sublinear in K), indicating stable aggregation despite a larger, noisier client pool. 1) Convergence Rate and Compression Quality: We empirically assess whether smartification preserves update directions by measuring cosine alignment between compressed and full gradients: cos(∠(∆comp , ∆full )) =

⟨∆comp , ∆full ⟩ . ∥∆comp ∥ ∥∆full ∥

12

D. Training Efficiency and Robustness Logistic Regression and Decision Trees train rapidly (< 2.5s), while Random Forest provides the best accuracyefficiency balance (12.3s training, 0.87 ms inference). SVM (18.7s) and KNN (3.21 ms inference, 412 MB memory) incur higher computational or memory costs, limiting scalability. Random Forest demonstrates strong stability (CV std < 0.3%, p < 0.001), with volumetric and temporal features (Flow Bytes/s, Flow Duration) contributing 52.7% of total importance. SMOTE substantially improves minority recall (Bot: 0.39 → 0.98) with minimal accuracy loss (0.4%), while PCA reduces dimensionality from 78 to 35 features (99.3% variance retained), lowering training time (−38%) and memory (−64%) with negligible performance impact. Under partial participation (C < 1), per-round bandwidth decreases, but convergence slows due to fewer client updates. IX. A BLATION S TUDY We performed a controlled ablation study XII to quantify the individual and joint contributions of EdgeDetect components across four axes: (i) classification performance, (ii) communication efficiency, (iii) privacy resilience, and (iv) convergence dynamics. Each component (smartification, homomorphic encryption, differential privacy, PCA, SMOTE, and FedProx) was selectively removed while keeping all other settings fixed (CIC-IDS2017, K = 50 clients, IID distribution, 5 runs, averaged).

(32)

Across all rounds, EdgeDetect achieves a mean cosine similarity 0.87±0.04, indicating that compression retains most directional information and explaining the near-parity in IID convergence despite 32× quantization. If the gradient direction cosine ≥ 0.8, then the expected descent holds:    E L W t+1 ≤ L W t − η cos(θ) ∥∇L(W t )∥22 + O(η 2 ). (33)

A. PCA: Detailed Attack Type Characterization Principal Component Analysis (PCA) was applied to reduce the original 78 high-dimensional network features to 35 uncorrelated components, retaining 99.3% of the variance (Table XIII). This transformation lowers computational overhead, mitigates noise, and enhances discriminative visualization between benign and attack traffic in the reduced-dimensional space. X. A BLATION S TUDY: C OMPONENT I MPACT A NALYSIS

C. Class-specific insights (concise) A. Impact of Gradient Smartification Volumetric attack classes exhibit near-linear separability in Removing gradient smartification (replacing binarization the PCA space: DoS/DDoS achieve F-1 scores of 0.989/0.987 with low mutual confusion (2.1%), indicating that PCA with full-precision gradients) while keeping encryption and DP preserves discriminative variance for rate- and volume-driven active: Communication Cost: Increases from 14.0 MB to 450.0 signatures. BENIGN traffic is identified reliably (F-1=0.989; MB per round (32.1× increase; 29.85 GB total communication). precision=0.992) with a 0.8% false positive rate, dominated by Accuracy: 98.2% vs. 98.0% (+0.2 pp improvement, statistically confusions with Port Scan (5 cases) and Brute Force (3 cases); insignificant at p ¿ 0.05). Convergence: 287 rounds to 98% false alarms toward DoS/DDoS remain < 0.1%. Application- (vs. 289 rounds), negligible difference. layer attacks are hardest due to overlap with legitimate flows: Conclusion: Smartification is a communication optimization Web Attack (F-1=0.939) and Bot (F-1=0.927) show the largest mechanism with a near-zero accuracy penalty. The modest confusion (11.2% mutual) and the highest false negatives, accuracy improvement (+0.2 pp) under full precision likely particularly Bot (8.1%; 43/530), consistent with encrypted reflects reduced quantization bias, but communication savings C&C and timing randomization. In contrast, volumetric false (32×) far outweigh this negligible gain. Table XVI provides negatives are rare (DoS: 1.2%, DDoS: 1.5%) and mostly a consolidated summary of the necessity and contribution of correspond to low-rate or short-duration phases. each component:

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

13

TABLE XII: Ablation Study: Component-wise Impact on Accuracy, Communication, Privacy, and Convergence Components Configuration

Accuracy Metrics

Smartif. HE DP PCA SMOTE Acc (%)

BASELINE COMPARISONS FedAvg (No Protection) signSGD (Binarization Only)

✗ ✓

F1

∆Acc

Communication Std

Privacy

/Round (MB) Ratio Total (GB) PSNR (dB)

Invert?

✗ ✗

✗ ✗

✓ ✓

✓ ✓

98.2 97.8

0.9790 — 0.0042 0.9754 -0.4 pp 0.0053

450.0 14.1

1.0× 31.9×

129.15 4.40

31.7 16.8

✓ Yes ✓ Partial

PROGRESSIVE REMOVAL (Ablation Track) Full EdgeDetect ✓ ✓ ✓ ✗ - Encrypt (Smartif only) - DP Noise ✓ ✓ - PCA (78 features) ✓ ✓ ✓ ✓ - SMOTE (Random US) - Smartif (Full Precision) ✗ ✓

✓ ✓ ✗ ✓ ✓ ✓

✓ ✓ ✓ ✗ ✓ ✓

✓ ✓ ✓ ✓ ✗ ✓

98.0 98.0 98.1 97.9 94.2 98.2

0.9789 0.9789 0.9791 0.9787 0.9341 0.9794

0.0 pp 0.0 pp +0.1 pp -0.1 pp -3.8 pp +0.2 pp

0.0048 0.0048 0.0045 0.0051 0.0067 0.0043

14.0 14.0 14.0 58.2 14.0 450.0

32.1× 32.1× 32.1× 7.7× 32.1× 1.0×

4.05 4.05 4.05 16.73 4.05 129.15

15.1 15.1 14.2 15.3 15.1 15.1

✗ No ✓ Vulnerable ✗ Protected ✗ Protected ✗ Protected ✗ Protected

MULTI-COMPONENT REMOVAL - Smartif - HE (Binarized) - Smartif - DP - HE - DP (Binarization) - PCA - SMOTE (Full)

0.0 pp +0.1 pp -0.2 pp -4.3 pp

✓ ✓ ✓ ✗

✗ ✓ ✗ ✓

✓ ✗ ✗ ✓

✓ ✓ ✓ ✗

✓ ✓ ✓ ✗

98.0 98.1 97.8 93.7

0.9789 0.9791 0.9754 0.9261

0.0048 0.0045 0.0053 0.0089

14.0 14.0 14.1 450.0

32.1× 32.1× 31.9× 1.0×

4.05 4.05 4.40 129.15

15.1 14.2 16.8 15.1

✓ Vulnerable ✗ Protected ✓ Vulnerable ✗ Protected

ALTERNATIVE CONFIGURATIONS FedProx instead of FedAvg ✓ Differential Privacy Only (DP-SGD) ✗ Secure Aggregation (SecAgg) ✓

✓ ✗ ✗

✓ ✓ ✗

✓ ✓ ✓

✓ ✓ ✓

98.4 93.8 98.0

0.9816 +0.4 pp 0.0041 0.9358 -4.2 pp 0.0062 0.9789 0.0 pp 0.0048

14.0 450.0 14.0

32.1× 1.0× 32.1×

3.79 129.15 4.05

15.1 18.9 31.7

✗ Protected ✗ Partial ✗ Protected

FEATURE ENGINEERING VARIANTS - No Temporal Features ✓ ✓ - No Entropy Features - All Original 78 Features ✓

✓ ✓ ✓

✓ ✓ ✓

✓ ✓ ✗

✓ ✓ ✓

96.3 97.1 97.9

0.9621 -1.7 pp 0.0058 0.9705 -0.9 pp 0.0054 0.9787 -0.1 pp 0.0051

14.0 14.0 58.2

32.1× 32.1× 7.7×

4.05 4.05 16.73

15.1 15.1 15.3

✗ Protected ✗ Protected ✗ Protected

Notes: All ablation experiments were conducted on CIC-IDS2017 with K = 50 clients, IID distribution, and 5 independent runs (averaged). Smartif = Gradient Smartification (median-threshold binarization); HE = Paillier Homomorphic Encryption; DP = Differential Privacy noise; PCA = Principal Component Analysis (35 components); SMOTE = Synthetic Minority Oversampling. ∆Acc = Accuracy change relative to Full EdgeDetect (0.0 pp = no difference, negative = worse, positive = better). Comm./Round = per-client per-round communication cost. Ratio = compression ratio vs. FedAvg. Total = total bandwidth to reach 98% target accuracy. PSNR = Peak Signal-to-Noise Ratio from gradient inversion attack (iDLG); higher = more vulnerable. Invert indicates gradient reconstruction success. ✓ = Present; ✗ = Absent.

TABLE XIII: PCA: Complete Attack Type Profiles Across Primary, Secondary, and Discriminative Feature Spaces Attack Type

Primary Separators (82.4% Var.)

Secondary Features (10.3% Var.)

Discriminative Components

Summary

PC2

PC3

PC4

PC5

PC6

PC13

PC23

PC24

PC26

PC27

PC31

∥x∥2

Class

Benign Traffic (Baseline) BENIGN1 -2.358 -0.055 -2.884 -0.070 BENIGN2 BENIGN3 -2.417 -0.057 -2.885 -0.070 BENIGN4 BENIGNavg -2.39 -0.06

0.577 0.911 0.615 0.912 0.70

0.734 1.763 0.851 1.765 1.03

3.730 8.846 4.304 8.852 6.43

0.235 0.620 0.276 0.619 0.44

1.638 6.053 2.132 6.056 3.97

-1.722 -5.607 -2.156 -5.609 -3.77

-0.070 0.296 -0.032 0.293 0.12

0.905 0.594 0.868 0.592 0.74

-0.148 0.283 -0.102 0.281 0.08

-0.219 -0.367 -0.235 -0.366 -0.30

4.14 10.31 4.83 10.32 7.40

Ref. Ref. Ref. Ref. Ref.

Volumetric attack classes exhibit near-linear separability in PCA space Attacks (DoS Family) 2.840 0.120 -0.920 -1.760 -8.850 -0.620 -6.060 DoS DDoS 1.510 -0.080 0.500 -0.290 0.540 -0.750 0.430

5.610 -0.290

-0.300 -0.390

-0.590 0.690

-0.280 0.600

0.370 -0.040

10.32 1.89

Extreme Extreme

Reconnaissance Attacks -0.450 Port Scan Brute Force -0.380

PC1

0.030 0.060

-0.210 -0.180

0.180 0.250

-0.650 -0.540

0.320 0.410

-0.120 -0.080

0.290 0.320

0.190 0.210

-0.330 -0.360

0.220 0.250

-0.010 0.020

0.87 0.81

Moderate Moderate

Application-Layer Attacks Web Attack -0.620 0.080 Bot -0.710 0.100

-0.350 -0.420

0.420 0.510

-0.890 -1.080

0.530 0.640

-0.150 -0.200

0.590 0.630

0.320 0.410

-0.480 -0.590

0.380 0.470

0.030 0.040

1.24 1.51

Ambiguous Ambiguous

Notes: Structure: Rows grouped by attack taxonomy. PC1–3 (82.4% variance) drive primary benign–attack separation; PC4–6 (10.3%) capture secondary variation. Key discriminative components (PC13, 23, 24, 26, 27, 31) enable multi-class differentiation. ∥x∥2 denotes the Euclidean norm over PC1–35. Class labels: Extreme (F1 > 0.98), Moderate (0.96–0.97), Ambiguous (< 0.94). BENIGNavg is computed over 10 samples; BENIGN1−4 illustrates intra-class variance. DoS exhibits a strong negative PC5 (−8.85), reflecting volumetric anomalies. Application-layer attacks overlap with BENIGN along PC1–3 and separate primarily via PC4–6.

Principal Component Analysis (PCA): Removing PCA increases communication from 14.0 MB to 58.2 MB (4.16×) Smartification (Gradient Binarization): Removing bina- and computation (+182%) with only 0.1 pp accuracy difference, rization increases per-round communication from 14 MB to revealing feature redundancy. SMOTE Balancing: Eliminating 450 MB (32×) with negligible accuracy change (98.0% → SMOTE reduces accuracy to 94.2% (-3.8 pp) and macro-F1 to 98.2%, p > 0.05), confirming its communication efficiency. 0.934 due to minority-class degradation, highlighting the need Homomorphic Encryption (HE): Disabling Paillier encryption for class balancing. Smartification + Encryption Synergy: enables gradient inversion (PSNR 31.7 dB; >95% label recov- Combined binarization and HE achieve inversion resistance ery) while accuracy remains 98.0%, indicating strong privacy (PSNR 15.1 dB; 14.3% label recovery ≈ random guessing) protection without performance cost. Differential Privacy with no accuracy loss. FedProx Integration: Adding FedProx (DP): With smartification + HE, DP yields marginal privacy improves heterogeneity robustness, raising accuracy to 98.4% gain and +0.1 pp accuracy change; standalone DP-SGD reduces and reducing total communication to 3.79 GB. accuracy by 4.2 pp, showing the privacy-utility trade-off. B. Ablation Study Key Findings

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

14

TABLE XIV: Principal Component Variance Decomposition and Attack Class Separation Metrics Attack Type

PC1–3 Norm

PC4–6 Norm

Disc. Norm†

Total Norm

Separation‡

Std Dev (PC1–5)

Distinctiveness

F1–Score

BENIGN (avg) DoS DDoS Port Scan Brute Force Web Attack Bot

0.84 2.93 1.58 0.23 0.21 0.37 0.45

1.09 1.07 0.54 0.39 0.38 0.57 0.68

0.74 0.55 0.75 0.25 0.23 0.36 0.44

7.40 10.32 1.89 0.87 0.81 1.24 1.51

— 9.17 6.84 0.31 0.26 0.42 0.51

3.59 4.44 0.59 0.39 0.36 0.55 0.66

Baseline Very High High Medium Medium Medium–Low Medium–Low

0.989 0.989 0.987 0.966 0.963 0.939 0.927

p Notes: Norms: ∥PC1–3∥ = PC12 + PC22 + PC32 (primary separation); ∥PC4-6∥ captures secondary variation; Disc. Norm = mean | · | over PC13, 23, 24, 26, 27, 31. Separation‡ = Euclidean distance from the BENIGN centroid in PC1–5 space (higher = stronger class separation). Std Dev (PC1–5) reflects dispersion in primary components. F1 = macro-averaged multi-class F1 (Sec. VII-J). Volumetric attack classes exhibit near-linear separability in PCA space attacks (DoS/DDoS) show maximal separation; application-layer attacks remain closest to the benign cluster.

TABLE XV: Detailed Principal Component Contributions: All Attack Types Across 35 Components (Selected Subset) PC1–11 (Variance Rank 1–11)

Type PC1 BENIGN DoS DDoS Port Scan Brute Force Web Attack Bot

PC2

PC3

PC4

PC5

PC6

PC7

PC8

PC15, PC20–35 (Key + Tail) PC9 PC10 PC11 PC15 PC20 PC23 PC24 PC26 PC27 PC29 PC31 PC32 PC34 PC35

-2.39 -0.06 0.70 1.03 6.43 0.44 -0.02 0.41 0.49 2.84 0.12 -0.92 -1.76 -8.85 -0.62 0.06 -1.11 -1.91 1.51 -0.08 0.50 -0.29 0.54 -0.75 -0.10 -0.73 1.15 -0.45 0.03 -0.21 0.18 -0.65 0.32 0.02 0.30 -0.35 -0.38 0.06 -0.18 0.25 -0.54 0.41 0.02 0.33 -0.30 -0.62 0.08 -0.35 0.42 -0.89 0.53 0.03 0.40 -0.48 -0.71 0.10 -0.42 0.51 -1.08 0.64 0.04 0.49 -0.59

0.97 2.76 0.56 -0.28 -0.22 -0.35 -0.43

-0.21 0.95 0.04 -0.02 -0.02 -0.03 -0.04

-0.80 4.73 0.61 -0.30 -0.25 -0.38 -0.47

-0.60 0.59 -0.11 0.06 0.08 0.11 0.14

-3.77 5.61 -0.29 0.23 0.26 0.47 0.56

0.12 -0.30 -0.39 0.19 0.21 0.32 0.41

0.74 -0.59 0.69 -0.33 -0.36 -0.48 -0.59

0.08 -0.28 0.60 0.22 0.25 0.38 0.47

0.80 -2.24 -0.80 0.19 0.17 0.26 0.32

-0.30 0.37 -0.04 -0.01 0.02 0.03 0.04

0.00 -0.01 0.01 0.00 0.00 0.00 0.00

0.02 -0.13 0.05 -0.02 -0.02 -0.04 -0.05

-0.04 0.19 0.00 0.00 0.00 0.01 0.01

Notes: Projection spans 35 principal components (22 most critical shown). PC1–11 capture dominant variance, while PC15 and PC20–35 include key discriminative components (notably PC24, PC26, and PC31). Pattern summary: DoS exhibits extreme deviation on PC5 (−8.85), reflecting volumetric anomalies; DDoS shows moderate displacement. Reconnaissance attacks cluster near the origin (0.2–0.4 across PC1–6). Application-layer attacks (Web Attack, Bot) shift primarily along PC4-6, indicating subtle evasion patterns. Tail components (PC32–35) remain near zero across classes, confirming effective dimensionality reduction at k = 35.

TABLE XVI: Ablation Study Summary: Component Necessity and Contribution Component Smartification (Binarization) Paillier Encryption Differential Privacy PCA SMOTE

Necessary?

Accuracy Impact

Communication Impact

Privacy Impact

Overhead

Recommendation

Yes (Comm.) Yes (Privacy) Optional Yes (Efficiency) Yes (Accuracy)

Negligible (-0.2 pp) None (0 pp) Negligible (+0.1 pp) Negligible (+0.1 pp) Critical (3.8 pp)

Critical (32×) None (0 MB) None (0 MB) Essential (4.16×) None (0 MB)

Important (↓16.8 dB) Critical Marginal None (0 dB) None (0 dB)

Low (+2.4%) Medium (+1,760% per round) Low (+4.1%) Medium (+182%) Low (-18%)

Keep Keep Optional Keep Keep

Notes: “Necessary” indicates significant degradation if removed (Comm. = communication, Privacy = gradient leakage, Accuracy = detection quality). Acc. Impact is measured in percentage points (pp). Communication impact reported relative to full-precision FedAvg. For bandwidth-unlimited deployments, smartification may be replaced by FedAvg. Encryption is recommended for sensitive environments. PCA and SMOTE are universally beneficial.

TABLE XVII: Unified Ablation Analysis: Privacy, Utility, Communication, and Efficiency Configuration Full EdgeDetect – No Smartification – No Encryption – No DP – No PCA (78 feat.) – No SMOTE

Acc. (%)

F1

Comm. (MB)

PSNR (dB)

Label Rec. (%)

Train (s)

Mem (MB)

Primary Impact

98.0 98.2 98.0 98.1 97.9 94.2

0.979 0.979 0.979 0.979 0.978 0.934

14.0 450.0 14.0 14.0 58.2 14.0

15.1 15.1 31.7 15.1 15.3 15.1

14.3 14.3 98.7 14.3 14.3 14.3

12.3 12.3 12.3 12.3 34.7 10.1

234 234 234 234 612 200

Secure + Efficient 32× Comm. increase Privacy collapse Marginal effect 4.16× Comm. increase Minority recall collapse

Notes: Acc. = test accuracy; F-1 = macro F1-score; Comm. = per-client per-round communication; PSNR = gradient reconstruction quality under iDLG (higher = more leakage); Label Rec. = attack-class recovery rate. Results averaged over 5 runs (K = 50). Removing encryption yields inversion success > 95%. Removing smartification increases communication from 14 MB to 450 MB per round (32×). Removing PCA raises total communication from 4.05 GB to 16.73 GB (+312%). Removing SMOTE reduces minority-class recall by up to 60%.

lowering gradient entropy and improving privacy. Combined with Paillier encryption, it retains 98.7% of centralized accuracy The results highlight three key insights for federated intrusion with complete inversion resistance. The framework remains detection in 6G-IoT. First, PCA reveals strong redundancy: 35 robust under poisoning (> 85% accuracy at 20% attackers) and components retain 99.3% variance with negligible performance efficient on edge devices (4.2 MB, 0.8 ms), though challenges loss, enabling efficient computation and communication. Secpersist in non-convex convergence, concept drift, and white-box ond, Random Forest outperforms SVM in stability–accuracy robustness. EdgeDetect thus establishes a strong privacy–utility– trade-off, achieving 98.0% accuracy and 97.9% macro F1 with efficiency trade-off for practical 6G-IoT deployment. very low variance (σ = 0.0017), indicating robustness for deployment. Third, imbalance handling is essential: SMOTE– XII. C ONCLUSION undersampling improves minority recall from 0.39 to 0.98, confirming class distribution as a first-order design factor. This paper introduced EdgeDetect, a privacy-preserving EdgeDetect introduces adaptive median-threshold smartification federated intrusion detection framework designed for resourcewith homomorphic encryption for federated IDS. Unlike constrained 6G-IoT environments. EdgeDetect employs gradisignSGD, it preserves gradient alignment (0.87 ± 0.04), achiev- ent smartification, a median-based binarization that compresses ing 96.9% communication reduction (450 MB→14 MB) while local updates to {+1, −1}, achieving a 32× communication XI. D ISCUSSION

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

reduction while maintaining convergence. Combined with Paillier homomorphic encryption, the framework ensures that only aggregated updates are revealed to the server, mitigating gradient inversion and honest-but-curious threats. Experiments on CIC-IDS2017 (2.8M flows, 7 attack classes) show that EdgeDetect achieves 98.0% accuracy and 97.9% macro F1, matching centralized performance while reducing per-round communication from 450 MB to 14 MB (96.9% reduction). Ablation analysis confirms that smartification enables efficient compression with negligible utility loss, encryption prevents gradient reconstruction (PSNR 15.1 dB vs. 31.7 dB undefended), and SMOTE significantly improves minorityclass recall. Under 5% poisoning and severe data imbalance, the system maintains 87% accuracy and 0.95 minority-class F1 (p < 0.001), demonstrating robustness for real-world deployment. Edge experiments on Raspberry Pi 4 further validate practicality, achieving a 4.2 MB memory footprint, 0.8 ms latency, and 12 mJ per inference with minimal accuracy degradation. Overall, EdgeDetect demonstrates that secure federated IDS can meet the strict privacy, efficiency, and reliability requirements of next-generation 6G-IoT edge networks. ACKNOWLEDGMENTS We thank the Canadian Institute for Cybersecurity for providing the CIC-IDS2017 dataset and the anonymous reviewers for their valuable feedback that improved this work. R EFERENCES [1] P. Kairouz, H. B. McMahan, B. Avent, A. Bellet, M. Bennis, A. N. Bhagoji, K. Bonawitz, Z. Charles, G. Cormode, R. Cummings et al., “Advances and open problems in federated learning,” Foundations and Trends in Machine Learning, vol. 14, no. 1–2, pp. 1–210, 2021. [2] Y. Liu, J. Zhang, and H. V. Poor, “Federated deep learning for intrusion detection with differential privacy,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 3291–3306, 2023. [3] Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated machine learning: Concept and applications,” ACM Transactions on Intelligent Systems and Technology, vol. 10, no. 2, pp. 1–19, 2019. [4] V. Mothukuri, P. Khare, R. M. Parizi, S. Pouriyeh, A. Dehghantanha, and G. Srivastava, “Federated learning-based anomaly detection for iot security attacks,” IEEE Internet of Things Journal, vol. 9, no. 4, pp. 2545–2554, 2021. [5] T. D. Nguyen, P. Rieger, M. Miettinen, and A.-R. Sadeghi, “Federated learning for intrusion detection system: Concepts, challenges and future directions,” Computer Networks, vol. 197, p. 108270, 2022. [6] X. Yin, Y. Zhu, and J. Hu, “A comprehensive survey of privacy-preserving federated learning: A taxonomy, review, and future directions,” ACM Computing Surveys, vol. 54, no. 6, pp. 1–36, 2023. [7] Y. Siriwardhana, P. Porambage, M. Liyanage, and M. Ylianttila, “Federated learning for 5g: A survey,” IEEE Communications Surveys & Tutorials, vol. 23, no. 3, pp. 1935–1962, 2021. [8] Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” Proceedings of the IEEE, vol. 107, no. 8, pp. 1738–1762, 2019, eXPERIMENTS: Edge computing evaluation framework. [9] F. Haddadpour, M. M. Kamani, A. Mokhtari, and M. Mahdavi, “Federated learning with compression: Unified analysis and sharp guarantees,” in International Conference on Artificial Intelligence and Statistics, 2021, pp. 2350–2358. [10] W. Y. B. Lim, N. C. Luong, D. T. Hoang, Y. Jiao, Y.-C. Liang, Q. Yang, D. Niyato, and C. Miao, “Federated learning in mobile edge networks: A comprehensive survey,” IEEE Communications Surveys & Tutorials, vol. 22, no. 3, pp. 2031–2063, 2020.

15

[11] Z. Chen, K. Zhang, M. Lu, Q. Zhu, and X. Zhang, “Privacy-preserving collaborative learning via automatic differential privacy budget allocation,” IEEE Transactions on Dependable and Secure Computing, vol. 21, no. 3, pp. 1456–1470, 2024. [12] J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar, “signsgd: Compressed optimisation for non-convex problems,” in International Conference on Machine Learning, 2018, pp. 560–569. [13] J. Chen and X. Ran, “Convergence of edge computing and deep learning: A comprehensive survey,” IEEE Communications Surveys & Tutorials, vol. 22, no. 2, pp. 869–904, 2020, eXPERIMENTS: Edge computing benchmarks. [14] O. Aouedi, K. Piamrat, G. Muller, and K. Singh, “Federated semisupervised learning for attack detection in industrial internet of things,” IEEE Transactions on Industrial Informatics, vol. 18, no. 5, pp. 3443– 3452, 2022. [15] K. Wei, J. Li, M. Ding, C. Ma, H. H. Yang, F. Farokhi, S. Jin, T. Q. Quek, and H. V. Poor, “Federated learning with differential privacy: Algorithms and performance analysis,” IEEE Transactions on Information Forensics and Security, vol. 15, pp. 3454–3469, 2020. [16] M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep learning with differential privacy,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 2016, pp. 308–318. [17] C. Ma, J. Li, M. Ding, B. Liu, K. Wei, J. Weng, and H. V. Poor, “Privacy-preserving federated learning based on multi-key homomorphic encryption,” International Journal of Intelligent Systems, vol. 37, no. 9, pp. 5880–5901, 2022. [18] S. Truex, N. Baracaldo, A. Anwar, T. Steinke, H. Ludwig, R. Zhang, and Y. Zhou, “A hybrid approach to privacy-preserving federated learning,” in Proceedings of the 12th ACM Workshop on Artificial Intelligence and Security, 2019, pp. 1–11. [19] W. Zhang, Y. Liu, T. Chen, and Q. Yang, “Secure and efficient federated learning via novel multi-key homomorphic encryption,” in USENIX Security Symposium, 2024, pp. 3421–3438. [20] K. Bonawitz, V. Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth, “Practical secure aggregation for privacy-preserving machine learning,” in Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 2017, pp. 1175–1191. [21] S. Rahman, I. Khalil, and M. Atiquzzaman, “Internet of things intrusion detection: Centralized, on-device, or federated learning?” IEEE Network, vol. 34, no. 6, pp. 310–317, 2020. [22] J. H. Bell, K. A. Bonawitz, A. Gašcón, T. Lepoint, and M. Raykova, “Secure single-server aggregation with (poly) logarithmic overhead,” in Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, 2020, pp. 1253–1269. [23] S. Kadhe, N. Rajaraman, O. O. Koyluoglu, and K. Ramchandran, “Fastsecagg: Scalable secure aggregation for privacy-preserving federated learning,” in Workshop on Federated Learning for User Privacy and Data Confidentiality, 2020. [24] C. Zhang, S. Li, J. Xia, W. Wang, F. Yan, and Y. Liu, “Batchcrypt: Efficient homomorphic encryption for cross-silo federated learning,” in Proceedings of the 2020 USENIX Annual Technical Conference, 2020, pp. 493–506. [25] D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “Qsgd: Communication-efficient sgd via gradient quantization and encoding,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 1709–1720. [26] J. Fang, Y. Wang, Y. Xu, and Q. Zhou, “Lightsecagg: A lightweight and versatile design for secure aggregation in federated learning,” Proceedings of Machine Learning and Systems, vol. 4, pp. 694–720, 2022. [27] J. So, B. Gürel, A. S. Amiri, B. Guler, and A. S. Avestimehr, “Turboaggregate: Breaking the quadratic aggregation barrier in secure federated learning,” IEEE Journal on Selected Areas in Information Theory, vol. 2, no. 1, pp. 479–489, 2021. [28] Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally, “Deep gradient compression: Reducing the communication bandwidth for distributed training,” in International Conference on Learning Representations, 2018. [29] A. Reisizadeh, A. Mokhtari, H. Hassani, A. Jadbabaie, and R. Pedarsani, “Fedpaq: A communication-efficient federated learning method with periodic averaging and quantization,” in International Conference on Artificial Intelligence and Statistics, 2020, pp. 2021–2031. [30] F. Sattler, S. Wiedemann, K.-R. Müller, and W. Samek, “Robust and communication-efficient federated learning from non-iid data,” IEEE Transactions on Neural Networks and Learning Systems, vol. 31, no. 9, pp. 3400–3413, 2019.

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

[31] X. Ma, L. Sun, and Y. Yao, “Efficient privacy-preserving federated learning with gradient compression,” IEEE Transactions on Information Forensics and Security, vol. 19, pp. 1123–1138, 2024. [32] W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “Terngrad: Ternary gradients to reduce communication in distributed deep learning,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 2055–2065. [33] D. Preuveneers, V. Rimmer, I. Tsingenopoulos, J. Spooren, W. Joosen, and E. Ilie-Zudor, “Distributed security framework for reliable threat intelligence sharing in federated deep learning,” Security and Communication Networks, vol. 2018, p. Article ID 6060253, 2018. [34] R. Zhao, Y. Wang, Z. Xue, T. Ohtsuki, B. Mao, N. Zhang, and H. Jiang, “Intelligent intrusion detection based on federated learning aided long short-term memory,” Physical Communication, vol. 42, p. 101157, 2020. [35] K. Bonawitz, H. Eichner, W. Grieskamp, D. Huba, A. Ingerman, V. Ivanov, C. Kiddon, J. Konečnỳ, S. Mazzocchi, B. McMahan et al., “Towards federated learning at scale: System design,” Proceedings of Machine Learning and Systems, vol. 1, pp. 374–388, 2019. [36] Q. Li, Y. Diao, Q. Chen, and B. He, “Federated learning on non-iid data silos: An experimental study,” in 2022 IEEE 38th International Conference on Data Engineering, 2020, pp. 965–978. [37] K. Cheng, T. Fan, Y. Jin, Y. Liu, T. Chen, D. Papadopoulos, and Q. Yang, “Secureboost: A lossless federated learning framework,” IEEE Intelligent Systems, vol. 36, no. 6, pp. 87–98, 2021. [38] H. Fereidooni, S. Marchal, M. Miettinen, A. Mirhoseini, H. Mollering, T. D. Nguyen, P. Rieger, A.-R. Sadeghi, T. Schneider, H. Yalame et al., “Safelearn: Secure aggregation for private federated learning,” in 2021 IEEE Security and Privacy Workshops, 2021, pp. 56–62. [39] H. Tang, S. Gan, C. Zhang, T. Zhang, and J. Liu, “Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression,” in International Conference on Machine Learning, 2019, pp. 6155–6165. [40] D. Rothchild, A. Panda, E. Ullah, N. Ivkin, I. Stoica, V. Braverman, J. Gonzalez, and R. Arora, “Fetchsgd: Communication-efficient federated learning with sketching,” in International Conference on Machine Learning, 2020, pp. 8253–8265. [41] P. Lu, Y. Wang, S. Li, H. Song, and D. Wang, “Adaptive gradient sparsification for efficient federated learning: An online learning approach,” IEEE Transactions on Neural Networks and Learning Systems, vol. 32, no. 12, pp. 5469–5481, 2020. [42] D. Basu, D. Data, C. Karakus, and S. Diggavi, “Qsparse-local-sgd: Distributed sgd with quantization, sparsification, and local computations,” IEEE Journal on Selected Areas in Information Theory, vol. 1, no. 1, pp. 217–226, 2020. [43] K. Mishchenko, E. Gorbunov, M. Takac, and P. Richtárik, “Distributed learning with compressed gradient differences,” arXiv preprint arXiv:1901.09269, 2019. [44] A. Koloskova, S. U. Stich, and M. Jaggi, “Decentralized deep learning with arbitrary communication compression,” in International Conference on Learning Representations, 2019. [45] S. Horváth, C.-Y. Ho, L. Horvath, A. N. Sahu, M. Canini, and P. Richtárik, “Natural compression for distributed deep learning,” Mathematical and Scientific Machine Learning, pp. 129–141, 2022. [46] H. Tang, X. Lian, T. Zhang, and J. Liu, “Communication-efficient distributed sgd with compressed sensing,” in International Conference on Machine Learning, 2021, pp. 10 259–10 269. [47] H. Jia, M. Yaghini, C. A. Choquette-Choo, N. Dullerud, A. Thudi, V. Chandrasekaran, and N. Papernot, “Proof-of-learning: Definitions and practice,” in 2021 IEEE Symposium on Security and Privacy, 2021, pp. 1039–1056. [48] S. P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. T. Suresh, “Scaffold: Stochastic controlled averaging for federated learning,” in International Conference on Machine Learning, 2020, pp. 5132–5143. [49] R. Xu, N. Baracaldo, Y. Zhou, A. Anwar, and H. Ludwig, “Hybridalpha: An efficient approach for privacy-preserving federated learning,” pp. 13–23, 2019. [50] J. Mills, J. Hu, and G. Min, “Communication-efficient federated learning for wireless edge intelligence in iot,” IEEE Internet of Things Journal, vol. 7, no. 7, pp. 5986–5994, 2019. [51] N. H. Tran, W. Bao, A. Zomaya, M. N. Nguyen, and C. S. Hong, “Federated learning over wireless networks: Optimization model design and analysis,” in IEEE INFOCOM 2019-IEEE Conference on Computer Communications, 2019, pp. 1387–1395. [52] Y. Deng, F. Lyu, J. Ren, H. Wu, Y. Zhou, Y. Zhang, and Y. Yang, “Adaptive scheduling for federated learning on resource-constrained edge devices,” IEEE Internet of Things Journal, vol. 7, no. 8, pp. 7942–7953, 2020.

16

[53] H. Wang, M. Yurochkin, Y. Sun, D. Papailiopoulos, and Y. Khazaeni, “Optimizing federated learning on non-iid data with reinforcement learning,” in IEEE INFOCOM 2020-IEEE Conference on Computer Communications, 2020, pp. 1698–1707. [54] H. Zhu, J. Xu, S. Liu, and Y. Jin, “Federated learning on non-iid data: A survey,” Neurocomputing, vol. 465, pp. 371–390, 2021. [55] D. Wu, F. Wang, Y. Cao, and J. Li, “Adaptive gradient quantization for privacy-preserving federated learning in iot networks,” IEEE Internet of Things Journal, vol. 11, no. 8, pp. 13 245–13 258, 2024. [56] M. S. Alam, M. R. Karim, and M. J. Hossain, “Intrusion detection using cnn-based image representation of network traffic,” IEEE Access, vol. 11, pp. 24 312–24 325, 2023. [57] N. Ghani, I. Ahmad, and M. K. Khan, “Explainable xgboost-based intrusion detection for network security,” Computers & Security, vol. 123, p. 102947, 2023. [58] D. Savić and M. Radovanović, “Lstm autoencoder-based network intrusion detection,” Journal of Network and Computer Applications, vol. 173, p. 102890, 2021. [59] G. Cerar and T. Zagar, “Anomaly detection using isolation forest for network intrusion detection,” Applied Sciences, vol. 10, no. 18, p. 6405, 2020. [60] S. Wang, T. Tuor, and T. Salonidis, “Adaptive federated learning in resource-constrained edge computing systems,” IEEE Journal on Selected Areas in Communications, vol. 40, no. 1, pp. 280–294, 2022. [61] C. Zhang, Y. Xie, and B. Li, “Federated learning with lstm for network intrusion detection,” IEEE Internet of Things Journal, vol. 9, no. 16, pp. 14 641–14 653, 2022. [62] M. Chen, W. Saad, and H. V. Poor, “Secure federated xgboost learning for iot intrusion detection,” IEEE Transactions on Information Forensics and Security, vol. 16, pp. 3674–3689, 2021.

A PPENDIX This appendix provides additional theoretical clarification of the proposed Gradient Smartification mechanism, its convergence properties relative to signSGD, and the adversarial threat model addressed by the combined binarization and encryption framework. TABLE XVIII: Hyperparameter configurations (nested 3-fold CV). Model

Configuration

LogReg SVM-RBF RF DT KNN GB MLP

C ∈ {0.1, 100}; solver={saga,sag}; ℓ2 C = 1.0; γ = 0.001 √ n ∈ {100, 200}; depth=20; m = ⌊ d⌋ depth∈ {6, 10, 15}; min split=5 k ∈ {3, 5, 7}; wt={uniform,dist} ν = 0.1; n = 100 [128,64]; drop=0.5; Adam(10−3 )

A. Relationship to signSGD Classical signSGD updates model parameters using the element-wise sign of stochastic gradients: W (r+1) = W (r) − η · sign(∇L(W (r) )).

(34)

In contrast, EdgeDetect applies a median-centered binarization:   (r) (r) (r) ∆i,bin = sign ∆i − θ , θ = median(|∆i |). (35) Unlike signSGD, which thresholds at zero, the proposed formulation suppresses low-magnitude gradient components whose absolute values fall below the median. This reduces stochastic noise and mitigates the influence of small-variance gradient coordinates common in high-dimensional IDS feature spaces.

EDGEDETECT: SECURE GRADIENT COMPRESSION FOR FEDERATED INTRUSION DETECTION

B. Key Distinction (r) ∆i

17

TABLE XIX: Computational Performance: Training Time and Inference Latency

Let’s = g + ϵ denote the true gradient g with Configuration Train Time (s) Inference (ms) Memory (MB) stochastic noise ϵ. Under zero-threshold binarization, small- Model noise perturbations may flip signs when |gj | is small. Median- Logistic Reg. C = 100, sag 2.4 0.12 45 rbf, C = 1, γ = 0.1 18.7 1.45 178 threshold binarization suppresses coordinates where |gj | < θ, SVM Random Forest T = 15, d = 8 12.3 0.87 234 reducing sign-flip probability and lowering gradient variance. Decision Tree d = 10 1.1 0.08 28 ∗ k = 7, distance-wt 0.3 3.21 412† Empirically, this improves convergence stability under hetero- KNN geneous client distributions. Notes: Benchmarked on Intel i7-9700K @ 3.6GHz, 32GB RAM, singleC. Convergence Sketch Assume: L(W ) is L-smooth, Stochastic gradients are unbi(r) ased: E[∆i ] = ∇L(W (r) ), Gradient variance is bounded: (r) E∥∆i − ∇L∥2 ≤ σ 2 . Under these conditions, signSGD achieves a convergence rate:   1 (r) 2 E[∥∇L(W )∥ ] = O √ . (36) r

threaded execution. Train Time includes hyperparameter search, crossvalidation, and final model fitting on the full training set (n = 12,000 for binary; n = 28,000 for multi-class). Inference was measured per sample on the test set. Memory = peak RAM consumption during training. ∗ KNN training is instantaneous (lazy learning) but requires † 412 MB to store all training instances for prediction.

F. Per-Class Performance Analysis Table XX reports detailed per-class results for the best Random Forest configuration (T = 15, depth=8). Errors are asymmetric across attack families, reflecting different separability in the PCA feature space.

Since Gradient Smartification preserves dominant gradient directions and discards only low-magnitude coordinates, the update remains directionally aligned with ∇L in expectation. Therefore, its convergence behavior asymptotically matches TABLE XX: Per-Class Performance Breakdown for Random signSGD under bounded noise assumptions. Full formal proof is Forest (Config 2) left for future work; empirical convergence curves in Section V Class True Pos. False Pos. False Neg. Precision Recall F1 support this claim. BENIGN 992 8 15 0.992 0.985 0.989 1) Proof Sketch of Theorem 1: We assume f is L-smooth. DoS 978 12 10 0.988 0.990 0.989 DDoS 975 14 11 0.986 0.989 0.987 By standard smoothness inequality, L f (wt+1 ) ≤ f (wt ) + ⟨∇f (wt ), wt+1 − wt ⟩ + ∥wt+1 − wt ∥22 . 2 Substituting the update rule wt+1 = wt − ηg̃t gives Lη 2 f (wt+1 ) ≤ f (wt ) − η⟨∇f (wt ), g̃t ⟩ + ∥g̃t ∥22 . 2 Taking expectation and applying Proposition 1,   E ⟨∇f (wt ), g̃t ⟩ ≥ γ∥∇f (wt )∥22 . Thus, Lη 2 E[f (wt+1 )] ≤ f (wt ) − ηγ∥∇f (wt )∥22 + d. 2 √ Choosing η = O(1/ T ) and summing over T iterations yields   1 min E∥∇f (wt )∥22 = O √ , t≤T γ T establishing convergence to a stationary point with a degradation factor 1/γ due to binarization. D. Bias and Stability of Gradient Smartification Median-threshold binarization introduces coordinate-wise bias: E[∆bin ] ̸= ∇L. (37) E. Computational Performance Analysis Table XIX quantifies training time and inference latency across all evaluated models, establishing practical feasibility for real-time intrusion detection deployment.

0.966 0.963 0.939 0.927

Port Scan Brute Force Web Attack Bot

935 928 885 863

42 48 78 94

23 24 37 43

0.957 0.951 0.919 0.902

0.976 0.975 0.960 0.953

Macro Avg.

6,556

296

163

0.956

0.975 0.966

Notes: Test set n = 7,000 (1,000 per class). True Pos./False Pos./False Neg. are computed one-vs-rest per class. Macro averages weight all classes equally.

1) Non-IID Data Distribution Analysis: Table XXI reports per-class F1 under increasing heterogeneity. Performance degrades smoothly as it α decreases, with minority/overlapping classes most affected. TABLE XXI: Per-Class F1-Scores Under Data Heterogeneity Attack Class BENIGN DoS DDoS Port Scan Brute Force Web Attack Bot Macro Avg. Accuracy

IID 0.989 0.989 0.987 0.966 0.963 0.939 0.927 0.966 98.0

α=10 0.987 0.988 0.986 0.964 0.961 0.936 0.923 0.964 97.8

α=1.0 0.983 0.985 0.982 0.957 0.952 0.924 0.908 0.956 96.8

α=0.5 0.978 0.981 0.976 0.948 0.941 0.908 0.889 0.946 95.7

α=0.1 0.971 0.974 0.968 0.934 0.923 0.881 0.854 0.929 94.2

Label Skew 0.984 0.987 0.981 0.961 0.956 0.929 0.918 0.959 97.1

Notes: α is the Dirichlet concentration parameter (smaller, α ⇒ higher heterogeneity). Label skew assigns each client 2–3 dominant classes with 70% probability. K = 50, C = 1.0, averaged over 5 runs.

Record · ID 18958 · SHA-256 b24b24ab6a6fb0be
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.