Conceptio › Archive › arXiv CS
arXiv CSopen access

An LLM-Assisted AutoML Framework for Intrusion Detection in IoT Networks

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

1

An LLM-Assisted AutoML Framework for Intrusion Detection in IoT Networks

arXiv:2609.23097v1 [cs.CR] 19 Sep 2026

Li Yang, Member, IEEE

Abstract—Internet of Things (IoT) systems are increasingly deployed in smart homes, transportation, energy systems, and critical infrastructure. This broad connectivity improves service intelligence, but also enlarges the attack surface of IoT networks. Machine Learning (ML)-based Intrusion Detection Systems (IDSs) are widely used to identify malicious network threats and protect IoT systems, but developing effective MLbased IDS models often requires human expertise and repeated manual decisions on many procedures, including data preprocessing, feature selection, model selection, and hyperparameter tuning. Automated Machine Learning (AutoML) reduces this burden by automating steps of the ML pipeline using optimization techniques, but conventional AutoML methods can consume substantial optimization time because they explore broad candidate model families and large hyperparameter spaces. This paper proposes a Large Language Model (LLM)-assisted AutoML framework for IoT intrusion detection. The proposed framework uses an LLM as a policy generator that converts dataset profiles into bounded and validated AutoML policies for automated data balancing, automated feature engineering, and Combined Algorithm Selection and Hyperparameter Optimization (CASH). Under an equal 10-trial budget, the proposed LLM-assisted policy achieves higher weighted test F1-score than traditional AutoML using the Tree-structured Parzen Estimator (TPE) on both datasets, reaching 99.680% on CICIDS2017 and 99.186% on IoTID20. Relative to the broader 30-trial Traditional AutoML-TPE baseline, the 10-trial proposed method reduces optimizer time by 63.7% and 49.9%, respectively, while achieving slightly higher F1-score. These results show that a bounded LLM policy can improve the quality of a low-budget AutoML search while retaining a clear efficiency advantage relative to a larger conventional search budget. Index Terms—AutoML, cybersecurity, Internet of Things, intrusion detection system, large language models, machine learning.

I. I NTRODUCTION

T

HE Internet of Things (IoT) refers to the integration of sensing, computing, and communication capabilities into physical devices that can collect and exchange data over networked environments [1], [2]. IoT has enabled many important applications, such as smart homes, intelligent transportation, smart grids, and critical infrastructure management [3]. However, the increased connectivity that supports IoT applications also increases cybersecurity risks. Resource-constrained IoT devices often have limited processing power and security mechanisms, making them vulnerable to various cyber-attacks, such as Denial of Service (DoS), botnets, scanning, and Li Yang is with the Faculty of Business and Information Technology, Ontario Tech University, Oshawa, ON L1G 0C5, Canada, and also with the Department of Electrical and Computer Engineering, Western University, London, ON N6A 3K7, Canada (e-mails: [email protected]; [email protected]).

Man-in-the-Middle (MITM) attacks [1], [4]. For example, in Electric Vehicle Charging Systems, connected chargers, backend servers, and smart grid interfaces improve charging intelligence, but they also expose charging operations and user information to cyber-attacks [5], [6]. Intrusion Detection Systems (IDSs) are essential security mechanisms for IoT systems because they analyze network traffic and detect malicious attacks that bypass preventive controls such as authentication and firewalls [7], [8]. Machine Learning (ML)-based IDSs have become a dominant approach because ML models can learn complex relationships between network-traffic features and attack classes and generalize these patterns to unseen samples [5], [9]. However, developing an effective IDS remains challenging. IDS datasets are usually high-dimensional, multi-class, and highly imbalanced, and model performance usually depends strongly on the ML development process, including data pre-processing, feature selection, model selection, and hyperparameter tuning [7]. Traditional ML development is constrained by human expertise and design bias, which can lead to suboptimal performance and insufficient generalizability. Automated Machine Learning (AutoML) has emerged as a promising approach to address these development challenges because it can automate key stages of the ML pipeline, including Automated Data Pre-processing (AutoDP), Automated Feature Engineering (AutoFE), automated model selection, and Hyperparameter Optimization (HPO) [7]. This makes AutoML attractive for autonomous cybersecurity, especially in future networks where zero-touch networking principles aim to reduce human intervention in monitoring, optimization, and protection [7]. The joint problem of automated model selection and HPO is commonly formulated as Combined Algorithm Selection and Hyperparameter Optimization (CASH), a central problem in AutoML [3], [10]. Nevertheless, conventional AutoML also has limitations. The AutoML search process can be computationally expensive due to broad search spaces, but defining an effective search space remains a difficult problem and often requires domain knowledge [7]. A generic AutoML process may spend many trials evaluating candidate models or hyperparameter regions that are unlikely to provide a favorable trade-off for a specific network security dataset. This challenge is especially important for IDS development in IoT systems, where efficient and resource-aware model construction is often needed [3]. Recent work has begun to use Large Language Models (LLMs) at specific stages of ML automation. Context-Aware Automated Feature Engineering (CAAFE) uses an LLM to generate context-aware feature transformations for tabular data

2

[11], while Large Language Models to Enhance Bayesian Optimization (LLAMBO) integrates LLM guidance into Bayesian optimization and hyperparameter search [12]. Broader studies have also examined LLMs as assistants for ML workflow design and AutoML [13], [14]. These results motivate the use of LLMs for search-space design, but unrestricted agents may also produce invalid code, unsupported operations, or inconsistent configurations [15]. For IDS development, we therefore restrict the LLM to a fixed AutoML policy schema. It proposes bounded choices from a compact dataset profile, while validated ML components perform data processing, model training, and evaluation. The central question in this work is whether such a bounded policy can outperform a conventional predefined AutoML policy under the same lowbudget setting, and whether the resulting compact search can remain competitive with a broader high-budget search. Therefore, this paper proposes an LLM-assisted AutoML framework for IoT intrusion detection. A local Llama 3.1 8-billion-parameter (8B) model [16] receives an aggregate dataset profile and returns a structured policy for AutoDPbased balancing, AutoFE, and CASH optimization. The policy is checked against predefined operators, parameter bounds, and fallback rules before it is executed. Model selection remains empirical: the validated pipeline is evaluated by stratified cross-validation on the training partition, and the hold-out set is used only for the final assessment. Local execution keeps the profiling information within the experimental environment and allows the same LLM version and decoding settings to be reused across experiments. In this design, the LLM shapes a compact AutoML policy but does not replace the deterministic training and validation procedures. Prior studies have examined LLM-assisted feature engineering and hyperparameter optimization, whereas this work focuses on a validated local policy interface that jointly coordinates AutoDP, AutoFE, and CASH for IoT IDS development. To the best of our knowledge, no prior IoT IDS framework has combined these three AutoML stages through a bounded LLM policy while leaving model fitting and scoring to deterministic and validation-driven procedures. The main contributions of this paper are summarized as follows: 1) An LLM-assisted AutoML framework1 is proposed for IoT intrusion detection, in which a local LLM generates bounded policies for AutoDP, AutoFE, and CASH optimization. The framework is designed to improve the quality of a low-budget AutoML search while reducing unnecessary exploration of a broad search space. 2) A validated policy-execution mechanism is developed to convert LLM-generated suggestions into executable AutoML configurations, preventing unsupported model families and invalid ranges. 3) A detailed AutoML pipeline is designed using partial adaptive oversampling, multi-stage feature selection with retention safeguards, and CASH based on Bayesian Optimization (BO) with the Tree-structured Parzen Estimator (TPE), denoted BO-TPE, using a com1 The complete code will be available upon paper acceptance at: https:// github.com/LiYangHart/LLM-Assisted-AutoML-For-Intrusion-Detection

pact gradient-boosted-tree search policy for the proposed branch. 4) The proposed framework is evaluated on public benchmark cybersecurity datasets, CICIDS2017 [17] and IoTID20 [18], and compared with traditional ML baselines, state-of-the-art AutoML and optimized ML models, and ablated LLM-assisted variants. The remainder of this paper is organized as follows. Section II reviews optimized and automated ML-based IDSs and recent LLM-assisted AutoML research. Section III presents the proposed LLM-assisted AutoML framework in detail. Section IV describes the experimental setup and discusses the results. Section V concludes the paper and outlines future research directions. II. R ELATED W ORK A. Optimized and Automated ML-based IDSs ML-based IDSs have been widely used for network and IoT security [7], [19]. Supervised learning methods distinguish benign and malicious traffic by learning decision boundaries from labeled network-traffic records [7]. Tree-based models are frequently adopted because network traffic datasets are usually tabular and nonlinear, while recent evidence also shows that tree-based models often remain highly competitive with deep learning on typical tabular data [20]. Jin et al. [21] proposed a real-time IDS that uses the Light Gradient Boosting Machine (LightGBM) with a parallel intrusion detection mechanism to improve detection efficiency on large-scale traffic data. Fernando et al. [22] introduced a new IoT intrusiondetection dataset and examined ML model generalization for IoT communication intrusion detection. Deep learning methods can also achieve strong detection performance, although tree-based models often remain competitive on tabular datasets and generally require less architectural design [20], [23]. Optimized ML-based IDSs improve conventional ML by tuning models, selecting features, or constructing ensembles. Elmasry et al. [24] proposed a double Particle Swarm Optimization (PSO) method to evolve the Long Short-Term Memory (LSTM) model for network intrusion detection. Naeem et al. [25] developed a multi-class vehicle-security framework that combines deep transfer learning-based Convolutional Neural Network (CNN) with Genetic Algorithm (GA)-based optimization. AutoML-based IDSs further reduce manual effort by automating the ML model optimization procedures. Khan et al. [26] proposed the Optimized Ensemble Intrusion Detection System (OE-IDS), an AutoML-based soft-voting ensemble model for network intrusion detection. Singh et al. [27] proposed Automated Machine Learning for Intrusion Detection (AutoML-ID), which uses AutoML to select and tune IDS models for wireless sensor networks. Yang and Shami [28] proposed an AutoML-based autonomous IDS framework that integrates automated data pre-processing, feature engineering, model selection, and hyperparameter optimization for intrusion detection. Although optimized ML-based and traditional AutoMLbased IDSs have improved detection performance, several limitations remain. Many optimized IDSs still rely on manually

3

TABLE I C OMPARISON OF RELATED IDS AND AUTO ML RESEARCH CATEGORIES . Category

Related Works

Main Contribution

Main Limitation / Novelty

Traditional ML-based IDSs

[21], [22] Detect attacks from network traffic using manually designed ML models.

Require manual pre-processing, model selection, and tuning; provide the IDS foundation.

Optimized ML-based IDSs

[24], [25] Improve IDS performance through feature selection, ensembles, or HPO.

Search design and objectives are usually manually defined; motivate optimized IDS development.

Traditional AutoMLbased IDSs

[27], [28] Automate model selection, HPO, and other ML pipeline stages.

Broad search spaces may increase runtime; provide the AutoML backbone.

LLM[11], [12] Use LLMs to suggest assisted ML workflow and ML/AutoML optimization choices.

Free-form LLM outputs require validation; motivate bounded policy generation.

Proposed framework

Integrates LLM guidance with validated AutoML execution for efficient and reproducible IDS development.

This paper

Uses local LLM policy generation for AutoDP, AutoFE, and CASH in IDS.

defined search spaces, while general AutoML methods may explore broad candidate spaces and consume considerable computational time. Moreover, most existing AutoML-based IDSs mainly rely on empirical search and provide limited support for using dataset characteristics to guide the AutoML process before optimization. These limitations motivate an LLM-assisted AutoML design in which a compact dataset profile is used to propose a targeted and validated search policy for IDS development. B. LLM-Assisted ML Workflow Design LLMs have recently been studied for data analysis, ML workflow construction, and AutoML search [13], [14]. Two direct examples are CAAFE, which uses an LLM to iteratively generate context-aware feature-engineering code for tabular data [11], and LLAMBO, which introduces LLM-based components for warm-starting, surrogate modeling, and candidate sampling in Bayesian optimization [12]. These methods show that LLMs can contribute useful prior information to individual stages of an AutoML process. However, free-form feature generation or direct LLM control of an optimization loop is difficult to validate in a security experiment, where data partitioning, pre-processing, and model scoring must remain reproducible. LLMs have also been integrated more directly into intrusion-detection and security-analysis workflows. IDSAgent uses an LLM-driven agent with specialized IDS tools, memory, and external knowledge to reason over network traffic and produce explainable intrusion decisions [29]. HuntGPT combines a Random Forest anomaly detector with explainableAI components and an LLM-based conversational interface

to present detected threats in an analyst-readable form [30]. These systems place the LLM in the detection or interpretation loop. In contrast, the proposed framework keeps the deployed IDS conventional: the LLM is used only during offline AutoML planning, does not inspect hold-out or live traffic, and does not produce final intrusion labels. Instead, it operates at the policy level, jointly guiding data balancing, feature selection, and CASH within a predefined and validated decision space, while deterministic ML procedures perform model fitting, validation, and final prediction. This design allows several interconnected AutoML stages to be adapted together rather than optimizing an isolated component, while also reducing unnecessary search and improving the efficiency of model development. The use of a local LLM further limits external data exposure and avoids dependence on online services. Overall, the proposed framework combines flexible LLM-assisted policy generation with a controlled and practical AutoML pipeline for IoT intrusion detection. Table I places this design relative to existing IDS and AutoML research. III. P ROPOSED LLM-A SSISTED AUTO ML F RAMEWORK A. System Overview The proposed framework uses LLM-guided AutoML planning to construct an optimized IDS with a smaller search cost. Rather than allowing the LLM to generate executable programs, its output is restricted to a bounded AutoML policy. All data processing, feature selection, model fitting, and scoring are performed by predefined operators. The LLM policy is returned as JavaScript Object Notation (JSON). Within deterministic AutoML execution, AutoDP can use the Synthetic Minority Over-sampling Technique (SMOTE), Adaptive Synthetic Sampling (ADASYN), random oversampling, or no balancing. AutoFE combines nearconstant feature removal, a Mutual Information (MI) prefilter, LightGBM gain-based importance selection, and Pearson correlation pruning. These operators and their policy-controlled parameters are shown explicitly in Fig. 1. Fig. 1 summarizes the workflow in three main stages. In Stage 1, the labeled traffic data are divided into an outer training partition and an independent hold-out set. Aggregate statistics from the training partition are provided to a local Llama 3.1 8B model [16], which returns a structured AutoML policy. The policy is validated against an admissible policy space, and unsupported choices are rejected or replaced by deterministic defaults. The hold-out set is excluded from policy fitting and model selection. Stage 2 performs deterministic AutoML execution within stratified cross-validation on the outer training partition. First, label encoding and z-score scaling are fitted on each foldtraining partition and applied to its paired validation fold. Second, AutoDP partially rebalances the fold-training data using the validated sampler, target class-count quantile, and perclass oversampling cap. Third, AutoFE applies near-constant removal, MI prefiltering, LightGBM gain-based selection, and Pearson correlation pruning with a final-feature retention safeguard. The fitted transformations and selected features are then

4

Network Traffic

Network data

1

Local LLM policy generation

Aggregate training statistics guide a bounded policy

CICIDS2017 IoTID20 Features + labels

Dataset profile

Local Llama 3.1 8B

Policy validation

Sample count Feature dimensions Class-count vector

Profile + fixed schema Temperature = 0 Structured JSON policy

Allowlisted operators and models Bounded ranges and trial budget Defaults and feasibility safeguards

Stratified split

Data partition

2

Training set / Test set

Validated component policies

Training partition

Deterministic AutoML execution

Stratified cross-validation within the training partition

Excluded from model selection and policy fitting

Hold-out test set

Preprocessing

AutoDP

AutoFE

Cached folds

Label encoding Z-score scaling

Partial balancing SMOTE / ADASYN Random / none

Near-constant removal MI prefilter + gain importance Pearson correlation pruning

Training / validation Fitted transforms Selected features

LLM-guided controls Sampling method Target quantile + oversampling cap

LLM-guided controls MI width + gain threshold Feature cap + retention floor Correlation threshold

Traditional AutoML Process

Reserved test data

Fit on fold-training data only; oversample training data only

CASH

Data / execution

Trial parameters

ML model selection LightGBM hyperparameter tuning by BO-TPE

Policy control

Candidate evaluation Mean weighted validation F1

LLM: bounded hyperparameter ranges

Validation-score feedback

Select the highest mean weighted F1

3

Final training and independent evaluation Refit selected pipeline

One-time hold-out test

IDS output

Full outer training set Selected features + best model

Accuracy + Precision + Recall + F1 Optimization and prediction time

Benign traffic or specific attack type

Apply fitted pipeline to test data

Fig. 1. Overview of the proposed LLM-assisted AutoML IDS.

cached with the paired training and validation folds. Fourth, CASH evaluates the validated LightGBM search policy with BO-TPE on these cached folds and selects the configuration with the highest mean weighted validation F1-score. In Stage 3, the selected pipeline is refitted on the full outer training set and evaluated once on the untouched hold-out set. Thus, the LLM provides bounded policy guidance, whereas data transformation, model fitting, and model selection remain deterministic or validation-driven. Table II summarizes the main components. Let D = {(xi , yi )}ni=1 denote a labeled traffic dataset, where xi ∈ Rd is a d-dimensional feature vector and yi ∈ C is its class label. The proposed framework searches for an IDS classifier f (x; θ) that achieves high weighted F1-score on unseen traffic while limiting unnecessary AutoML evaluations. The LLM-assisted AutoML policy is represented as:

per-class oversampling factor rmax . The vector ϕ denotes the feature-selection controls, A is the validated candidate model set, Ω is the corresponding conditional hyperparameter space, T is the optimization budget, and Π is the admissible policy space. The selected IDS model is obtained by maximizing the mean weighted F1-score under stratified cross-validation: K  1 X w (k) (k) F1 a, λ, Dtrain , Dvalid , (2) a∗ , λ∗ ∈ argmax a∈A,λ∈Ωa K

π = {β, ϕ, A, Ω, T } ∈ Π,

B. LLM Policy Generation and Validation The LLM-assisted stage begins by summarizing the outer training partition Dtrain into a compact dataset profile P. The

(1)

where β = {b, q, rmax } contains the AutoDP controls: balancing method b, target class-count quantile q, and maximum

k=1

(k) (k) where Dtrain and Dvalid are the training and validation folds

in the k-th cross-validation split, respectively. Equation (2) clarifies the role of the LLM: the LLM defines the AutoML policy, while the final model and its hyperparameters are selected according to measured validation performance.

5

TABLE II S UMMARY OF THE PROPOSED LLM- ASSISTED AUTO ML FRAMEWORK . Stage

Objectives

LLM-Assisted Decision

Dataset profiling

Compute sample count, feature types and dimensions, and class counts from the training partition.

Provides aggregate evidence for AutoML policy generation.

ranges. The raw JSON is never executed directly; it first passes through the validation layer described below. The generated policy object is defined as   G = {b, q, rmax },    AutoDP  G = GAutoFE = {τcorr , γimp , mimp , kMI , mmin } , , (4)     GCASH = {A, Ω, T }

where b, q, and rmax denote the balancing method, target quantile, and per-class oversampling cap, respectively. The AutoFE controls are the correlation threshold τcorr , cumulative AutoDP Partially rebalance Selects the sampler, target gain-importance threshold γimp , feature cap mimp , mutualbalancing minority classes within class-count quantile, and per-class information width kMI , and minimum final-feature count the training partition. oversampling cap. mmin . The variables A, Ω, and T denote the validated candiAutoFE Remove weak, irrelevant, Selects the mutual-information and redundant features width, importance threshold, date set, conditional hyperparameter space, and optimization while preserving a feature cap, retention floor, and budget. minimum feature set. correlation threshold. The validation layer converts the JSON response into an exCASH Optimize the validated Provides bounded search-space ecutable policy by applying an allowlist and numeric bounds. model family and its guidance; the proposed branch hyperparameters. uses a compact The balancing method is restricted to SMOTE, ADASYN, LightGBM-centered space and a random oversampling, or no balancing [31], [32]; the target small trial budget. quantile and oversampling factor are clipped to bounded ranges. AutoFE thresholds and feature counts are likewise constrained, including the minimum final-feature count. Model profile includes the number of samples, feature dimensionality, names are limited to LightGBM, random forest, and extra numbers of numerical and categorical features, and the classtrees, and invalid or missing fields are replaced by determincount vector. It is expressed as istic defaults. For the proposed low-budget CASH branch, the P(Dtrain ) = {n, d, dnum , dcat , cy }, cy = [|{i : yi = c}|]c∈C , executable policy uses LightGBM as the model family and (3) a compact predefined search envelope. If the LLM provides where n is the number of samples, d is the total number of model-specific LightGBM ranges, their intersection with this input features, dnum and dcat are the numbers of numerical envelope is used; an invalid or missing intersection leaves the and categorical features, respectively, cy is the class-count predefined range unchanged. Additional feasibility checks are applied during execution. vector, and C denotes the set of traffic classes. The classThe number of neighbors for SMOTE or ADASYN is reduced count vector characterizes the degree of imbalance, while when a sampled class is small, and random oversampling the dimensionality and feature-type counts provide evidence is used if synthetic sampling cannot be performed. AutoFE for feature-selection and search-space decisions. The profile also enforces feature-retention safeguards before and after contains aggregate information only; raw traffic records are correlation pruning. These rules prevent malformed LLM outnot supplied to the LLM. put or small-class geometry from changing the experimental A local Llama 3.1 8B model is used as the policy generator protocol. [16]. The model receives the aggregate profile together with a fixed JSON schema and is queried with temperature set to zero. No task-specific fine-tuning is performed. The prompt C. LLM-Assisted Automated Data Balancing explicitly limits the LLM to policy generation: it does not Before data balancing and model learning, categorical traffic access raw samples, score candidate models, or produce final attributes are converted into numerical values using label IDS labels. The 8B variant was selected because it can be encoding [33]. For each categorical feature, an encoder is fitted executed locally on the experimental workstation while follow- on the active training partition and then applied to the paired ing structured instructions. Local execution keeps the profile validation or hold-out partition. Previously unseen categories and experiment metadata within the experimental environment, are mapped to a dedicated unknown token. This transformation avoids dependence on an external LLM service, and allows a allows tree-based models and feature selectors to process catfixed model version and decoding configuration to be used egorical fields such as addresses, ports, or protocol identifiers throughout the study. without expanding the feature space through one-hot encoding. The returned policy has three components: AutoDP, Aut- After encoding, all numerical features are normalized by zoFE, and CASH. AutoDP specifies the sampler together with score scaling. For a feature value x, the normalized value is the target class-count quantile and a per-class oversampling [33]: x − µtr cap. AutoFE specifies the Pearson correlation threshold, cumu, (5) x′ = σtr lative gain-importance threshold, maximum retained-feature count, mutual-information prefilter width, and minimum final- where µtr and σtr are the mean and standard deviation estifeature count. CASH specifies candidate model information, mated from the active training partition. The same parameters a trial budget, and optional model-specific hyperparameter are then applied to the paired validation or hold-out data. Encoding and scaling

Encode categorical predictors and scale numerical features.

Uses a fixed process for consistency and does not participate in model scoring.

6

Numerical features are standardized before balancing so that distance-based oversampling is not dominated by features with larger numerical ranges. Applying the same label encoding and z-score normalization protocol to all comparison methods ensures that the observed differences come from balancing, feature selection, and CASH policies rather than inconsistent pre-processing. Network intrusion datasets are often highly imbalanced because normal traffic and frequent attack categories may dominate the data, whereas rare attacks contain far fewer samples [5]. The AutoDP module computes the class-count vector c on the active training partition and the imbalance ratio max(c) . (6) ρ= min(c)

width, cumulative importance threshold, feature cap, minimum final-feature count, and correlation threshold. The stages serve different purposes: variance filtering removes nearly constant inputs, mutual information provides a fast relevance screen, LightGBM gain importance provides model-aware ranking, and correlation pruning removes redundant retained features. The first step removes near-constant features using a variance threshold of 10−8 . Let Fv denote the feature set remaining after variance filtering, and let dv = |Fv |. The second step applies a mutual-information prefilter. Let kMI denote the LLM-selected mutual-information width and mimp denote the maximum number of important features. The prefilter width is defined as:

No resampling is performed when ρ ≤ 1.5. Otherwise, AutoDP performs partial balancing rather than forcing every minority class to the majority count. For a validated target quantile q, the common target is

If dv > mMI , only the top mMI features ranked by mutual information with the class label are retained. The third step trains a LightGBM classifier on the training partition and ranks features using gain-based importance. Let qj denote the raw LightGBM gain importance of feature j. The normalized importance score is defined as [28]: qj , j ∈ FMI , (10) Ij = P r∈FMI qr

tq = min (max(c), max (Qq (c), min(c) + 1)) ,

(7)

where Qq (c) is the q-quantile of the class-count distribution. In the conventional AutoDP setting, q = 0.75 and no perclass growth cap is applied. In the LLM-assisted setting, the policy also provides rmax ; therefore, a class j with count cj is assigned tj = min (tq , ⌊rmax cj ⌋) , (8) and it is oversampled only when cj < tj . This cap is particularly useful for very small classes because it avoids a large increase in the training-set size. The policy chooses among SMOTE, ADASYN, random oversampling, and no balancing. SMOTE interpolates between neighboring minority samples [31], whereas ADASYN allocates more synthetic samples to locally difficult minority regions [32]. For SMOTE and ADASYN, the neighborhood size is reduced when a sampled class contains only a few examples. If fewer than two samples are available, or if synthetic sampling otherwise fails, random oversampling is used as a deterministic fallback. All resampling is fitted on fold-training data only; validation and hold-out samples are never used to construct synthetic data. D. LLM-Assisted Automated Feature Engineering Feature engineering aims to improve the feature set used by ML models. In network intrusion detection, raw datasets often contain weakly informative, redundant, or near-constant features. Feature selection, a major component of feature engineering, removes unnecessary features while retaining those important for intrusion detection [3], [7]. This step is important for IoT IDS development because fewer features can reduce data complexity, model execution time, and the risk of overfitting. The proposed AutoFE process has four stages: near-constant feature removal, mutual-information prefiltering, LightGBM gain-based cumulative importance selection, and Pearson correlation pruning. All selectors are fitted on the active training partition only. The LLM controls the mutual-information

mMI = min {max (3mimp , kMI , 60) , dv } .

(9)

where FMI is the feature set retained after mutual-information prefiltering. Features are sorted in descending order of Ij , and the smallest subset whose cumulative importance reaches the LLM-selected threshold γimp is retained:   Fimp = {j1 , j2 , . . . , jm∗ } ,  ( ) m X (11) ∗  Ijℓ ≥ γimp .   m = min m : ℓ=1

where Ij1 ≥ Ij2 ≥ · · · after sorting. In implementation, the cumulative threshold is evaluated with a single LightGBM fit, and the retained count is also bounded by mimp . To avoid an overly aggressive reduction, the importance stage retains at least mimp floor = min (mimp , max (30, ⌈0.8mimp ⌉ , mmin ))

(12)

features when the importance selector is active. If the MIreduced feature set already contains no more than the requested feature cap, the importance-selection step is skipped. This single-fit selector avoids the repeated model fitting required by recursive elimination. The final step removes redundant features using the absolute Pearson correlation matrix estimated from the training partition. For two selected features p and q, the absolute correlation is computed as Rpq = |corr (xp , xq )| .

(13)

If Rpq > τcorr , one of the two features is removed. The LLMassisted policy includes a retention floor mmin . If correlation pruning leaves fewer than mmin features, features are restored from the pre-correlation subset until the floor is met; if no feature remains, the pre-correlation subset is retained. This safeguard is applied after the MI and gain-importance stages

7

so that correlation pruning cannot collapse an already compact feature set. Thus, the LLM-assisted AutoFE policy adapts five controls to the dataset profile: kMI , γimp , mimp , mmin , and τcorr . The deterministic bounds are intentionally conservative for highdimensional data, so the LLM can reduce redundant inputs without selecting an impractically small final feature set. E. LLM-Assisted Automated Model Learning and CASH Optimization After pre-processing, balancing, and AutoFE, the framework performs BO-TPE-based model optimization. LightGBM, Random Forest (RF), and Extra Trees (ET) are retained in the conventional ML and AutoML comparisons because they are strong and efficient models for tabular network-flow data. The proposed low-budget LLM-assisted CASH branch uses LightGBM as its executable model family and optimizes it within a compact validated hyperparameter space. This keeps the model family fixed across LLM-assisted trials while allowing the LLM policy to guide the search bounds and budget. RF constructs an ensemble of decision trees trained on bootstrapped samples and random subsets of features [22]. Its prediction is obtained by aggregating the votes of individual trees. RF is robust, relatively insensitive to feature scaling, and effective for high-dimensional tabular data. It is used as a strong plain ML and standard AutoML baseline because it often performs well on IDS datasets with limited tuning. ET further randomizes both feature selection and split thresholds [22]. This stronger randomization can reduce variance and training cost, but it may require a sufficiently large ensemble to achieve the best performance. In the experiments, ET provides an efficiency-oriented baseline and helps evaluate whether the additional randomness improves or degrades IDS classification under the selected datasets. LightGBM [34] is a gradient-boosted decision-tree algorithm designed for efficient large-scale learning. It grows boosted trees sequentially, where each new tree corrects the errors of the current ensemble. LightGBM improves efficiency through histogram-based binning, Gradient-based One-Side Sampling (GOSS), and Exclusive Feature Bundling (EFB). These properties make it especially suitable for tabular network traffic with many numerical features. In the proposed LLM-assisted branch, LightGBM is selected as the final model family because it provides a favorable balance between detection performance and search efficiency. The learning objective of a gradient-boosted tree ensemble can be written as [34]: ŷi =

M X

fm (xi ),

fm ∈ F,

(14)

m=1

L=

n X i=1

ℓ(yi , ŷi ) +

M X

R(fm ).

(15)

m=1

Here, fm is an individual tree, ℓ(·) is the classification loss, and R(·) is a regularization term that penalizes overly complex trees. In LightGBM, important hyperparameters include the

number of estimators, maximum tree depth, learning rate, number of leaves, minimum child samples, subsampling ratio, feature subsampling ratio, and L1/L2 regularization terms. In RF and ET, important hyperparameters include the number of trees, maximum depth, minimum split size, minimum leaf size, and maximum feature ratio. BO-TPE is selected as the HPO method because it is well suited to mixed and conditional search spaces [35]. Unlike grid search, BO-TPE does not exhaustively evaluate every combination. Unlike random search, it uses previous observations to guide future trials. Compared with Gaussian-process Bayesian optimization, TPE handles categorical, integer, and conditional parameters more naturally. This is important for CASH because each model family has a different hyperparameter space. TPE models the density of good and poor configurations separately. Let y be the validation loss or negative validation score, and let y ∗ be a quantile threshold that separates promising configurations from weaker ones. TPE estimates ( l(λ), y < y ∗ , p(λ|y) = . (16) g(λ), y ≥ y ∗ New configurations are sampled to increase the ratio l(λ)/g(λ), which favors regions that are more likely under good trials than under poor trials. In this work, Optuna implements the BO-TPE optimizer [36]. The objective function is the mean weighted F1-score under stratified cross-validation. K

J(a, λ) =

1 X w (k) (k) F1 (a, λ, Dtr , Dval ). K

(17)

k=1

The LLM-assisted CASH stage is configured before BOTPE begins. The prompt allows bounded model-specific ranges, but the executable policy fixes LightGBM for the proposed branch to keep the optimization cost stable. A predefined compact LightGBM envelope provides the safe search space; when the LLM returns valid LightGBM bounds, they are intersected with that envelope, and missing or invalid bounds leave the predefined ranges unchanged. The proposed branch uses a 10-trial budget, whereas the standard CASH baseline searches LightGBM, RF, and ET with 30 trials. BOTPE and cross-validation, rather than the LLM, determine the final hyperparameters from measured validation F1-score. F. Proposed Framework Summary and Deployment Algorithm 1 summarizes the complete training and evaluation procedure. The framework first creates an independent outer train–test split, constructs the dataset profile from the training partition, and converts the local LLM output into a validated policy for AutoDP, AutoFE, and CASH. The deterministic AutoML stage then prepares the stratified crossvalidation folds by fitting encoding and scaling, applying partial AutoDP balancing to fold-training data only, and fitting the four-stage AutoFE process before caching the prepared folds. BO-TPE evaluates the validated CASH space on these cached folds and selects the configuration with the highest mean weighted validation F1-score. Finally, the selected pipeline is refitted on the full outer training partition, and the hold-out test set is used once for independent evaluation.

8

Algorithm 1 Proposed LLM-Assisted AutoML IDS Framework 1: Input: Labeled dataset D = {X, Y }; local LLM MLLM ; admissible policy space Π; number of CV folds K. 2: Output: Optimized IDS M∗ ; validated policy π; selected features F ∗ ; best configuration (a∗ , λ∗ ); evaluation results R. 3: // Step 1: LLM-assisted AutoML policy generation and validation 4: Split D → {Dtrain , Dtest } using a stratified outer split; reserve Dtest for final evaluation. 5: P ← {n, d, dnum , dcat , cy } from Dtrain only. 6: G ← MLLM (P). {LLM proposes bounded AutoDP, AutoFE, and CASH settings} 7: π ← Validate(G, Π). {Deterministic validation corrects invalid or missing settings} 8: Extract β = {b, q, rmax }, ϕ = {kMI , γimp , mimp , τcorr , mmin }, and {A, Ω, T } from π. 9: // Step 2: LLM-guided, deterministic AutoDP and AutoFE (k) (k) 10: Create K stratified folds {(Dtr , Dval )}K k=1 from Dtrain . 11: for k = 1, . . . , K do (k) 12: Fit label encoding and z-score scaling on Dtr ; transform both fold partitions. {Fixed preprocessing; no LLM decision} (k) (k) 13: ρ(k) ← maxj cj / minj cj . (k) 14: if ρ > 1.5 and b ̸= None then 15: t(k) ← Targets(c(k) , q, rmax ). (k) (k) 16: Dtr ← AutoDP(Dtr , b, t(k) ). {LLM selects bounded balancing policy; operator performs resampling} 17: Adjust neighbor count when needed; use random oversampling if SMOTE/ADASYN is infeasible. 18: end if (k) (k) 19: Fv ← VarianceFilter(Dtr ). (k) (k) 20: FMI ← TopMI(Fv , kMI ). (k) (k) 21: Fimp ← GainSelect(FMI , γimp , mimp ). (k)

F (k) ← CorrPrune(Fimp , τcorr , mmin ). {LLM sets AutoFE controls; feature selection is fitted on training data} 23: Apply F (k) to the training and validation partitions and cache the prepared fold V (k) . 24: end for 25: // Step 3: LLM-constrained CASH optimization using BO-TPE 26: Initialize BO-TPE over {(a, Ωa ) : a ∈ A} with trial budget T . {LLM defines the bounded search policy; BO-TPE performs empirical search} 27: for t = 1, . . . , T do 28: (at , λt ) ← TPE(A, Ω). 29: for k = 1, . . . , K do 30: Train (at , λt ) on the cached training data in V (k) . (k) 31: st ← F1w (at , λt ; V (k) ) on the paired validation data. 32: end for P (k) K 1 33: s̄t ← K k=1 st ; update BO-TPE using s̄t . 34: end for 35: (a∗ , λ∗ ) ← arg maxt=1,...,T s̄t . {Final configuration is selected by measured CV performance, not by the LLM} 36: // Step 4: Final pipeline fitting and independent evaluation 37: Fit encoding and z-score scaling on Dtrain and apply them to both outer partitions. 38: Apply AutoDP(·; β) to Dtrain only. 39: F ∗ ← AutoFE(Dtrain ; ϕ); apply F ∗ to both outer partitions. 40: M∗ ← Train(a∗ , λ∗ , Dtrain , F ∗ ). 41: R ← Evaluate(M∗ , Dtest ). {LLM has no role in model scoring or hold-out evaluation} 42: return M∗ , π, F ∗ , (a∗ , λ∗ ), and R. 22:

The proposed framework provides four main advantages: 1) It improves search efficiency by adapting a compact AutoML policy to the dataset profile rather than exploring a broad, generic search space. 2) It supports reproducible execution because each generated policy is validated against predefined ranges and corrected using deterministic defaults when necessary. 3) It supports local cybersecurity applications because Llama 3.1 8B can operate without transmitting exper-

imental information to an external LLM service. 4) It preserves experimental rigor by limiting the LLM to policy generation, while model selection remains based on cross-validation and the hold-out test set is reserved for final evaluation. Traditional CASH baselines spend most of their runtime repeatedly training models across a broad candidate space. In contrast, the proposed method performs policy generation and validation once, followed by fewer BO-TPE trials within a narrower conditional space [36]. This shifts the computational effort from broad exploration to guided optimization, while cross-validation remains responsible for selecting the final model. The framework is also suitable for constrained IoT environments. Model training and LLM-assisted planning can be performed on a gateway, server, or cloud platform, while the optimized tree-based IDS can be deployed close to the traffic source. Tree ensembles provide efficient prediction for tabular network traffic [20], [34], and the LLM is required only during AutoML planning rather than during IDS inference. This deployment model is deliberately different from direct LLM-based IDS agents [29], [30]: after model development, the LLM is not part of the detection loop and does not inspect live traffic or generate operational alerts. A further practical advantage is that the generated policy can be cached and audited. The stored policy records the selected balancing method, AutoFE thresholds, candidate model family, hyperparameter ranges, and trial budget. This makes the AutoML configuration easier to reproduce than an iterative manual tuning process and allows researchers to inspect the LLM-generated decisions before expensive model search is executed. Such inspection is useful in cybersecurity applications, where an unreasonable policy can be identified before it affects the training pipeline. Overall, the framework uses the LLM as a bounded source of prior knowledge while leaving model fitting, selection, and evaluation to the validated AutoML pipeline. This separation allows the LLM to guide a more focused search without becoming part of the deployed detector, while keeping the development process efficient, inspectable, and reproducible. IV. P ERFORMANCE E VALUATION A. Experimental Setup All experiments in this work were conducted on a Lenovo Legion 5 machine with an Intel Core Ultra 7 255HX Central Processing Unit (CPU), an NVIDIA RTX 5070 Graphics Processing Unit (GPU), and 32 GB of Random-Access Memory (RAM), representing an IoT server machine. The experiments use Python-based ML libraries, including Scikitlearn for classical ML procedures, Imbalanced-learn [37] for oversampling, LightGBM for LightGBM model development, Optuna for BO-TPE optimization, and Ollama with Llama 3.1 8B for local LLM-based AutoML policy generation. A fixed random seed of 0 is used for train-test splitting, stratified cross-validation, model construction, sampling, and HPO. This improves reproducibility across the main comparison and ablation studies.

9

The proposed framework was evaluated on two public benchmark intrusion detection datasets, CICIDS2017 [17] and IoTID20 [18]. CICIDS2017 is a widely used benchmark cybersecurity dataset that contains modern network traffic with benign samples and multiple attack types such as DoS, brute force, web attacks, botnet, port scan, and infiltration attacks. IoTID20 is an IoT-focused benchmark dataset containing benign traffic and several IoT attack categories, including Mirai, DoS, scanning, and MITM attacks. These two datasets are selected because they jointly represent general network and IoT-specific intrusion detection scenarios, and both contain various types of attacks, heterogeneous network-flow features, and highly imbalanced class distributions. Given the focus of this work on resource-aware IDS development for IoT and edge environments, representative subsets of CICIDS2017 and IoTID20 were used for model development and evaluation. The use of representative subsets reduces the computational and memory burden of repeated fold-wise preprocessing, AutoFE, and CASH optimization, while retaining the evaluated normal/benign traffic and attack categories as well as their pronounced class imbalance. This is also consistent with practical IoT settings, where gateways and edge systems may have limited storage and computing resources and may not retain or repeatedly process all historical traffic records [38], [39]. The same fixed subsets are used by all compared methods to maintain a consistent evaluation protocol. The final CICIDS2017 subset contains 26,800 samples, while the IoTID20 subset contains 62,578 samples. Table III shows that both evaluated subsets remain highly imbalanced, which is important for evaluating the AutoDP component. In CICIDS2017, benign traffic accounts for 68.004% of the subset, whereas infiltration and brute-force samples represent only 0.134% and 0.358%, respectively. In IoTID20, Mirai traffic accounts for 66.062% of the subset, while the other classes are substantially smaller. These distributions provide a useful setting for examining whether the proposed partial-balancing policy can improve minority-class representation without unnecessarily expanding every class to the majority size. Each dataset is divided into an outer training set and a hold-out test set using an 80/20 stratified split, with the resulting sample counts also reported in Table III. Model selection and HPO use the outer training set only. To keep repeated CASH evaluation computationally manageable, a fixed stratified 12,000-sample search subset is drawn from the outer training partition and used for 3-fold stratified crossvalidation. The same search-subset rule is used for the plain and AutoML search comparisons, and the hold-out set is never used for this sampling or for model selection. Within each fold, label encoding and z-score normalization are fitted on the foldtraining data, AutoDP resamples only the fold-training data under its partial-balancing policy, and AutoFE is fitted on the resulting training fold and then applied to the paired validation fold. After HPO, the selected pipeline is fitted on the complete outer training set and evaluated once on the hold-out test set. Accuracy and weighted precision, recall, and F1-score are used to evaluate the proposed IDSs comprehensively. The F1-

score balances precision and recall and is particularly useful when false positives and false negatives both matter. Efficiency is evaluated using optimizer time and hold-out prediction time per sample. Optimizer time measures the model-search procedure itself; the one-time LLM policy-generation latency is tracked separately and included in the total wall-clock time rather than in the optimizer-time column. The LLM calls used in the reported experiments required 6.88 s for CICIDS2017 and 4.99 s for IoTID20, produced valid JSON outputs, and did not trigger a fallback. The conventional CASH comparisons use Random Search (RS) and BO-TPE as their optimization strategies. B. Experimental Results and Discussion Table IV compares the proposed LLM-AutoML framework with optimized deep-learning models [24], [25], existing AutoML-based IDSs [26], [27], plain tree-based learners [21], [22], and traditional AutoML baselines [28]. For a fair comparison, all methods were reproduced and evaluated under the same experimental environment, dataset partitions, preprocessing procedure, and evaluation protocol. Traditional AutoMLTPE is evaluated with both 10 and 30 trials: the 10-trial setting provides the direct budget-matched comparison with the proposed method, while the 30-trial setting represents the broader conventional search used as the secondary efficiency reference. On CICIDS2017, the proposed 10-trial LLM-AutoML method achieves the highest test performance, with 99.683% accuracy, 99.685% precision, 99.683% recall, and 99.680% weighted F1-score. Under the matched 10-trial budget, Traditional AutoML-TPE reaches 99.532% F1-score, so the proposed policy improves the test F1-score by 0.149 percentage points. Its optimizer time is 46.55 s compared with 37.41 s for the matched 10-trial TPE run, showing that the equalbudget benefit is primarily better search quality rather than lower raw optimizer time. The broader 30-trial Traditional AutoML-TPE baseline reaches 99.624% F1-score and requires 128.41 s. Relative to that baseline, the proposed 10-trial method improves F1-score by 0.056 percentage points and reduces optimizer time by 63.7%. Compared with Traditional AutoML-RS, it improves F1-score by 0.167 percentage points and reduces optimizer time by 58.4%. An additional 30-trial LLM-assisted run reached the same 99.680% test F1-score as the 10-trial proposed method while requiring 111.40 s of optimizer time, indicating that the selected compact policy had already converged to the reported hold-out result within the smaller budget. The same pattern is observed on IoTID20. The proposed method obtains 99.185% accuracy, 99.187% precision, 99.185% recall, and 99.186% weighted F1-score. Under the equal 10-trial budget, Traditional AutoML-TPE reaches 98.844% F1-score, giving the proposed method a 0.341percentage-point advantage, although optimizer time increases from 29.66 s to 36.28 s. Compared with the 30-trial TPE baseline, which reaches 99.170% F1-score in 72.36 s, the proposed method improves F1-score by 0.016 percentage points and reduces optimizer time by 49.9%. It also reduces

10

TABLE III C LASS DISTRIBUTIONS OF THE REPRESENTATIVE CICIDS2017 AND I OTID20 SUBSETS USED IN THE EXPERIMENTS . Class

Samples

Distribution (%)

Training samples for Cross Validation

Hold-Out Test samples

Benign

18,225

68.004

14,580

3,645

DoS

3,042

11.351

2,433

609

Web Attack

2,180

8.134

1,744

436

Botnet

1,966

7.336

1,573

393

Port Scan

1,255

4.683

1,004

251

Brute Force

96

0.358

77

19

Infiltration

36

0.134

29

7

Total

26,800

100.000

21,440

5,360

Mirai

41,340

66.062

33,072

8,268

Scan

7,620

12.177

6,096

1,524

DoS

6,047

9.663

4,837

1,210

Normal

4,031

6.442

3,225

806

MITM ARP Spoofing

3,540

5.657

2,832

708

Total

62,578

100.000

50,062

12,516

Dataset

CICIDS2017

IoTID20

TABLE IV C LASSIFICATION PERFORMANCE AND EFFICIENCY COMPARISON ON CICIDS2017 AND I OTID20. Dataset

CICIDS2017

IoTID20

Method

Best ML model

Test Accuracy (%)

Test Precision (%)

Test Recall (%)

Test F1 (%)

Optimization time (s)

Test time per sample (ms)

Features

PSO-LSTM [24]

LSTM

95.951

95.997

95.951

95.956

1110.95

0.0629

23

GA-CNN [25]

CNN

96.679

96.809

96.679

96.704

960.11

0.0703

77

OE-IDS [26]

Soft Voting

99.384

99.387

99.384

99.378

382.67

0.0582

15

AutoML-ID [27]

Boosting Ensemble

99.515

99.516

99.515

99.515

292.91

0.0124

77

LightGBM [21]

LightGBM

99.422

99.430

99.422

99.420

3.60

0.0053

77

RF [22]

Random Forest

99.571

99.578

99.571

99.569

2.35

0.0091

77

ET [22]

Extra Trees

98.433

98.514

98.433

98.450

1.11

0.0109

77

Traditional AutoML-RS

LightGBM

99.515

99.523

99.515

99.514

111.80

0.0079

49

Traditional AutoML-TPE - 10 trials [28]

LightGBM

99.534

99.539

99.534

99.532

37.41

0.0145

49

Traditional AutoML-TPE - 30 trials [28]

LightGBM

99.627

99.630

99.627

99.624

128.41

0.0208

49

Proposed LLM-AutoML - 10 trials

LightGBM

99.683

99.685

99.683

99.680

46.55

0.0198

41

PSO-LSTM [24]

LSTM

95.262

95.228

95.262

95.121

1281.44

0.0203

28

GA-CNN [25]

CNN

98.154

98.167

98.154

98.155

1480.58

0.0401

81

OE-IDS [26]

Soft Voting

97.643

97.649

97.643

97.616

294.70

0.0266

19

AutoML-ID [27]

Boosting Ensemble

99.017

99.022

99.017

99.019

239.80

0.0403

81

LightGBM [21]

LightGBM

98.778

98.809

98.778

98.787

5.12

0.0070

81

RF [22]

Random Forest

99.073

99.080

99.073

99.075

1.94

0.0050

81

ET [22]

Extra Trees

98.634

98.670

98.634

98.645

1.32

0.0066

81

Traditional AutoML-RS

Random Forest

99.081

99.101

99.081

99.087

107.25

0.0076

59

Traditional AutoML-TPE - 10 trials [28]

LightGBM

98.834

98.869

98.834

98.844

29.66

0.0149

59

Traditional AutoML-TPE - 30 trials [28]

Random Forest

99.169

99.173

99.169

99.170

72.36

0.0059

59

Proposed LLM-AutoML - 10 trials

LightGBM

99.185

99.187

99.185

99.186

36.28

0.0196

19

optimizer time by 66.2% relative to Traditional AutoML-RS while improving F1-score by 0.098 percentage points. As on CICIDS2017, the auxiliary 30-trial LLM-assisted run did not improve the test F1-score beyond 99.186%, but increased optimizer time to 92.63 s. This result supports the use of the smaller LLM-guided budget for the final method. The optimized IDS and plain-model comparisons provide additional context. On CICIDS2017, the plain RF baseline is already strong at 99.569% F1-score, illustrating that the dataset is favorable to tree-based classification. The proposed method nevertheless reaches a higher F1-score while reducing the final feature set from 77 raw features to 41. On

IoTID20, the strongest plain RF baseline reaches 99.075% F1-score, whereas the proposed method reaches 99.186% with only 19 retained features. The proposed model also remains substantially less expensive to optimize than PSO-LSTM, GACNN, OE-IDS, and AutoML-ID in the reported comparison. These results suggest that the main contribution is not a large absolute increase over already strong tree baselines, but a more compact and controlled AutoML process that reaches a strong final model with a small search budget. The feature-selection results in Figs. 2 and 3 further show how the final representations differ between the two datasets. CICIDS2017 retains 41 features. Its three

11

Fig. 2. Selected features and normalized importance scores of the final LLM-assisted IDS on CICIDS2017.

Fig. 3. Selected features and normalized importance scores of the final LLM-assisted IDS on IoTID20.

largest normalized importance values are associated with Init Win bytes forward, Init Win bytes backward, and Flow IAT Min, which together account for approximately 35.0% of the total importance of the retained set. The remaining importance is distributed across timing, packet-length, and flag-related features, indicating that the final CICIDS2017 model still uses a relatively broad representation. IoTID20 is reduced more aggressively to 19 features. Src Port, Dst IP, and Init Bwd Win Byts are the three most important retained features and together account for approximately 36.7% of the normalized importance. The other retained variables include flow-rate, inter-arrival-time, packet-length, protocol, and acknowledgement-related information. These figures are consistent with the feature counts in Table IV and show that the LLM-assisted AutoFE policy does not impose the same reduction level on both datasets. The inference-time results indicate a modest cost for the final tuned LightGBM configuration. On CICIDS2017, the proposed model requires 0.0198 ms per test sample, which is slower than the plain tree-based baselines and Traditional AutoML-RS, but comparable to the 30-trial Traditional AutoML-TPE model at 0.0208 ms. On IoTID20, the proposed model requires 0.0196 ms per sample, which is higher than the plain tree-based and Traditional AutoML baselines but remains

below 0.1 ms per sample. Thus, the proposed method trades a small amount of per-sample inference time for higher test F1score and a substantially smaller final feature set, particularly on IoTID20. Table V expands the ablation study to the same performance and efficiency measures used in the main comparison. All variants use a 10-trial CASH budget, so the table isolates how the LLM-controlled AutoDP, AutoFE, and CASH components interact under a fixed search budget. On CICIDS2017, A1 (LLM AutoDP only) is effectively unchanged from A0 in F1-score, while A2 (LLM AutoFE only) increases F1-score from 99.532% to 99.569% and reduces the feature count from 49 to 40. A3 (LLM CASH only) provides the largest individual performance increase, reaching 99.606% F1-score. A4, which combines LLM AutoDP and AutoFE but retains the traditional CASH policy, drops slightly to 99.513% F1score. The full A5 method then reaches 99.680%, clearly above every partial variant. This non-additive pattern indicates that the transformed data and feature space benefit from a CASH policy that is aligned with them; the gains from the three controls should therefore be interpreted as coordinated rather than independent additive effects. IoTID20 shows an even clearer separation of roles. A1 produces the same 98.844% F1-score as A0, indicating that the

12

TABLE V A BLATION STUDY OF THE LLM- CONTROLLED COMPONENTS UNDER A 10- TRIAL CASH BUDGET. Dataset

CICIDS2017

IoTID20

Variant

Test Accuracy (%)

Test Precision (%)

Test Recall (%)

Test F1 (%)

Optimization time (s)

Test time per sample (ms)

Features

A0: Traditional AutoML - 10 trials

99.534

99.539

99.534

99.532

37.41

0.0145

49

A1: LLM AutoDP only

99.534

99.540

99.534

99.532

33.79

0.0137

48

A2: LLM AutoFE only

99.571

99.576

99.571

99.569

34.47

0.0178

40

A3: LLM CASH only

99.608

99.612

99.608

99.606

54.07

0.0160

49

A4: LLM AutoDP + AutoFE

99.515

99.521

99.515

99.513

33.60

0.0136

41

A5: Full proposed LLM-AutoML (LLM AutoDP + AutoFE + CASH)

99.683

99.685

99.683

99.680

46.55

0.0198

41

A0: Traditional AutoML - 10 trials

98.834

98.869

98.834

98.844

29.66

0.0149

59

A1: LLM AutoDP only

98.834

98.869

98.834

98.844

28.69

0.0155

59

A2: LLM AutoFE only

98.866

98.897

98.866

98.875

21.99

0.0140

19

A3: LLM CASH only

99.153

99.156

99.153

99.154

65.31

0.0364

59

A4: LLM AutoDP + AutoFE

98.866

98.897

98.866

98.875

34.18

0.0200

19

A5: Full proposed LLM-AutoML (LLM AutoDP + AutoFE + CASH)

99.185

99.187

99.185

99.186

36.28

0.0196

19

LLM AutoDP policy alone does not change the final detection result in this run. A2 improves F1-score slightly to 98.875% while reducing the feature set from 59 to 19. A3 provides the dominant individual performance gain and reaches 99.154% F1-score, although its optimizer time is the highest among the partial variants at 65.31 s. A4 retains the compact 19-feature representation but does not improve beyond A2. When all three components are coordinated in A5, F1-score rises to 99.186%, the highest result in the ablation table. The ablation therefore suggests that CASH guidance contributes most directly to predictive performance, whereas AutoFE contributes strongly to compactness and becomes most effective when coupled with the final LLM-guided search. Taken together, the equal-budget and ablation results refine the efficiency claim of this work. The LLM-assisted policy is not faster than Traditional AutoML-TPE when both are restricted to exactly 10 trials; it incurs modest additional optimizer cost while reaching a clearly stronger test result on both datasets. Its efficiency advantage appears when the proposed compact 10-trial search is compared with the broader 30-trial conventional search, where it reduces optimizer time by approximately 63.7% on CICIDS2017 and 49.9% on IoTID20 while maintaining slightly higher F1-score. The results therefore support the bounded-policy architecture as a way to improve low-budget search quality and reduce the amount of broader AutoML exploration required, while final model selection remains governed by cross-validation and independent hold-out testing. V. C ONCLUSION AutoML has emerged as a promising approach to autonomous cybersecurity for IoT and zero-touch networks, but the cost of broad model and hyperparameter search remains an important limitation. This paper proposed an LLM-assisted AutoML framework for IoT intrusion detection in which a local Llama 3.1 8B model generates bounded and validated policies for data balancing, feature selection, and CASH optimization. The LLM does not train or evaluate the IDS model; its role is limited to policy generation, while the resulting configurations are executed and assessed through deterministic

ML procedures, cross-validation, and an independent hold-out test set. The final experiments clarify both the effectiveness and the efficiency of this design. Under the same 10-trial budget, the proposed policy increases weighted test F1-score from 99.532% to 99.680% on CICIDS2017 and from 98.844% to 99.186% on IoTID20 relative to Traditional AutoMLTPE. The equal-budget LLM-assisted search requires modestly more optimizer time, but it reaches a stronger low-budget result. Relative to the broader 30-trial Traditional AutoMLTPE baseline, the proposed 10-trial method reduces optimizer time by 63.7% on CICIDS2017 and 49.9% on IoTID20 while maintaining slightly higher F1-score. The ablation study further shows that the three LLM-controlled stages interact non-additively: CASH guidance contributes most directly to detection performance, whereas AutoFE substantially reduces the retained feature set, and their coordinated use produces the strongest final result. These findings support the boundedpolicy architecture as a practical way to improve the quality of a compact AutoML search without placing the LLM in the deployed IDS decision loop. Future work will extend the framework to online and continual IDS learning, evaluate additional local LLMs, and incorporate multi-objective deployment constraints such as latency, memory, and energy consumption. VI. ACKNOWLEDGMENT This work was supported in part by the Natural Sciences and Engineering Research Council of Canada (NSERC) under Discovery Grant RGPIN-2025-05840, and in part by the Department of National Defence (DND)/NSERC Discovery Grant Supplement DGDND-2025-05840. R EFERENCES [1] A. Sharma and K. Bhushan, “A comprehensive survey on IoT security: Challenges, security issues, and countermeasures,” Comput. Sci. Rev., vol. 59, Art. no. 100839, Feb. 2026. [2] N. Mishra and S. Pandya, “Internet of Things applications, security challenges, attacks, intrusion detection, and future visions: A systematic review,” IEEE Access, vol. 9, pp. 59353–59377, 2021. [3] L. Yang and A. Shami, “IoT data analytics in dynamic environments: From an automated machine learning perspective,” Eng. Appl. Artif. Intell., vol. 116, Art. no. 105366, Nov. 2022.

13

[4] M. Bagaa, T. Taleb, J. Bernal Bernabe, and A. Skarmeta, “A machine learning security framework for IoT systems,” IEEE Access, vol. 8, pp. 114066–114077, 2020. [5] W. Zaheer, C. H. Nwokoye, S. N. Afrasiabi, K. El-Khatib, and L. Yang, “A comprehensive survey on online AutoML and adversarial robustness for IoT and EV charging network security,” Sensors, vol. 26, no. 12, Art. no. 3886, Jun. 2026. [6] F. Dehrouyeh, L. Yang, F. B. Ajaei, and A. Shami, “On TinyML and cybersecurity: Electric vehicle charging infrastructure use case,” IEEE Access, vol. 12, pp. 108703–108730, 2024. [7] L. Yang, M. El Rajab, A. Shami, and S. Muhaidat, “Enabling AutoML for zero-touch network security: Use-case driven analysis,” IEEE Trans. Netw. Serv. Manag., vol. 21, no. 3, pp. 3555–3582, 2024. [8] M. S. A. Muthanna, R. Alkanhel, A. Muthanna, A. Rafiq, and W. A. M. Abdullah, “Towards SDN-enabled, intelligent intrusion detection system for Internet of Things (IoT),” IEEE Access, vol. 10, pp. 22756–22768, 2022. [9] A. Fatani, M. Abd Elaziz, A. Dahou, M. A. A. Al-Qaness, and S. Lu, “IoT intrusion detection system using deep learning and enhanced transient search optimization,” IEEE Access, vol. 9, pp. 123448–123464, 2021. [10] L. Yang and A. Shami, “On hyperparameter optimization of machine learning algorithms: Theory and practice,” Neurocomputing, vol. 415, pp. 295–316, Nov. 2020. [11] N. Hollmann, S. Müller, and F. Hutter, “Large language models for automated data science: Introducing CAAFE for context-aware automated feature engineering,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, 2023. [12] T. Liu, N. Astorga, N. Seedat, and M. van der Schaar, “Large language models to enhance Bayesian optimization,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2024. [13] Y. Gu, H. You, J. Cao, M. Yu, H. Fan, and S. Qian, “Large language models for constructing and optimizing machine learning workflows: A survey,” ACM Trans. Softw. Eng. Methodol., Oct. 2025, doi: 10.1145/3773084. [14] A. Tornede et al., “AutoML in the age of large language models: Current challenges, future opportunities and risks,” Trans. Mach. Learn. Res., pp. 1–31, 2024. [15] J. Xu et al., “Large language models synergize with automated machine learning,” Trans. Mach. Learn. Res., 2024. [16] A. Grattafiori et al., “The Llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. [17] I. Sharafaldin, A. H. Lashkari, and A. A. Ghorbani, “Toward generating a new intrusion detection dataset and intrusion traffic characterization,” in Proc. Int. Conf. Inf. Syst. Secur. Privacy (ICISSP), 2018, pp. 108–116. [18] I. Ullah and Q. H. Mahmoud, “A scheme for generating a dataset for anomalous activity detection in IoT networks,” in Advances in Artificial Intelligence, Lecture Notes in Computer Science, vol. 12109. Cham, Switzerland: Springer, 2020, pp. 508–520, doi: 10.1007/978-3030-47358-7 52. [19] L. Yang, A. Moubayed, and A. Shami, “MTH-IDS: A multitiered hybrid intrusion detection system for Internet of Vehicles,” IEEE Internet Things J., vol. 9, no. 1, pp. 616–632, Jan. 2022. [20] L. Grinsztajn, E. Oyallon, and G. Varoquaux, “Why do tree-based models still outperform deep learning on typical tabular data?” Adv. Neural Inf. Process. Syst., vol. 35, pp. 507–520, 2022. [21] D. Jin, Y. Lu, J. Qin, Z. Cheng, and Z. Mao, “SwiftIDS: Real-time intrusion detection system based on LightGBM and parallel intrusion detection mechanism,” Comput. Secur., vol. 97, Art. no. 101984, Oct. 2020. [22] F. Gutiérrez-Portela, H. B. Arteaga-Arteaga, F. Almenares-Mendoza, L. Calderón-Benavides, H.-G. Acosta-Mesa, and R. Tabares-Soto, “Enhancing intrusion detection in IoT communications through ML model generalization with a new dataset (IDSAI),” IEEE Access, vol. 11, pp. 70542–70559, 2023. [23] L. Yang and A. Shami, “A transfer learning and optimized CNN based intrusion detection system for Internet of Vehicles,” in Proc. IEEE Int. Conf. Commun. (ICC), 2022, pp. 2774–2779. [24] W. Elmasry, A. Akbulut, and A. H. Zaim, “Evolving deep learning architectures for network intrusion detection using a double PSO metaheuristic,” Comput. Netw., vol. 168, Art. no. 107042, Feb. 2020. [25] H. Naeem, F. Ullah, O. Krejcar, D. Li, and D. Vasan, “Optimizing vehicle security: A multiclassification framework using deep transfer learning and metaheuristic-based genetic algorithm optimization,” Int. J. Crit. Infrastruct. Prot., vol. 49, Art. no. 100745, 2025. [26] M. A. Khan, N. Iqbal, Imran, H. Jamil, and D. H. Kim, “An optimized ensemble prediction model using AutoML based on soft voting classifier

for network intrusion detection,” J. Netw. Comput. Appl., vol. 212, Art. no. 103560, Mar. 2023. [27] A. Singh, J. Amutha, J. Nagar, S. Sharma, and C. C. Lee, “AutoMLID: Automated machine learning model for intrusion detection using wireless sensor network,” Sci. Rep., vol. 12, Art. no. 9074, May 2022. [28] L. Yang and A. Shami, “Towards autonomous cybersecurity: An intelligent AutoML framework for autonomous intrusion detection,” in Proc. Workshop Autonomous Cybersecurity (AutonomousCyber), ACM SIGSAC Conf. Comput. Commun. Secur. (CCS), 2024, pp. 68–78, doi: 10.1145/3689933.3690833. [29] Y. Li, Z. Xiang, N. D. Bastian, D. Song, and B. Li, “IDS-Agent: An LLM agent for explainable intrusion detection in IoT networks,” in NeurIPS 2024 Workshop on Open-World Agents, 2024. [30] T. Ali, P. Kostakos, and S. Sheikhi, “HuntGPT: Integrating machine learning-based anomaly detection and explainable AI with large language models (LLMs),” Telecom, vol. 7, no. 3, Art. no. 73, Jun. 2026. [31] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” J. Artif. Intell. Res., vol. 16, pp. 321–357, 2002. [32] H. He, Y. Bai, E. A. Garcia, and S. Li, “ADASYN: Adaptive synthetic sampling approach for imbalanced learning,” in Proc. Int. Joint Conf. Neural Netw. (IJCNN), 2008, pp. 1322–1328. [33] F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011. [34] G. Ke et al., “LightGBM: A highly efficient gradient boosting decision tree,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 3146–3154. [35] S. Watanabe, “Tree-structured Parzen estimator: Understanding its algorithm components and their roles for better empirical performance,” arXiv preprint arXiv:2304.11127, 2023. [36] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization framework,” in Proc. ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., 2019, pp. 2623–2631. [37] G. Lemaı̂tre, F. Nogueira, and C. K. Aridas, “Imbalanced-learn: A Python toolbox to tackle the curse of imbalanced datasets in machine learning,” J. Mach. Learn. Res., vol. 18, no. 17, pp. 1–5, 2017. [38] L. Yang, D. M. Manias, and A. Shami, “PWPAE: An ensemble framework for concept drift adaptation in IoT data streams,” in Proc. IEEE Global Commun. Conf. (GLOBECOM), 2021, pp. 1–6. [39] L. Yang and A. Shami, “A multi-stage automated online network data stream analytics framework for IIoT systems,” IEEE Trans. Ind. Informat., vol. 19, no. 2, pp. 2107–2116, Feb. 2023.

Li Yang (Member, IEEE) is an Assistant Professor in the Faculty of Business and Information Technology at Ontario Tech University, and an Adjunct Research Professor in the Department of Electrical and Computer Engineering at Western University. He received his Ph.D. in Electrical and Computer Engineering from Western University in 2022. He was awarded the Faculty of Business and Information Technology (FBIT) Early Career Research Chair in Cybersecurity for Business, the FBIT Research Excellence Award, and the Ontario Tech University Research Excellence Award in 2026. Li Yang has been named by the IEEE Computer Society as one of the recipients of Computing’s Top 30 Early Career Professionals for 2025. Li Yang is also included in Stanford University/Elsevier’s List of the World’s Top 2% Scientists, and he was ranked among the world’s top 0.5% of researchers in ’Networking & Telecommunications’ in 2024 and 2025. He is currently an Associate Editor of IEEE Transactions on Industrial Informatics and IEEE Transactions on Network and Service Management. His research interests include cybersecurity, machine learning, deep learning, AutoML, Tiny Machine Learning (TinyML), LLMs, model optimization, network data analytics, IoT, intrusion detection, anomaly detection, concept drift, continual learning, and adversarial machine learning.

Record · ID 1028623 · SHA-256 f68395d4c7c445ba
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.