ConceptioArchivearXiv CS
arXiv CSopen access

SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

Preprint. Under review.

SMETA-ZSL: Semantic Meta-Alignment for Zero-Shot Threat Classification Ivan Alejandro Montoya Sanchez, Anantaa Kotal, Aritran Piplai The University of Texas at El Paso 500 W. University Avenue El Paso, TX 79968, USA [email protected], {akotal,apiplai}@utep.edu

arXiv:2607.09936v1 [cs.LG] 10 Jul 2026

Abstract Cybersecurity systems must adapt rapidly to emerging threats. However, labeled data for new threat categories is unavailable when those threats first appear. Generalized zero-shot learning offers a natural solution by enabling recognition of unseen classes through auxiliary semantic knowledge rather than labeled examples. Large language models are particularly promising in this setting because they can convert unstructured CTI reports into semantic prototypes for emerging threats. However, applying language-driven zero-shot learning to cybersecurity is difficult due to strong semantic overlap between threat descriptions, heterogeneity between behavioral attributes and text, severe class imbalance, and open-set conditions where unseen threats are unknown during training. We propose SMETA-ZSL, that learns semantic prototypes from overlapping language descriptions through contrastive finetuning, aligns behavioral features through episodic meta-learning and knowledge distillation, and performs adaptive routing for generalization across seen-unseen classes. Across 7 benchmarks, SMETA-ZSL delivers the strongest overall generalized zeroshot performance under the strictest inductive setting, surpassing prior methods by 10.8 points on average, with gains up to 18.1 points. Github: https://github.com/Security-And-Intelligence-Lab-UTEP/SMETA-ZSL

1

Introduction

Cyber defense systems rely on large amounts of raw data, such as malware samples or logs, to train and update models, yet this data is often difficult to obtain, slow to analyze, and not always available when new threats emerge. Consider the task of malware classification: the continuous emergence of new variants and families (Tang et al., 2023) necessitates frequent retraining of classifiers, imposing significant computational overhead (Li et al., 2024a) and making it difficult to obtain timely, labeled data for adapting models to new threats (Barros et al., 2022; Aurna et al., 2025). In contrast, Cyber Threat Intelligence (CTI) reports are regularly published by analysts and provide timely natural language descriptions of attacker behavior. However, most defense systems do not operate over natural language, and therefore cannot directly use CTI to adapt or update their behavior, leaving valuable intelligence underutilized. Large Language Models (LLMs) offer a way to process CTI and extract information that can be used by downstream systems. Prior work (Chakraborty et al., 2026; Mitra et al., 2025; Bertiger et al., 2025; Schwartz et al., 2025) has largely focused on using LLMs to generate rule-based defenses. In contrast, leveraging CTI to update data-driven, black-box machine learning systems is significantly more difficult, as it requires converting textual descriptions into signals that can influence model behavior without access to raw data. This gap creates an important language modeling problem: can language models convert unstructured threat reports into representations that are useful for updating non-linguistic machine learning systems, even when no labeled examples of the new threat exist? Directly using LLMs for zero-shot malware or attack detection is challenging because cybersecurity artifacts are often too large to analyze in full, making inference costly and frequently exceeding practical context-window limits (Qian et al., 2025). Truncating these artifacts into abstract 1

Preprint. Under review.

features degrades performance (Zhou et al., 2024). Instead, the goal is to extract and transfer knowledge from CTI using language models to adapt existing cyber-defense systems. We formulate this problem as a Generalized Zero-Shot Learning (GZSL) setting, where a downstream machine learning model must recognize both seen and unseen classes, with Cyber Threat Intelligence (CTI) providing semantic information for unseen threats. While GZSL with semantic auxiliary information has been extensively explored in generalized settings, such as the vision-language domain (Chen et al., 2023; Ma & Hu, 2020; Rao et al., 2024; Lei et al., 2024), its application to cybersecurity remains an open and largely unsolved challenge, despite being of critical practical importance. The unique challenges in this setup are as follows. 1. Semantic ambiguity of class prototypes: Unlike generalized domains with distinguishable semantic features, cybersecurity relies on CTI reports whose descriptions often overlap heavily across malware families. Different families frequently share APIs, behavioral patterns, and attack techniques, making semantic prototypes weakly separable and reducing transfer to unseen classes. Figure 1 illustrates this challenge. 2. Cross-modal heterogeneity: Malware features are derived from behavioral observations such as API calls, system traces, and network traffic . These features differ substantially from the sparse, high-level semantics of CTI reports, making alignment between feature and semantic spaces difficult. 3. Class imbalance: Seen class dominance is already a major challenge in GZSL . In cybersecurity, this is amplified by severe imbalance, where benign samples dominate and emerging malware families are rare , biasing predictions toward frequent seen classes. 4. Open-set classification: In practice, the number and identity of unseen malware families are unknown during training. Models must therefore distinguish between seen and unseen threats without assuming a fixed set of candidate unseen classes in advance. We propose SMETA-ZSL, a semantically grounded meta-GZSL framework designed for the realistic cyber threat setting in which unseen classes are not predefined during training, unlabeled unseen instances are unavailable, and adaptation must rely solely on naturallanguage CTI reports rather than raw artifacts. Unlike prior methods that assume closed sets, access to raw unseen-class data, or predefined unseen prototypes, SMETA-ZSL operates under a stricter open-set, class-inductive, instance-inductive setting. To address the resulting challenges, our framework combines: (1) a contrastively fine-tuned LLM encoder that produces discriminative semantic prototypes from overlapping CTI descriptions; (2) a cross-modal alignment framework that uses episodic meta-learning to explicitly simulate unseen-class emergence during training; and (3) a parameter-free gating mechanism for adaptive inference. Across 7 benchmark datasets, SMETA-ZSL delivers the strongest overall generalized zero-shot performance under the strictest inductive setting.

Figure 1: CTI reports for two distinct malware families: Mobidash and Dowgin, share substantial lexical overlap in behavioral descriptions, illustrating why cybersecurity class prototypes are weakly separable in semantic space.

2

Background

ZSL addresses the problem of recognizing instances from classes never observed during training (Pourpanah et al., 2022; Wang et al., 2019). In GZSL, the classifier must operate over the union of both seen and unseen classes simultaneously (Li et al., 2024b; Verma et al., 2020), compounding the difficulty with seen-class dominance — where the posterior probability of seen classes systematically overwhelms that of unseen classes at inference 2

Preprint. Under review.

time (Kwon & Al Regib, 2022; Bhat et al., 2025; Xian et al., 2017). Since unseen classes are never observed during training, ZSL methods rely on auxiliary semantic information to bridge the gap. Drawing inspiration from human cognition, where unfamiliar concepts are recognized through background knowledge relating them to familiar ones (Fu et al., 2015; Romera-Paredes & Torr, 2015; Tang et al., 2024; Li et al., 2024b; Sanchez et al., 2025), existing approaches broadly fall into three families: learning a joint embedding space between modalities and semantic representations (Ali & Khan, 2023; Nawaz et al., 2022; Chen et al., 2023; Lei et al., 2024), synthetically generating unseen-class data from semantic descriptions (Mishra et al., 2018; Ma & Hu, 2020; Tang et al., 2024; Wu & Bergman, 2025b; Marszałek et al.), and learning semantic prototypes (Wang et al., 2025a; Fu et al., 2017; Wang et al., 2021; Rao et al., 2024; Fan et al., 2026; Gill et al., 2026). Among these, the semantic prototype approach is particularly well-suited for settings where class-level descriptions are available but instance-level data for unseen classes is absent. ZSL has also been adapted to domainspecific settings, including cybersecurity, where specialized frameworks address malware detection under data scarcity (Barros et al., 2022; Wang et al., 2025b; Aurna et al., 2025).

3

Methodology

We propose SMETA-ZSL, a semantically grounded GZSL framework for tabular data classification. Our core hypothesis is that language models, when appropriately fine-tuned, can serve as a reliable semantic bridge between overlapping natural language descriptions and fine-grained behavioral attributes, enabling zero-shot recognition of unseen classes through episodic meta-training that explicitly aligns heterogeneous modalities. Figure 2 gives an overview of our proposed framework.

Figure 2: Overview of our proposed SMETA-ZSL framework that uses semantic prototype for GZSL, when semantic descriptions are overlapping and at a higher level of abstraction. 3.1

Preliminaries

In Zero-shot learning there are two disjoint sets of classes: the seen classes S , for which labeled training instances are available, and the unseen classes U , for which no labeled instances exist at training time. Specifically in GZSL, the classifier f (·) : X → S ∪ U must operate over the union of both seen and unseen classes. Each class c ∈ S ∪ U is represented by a prototype vector t ∈ T ⊆ R M , encoded via a prototyping function π (·) : S ∪ U → T . This yields prototype sets Ts = {tis }iN=s1 and Tu = {tiu }iN=u1 for seen and unseen classes respectively, serving as the semantic bridge through which knowledge is transferred from S to U . Given a test instance xi , a projection network f θ : X → T maps it into the shared semantic space, and the predicted class is determined by nearest-prototype matching: f θ ( xi ) · t c ŷi = arg max (1) c∈S∪U ∥ f θ ( xi )∥∥ tc ∥ The cybersecurity setting introduces two additional structural constraints that go beyond standard GZSL. First, we operate under the open-set assumption: unlike closed-set GZSL where S ∪ U is fully predefined, only S is known at training time. The unseen classes U , including their cardinality |U | = Nu and identities {ciu }iN=u1 , are revealed only at inference, reflecting the real-world scenario where novel malware families cannot be anticipated in advance. Second, we adopt the Class-Inductive, Instance-Inductive (CIII) setting (Pourpanah et al., 2022), under which the model is trained using only Dtr and Ts , without access to Tu 3

Preprint. Under review.

during training. The unseen prototypes are made available only at inference, serving as the semantic bridge through which the model generalizes to novel unseen classes. In the cybersecurity domain, this bridge is constructed from cyber threat intelligence (CTI) reports, unstructured expert-authored documents describing behavioral signatures, tactics, and techniques of each threat class. Our goal is to ensure that knowledge transfers reliably from S to U despite the semantic overlap and cross-modal heterogeneity these reports introduce, a challenge addressed directly by SMETA-ZSL. 3.2

Discriminative Semantic Prototype Learning with LLM

LLM for Semantic Prototyping: LLMs are natural candidates for semantic prototype construction: given a natural language description of a class, a pretrained encoder can map the text into a dense vector that captures its semantic content. These vectors serve as class prototypes in T , providing a representation of each class However, naively encoding CTI reports with a pretrained language model yields weakly separable prototypes, as threat descriptions across different classes frequently share structural vocabulary and domainspecific terminology, causing their embeddings to cluster around a shared centroid rather than reflect class-discriminative structure. Supervised Contrastive Objective: Let zi ∈ R M denote the L2-normalized embedding of description sample i produced by the language model. To enforce intra-class compactness and inter-class separation in the language model’s representation space, we optimize it with a Supervised Contrastive loss (Khosla et al., 2020), which pulls descriptions of the same class together while pushing descriptions of different classes apart: exp(zi · z p /τ ) −1 LSupCon = ∑ (2) ∑ log ∑ | P ( i )| a∈ A(i ) exp( zi · z a /τ ) i∈ I p ∈ P (i ) where P(i ) is the set of positive samples whose descriptions share the same class label as i, A(i ) is the set of all other description samples in the batch, and τ is a temperature hyperparameter. Figure 3 illustrates the effect of contrastive finetuning on the geometry of the learned semantic space. Isotropy Regularization: Optimizing solely on LSupCon is insufficient when class descriptions share substantial structural overlap. The encoder tends to learn the general structure of descriptions rather than class-discriminative signatures, causing embeddings to collapse toward a shared mean. To explicitly counteract this, we introduce an isotropy regularization term that penalizes the mean squared cosine similarity of each embedding zi to the L2-normalized batch mean z̄, discouraging the semantic space from collapsing toward a centroid driven by shared boilerplate vocabulary: 1 N (zi · z̄)2 (3) N i∑ =1 The combined training objective for the semantic encoder is Lsem = LSupCon + γLIso , where γ ≥ 0 controls the strength of the isotropy regularization relative to the contrastive objective.

LIso =

Prototype Construction: Once the encoder is trained, a class prototype tc ∈ T is computed for each class c ∈ S ∪ U as the L2-normalized mean of its constituent embeddings, where Sc denotes the set of description samples belonging to class c. This yields the prototype sets Ts and Tu , which serve as fixed semantic anchors for all subsequent alignment and inference. 3.3

Cross-Modal Alignment via Meta Knowledge Distillation

Even with discriminative semantic prototypes, aligning X with T remains non-trivial due to the granularity mismatch between precise instance-level behavioral observations and coarse, abstract semantic descriptions. We address this through a novel framework that combines meta-knowledge distillation (Pan et al., 2021) with episodic meta-learning, where the zero-shot condition is explicitly simulated at every training step. Two-Tower Architecture: The framework operates across two representation spaces. The Semantic Tower encodes class descriptions into T , producing the fixed class prototypes {tc } 4

Preprint. Under review.

(a) Semantic embeddings with Pretrained LLM

(b) Semantic embeddings with Contrastively Finetuned LLM

Figure 3: Comparison of semantic embedding quality with and without contrastive finetuning for datapoints in CIC-AndMal (Canadian Institute for Cybersecurity (CIC) and Canadian Centre for Cyber Security (CCCS), 2020) dataset.

derived in Section 3.2. The Behavioral Tower consists of a projection network f θ : X → T that learns to map behavioral feature vectors into the shared semantic space, producing ẑi = f θ ( xi ) ∈ R M for each instance xi ∈ X . Episodic Meta-Training: To prevent f θ from overfitting to seen-class alignment and to explicitly enforce generalization to unseen classes, training is structured as a series of metalearning episodes that simulate the zero-shot condition at every step. At each episode, the set of available training classes Ctrain is randomly partitioned into a support set Csup and a query set Cqry , where Ctrain = Csup ∪ Cqry and Csup ∩ Cqry = ∅. The query set classes are treated as proxy-unseen classes, known training classes deliberately withheld at each episode to simulate novel conditions, forcing f θ to generalize its cross-modal alignment beyond the classes it is currently supervised on. Dual-Objective Distillation Loss: The loss function balances exact knowledge transfer on support classes with generalization on withheld query classes. Let si,c = ẑi · tc /∥ẑi ∥∥tc ∥ denote the cosine similarity between the projected instance and prototype tc . For the support set, the student is supervised via a distillation objective that combines soft knowledge transfer through KL divergence against the teacher’s similarity distribution with a hard cross-entropy target:      s ŝi,c Lsup = LKL σ i,c σ + LCE (ẑi , yi ) (4) T T where T is the distillation temperature, ŝi,c are the teacher’s similarity logits, and σ (·) denotes the softmax function. The KL divergence transfers the teacher’s soft relational structure across classes, while the cross-entropy term anchors the student to the correct hard label. For the query set, no distillation supervision is provided and the student is evaluated purely on prototype matching, acting as a generalization regularizer: Lqry = LCE (ẑi , yi ), i ∈ Cqry . The combined training objective is: Lalign = Lsup + λLqry (5) where λ ≥ 0 scales the generalization penalty relative to the distillation objective. 3.4

Generalized Inference via Adaptive Confidence Gating

Z-Score Confidence Estimation: The central observation motivating the gating mechanism is that a seen-class instance produces a similarity distribution over seen prototypes with one sharply dominant score, while an unseen-class instance produces a flatter, less discriminative distribution, as empirically confirmed in Figure 4. Rather than comparing the top seen-class score against a fixed threshold, we evaluate how statistically dominant it is relative to the per-sample distribution of all seen-class similarities. Given a test instance xi , we compute ẑi = f θ ( xi ) and the cosine similarity si,c = ẑi · tc /∥ẑi ∥∥tc ∥ against every seen prototype tc ∈ Ts . The per-sample mean µi and standard deviation σi of this seen-class similarity 5

Preprint. Under review.

Figure 4: Distribution of per-sample σi (standard deviation of seen-class cosine similarities) stratified by ground truth label.

distribution are then computed, and the Z-score of the maximum seen-class similarity smax = maxc∈S si,c measures how many standard deviations it lies above the mean: i smax − µi Zi = i (6) max(σi , ϵ) where ϵ ensures numerical stability when all seen-class similarities are nearly identical. Routing Decision: Zi quantifies how statistically dominant the best seen-class match is for a given sample. A high Zi indicates that one seen prototype produces a sharply dominant response, suggesting the instance belongs to a seen class. A low Zi indicates a flat similarity distribution, suggesting the instance is out-of-distribution with respect to S and likely belongs to an unseen class. The routing decision is formalized as:  arg maxc∈S si,c if Zi ≥ τ ŷi = (7) arg maxc∈U si,c if Zi < τ where τ is a confidence threshold optimized on the validation set to maximize the harmonic mean between seen and unseen class accuracy. This mechanism requires no additional learned parameters and operates entirely on the cosine similarity scores already computed during inference.

4

Experiments

Baselines: We evaluate nine methods capable of zero-shot tabular data classification — a critical requirement in cybersecurity — spanning three paradigms: generative (Mishra et al., 2018; Verma et al., 2020), cybersecurity-specific (Wang et al., 2025b; Aurna et al., 2025; Barros et al., 2022), and LLM-based (Shi et al., 2024; Wang et al., 2025a; Ali & Khan, 2023; Yun et al., 2024). We additionally evaluate TabPFN (Hollmann et al., 2023), APT (Wu & Bergman, 2025a), and ZEUS (Marszałek et al.), which generalise to novel datasets with minimal oversight but still require at least one sample from each unseen class at inference time. As such, these models do not strictly satisfy our GZSL formulation and are evaluated in the one-shot setting. Finally, HistGradBoost (Pedregosa et al., 2011) and XGBoost (Chen & Guestrin, 2016) serve as standard tabular classification baselines. Table 1 highlights four axes along which methods differ from SMETA-ZSL. ZET-LLM (Shi et al., 2024), ProtoLLM (Wang et al., 2025a), and SMELL (Barros et al., 2022) require architectural adaptation for zero-shot inference, while ZEUS, TabPFN, and APT do not support it at all and are therefore included only in the one-shot comparison. Only TabPFN and APT require unseen class prototypes at training time. CLIP-decoder (Ali & Khan, 2023), P2T (Yun et al., 2024), ZEUS, and TZSL (Wang et al., 2025b) require unlabelled unseen instances during training. Only CLIP-decoder and P2T support open-set recognition among the baselines. SMETA-ZSL is the only method satisfying all four criteria simultaneously. As baselines differ in which assumptions they relax, we evaluate each under its native setting while holding SMETA-ZSL to the strictest assumption throughout. 6

Preprint. Under review.

Table 1: Requirements and capabilities of baseline methods. Checkmarks indicate whether each method supports zero-shot classification natively, trains without unseen prototypes or unlabelled instances, and handles open-set recognition.

Method

Architecture

Learning Principle

Zero-Shot Classification

Trains w/o Unseen Prototypes

Trains w/o Unlabelled Unseen Instances

Open-Set

CVAE-ZSL (Mishra et al. (2018)) MZSL (Verma et al. (2020)) CLIP-decoder (Ali & Khan (2023)) P2T (Yun et al. (2024)) ZET-LLM (Shi et al. (2024)) ProtoLLM (Wang et al. (2025a)) ZEUS (Marszałek et al.) TabPFN (Hollmann et al. (2023)) APT (Wu & Bergman (2025a))

VAE GAN + MLP ViT LLM LLM + MLP LLM LLM LLM LLM

Generative, prototype Generative, meta-learning Transfer learning Transfer learning Feature extraction Prototype, transfer learning Federated learning Transductive Metric learning

✓ ✓ ✓ ✓ ✓(Modified) ✓(Modified) ✗ ✗ ✗

✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✗

✓ ✓ ✗ ✗ ✓ ✓ ✗ ✓ ✓

✗ ✗ ✓ ✓ ✗ ✗ ✗ ✗ ✗

SMELL (Barros et al. (2022)) FL-ZSL (Aurna et al. (2025)) TZSL (Wang et al. (2025b))

MLP MLP VQ-VAE

Metric learning Continual learning Transfer learning

✓(Modified) ✓ ✓

✓ ✓ ✓

✓ ✓ ✗

✗ ✗ ✗

SMETA-ZSL

LLM + MLP

Prototype, meta-learning

ZSL for Cybersecurity

Datasets: We evaluate SMETA-ZSL on seven benchmark datasets spanning four cybersecurity domains CIC-AndMal-2020 (Canadian Institute for Cybersecurity (CIC) and Canadian Centre for Cyber Security (CCCS) (2020)), BODMAS (Yang et al. (2021)), APIGRAPH (Zhang et al. (2020)), AVASTCTU (Bošanskỳ et al. (2022)) and three general-domain tabular datasets GOODREADS (Wan & McAuley (2018); Wan et al. (2019)), PETFINDER (PetFinder.my (2018)), FAKEDDIT (Nakamura et al. (2020)). Dataset details including the class distributions are provided in Table 2. Table 2: Overview of evaluation datasets spanning cybersecurity and general-domain benchmarks Dataset

Features

CIC-AndMal-2020 BODMAS APIGRAPH AVASTCTU GOODREADS PETFINDER FAKEDDIT

Dynamic behavioral features (API, memory, network, logcat) Static PE features (2,381-dim) Android API calls Dynamic sandbox execution logs (CAPEv2) Book metadata (page count, publication year, rating) Pet attributes (age, breed, gender) Post metadata (score, upvote ratio)

Size

Semantic Source

Total Classes

Seen

Unseen

400,000 134,435 322,594 ∼49,000 10,000 15,000 420,000

CTI reports CTI reports CTI reports CTI reports Book description Adoption profile Title/OCR text

26 44 74 10 8 5 6

22 40 69 7 6 3 4

4 4 5 3 2 2 2

Implementation Details: Semantic prototypes are encoded using LLaMA-3.1-8B, loaded in 4-bit precision via bitsandbytes and fine-tuned with LoRA to embed per-class behavioral descriptions into 4,096-dimensional prototype vectors. The projection network f θ is a twoblock residual MLP with hidden size 1,024, GELU activations, BatchNorm, and dropout (p=0.1), mapping tabular features into T with L2-normalized outputs. All hyperparameters are tuned on the validation set; sensitivity analysis is provided in Appendix 9. CTI Report Collection: CTI reports were collected from two open-source repositories: the ORKL Community CTI Library (ORKL (2024)), a searchable corpus of public threat intelligence reports, and the APT REPORT archive (CyberMonitor (2024)), a communitymaintained collection of vendor and government CTI documents. To ensure semantic coverage across all malware families, particularly for unseen classes with limited report availability, we employed a structured augmentation pipeline: a LLM was prompted to generate synthetic CTI variants conditioned on relevant MITRE ATT&CK technique descriptors (TTPs) corresponding to each malware family (The MITRE Corporation (2024)). LLM-based generation of CTI has been explored as a viable approach to address report scarcity in threat intelligence workflows (Ranade et al. (2021)). Evaluation Protocol: For finetuning of the LLM we use a 90/10 split for training and validation. Unseen classes are selected previously and held out during the whole fine-tuning process. During the meta-learning stage, we hold out 15% of the seen samples as validation and 15% for test.The number of unseen families varies between 3 and 2 depending on the dataset label size. We follow the standard GZSL evaluation protocol, reporting accuracy on seen classes (S), accuracy on unseen classes (U), and their harmonic mean (H) as the primary metric. The harmonic mean is used as the primary ranking criterion as it penalizes 7

Preprint. Under review.

methods that trivially favour either seen or unseen classes. Results are averaged over five random splits and reported with standard deviation.

5

Results

GZSL Baseline Comparison: Table 3 reports GZSL results across 5 runs for seven tabular benchmarks spanning both cybersecurity and general-domain datasets. On average, SMETAZSL improves the harmonic mean by approximately 10.8 points over the strongest baseline on the datasets where it leads, with the largest gains observed on CIC-AndMal (+18.1 over MZSL), GOODREADS (+11.6 over ProtoLLM), AVASTCTU (+12.5 over FL-ZSL) and APIGRAPH (+19.4 over TZSL). On FAKEDDIT, ProtoLLM leads (47.80 vs. 44.45), where prototype-based alignment appears particularly well-suited to the dataset’s multimodal label structure. SMETA-ZSL is the only method that performs consistently well across all seven datasets under strict inductive assumptions. Figure 5 further illustrates this trend: SMETA-ZSL attains the highest average unseen accuracy across all seven benchmarks while maintaining competitive seen accuracy. The full results for seen, unseen and mean accuracy across all 5 runs are given in Appendix. Table 3: GZSL results across seven tabular datasets. SMETA-ZSL is evaluated under the strictest inductive assumptions; baselines under their native settings. Best results (in blue) indicate the highest harmonic mean between seen and unseen class accuracy. Method

CIC-AndMal

BODMAS

APIGRAPH

AVASTCTU

GOODREADS

PETFINDER

FAKEDDIT

CVAE-ZSL MZSL CLIP P2T ZET-LLM ProtoLLM SMELL FL-ZSL TZSL

17.56 ± 2.61 39.70 ± 5.89 14.34 ± 3.72 0.00 ± 0.00 13.03 ± 1.20 38.15 ± 0.67 28.35 ± 1.48 29.43 ± 3.27 10.94 ± 4.88

30.20 ± 2.50 42.23 ± 4.97 43.15 ± 8.39 0.00 ± 0.00 32.32 ± 8.36 37.01 ± 0.19 27.60 ± 2.18 48.12 ± 1.76 41.20 ± 9.47

8.22 ± 2.32 23.21 ± 2.62 21.49 ± 4.55 0.47 ± 0.49 12.67 ± 1.57 21.59 ± 0.16 12.32 ± 5.45 27.16 ± 2.53 30.80 ± 3.59

24.85 ± 7.23 22.27 ± 14.74 39.14 ± 6.46 0.48 ± 0.10 0.01 ± 0.01 42.26 ± 4.15 36.71 ± 1.09 44.46 ± 5.54 29.67 ± 23.09

11.89 ± 1.72 22.50 ± 0.72 18.76 ± 0.90 7.81 ± 0.93 20.48 ± 0.55 24.34 ± 0.25 24.21 ± 1.54 21.89 ± 1.02 0.00 ± 0.00

1.46 ± 1.95 30.96 ± 1.05 32.24 ± 1.25 2.52 ± 0.68 23.14 ± 0.78 25.69 ± 0.91 3.15 ± 0.59 31.41 ± 0.42 6.13 ± 0.37

14.33 ± 4.05 38.16 ± 2.19 36.56 ± 14.66 0.55 ± 0.35 1.05 ± 0.19 47.80 ± 0.20 21.10 ± 5.52 21.20 ± 4.65 17.58 ± 1.78

SMETA-ZSL

57.78 ± 1.02

50.20 ± 9.10

50.19 ± 3.61

57.00 ± 1.22

35.92 ± 2.50

33.36 ± 0.58

44.45 ± 4.28

Figure 5: Average seen and unseen class accuracy across all benchmark datasets. Our proposed SMETA-ZSL achieves the highest unseen accuracy among all compared methods while maintaining competitive seen accuracy. One-Shot Baseline Comparison: Table 4 compares SMETA-ZSL (zero-shot) against few-shot and standard tabular baselines. SMETA-ZSL achieves the best harmonic mean on five of seven benchmarks, with margins of +15.0 on GOODREADS, +10.4 on PETFINDER, and +7.1 on APIGRAPH. On BODMAS and AVASTCTU, HistGradBoost and ZEUS respectively benefit from labelled support examples unavailable to SMETA-ZSL. Despite strictly zeroshot assumptions, SMETA-ZSL remains the most consistently strong method across the suite. Fully supervised upper-bound results are in Appendix 8. Effect of Embedding Model: Table 5 reports the effect of substituting the semantic encoder on GZSL performance, holding the remainder of SMETA-ZSL fixed. LLaMA3.1-8B is the best-performing encoder overall, achieving the highest scores on both CIC-AndMAL (56.70 8

Preprint. Under review.

Table 4: One-shot classification results across seven tabular benchmarks. SMETA-ZSL is compared against methods designed for few-shot tabular generalisation and standard tabular classifiers. Best results in blue. Method

CIC-AndMal

BODMAS

APIGRAPH

AVASTCTU

GOODREADS

PETFINDER

FAKEDDIT

TabPFN APT ZEUS HistGradBoost XGBoost

36.28 ± 0.38 51.92 ± 1.10 47.29 ± 0.69 37.21 ± 1.09 24.57 ± 7.60

49.09 ± 5.41 49.44 ± 0.72 58.72 ± 4.22 59.83 ± 3.18 27.85 ± 17.82

4.91 ± 6.61 11.54 ± 0.64 43.12 ± 0.93 38.93 ± 0.96 1.67 ± 1.08

42.14 ± 1.31 65.52 ± 0.66 79.87 ± 3.31 61.70 ± 2.11 9.26 ± 10.32

0.00 ± 0.00 20.57 ± 0.74 20.97 ± 3.77 0.02 ± 0.02 0.00 ± 0.00

0.00 ± 0.00 14.82 ± 1.40 23.00 ± 1.97 0.57 ± 0.06 0.00 ± 0.00

12.66 ± 8.36 10.59 ± 3.07 40.84 ± 3.50 1.04 ± 0.15 0.00 ± 0.00

SMETA-ZSL

56.7 ± 1.6

50.19 ± 3.61

50.20 ± 9.10

57.00 ± 1.22

35.92 ± 2.50

33.36 ± 0.58

44.45 ± 4.28

± 1.6) and BODMAS (50.19 ± 3.61). It also exhibits relatively low variance, suggesting stable semantic prototypes across splits. Table 5: Effect of language model on SMETA-ZSL Dataset

Qwen-3-4B

Gemma-3-4B

LLaMA3.1-3B

LLaMA3.1-8B

Mistral-Nemo-Base-12B

Qwen-14B

CIC-AndMAL BODMAS

55.45 ± 2.03 33.25 ± 2.76

54.34 ± 3.57 48.71 ± 2.29

45.61 ± 4.93 41.68 ± 1.96

56.70 ± 1.6 50.19 ± 3.61

37.14 ± 6.79 33.25 ± 2.76

48.58 ± 3.36 48.30 ± 5.02

Ablation Studies: Table 6 reports each component’s contribution via systematic ablation. Removing LLM semantic embeddings drops the harmonic mean by 9.3 points on average, confirming prototype quality as the primary performance driver. Replacing episodic metalearning with standard training degrades unseen-class accuracy substantially (−5.11 on APIGRAPH, −4.88 on AVASTCTU), validating that simulating the zero-shot condition during training is critical. Removing knowledge distillation yields a more moderate but consistent drop, indicating soft relational transfer provides complementary signal beyond hard label supervision. Few-Shot Setting: Appendix 16 examines how SMETA-ZSL performs as labeled samples per unseen class increase from zero to five. Performance generally continues to improve with more shots, with gains of up to 18.8 points on AVASTCTU at K = 5, though with diminishing returns beyond K = 2. Table 6: Ablation study across cybersecurity datasets. Each row removes or replaces one component of SMETA-ZSL. Configuration Full SMETA-ZSL w/o LLM semantic embeddings (random init) w/o meta-learning episodes w/o knowledge distillation Replace LLM with static GloVe vectors

CIC-AndMAL

BODMAS

APIGRAPH

AVASTCTU

56.7 36.12 54.63 55.78 22.99

59.10 48.71 59.11 57.41 33.95

50.19 46.12 45.08 44.55 39.84

57.00 54.92 52.12 52.12 49.33

Effect of Synthetic CTI: Our problem space has a core data scarcity issue: the number of real CTI reports per malware family is too small to provide sufficient within-class diversity for optimizing the supervised contrastive objective. Large-scale paired malware-CTI data is difficult to obtain in practice, as organizations often release malware artifacts or CTI descriptions, but rarely both Johnson et al. (2016); MISP Project (2024). This is precisely the premise of our problem setting: if the malware artifact is available, its corresponding attack description may not be, and conversely, if a CTI report describing an attack is released, the real malware sample may not be shared. To verify that training-time synthetic CTI does not confer an unfair advantage at test time, we ablate synthetic CTI entirely for unseen classes (i.e., classes present only at inference), constructing unseen-class prototypes from real CTI descriptions only, while keeping seenclass training unchanged. As shown in Table 7, unseen accuracy (U) and harmonic mean (H) are statistically indistinguishable with and without synthetic CTI across both datasets, 9

Preprint. Under review.

Table 7: Ablation on synthetic CTI usage for unseen-class prototype construction. S: seen accuracy, U: unseen accuracy, H: harmonic mean. Dataset

Unseen Prototype Source

S

U

H

CIC-AndMal

Real CTI only Real + Synthetic CTI

53.75 53.75

50.00 50.67

51.81 52.16

BODMAS

Real CTI only Real + Synthetic CTI

66.75 66.56

35.00 33.50

45.92 44.57

confirming that inference-time performance is driven by learned cross-modal alignment rather than synthetic CTI influencing prototype construction. Reproducibility Statement: All experiments reporting mean ± std are repeated over 5 runs with different seeds. Per seed result reported in Appendix. Hyperparameters are listed in Appendix 9. Semantic prototypes were generated with Llama-3.1-8B and are released as fixed .npy files to avoid non-determinism. Training was performed on 3× NVIDIA RTX ADA 48GB GPUs. Code, checkpoints, and preprocessing scripts are available at https: //github.com/Security-And-Intelligence-Lab-UTEP/SMETA-ZSL/blob/main/README.md

6

Conclusion

We presented SMETA-ZSL, a framework for a realistic but underexplored cyber-defense setting in which new threat classes emerge without labeled behavioral examples, unlabeled unseen instances, or prior knowledge of unseen classes. In practice, the only available supervision often comes from natural-language CTI reports. We show that language models can act as semantic adaptation interfaces for downstream cyber-defense systems by converting CTI into discriminative semantic prototypes that transfer to structured behavioral data. Across seven benchmarks spanning four cybersecurity and three general-domain datasets, SMETA-ZSL consistently outperforms prior methods under stricter open-set assumptions, improving baseline by an average of 10.8 points.

References Muhammad Ali and Salman Khan. Clip-decoder: Zeroshot multilabel classification using multimodal clip aligned representations. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4675–4679, 2023. Nahid Ferdous Aurna, Yuzo Taenaka, and Youki Kadobayashi. A feedback-driven federated zero-shot learning framework for adaptive detection of evolving banking malware. IEEE Access, 2025. Pedro H Barros, Eduarda TC Chagas, Leonardo B Oliveira, Fabiane Queiroz, and Heitor S Ramos. Malware-smell: A zero-shot learning strategy for detecting zero-day vulnerabilities. Computers & Security, 120:102785, 2022. Anna Bertiger, Bobby Filar, Aryan Luthra, Stefano Meschiari, Aiden Mitchell, Sam Scholten, and Vivek Sharath. Evaluating llm generated detection rules in cybersecurity. arXiv preprint arXiv:2509.16749, 2025. S Divakar Bhat, Amit More, Mudit Soni, and Bhuvan Aggarwal. Pc-gzsl: Prior correction for generalized zero shot learning. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 7173–7183. IEEE, 2025. Branislav Bošanskỳ, Daniel Kouba, Ondřej Manhal, Tomáš Sick, Viliam Lisy, Jakub Kroustek, and Petr Somol. Avast-ctu public cape dataset. arXiv preprint arXiv:2209.03188, 2022. Canadian Institute for Cybersecurity (CIC) and Canadian Centre for Cyber Security (CCCS). CCCS-CIC-AndMal-2020 Dataset. https://www.unb.ca/cic/datasets/andmal2020.html, 2020. 10

Preprint. Under review.

Arjun Chakraborty, Sandra Ho, Adam Cook, and Manuel Meléndez. Cti-realm: Benchmark to evaluate agent performance on security detection rule generation capabilities. arXiv preprint arXiv:2603.13517, 2026. Tianqi Chen and Carlos Guestrin. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794. ACM, 2016. doi: 10.1145/2939672.2939785. Zhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng, Wen Zhang, Yin Fang, Jeff Z Pan, and Huajun Chen. Duet: Cross-modal semantic grounding for contrastive zero-shot learning. In Proceedings of the AAAI conference on artificial intelligence, volume 37, pp. 405–413, 2023. CyberMonitor. APT CyberCriminal Campagin Collections. CyberMonitor/APT CyberCriminal Campagin Collections, 2024.

https://github.com/

Jinfu Fan, Jiangnan Li, Xiaowen Yan, Xiaohui Zhong, Wenpeng Lu, and Linqing Huang. Clip-driven zero-shot learning with ambiguous labels. arXiv preprint arXiv:2603.05053, 2026. Zhenyong Fu, Tao Xiang, Elyor Kodirov, and Shaogang Gong. Zero-shot object recognition by semantic manifold distance. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2635–2644, 2015. Zhenyong Fu, Tao Xiang, Elyor Kodirov, and Shaogang Gong. Zero-shot learning on semantic class prototype graph. IEEE transactions on pattern analysis and machine intelligence, 40(8):2009–2022, 2017. Naveen Gill et al. Llm-fs: Zero-shot feature selection for effective and interpretable malware detection. arXiv preprint arXiv:2602.09634, 2026. Noah Hollmann, Samuel Müller, Katharina Eggensperger, and Frank Hutter. TabPFN: A transformer that solves small tabular classification problems in a second. In Proceedings of the 11th International Conference on Learning Representations (ICLR), 2023. URL https: //arxiv.org/abs/2207.01848. Chris Johnson, Lee Badger, David Waltermire, Julie Snyder, Clem Skorupka, et al. Guide to cyber threat information sharing. NIST special publication, 800(150):35, 2016. Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. Advances in neural information processing systems, 33:18661–18673, 2020. Gukyeong Kwon and Ghassan Al Regib. A gating model for bias calibration in generalized zero-shot learning. IEEE Transactions on Image Processing, 2022. Jiaming Lei, Lin Li, Chunping Wang, Jun Xiao, and Long Chen. Seeing beyond classes: Zero-shot grounded situation recognition via language explainer. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 1602–1611, 2024. Adrian Shuai Li, Arun Iyengar, Ashish Kundu, and Elisa Bertino. Revisiting concept drift in windows malware detection: Adaptation to real drifted malware with minimal samples. arXiv preprint arXiv:2407.13918, 2024a. Yapeng Li, Yong Luo, Zengmao Wang, and Bo Du. Improving generalized zero-shot learning by exploring the diverse semantics from external class names. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23344–23353, 2024b. Peirong Ma and Xiao Hu. A variational autoencoder with deep embedding model for generalized zero-shot learning. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 11733–11740, 2020. 11

Preprint. Under review.

Patryk Marszałek, Tomasz Kuśmierczyk, Witold Wydmański, Jacek Tabor, and Marek Śmieja. Zeus: Zero-shot embeddings for unsupervised separation of tabular data. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. Ashish Mishra, Shiva Krishna Reddy, Anurag Mittal, and Hema A Murthy. A generative model for zero shot learning using conditional variational autoencoders. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 2188–2196, 2018. MISP Project. MISP – open source threat intelligence platform & open standards for threat information sharing. https://www.misp-project.org/, 2024. Shaswata Mitra, Azim Bazarov, Martin Duclos, Sudip Mittal, Aritran Piplai, Md Rayhanur Rahman, Edward Zieglar, and Shahram Rahimi. Falcon: Autonomous cyber threat intelligence mining with llms for ids rule generation. arXiv preprint arXiv:2508.18684, 2025. Kai Nakamura, Sharon Levy, and William Yang Wang. r/fakeddit: A new multimodal benchmark dataset for fine-grained fake news detection. In Proceedings of the 12th Language Resources and Evaluation Conference, pp. 6149–6158, 2020. Shah Nawaz, Jacopo Cavazza, and Alessio Del Bue. Semantically grounded visual embeddings for zero-shot learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4589–4599, 2022. ORKL. The ORKL Community CTI Library. https://orkl.eu/, 2024. Accessed: 2026-03-31. Haojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang, Yaliang Li, and Jun Huang. Meta-kd: A meta knowledge distillation framework for language model compression across domains. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 3026–3036, 2021. Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825– 2830, 2011. PetFinder.my. Petfinder.my adoption prediction. petfinder-adoption-prediction, 2018.

https://www.kaggle.com/c/

Farhad Pourpanah, Moloud Abdar, Yuxuan Luo, Xinlei Zhou, Ran Wang, Chee Peng Lim, XiZhao Wang, and QM Jonathan Wu. A review of generalized zero-shot learning methods. IEEE transactions on pattern analysis and machine intelligence, 45(4):4051–4070, 2022. Xingzhi Qian, Xinran Zheng, Yiling He, Shuo Yang, and Lorenzo Cavallaro. Lamd: Contextdriven android malware detection and classification with llms. In 2025 IEEE Security and Privacy Workshops (SPW), pp. 126–136. IEEE, 2025. Priyanka Ranade, Aritran Piplai, Sudip Mittal, Anupam Joshi, and Tim Finin. Generating fake cyber threat intelligence using transformer-based models. In 2021 International Joint Conference on Neural Networks (IJCNN), pp. 1–9, 2021. doi: 10.48550/arXiv.2102.04351. Zhijie Rao, Jingcai Guo, Xiaocheng Lu, Jingming Liang, Jie Zhang, Haozhao Wang, Kang Wei, and Xiaofeng Cao. Dual expert distillation network for generalized zero-shot learning. arXiv preprint arXiv:2404.16348, 2024. Bernardino Romera-Paredes and Philip Torr. An embarrassingly simple approach to zeroshot learning. In International conference on machine learning, pp. 2152–2161. PMLR, 2015. Ivan Montoya Sanchez, Shaswata Mitra, Aritran Piplai, and Sudip Mittal. Semantic-aware contrastive fine-tuning: Boosting multimodal malware classification with discriminative embeddings. In 2025 International Joint Conference on Neural Networks (IJCNN), pp. 1–8. IEEE, 2025. 12

Preprint. Under review.

Yuval Schwartz, Lavi Ben-Shimol, Dudu Mimran, Yuval Elovici, and Asaf Shabtai. Llmcloudhunter: Harnessing llms for automated extraction of detection rules from cloud-based cti. In Proceedings of the ACM on Web Conference 2025, pp. 1922–1941, 2025. Zhiyi Shi, Junsik Kim, Davin Jeong, and Hanspeter Pfister. Surprisingly simple: Large language models are zero-shot feature extractors for tabular and text data. 2024. Bowen Tang, Jing Zhang, Long Yan, Qian Yu, Lu Sheng, and Dong Xu. Data-free generalized zero-shot learning. In Proceedings of the AAAI conference on artificial intelligence, volume 38, pp. 5108–5117, 2024. Lihong Tang, Xiao Chen, Sheng Wen, Li Li, Marthie Grobler, and Yang Xiang. Demystifying the evolution of android malware variants. IEEE Transactions on Dependable and Secure Computing, 21(4):3324–3341, 2023. The MITRE Corporation. MITRE ATT&CK. https://attack.mitre.org/, 2024. Accessed: 2026-03-31. Vinay Kumar Verma, Dhanajit Brahma, and Piyush Rai. Meta-learning for generalized zero-shot learning. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 6062–6069, 2020. Mengting Wan and Julian McAuley. Item recommendation on monotonic behavior chains. In Proceedings of the 12th ACM Conference on Recommender Systems, pp. 86–94, 2018. Mengting Wan, Rishabh Misra, Ndapa Nakashole, and Julian McAuley. Fine-grained spoiler detection from large-scale review corpora. arXiv preprint arXiv:1905.13416, 2019. Chaoqun Wang, Shaobo Min, Xuejin Chen, Xiaoyan Sun, and Houqiang Li. Dual progressive prototype network for generalized zero-shot learning. Advances in Neural Information Processing Systems, 34:2936–2948, 2021. Peng Wang, Dongsheng Wang, He Zhao, Hangting Ye, Dandan Guo, and Yi Chang. Llm empowered prototype learning for zero and few-shot tasks on tabular data. arXiv preprint arXiv:2508.09263, 2025a. Ping Wang, Hao-Cyuan Li, Hsiao-Chung Lin, Wen-Hui Lin, and Nian-Zu Xie. A transductive zero-shot learning framework for ransomware detection using malware knowledge graphs. Information, 16(6):458, 2025b. Wei Wang, Vincent W Zheng, Han Yu, and Chunyan Miao. A survey of zero-shot learning: Settings, methods, and applications. ACM Transactions on Intelligent Systems and Technology (TIST), 10(2):1–37, 2019. Yulun Wu and Doron L. Bergman. Zero-shot meta-learning for tabular prediction tasks with adversarially pre-trained transformer. In Proceedings of the 42nd International Conference on Machine Learning. PMLR, 2025a. URL https://arxiv.org/abs/2502.04573. ICML. Yulun Wu and Doron L Bergman. Zero-shot meta-learning for tabular prediction tasks with adversarially pre-trained transformer. In International Conference on Machine Learning, pp. 67111–67127. PMLR, 2025b. Yongqin Xian, Bernt Schiele, and Zeynep Akata. Zero-shot learning-the good, the bad and the ugly. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4582–4591, 2017. Limin Yang, Arridhana Ciptadi, Ihar Laziuk, Ali Ahmadzadeh, and Gang Wang. Bodmas: An open dataset for learning based temporal analysis of pe malware. In 4th Deep Learning and Security Workshop (DLS), 2021. Sukmin Yun, Jaehyun Nam, Woomin Song, Seong Hyeon Park, Jihoon Tack, Jaehyung Kim, Kyu Hwan Oh, and Jinwoo Shin. Tabular transfer learning via prompting llms. In Conference on Language Modeling (COLM), pp. 1–18. COLM Organizing Committee, 2024. 13

Preprint. Under review.

Xueqiang Zhang, Yating Zhang, Ming Zhong, Dandan Ding, Yinzhi Cao, Yin Zhang, Min Zhang, and Min Yang. Enhancing state-of-the-art classifiers with api semantics to detect evolved android malware. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, pp. 757–770, 2020. Xin Zhou, Kisub Kim, Bowen Xu, DongGyun Han, and David Lo. Out of sight, out of mind: Better automatic vulnerability repair by broadening input ranges and sources. In Proceedings of the IEEE/ACM 46th international conference on software engineering, pp. 1–13, 2024.

A

Appendix

Table 8: Fully supervised ceiling performance across seven tabular benchmarks. Each method is trained and evaluated on all classes with full label access, representing the upper bound that zero-shot methods aspire to approach. Results are mean ± standard deviation over five runs. Method TabPFN APT ZEUS HistGradBoost XGBoost MLP

CIC-AndMal-2020

BODMAS

APIGRAPH

AVASTCTU

GOODREADS

PETFINDER

FAKEDDIT

85.86 ± 1.05 66.85 ± 0.81 56.40 ± 1.90 83.76 ± 1.44 83.27 ± 1.70 77.15 ± 0.15

94.92 ± 0.33 56.26 ± 0.53 82.84 ± 1.36 93.62 ± 0.81 94.06 ± 0.57 90.72 ± 0.59

77.08 ± 1.15 15.52 ± 1.09 57.59 ± 0.69 78.83 ± 1.27 79.63 ± 1.25 77.88 ± 1.03

99.35 ± 0.30 98.55 ± 0.21 93.26 ± 1.07 98.78 ± 0.16 98.81 ± 0.15 98.67 ± 0.15

58.91 ± 0.66 41.97 ± 0.39 32.81 ± 2.28 58.69 ± 0.86 59.33 ± 0.67 47.71 ± 1.28

63.75 ± 0.62 61.97 ± 1.19 38.64 ± 2.26 36.14 ± 0.45 39.53 ± 0.94 36.08 ± 0.97

99.99 ± 0.02 99.99 ± 0.02 91.16 ± 1.08 99.99 ± 0.01 99.99 ± 0.01 98.91 ± 0.68

Table 9: Sensitivity of SMETA to key hyperparameters across cybersecurity datasets. H = Harmonic Mean (%). One parameter is varied at a time while others are fixed to their default values (marked with ⋆). Stable H across values indicates low sensitivity to that parameter. Hyperparameter

Value

CIC-AndMAL

BODMAS

APIGRAPH

Query Loss λqry

0.1 0.5 1.0 2.0⋆ 3.0 4.0

53.7 50.6 46.7 44.8 42.9 42.4

64.1 59.2 50.6 47.9 47.8 47.8

33.3 29.3 33.7 32.6 33.2 31.7

KD Temp tkd

1.0 2.0 3.0⋆ 4.0 5.0

44.6 45.0 44.8 44.8 44.5

48.5 48.2 47.9 48.2 48.0

32.6 32.6 32.6 32.4 32.4

Alpha α

0.1 0.3 0.5 0.7⋆ 0.9

48.3 46.8 45.9 44.8 42.5

53.6 51.0 50.1 47.9 51.3

33.6 33.4 32.3 32.6 32.4

Ratio (S, U)

(15, 3) (13, 5)⋆ (10, 8) (8, 10) (5, 13)

35.9 44.8 51.7 53.0 54.4

36.9 47.9 46.9 61.0 62.4

33.9 32.6 35.2 30.5 33.4

14

Preprint. Under review.

Table 10: GZSL results on CIC-AndMal-2020. Seen = Seen Acc (%), Unseen = Unseen Acc (%), Mean = H-Mean (%). Best Mean in bold. Class

CVAE-ZSL

MZSL

CLIP-Decoder

P2T

ZET-LLM

ProtoLLM

SMELL

FL-ZSL

TZSL

SMETA-ZSL

47.87 33.05 39.10

19.89 56.26 29.39

61.88 22.97 33.50

11.29 25.87 15.72

66.59 49.54 56.81

43.97 32.83 37.59

21.05 55.90 30.58

61.29 20.53 30.76

8.00 17.40 10.96

66.12 53.71 59.27

45.91 33.27 38.58

18.81 56.17 28.19

62.71 17.05 26.81

10.00 6.03 7.53

66.94 51.62 58.29

44.65 32.40 37.55

17.66 55.10 26.74

61.76 16.01 25.43

11.65 24.13 15.71

66.12 50.93 57.54

44.52 33.04 37.93

17.35 59.12 26.82

60.47 20.53 30.66

2.71 20.88 4.79

67.06 49.54 56.98

45.38±1.56 32.92±0.33 38.15±0.67

18.95±1.38 56.51±1.37 28.35±1.48

61.62±0.82 19.42±2.84 29.43±3.27

8.73±3.66 18.86±7.87 10.94±4.88

66.56±0.44 51.07±1.73 57.78±1.02

Run 1 (Seed 42) Seen Unseen Mean

48.23 9.10 15.32

57.76 27.73 37.47

24.59 6.50 10.28

5.41 0.00 0.00

56.22 6.54 11.72

Seen Unseen Mean

48.69 11.06 18.02

60.12 33.29 42.86

18.59 21.69 20.02

5.41 0.00 0.00

Seen Unseen Mean

47.18 8.57 14.51

59.06 42.81 49.64

18.82 9.40 12.54

5.29 0.00 0.00

Seen Unseen Mean

49.14 11.99 19.28

70.47 23.20 34.91

19.29 13.23 15.69

5.42 0.00 0.00

Seen Unseen Mean

48.61 13.13 20.68

58.94 23.55 33.65

18.47 10.21 13.15

5.41 0.00 0.00

Seen Unseen Mean

48.37±0.74 10.77±1.92 17.56±2.61

61.27±4.66 30.12±7.32 39.70±5.89

19.95±2.61 12.20±5.82 14.34±3.72

5.39±0.05 0.00±0.00 0.00±0.00

Run 2 (Seed 123) 56.25 8.40 14.62 Run 3 (Seed 456) 56.13 6.92 12.32 Run 4 (Seed 789) 56.57 7.94 13.93 Run 5 (Seed 2025) 56.07 7.09 12.58 Average ± Std 56.25±0.19 7.38±0.77 13.03±1.20

Table 11: GZSL results on BODMAS. Seen = Seen Acc (%), Unseen = Unseen Acc (%), Mean = H-Mean (%). Best Mean in bold. Class

CVAE-ZSL

MZSL

CLIP-Decoder

P2T

ZET-LLM

ProtoLLM

SMELL

FL-ZSL

TZSL

SMETA-ZSL

72.64 25.00 37.20

21.46 27.00 23.92

64.69 41.50 50.56

57.56 58.67 58.11

83.25 50.00 62.48

70.20 24.83 36.69

43.96 21.67 29.03

64.06 37.00 46.91

45.00 32.50 37.74

76.31 32.00 45.09

72.98 24.83 37.06

47.56 22.17 30.24

65.00 36.17 46.47

46.00 31.17 37.16

82.31 44.00 57.35

71.83 25.00 37.09

37.01 20.83 26.66

64.06 40.17 49.38

38.88 34.00 36.27

74.00 31.00 43.70

71.29 25.00 37.02

32.01 25.17 28.18

63.56 37.67 47.30

42.13 32.50 36.69

74.31 29.67 42.40

71.79±1.11 24.93±0.09 37.01±0.19

36.40±9.21 23.37±2.33 27.60±2.18

64.28±0.57 38.50±2.25 48.12±1.76

45.91±7.08 37.77±11.73 41.20±9.47

78.04±4.43 37.33±9.11 50.20±9.10

Run 1 (Seed 42) Seen Unseen Mean

30.50 37.17 33.50

92.38 24.17 38.31

62.44 33.50 43.60

3.31 0.00 0.00

Seen Unseen Mean

23.75 32.00 27.26

92.75 24.00 38.13

60.44 19.50 29.49

3.50 0.00 0.00

Seen Unseen Mean

30.44 32.67 31.51

90.56 30.33 45.45

64.69 34.83 45.28

3.50 0.00 0.00

Seen Unseen Mean

28.75 32.33 30.44

92.69 24.50 38.76

57.38 48.33 52.47

3.50 0.00 0.00

Seen Unseen Mean

25.25 32.17 28.29

91.94 34.83 50.52

63.25 34.83 44.93

3.50 0.00 0.00

Seen Unseen Mean

27.74±3.08 33.27±2.19 30.20±2.50

92.06±0.80 27.57±4.34 42.23±4.97

61.64±2.84 34.20±10.21 43.15±8.39

3.46±0.08 0.00±0.00 0.00±0.00

85.31 18.67 30.63 Run 2 (Seed 123) 86.31 28.00 42.28 Run 3 (Seed 456) 85.06 12.00 21.03 Run 4 (Seed 789) 85.56 17.50 29.06 Run 5 (Seed 2025) 86.50 24.83 38.59 Average ± Std 85.75±0.63 20.20±6.31 32.32±8.36

15

Preprint. Under review.

Table 12: GZSL results on APIGRAPH. Seen = Seen Acc (%), Unseen = Unseen Acc (%), Mean = H-Mean (%). Best Mean in bold. Class

CVAE-ZSL

MZSL

CLIP-Decoder

P2T

ZET-LLM

ProtoLLM

SMELL

FL-ZSL

TZSL

SMETA-ZSL

43.77 14.58 21.87

53.83 4.23 7.85

53.08 19.04 28.03

32.23 44.55 37.40

46.32 57.92 51.48

41.29 14.48 21.44

49.46 5.46 9.84

52.53 19.76 28.72

29.22 28.50 28.86

39.73 64.21 49.09

42.76 14.38 21.52

54.86 13.52 21.70

52.53 14.97 23.30

29.22 33.41 31.18

48.33 65.85 55.74

41.17 14.63 21.59

49.15 6.83 11.99

53.23 17.25 26.05

37.69 20.84 26.84

40.35 54.37 46.32

42.18 14.48 21.55

57.59 5.60 10.21

52.33 20.72 29.68

31.93 27.78 29.71

38.09 66.12 48.33

42.23±1.05 14.51±0.09 21.59±0.16

52.98±3.63 7.13±3.69 12.32±5.45

52.74±0.39 18.35±2.28 27.16±2.53

32.06±3.09 31.02±7.86 30.80±3.59

42.56±4.48 61.69±5.27 50.19±3.61

Run 1 (Seed 42) Seen Unseen Mean

23.38 7.36 11.20

71.13 12.34 21.02

58.40 10.54 17.86

1.45 0.00 0.00

34.62 6.80 11.37

Seen Unseen Mean

25.39 4.45 7.58

71.23 17.25 27.77

57.29 16.77 25.94

2.23 0.15 0.28

Seen Unseen Mean

23.56 4.75 7.91

73.63 14.73 24.55

57.94 15.69 24.69

2.58 0.00 0.00

Seen Unseen Mean

21.46 2.81 4.97

74.89 12.46 21.36

57.29 14.73 23.44

1.35 0.61 0.84

Seen Unseen Mean

22.86 5.96 9.46

74.19 12.46 21.33

56.84 8.98 15.51

1.99 0.91 1.25

Seen Unseen Mean

23.33±1.42 5.07±1.71 8.22±2.32

73.01±1.55 13.84±1.92 23.21±2.62

57.55±0.61 13.34±3.39 21.49±4.55

1.92±0.47 0.33±0.36 0.47±0.49

Run 2 (Seed 123) 33.92 6.70 11.19 Run 3 (Seed 456) 34.20 7.50 12.30 Run 4 (Seed 789) 34.27 9.50 14.88 Run 5 (Seed 2025) 34.43 8.50 13.63 Average ± Std 34.29±0.26 7.80±1.19 12.67±1.57

Table 13: GZSL results on AVASTCTU. Seen = Seen Acc (%), Unseen = Unseen Acc (%), Mean = H-Mean (%). Best Mean in bold. Class

CVAE-ZSL

MZSL

CLIP-Decoder

P2T

ZET-LLM

ProtoLLM

SMELL

FL-ZSL

TZSL

SMETA-ZSL

90.07 31.40 46.57

81.18 25.40 38.69

96.45 30.60 46.46

52.50 58.62 55.39

96.68 42.07 58.63

90.48 28.68 43.56

81.13 23.42 36.34

95.61 26.23 41.17

18.50 39.62 25.22

88.28 40.83 55.84

91.73 27.55 42.37

80.90 23.17 36.02

95.84 29.72 45.37

35.54 11.62 17.51

97.43 39.87 56.58

89.74 28.65 43.43

81.18 22.77 35.56

96.40 35.58 51.98

83.65 35.90 50.24

95.89 39.57 56.02

90.30 22.00 35.38

81.13 23.90 36.92

95.14 23.23 37.35

0.00 41.92 0.00

97.43 41.20 57.91

90.46±0.76 27.66±3.47 42.26±4.15

81.10±0.10 23.73±0.91 36.71±1.09

95.89±0.55 29.07±4.67 44.46±5.54

38.04±32.11 37.53±16.90 29.67±23.09

95.14±3.89 40.71±1.01 57.00±1.22

Run 1 (Seed 42) Seen Unseen Mean

46.18 22.33 30.11

99.44 14.90 25.92

65.30 35.73 46.19

14.29 0.33 0.65

99.48 0.00 0.00

Seen Unseen Mean

40.81 24.07 30.28

99.44 0.78 1.55

62.03 34.67 44.48

14.29 0.18 0.36

Seen Unseen Mean

31.09 13.92 19.23

99.39 11.35 20.37

62.12 19.78 30.01

14.29 0.23 0.46

Seen Unseen Mean

37.53 9.37 14.99

99.44 30.72 46.94

63.80 26.48 37.43

14.29 0.27 0.52

Seen Unseen Mean

29.79 29.52 29.65

99.67 9.05 16.59

37.27 37.88 37.58

14.29 0.20 0.39

Seen Unseen Mean

37.08±6.82 19.84±8.10 24.85±7.23

99.48±0.10 13.36±9.84 22.27±14.74

58.10±11.72 30.91±7.57 39.14±6.46

14.29±0.00 0.24±0.05 0.48±0.10

Run 2 (Seed 123) 99.48 0.00 0.00 Run 3 (Seed 456) 99.48 0.00 0.00 Run 4 (Seed 789) 99.48 0.00 0.00 Run 5 (Seed 2025) 99.48 0.00 0.00 Average ± Std 99.48±0.00 0.00±0.00 0.01±0.01

16

Preprint. Under review.

Table 14: GZSL results on GOODREADS. Seen = Seen Acc (%), Unseen = Unseen Acc (%), Mean = H-Mean (%). Best Mean in bold. Class

CVAE-ZSL

MZSL

CLIP-Decoder

P2T

ZET-LLM

ProtoLLM

SMELL

FL-ZSL

TZSL

SMETA-ZSL

23.92 24.75 24.33

21.28 36.18 26.80

34.53 17.65 23.36

16.67 0.00 0.00

34.03 47.50 39.65

24.27 24.77 24.52

21.64 22.33 21.98

32.75 16.52 21.96

16.67 0.00 0.00

25.06 50.77 33.55

23.36 24.78 24.05

22.17 25.93 23.90

34.36 15.30 21.17

16.67 0.00 0.00

31.56 43.73 36.66

24.45 24.88 24.67

20.81 28.20 23.94

36.33 16.05 22.26

16.67 0.00 0.00

31.47 35.80 33.50

23.54 24.82 24.16

20.92 29.32 24.41

31.78 15.37 20.72

16.67 0.00 0.00

30.67 44.30 36.24

23.91±0.47 24.80±0.05 24.34±0.25

21.36±0.50 28.39±4.57 24.21±1.54

33.95±1.76 16.18±0.97 21.89±1.02

16.67±0.00 0.00±0.00 0.00±0.00

30.56±3.32 44.42±5.59 35.92±2.50

Run 1 (Seed 42) Seen Unseen Mean

6.83 36.53 11.51

25.28 18.72 21.51

12.50 31.00 17.82

17.89 4.58 7.30

14.56 31.28 19.87

Seen Unseen Mean

7.72 45.02 13.18

26.89 20.03 22.96

13.06 34.58 18.96

26.69 4.00 6.96

Seen Unseen Mean

6.58 42.53 11.40

16.75 31.13 21.78

14.19 33.23 19.89

29.67 4.40 7.66

Seen Unseen Mean

8.67 34.92 13.89

27.11 20.53 23.37

13.75 32.05 19.24

28.53 4.33 7.52

Seen Unseen Mean

5.33 42.93 9.49

15.56 43.35 22.90

12.86 29.33 17.88

21.97 6.15 9.61

Seen Unseen Mean

7.03±1.25 40.39±4.40 11.89±1.72

22.32±5.09 26.75±9.41 22.50±0.72

13.27±0.69 32.04±2.02 18.76±0.90

24.95±4.40 4.69±0.75 7.81±0.93

Run 2 (Seed 123) 15.81 32.50 21.27 Run 3 (Seed 456) 15.36 29.18 20.13 Run 4 (Seed 789) 15.50 29.78 20.39 Run 5 (Seed 2025) 15.06 33.45 20.77 Average ± Std 15.26±0.48 31.24±1.79 20.48±0.55

Table 15: GZSL results on PETFINDER. Seen = Seen Acc (%), Unseen = Unseen Acc (%), Mean = H-Mean (%). Best Mean in bold. Class

CVAE-ZSL

MZSL

CLIP-Decoder

P2T

ZET-LLM

ProtoLLM

SMELL

FL-ZSL

TZSL

SMETA-ZSL

27.94 25.72 26.78

53.71 1.90 3.67

44.30 24.19 31.29

3.91 25.08 6.76

29.01 39.33 33.39

24.15 25.66 24.88

51.13 1.71 3.30

44.24 24.51 31.54

3.71 17.12 6.10

30.79 38.59 34.26

27.41 25.53 26.44

52.72 1.02 1.99

44.50 25.07 32.07

4.04 11.57 5.99

27.48 40.18 32.64

25.50 25.63 25.57

50.93 1.81 3.50

44.24 23.92 31.05

3.38 19.87 5.77

28.15 40.89 33.34

23.98 25.59 24.76

52.98 1.69 3.28

43.58 24.14 31.07

4.30 10.03 6.02

29.27 38.24 33.16

25.80±1.82 25.63±0.07 25.69±0.91

52.29±1.08 1.63±0.31 3.15±0.59

44.17±0.35 24.37±0.45 31.41±0.42

3.87±0.35 16.73±6.15 6.13±0.37

28.94±1.26 39.45±1.10 33.36±0.58

Run 1 (Seed 42) Seen Unseen Mean

1.95 48.47 3.76

36.95 27.87 31.77

32.52 36.24 34.28

28.20 1.64 3.09

25.40 19.85 22.28

Seen Unseen Mean

0.05 49.44 0.11

27.15 34.05 30.21

31.79 31.47 31.63

31.87 0.78 1.52

Seen Unseen Mean

0.00 49.48 0.00

29.74 29.99 29.86

31.52 30.74 31.13

38.23 1.72 3.29

Seen Unseen Mean

1.78 49.86 3.43

34.97 30.56 32.62

35.43 30.08 32.53

30.65 1.43 2.73

Seen Unseen Mean

0.00 50.27 0.00

25.23 38.02 30.33

32.85 30.53 31.65

34.18 1.01 1.95

Seen Unseen Mean

0.76±1.01 49.51±0.67 1.46±1.95

30.81±4.49 32.10±3.57 30.96±1.05

32.82±1.55 31.81±2.53 32.24±1.25

32.63±3.40 1.32±0.36 2.52±0.68

Run 2 (Seed 123) 27.05 21.45 23.93 Run 3 (Seed 456) 26.30 19.37 22.31 Run 4 (Seed 789) 27.15 20.85 23.59 Run 5 (Seed 2025) 25.78 21.75 23.60 Average ± Std 26.34±0.77 20.65±1.02 23.14±0.78

Table 16: SMETA-ZSL performance under varying numbers of labeled samples per unseen family (K-shot) across cybersecurity datasets. K = 0 is the zero-shot setting; S = Seen Acc (%), U = Unseen Acc (%), H = Harmonic Mean (%). CIC-AndMAL-2020

BODMAS

AVASTCTU

K

S

U

H

S

U

H

S

U

H

0 (zero-shot) 1 2 5

66.56 57.41 61.53 57.53

51.07 64.45 61.36 61.88

57.78 60.73 61.44 59.62

77.75 80.69 78.94 80.69

47.67 55.28 66.50 65.13

59.10 65.61 72.19 72.08

99.58 95.80 95.75 95.84

34.28 54.96 63.76 68.17

51.01 69.85 76.55 79.67

17

Record · ID 363204 · SHA-256 35753d6074ef4dc7
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.