arXiv:2609.16433v1 [cs.CR] 14 Sep 2026
Evaluating the NIST Bugs Framework Against CWE as a Successor for Automated Vulnerability Classification MD NAZMUL HOQUE, The University of Alabama, USA SHASWATA MITRA, The University of Alabama, USA SUBASH NEUPANE, The University of Alabama, USA SUDIP MITTAL, The University of Alabama, USA SHAHRAM RAHIMI, The University of Alabama, USA Vulnerability classification based on root cause weaknesses is essential for numerous cybersecurity activities, where the Common Weakness Enumeration (CWE) serves as a public repository of such flaws. However, its overlapping entries create a non-orthogonal structure. The result is the same vulnerability being mapped to multiple weaknesses, complicating Root Cause Analysis (RCA) and triage. To address this, NIST Special Publication 800-231 introduces the Bugs Framework (BF), which organizes vulnerabilities into ⟨cause, operation, consequence⟩ triples and links such triples into a causal chain, so that a vulnerability carries its root cause and its sink together instead of a single terminal label. To date, however, BF has been specified but not evaluated regarding its performance against the challenges to automated classification. Additionally, the evidence required for adoption has not been investigated empirically. In this paper, we evaluate BF as a classification target and a complement to CWE using a systematically screened corpus of automated Common Vulnerabilities and Exposures (CVEs) linked to CWE research. We assess the reproducibility of CVE-to-BF classification through two evaluations. The first is qualitative: an anonymized inter-rater study in which 2 subject-matter experts (SMEs) independently mapped 13 CVEs onto the four BF axes. Annotators showed strong agreement on the cause and operation axes, while the attribute axis indicated fair agreement. We also tested our automated framework1 across two large language model (LLM) deployments under different budgets for reproducibility analysis. Despite limitations, such as evidence availability and the absence of retrievable fix commits for closed-source software, our findings support the claim that BF is a more structured and automation-friendly framework than CWE. Our exploration reveals specific gaps in BF, including under-specified guidance on attributes. CCS Concepts: • Security and privacy → Software and application security; • Information systems → Information retrieval. Additional Key Words and Phrases: Bugs Framework (BF), Common Weakness Enumeration (CWE), Vulnerability Classification, Root Cause Analysis (RCA), Vulnerability Chain Inference, Cybersecurity, Large Language Model (LLM), Artificial Intelligence (AI)
1
Introduction
Vulnerability classification is the semantic layer on which much of modern security operations depends. When this layer captures only the surface form of a weakness, and not the causal chain that produces it, remediation can miss the real defect at serious cost. The 2017 Equifax breach shows the operational impact of this distinction. Attackers exploited a known Apache Struts flaw (CVE-2017-5638) [57] that had been publicly disclosed and flagged by US-CERT months earlier, but the vulnerable component was never patched. The National Vulnerability Database labels that flaw CWE-755 (Improper Handling of Exceptional Conditions) [60], a generic weakness class that 1 Experiment and Framework Code: github.com/shaswata09/cve2bf
Authors’ Contact Information: Md Nazmul Hoque, [email protected], The University of Alabama, Tuscaloosa, Alabama, USA; Shaswata Mitra, [email protected], The University of Alabama, Tuscaloosa, Alabama, USA; Subash Neupane, [email protected], The University of Alabama, Tuscaloosa, Alabama, USA; Sudip Mittal, [email protected], The University of Alabama, Tuscaloosa, Alabama, USA; Shahram Rahimi, [email protected], The University of Alabama, Tuscaloosa, Alabama, USA.
1:2
Hoque et al.
names the failure mode, not the root cause an engineer must locate and fix. Intruders operated undetected for 76 days and exposed the personal records of roughly 148 million consumers [83, 84]. A published technical analysis attributes the breach to the gap between the disclosed weakness and the specific vulnerable library inside a legacy application [32]. Equifax later settled with United States regulators for at least $575 million [82]. The quality of vulnerability classification is therefore an operational concern, not only a theoretical one. CVE-2014-0160 (Heartbleed) one disclosed vulnerability current practice
(a) Current target: one flat CWE label
CWE-125 Out-ofbounds Read assigned by NVD (the sink)
CWE-119 parent class, also defensible CWE-20 root-cause reading, also defensible
flat, sink-named, unordered: no canonical choice among the labels (failures F1 to F4; the survey corpus meets them as a performance ceiling)
evaluated in this paper
(b) Evaluated target: BF causal weakness chain DVR (Missing Code, Verify) → Inconsistent Value
MAD (Wrong Size, Reposition) → Overbound Pointer
root cause to sink, ordered; one class per operation (disjoint operation sets)
MUS (Overbound Pointer, Read) → Buffer Over-Read
IEX failure
The CWE label admits defensible alternatives. The full BF chain is in Figure 14. Fig. 1. The same CVE mapped to both targets: a flat sink-level CWE label (left) against the ordered BF weakness chain (right).
CVE records [39] identify such disclosed vulnerabilities, but downstream systems require additional metadata: severity, affected platforms, exploitability, cause of the weakness, likely consequences, and defensive relevance. The volume of disclosures continues to grow at approximately 263% between 2020 and 2025 [61], with more than 40,000 CVEs published in 2024 and over 48,000 in 2025. Each disclosure must be enriched with a severity score (CVSS) [18, 33], a weakness label (CWE [40]), and context that maps the disclosure into adversary models such as MITRE ATT&CK [51] and D3FEND [50]. In practice, CWE has become the principal target of weakness for this enrichment because of its breadth and community adoption. The same breadth, however, exposes a central problem for trustworthy automation. CWE entries overlap, sit at different abstraction levels, and often describe the final manifestation of a defect instead of the defect’s root cause [6], so the same CVE admits several defensible labels and independent annotators or classifiers legitimately diverge. The operational cost of this divergence is directly observable: tool-assisted CVE classification becomes less definitive, cross-report correlation weakens, and remediation priorities are obscured (Figure 1 previews the pattern, alongside the alternative target this paper evaluates). This issue is addressed in NIST Special Publication 800-231 [6]
1:3
by introducing the Bugs Framework (BF), which represents a weakness as a structured relation among a cause, an operation, and a consequence. It organizes weakness classes around disjoint operation sets: the specification requires that no two BF classes share an operation. However, beyond theoretical declarations, applied research on the BF specification remains Research Questions (RQ) RQ-1: What are the structural limitations of CWE that automated CVE-to-CWE classification repeatedly encounters, and how can those limitations be addressed in a successor framework and tested against? RQ-2: Does the BF specification, by design, address each of the operationally defined CWE limitations, and can the correspondence be traced directly to the specification language and to published BF chains? RQ-3: To what degree is CVE-to-BF classification reproducible in practice, both across independent SMEs and repeated runs in an automated pipeline under a constant configuration? scarce compared against the two decades of CVE-to-CWE work. Therefore, a literature-grounded evaluation of the NIST Bugs Framework is both timely and necessary before the community invests in CVE-to-BF automation. We therefore ask the following three research questions (RQs). To answer these questions, we conducted the research as a sequence of methodological steps, reported in Section 2: a systematic literature review with a grounded characterization schema for RQ1, a specification-level correspondence analysis for RQ2, and a two-study empirical evaluation for RQ3. Those steps produced the contribution (C𝑥 ) of this work, a literature-grounded case study that evaluates NIST SP 800-231 as a classification target, in five parts: C1. Corpus and schema. We assemble and characterize a systematically screened corpus of automated CVE-to-CWE research using the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) procedure across six databases, and report the corpus according to a defensible five-axis characterization schema. C2. Failure analysis. We read the corpus against four operationally defined CWE failures, namely a non-orthogonal hierarchy, an intractable target space, sink-only labeling, and the absence of a causal chain, and show that the surveyed systems repeatedly accommodate those failures rather than overcome them. C3. Structural correspondence. We demonstrate how BF addresses each failure by construction and support the claim with verbatim specification language. C4. Empirical evaluation. We report two complementary studies: an anonymous inter-rater reliability study that assesses reproducibility per BF axis, and an automated determinism study that executes a reproducible, evidence-based CVE-to-BF pipeline and measures chain-level determinism across repeated runs of 2 deployments and budget combinations under baseline configurations. C5. Case studies. We present case studies, one per BF class type, that tie documented CWE assignments and reassignments to the structural failures they expose and to the BF properties that address them. In the remainder of the manuscript, Section 2 provides the survey protocol, the characterization schema, and the design of the evaluations; Section 3 defines the classification properties at issue and the four structural failures, worked on a single CVE; and Section 4 situates vulnerability classification within the Cyber Threat Intelligence (CTI) pipeline. The research questions are
1:4
Hoque et al.
answered in turn: Section 5 synthesizes the corpus (RQ1), Section 6 traces each documented CWE failure to the corresponding BF specification language (RQ2), Section 7 presents one case study per BF class type, and Section 8 reports the inter-rater and determinism studies (RQ3). Section 9 states the boundary conditions of every claim, and Section 10 concludes; the appendices carry the acronym table, the threats to validity, and research protocol. 2
Methodology
This section reports how the research was conducted, one methodological step per research question. To answer RQ1, we collected a corpus of CVE-to-CWE research under a registered systematic procedure, characterized each included paper with a schema derived from a citable taxonomydevelopment method, and let the recurring limitation patterns emerge from that characterization, not from any prior position. To answer RQ2, we traced each limitation to the verbatim language of NIST SP 800-231 and to published BF chains (Section 6). The correspondence between the limitations the corpus surfaced and the properties the BF specification formalizes is an empirical finding of this procedure, not its premise. To answer RQ3, we designed two studies: an anonymous inter-rater reliability study and an automated determinism study, both conducted over a curated, evidencebased pipeline, whose apparatus and metrics are specified along with their results (Section 8). The corpus procedure follows the PRISMA reporting structure [63], and the review protocol follows the guidelines of Kitchenham et al. [27]. The characterization schema is constructed using the systematic-mapping classification practice of Petersen et al. [66] and the taxonomy-development method of Nickerson et al. [54], so that every axis of the schema is grounded in a named, citable methodology rather than ad hoc judgment. 2.1
Corpus Collection: A PRISMA Procedure
PHASE 1: SOURCING
PHASE 2: PRISMA FILTERING
PHASE 3: SYNTHESIS
Identification
Screening
Eligibility
Included Corpus
Per-database keyword search across six databases: IEEE Xplore, ACM Digital Library, Springer Link, arXiv, Google Scholar, and Semantic Scholar.
Title, abstract, and (where publicly available) introduction assessed against the inclusion criteria.
Full-text assessment against scale, method, and evaluation criteria.
Studies retained: 𝑁 = 20
Year filter: January 2018 to April 2026 Venue filter: peerreviewed or citable preprint Pooled records: 660
Exclusion criteria: • Downstream-only target • Non-learned lookup table • Incidental CVE use • Grey literature
Exclusion criteria: • Fewer than 500 evaluation CVEs • Undisclosed class composition • Synthetic or untraceable corpus • Manual intervention at inference
Method families covered: • Classical machine learning • Transformer classifiers • Retrieval systems • Knowledge graphs • Correction pipelines • LLM-based methods
Sourcing across six databases, filtering through Screening and Eligibility, and synthesis on the retained corpus grouped by method family. Solid arrows denote flow; dashed backgrounds group the phases. Fig. 2. PRISMA-style corpus construction for the 20-paper survey.
1:5
The full search protocol and the anchor list are reproduced in Appendix C. The search used keyword queries across six databases (IEEE Xplore, the ACM Digital Library, SpringerLink, arXiv, Google Scholar, and Semantic Scholar). The master tracking sheet was created by merging six database exports and deduplicating records based on persistent identifiers and normalized titles. To be included, a work needed to use a CVE artifact, description text, NVD metadata, or scanner output as input and generate a CWE label as a primary output or an independently evaluated stage. Predictions had to be produced using learned, retrieval-based, or generative methods without manual intervention, and had to utilize at least 500 distinct CVE records from a named public registry with disclosed class composition. The publication must have been peer-reviewed or a citable preprint with a reimplementable method. Exclusions included downstream-only targets, non-learned lookup tables, incidental use of CVEs, untraceable or synthetic corpora, undisclosed evaluation details, grey literature, and non-substantive duplication. Two criteria allowed the inclusion of a corpus with fewer than 500 records if it occupied a unique design-space cell and maintained a downstream pipeline, provided the CWE stage was evaluated independently. After screening, 20 papers (Table 2a) met all inclusion criteria and form the final corpus. 2.2
Manual Data Extraction
The authors performed abstract screening and full-text data extraction using a pre-defined form, following Kitchenham’s guidelines [27]. They applied inclusion, exclusion, and safety-valve criteria to each title, abstract, and available introduction, documenting the decisions along with the criteria met or failed. Records labeled as borderline or unclear were set aside for full-text eligibility assessment and discussed among the authors before making a decision. For eligible papers, they recorded properties based on five schema axes and retained relevant excerpts for traceability. To minimize transcription errors, a second author verified a sample of the records against the original papers, resolving any discrepancies through joint review. 2.3
Characterization Schema
Each included paper is characterized along five axes that together form the Characterization Schema used throughout the manuscript: input artifact, target form, method family, evaluation setting, and stated limitation. The schema was derived with the method of Nickerson et al. [54]: the meta-characteristic is “properties of a CVE-to-CWE system that determine how its evidence bears on the suitability of the classification target,” and the five axes were obtained through the empirical-to-conceptual iteration that the method prescribes. The schema is designed to support synthesis, not only inventory. When many systems narrow the target space, the pattern is interpreted as evidence of target-space intractability. When systems report parent-child or sibling errors, the pattern is interpreted as evidence of non-orthogonality (the terms are defined operationally in Section 3). When repair systems identify invalid or generic labels, the pattern is interpreted as evidence that CWE assignments often fail to encode root cause. The schema is reported in the integrated Tables 2a and 2b, which give both the methodological dimensions and the headline performance with its test-set conditions, and the same schema positions the 20 papers in the visual taxonomy of Figure 7. Performance is reported beside those conditions because a number is interpretable only against its label-space size, class composition and evaluation metric, and the table is the evidentiary basis for the synthesis in Section 5 and for the failure analysis in Section 6.
1:6
3
Hoque et al.
Preliminaries
This section defines each of the terms the comparison relies on operationally. Two terms describe properties a classification scheme may or may not have; four name the structural failures of CWE; the last composes them. 3.1 Operational Definitions Orthogonality of a classification scheme. A taxonomy is orthogonal when its categories partition the space of classified objects without semantic overlap: for any object 𝑂 there exists exactly one category 𝐶 such that 𝑂 belongs to 𝐶. That assignment is invariant under reasonable refinements of analysis [54, 69]. BF realizes this property by associating each class with a disjoint operation set, so that for any two classes 𝑐𝑖 , 𝑐 𝑗 with 𝑖 ≠ 𝑗 one has ops(𝑐𝑖 ) ∩ ops(𝑐 𝑗 ) = ∅ [6]. CWE Structural Failures ID
Failure
F1
Non-orthogonal Labels at different abstraction levels co-apply to one defect, so no hierarchy single label is canonical. Example: near-identical descriptions receive the class-level CWE-74 (CVE-2019-5404) and its base-level descendant CWE-94 (CVE-2018-0461) [64].
F2
Intractable target space
F3
Sink-only label- The assigned label names the final manifestation (sink), not the uping stream root cause. Example: Heartbleed (CVE-2014-0160) is recorded as CWE-125, the out-of-bounds read at the sink, while the missing length verification goes unrecorded [6, 62].
F4
Absence of a causal chain
Definition and documented example
A large, long-tailed label set that supervised classifiers cannot learn across, forcing restriction to a subset. Example: about 70% of CWE classes have fewer than 100 mapped CVEs and only 10% have more than 500 [16].
Weaknesses are recorded as an unordered label set with no perinstance causal order. Example: the CVE-2018-5907 record first carried the set {CWE-20, CWE-119, CWE-190}, then was remapped to CWE-190 alone [6, 59].
Table 1. The four structural CWE failures: notation, definition, and one documented example of each.
Determinism of a mapping procedure. A mapping procedure is deterministic when repeated independent classification of the same object by qualified experts converges on the same category above an agreed agreement threshold [14, 30]. Determinism is a property of the procedure and its target jointly: an orthogonal target removes the principal source of inter-annotator divergence. Table 1 fixes the notation for the four structural CWE failures, F1 through F4, and pairs each with one documented example. The following paragraphs provide the operational definition and test for each failure. Non-orthogonal hierarchy (CWE failure F1). This failure denotes a hierarchical taxonomy in which labels at different abstraction levels can simultaneously apply to the same defect. The operational test is that the same CVE plausibly maps to two or more CWE entries that are ancestors and descendants, or siblings, in the CWE tree. Pan et al. [64] document the operational test directly on the two-CVE pair recorded in Table 1.
1:7
Intractable target space (CWE failure F2). This failure is structural in origin but automation-facing in effect: it denotes a label space whose cardinality is large and whose label frequency is skewed enough that supervised classifiers cannot reliably learn the long tail, thereby forcing systems to restrict themselves to a subset and trade coverage for accuracy. What makes it specific to CWE is the source of the skew: CWE encodes specificity by enumeration and assigns each weakness its own flat entry, so the catalog grows beyond 900 interdependent entries [40]. Each rare type receives its own sparsely populated label. In operational terms, the test is that surveyed systems restrict to a fraction of the available labels and that accuracy degrades sharply on out-of-restriction labels. V2W-BERT quantifies the skew as recorded in Table 1 [16], and the corpus-wide pattern in Table 2a shows the majority of the 20 systems declining the full CWE label space and operating instead on CWE-1003, a frequent-class subset, or a custom-curated view. Sink-only labeling (CWE failure F3). This failure denotes a labeling convention in which the assigned category names the defect’s final manifestation, the sink, rather than the upstream cause, the source. The operational test for F3 follows directly. Two CVEs that share a sink but differ in their upstream causes receive the same sink-level CWE label, yet they would require different mitigations under that label. Within the corpus, FixV2W finds this loss to be widespread: more than half of all CVEs carry invalid or insufficiently detailed CWE mappings [25, 74]. Absence of a per-instance causal chain (CWE failure F4). This failure denotes a taxonomy that records a defect’s weaknesses as an unordered identifier set, with no per-instance mechanism that fixes the causal order in which one weakness creates the conditions for the next. Consider Heartbleed (CVE-2014-0160), where a missing bounds check permits an out-of-bounds read, yet the NVD records the defect under a single buffer-error weakness [62]. The documented CVE-2018-5907 case in Table 1, whose narrated chain runs from improper input validation through integer overflow to a memory-buffer error, shows this two-step loss in the record itself: the record first carried the three weaknesses as an unordered set, with no causal order among them, and was then remapped to the single label CWE-190 [6, 59]. Structural inconsistency of the CVE-CWE classification system. We define the structural inconsistency of the CVE-CWE classification system, as enriched and published in the NVD, as the simultaneous presence of failures F1 through F4 in the target taxonomy and the mapping procedure. However, we restrict the claim to the operational classification: CWE supplies mechanisms for both root-cause mapping and named weakness chains (CWE-709). However, a per-CVE NVD record encodes neither, so the four failures persist at the point of use. 3.2
Demonstration of Multiple Defensible CWE Mappings
CVE-2021-3156, the Sudo heap-based buffer overflow analyzed in Section 7.1, admits several defensible CWE mappings at once: original record characterizes the defect as an off-by-one error (CWE-193) [45, 58], current primary assignment is the heap-based buffer overflow CWE-122 [42, 58], and class-level CWE-119 together with base-level CWE-787 and CWE-120 plausibly apply to the same defect at other hierarchy levels (see Figure 1 for end to end ambiguity preview). 3.3
Overview of the Bugs Framework (BF)
NIST SP 800-231 defines BF as “a classification of security bugs and related faults that features a formal language for the unambiguous specification of software and hardware security weaknesses and vulnerabilities” [6]. The specification states four properties: structured, orthogonal, multidimensional, and context-free. The comparison developed in this paper rests on the last two in particular. “Structured means that a weakness is expressed as a ⟨cause, operation⟩→consequence triple with a
1:8
Hoque et al.
® Pooled, deduplicated candidate records (𝑁 = 660) Does the work take a CVE artifact as input and produce a CWE label as output or as an evaluated intermediate stage?
No
q Excluded: downstream-only target, non-learned lookup table, or incidental CVE use
No
q Excluded: manual intervention at inference
No
q Excluded: sub-500 corpus, undisclosed composition, or synthetic or untraceable corpus
No
q Excluded: grey literature or non-substantive duplication
è Safety valve B: downstream pipeline retained when its CWE stage is evaluated in its own right
Yes ® Is the prediction produced by a learned, retrieval-based, or generative method without manual intervention at inference? Yes ® Does the evaluation employ at least 500 distinct CVE records from a named public registry with disclosed class composition?
è Safety valve A: sub-500 corpus retained when it occupies a designspace cell no other work occupies
Yes ® Is the venue peer-reviewed, or is the work a citable preprint with a reimplementable method? Yes ¥ Included corpus 𝑁 = 20 papers Each internal node is one inclusion test; a Yes edge descends toward inclusion and a No edge exits to an exclusion category. The two dashed safety valves are the documented overrides (criteria of Section 2). Fig. 3. Decision-tree rendering of the screening and eligibility procedure.
precise causal relation” [6]. “Orthogonal means that the intersection of the sets of operations of any two BF classes is the empty set” [6]. “Multidimensional means that weaknesses are organized
1:9
BUGS FRAMEWORK (BF) Type categories
BF Weakness Type
BF Failure Type
Class types
Classes
BF_INP I/O Check DVL Data Validation DVR Data Verification
BF_MEM Memory MAD Memory Addressing MMN Memory Management MUS Memory Use
BF_DAT Data Type DCL Declaration NRS Name Resolution TCV Type Conversion TCM Type Computation
BF_FLR Failure IEX Information Exposure ACE Arbitrary Code Execution DOS Denial of Service TPR Data Tampering
Fig. 4. BF taxonomy structure up to the class level. The figure separates the BF weakness branch from the BF failure branch, with each class type expanded into its constituent classes.
not only by their operations but also by their causes, consequences, and operation and operand attributes” [6]. “Context-free means an operation cannot have different meanings depending on the language or domain” [6]. A weakness is an instance of a BF class consisting of one cause, one operation, one consequence, and operation and operand attributes [6]. A vulnerability is “a chain of weaknesses linked by causality via a consequence↷cause propagation that eventually enables a security failure” [6]. The classes used in this manuscript follow the taxonomy of NIST SP 800-231 [6]: the Input/Output Check (INP) classes Data Validation (DVL) and Data Verification (DVR); the Memory (MEM) classes Memory Addressing (MAD), Memory Management (MMN), and Memory Use (MUS); the Data Type (DAT) classes Type Conversion (TCV) and Type Computation (TCM); and the Failure (FLR) classes Information Exposure (IEX) and Arbitrary Code Execution (ACE), a subset of the full SP 800-231 taxonomy selected for the vulnerability classes this study analyzes. Figure 4 renders the full class taxonomy specified in NIST SP 800-231 in detail. 3.4
The BF Research Program as a Design Response
NIST SP 800-231 [6] also defines BF as a classification of security bugs and related faults with multi-dimensional weakness and failure taxonomies and a formal language. The antecedent BF paper [7] establishes the ⟨cause, operation⟩→consequence triple as the unit of weakness specification. Galhardo et al. [19] measurements sit alongside the BF program as a metric-design contribution. 4
CTI Frameworks and Defensive Operations Pipeline
The frameworks introduced in Section 3 do not operate as isolated artifacts; rather, they compose a directed CTI [34] pipeline in which the output of one stage becomes the input of the next.
1:10
Hoque et al.
Within this pipeline, a disclosed vulnerability is entered as a CVE record and assigned a CWE weakness label and a CVSS severity score through NVD enrichment. It then gains an optional Exploit Prediction Scoring System (EPSS) probability of in-the-wild exploitation [22], maps through CAPEC and ATT&CK to adversary behavior, and finally links to a defensive countermeasure in MITRE D3FEND [24, 50]. As depicted in Figure 5, this end-to-end flow runs from CVE disclosure on the left to a D3FEND countermeasure on the right. However, each stage is downstream of, and bounded by, the fidelity of the one preceding it; NIST itself documents that CWE and CVE deficiencies propagate into the NVD and can yield imprecise or wrong weakness assignments [6], and the structural failures of the CVE-to-CWE step therefore constrain every artifact that the later stages can recover. This bottleneck is the central concern of the present analysis. CWE
CVE disclosure
weakness label
CVSS / EPSS severity, exploit
CAPEC attack pattern
ATT&CK
D3FEND
tactic, technique
countermeasure
inverse mapping: defense to posture Fig. 5. End-to-end CTI pipeline, from a CVE through to a D3FEND countermeasure.
4.1
Operational Workflow of the CTI Pipeline
A typical enterprise vulnerability-management workflow applies these stages in sequence. After a CVE is disclosed, NVD analysts, or increasingly an automated classifier of the kind surveyed in Section 5, assign a CWE label and a CVSS vector. The Stakeholder-Specific Vulnerability Categorization (SSVC) methodology [75] then supplements this severity signal with decision trees over stakeholder decision points to yield a remediation priority [35, 65]. At the same time, EPSS contributes a complementary probabilistic estimate of exploitation. Bridging weakness to adversary, Common Attack Pattern Enumeration and Classification (CAPEC) [38] maps each weakness to the attacker methodologies that exploit it, and ATT&CK [51, 76] supplies the operational tactics, techniques, and procedures [78] against which detections are authored. However, the severity signal supplied at this stage is a coarse first-order filter rather than a causal account of the defect: CVSS Base scores correlate poorly with observed exploitation [5], which is the gap that EPSS is designed to narrow. The ATT&CK situates a vulnerability within a broader family of adversary-modeling. Among these, the Cyber Kill Chain [21] and the Diamond Model [10] structure incident reasoning. Neither yields a machine-readable weakness label, and ATT&CK itself sits at a coarser granularity than CWE. The design documentation states that ATT&CK does not enumerate attack vectors against software and instead defers to CAPEC and CWE, so a single technique, such as Exploit Public-Facing Application (EPFA), subsumes vulnerabilities arising from multiple distinct underlying weaknesses. 4.2
Bidirectional Mapping between ATT&CK and D3FEND
D3FEND is MITRE’s defensive analog to ATT&CK. Where ATT&CK enumerates the techniques an adversary employs, D3FEND enumerates the digital artifacts and countermeasures that defenders deploy to disrupt those techniques [24, 50]. The two knowledge graphs are explicitly cross-linked. By construction, this bidirectional mapping can produce a defensive recommendation only as specific as the upstream CWE and CAPEC labels permit. For instance, a vulnerability tagged only with the broad CWE-787 (Out-of-bounds Write) yields correspondingly coarse D3FEND countermeasures,
1:11
even where the root cause would call for precise control over a single length field. The same bound applies on the failure side, where mitigations aimed at a denial-of-service failure, such as adaptive proof-of-work admission control [11], are selected against the failure class rather than against the weakness chain that produced it. Consequently, the pipeline inherits the ambiguity introduced at the classification stage, a propagation that NIST documents for the CWE-to-NVD step itself [6], and each subsequent mapping can only preserve, never restore, the causal precision lost there. This is the same sink-only and non-orthogonal behavior (see Section 3) that the Heartbleed study (see Section 7) specifies. Replacing the CWE label with a BF causal chain removes this bound at its source. 5
CVE-to-CWE Automation Corpus
During 2020 to 2025, the automated mapping of CVE records to CWE labels has advanced from shallow lexical classifiers toward fine-tuned transformer encoders, retrieval over learned embeddings, knowledge-graph inference, post-prediction label repair, and large language model (LLM) in-context inference. To organize this body of work on common axes, we apply the Characterization Schema of Section 2.3 to the 20-paper corpus summarized in the integrated Tables 2a and 2b. Amid this heterogeneity, a single pattern recurs: the reported performance ceiling for CVE-to-CWE classification is set less by model capacity than by the structure of the target taxonomy, a thesis that this section develops by reading the corpus against the four structural failures defined in Section 3.1. 5.1
Methods and Reported Performance
Reported accuracy
unreachable: capped by CWE label-space ambiguity performance ceiling
restrict or repair the target labels
long tail, sparse classes, parent-child and sibling overlap
CWE label space covered (long tail admitted →) Fig. 6. Conceptual performance ceiling created by CWE label-space ambiguity.
The strongest headline numbers in the corpus are obtained under conditions that narrow the task. The LLM-based and LLM-hybrid entries in the corpus, Text2Weak [72], Key Term Extraction [71], LLM-VulnClass [77], and LLM-Triage [23], show that modern LLMs help most at the term-extraction and label-ranking stages and help least at single-label discrimination across the full hierarchy. Across Table 2b, the reported figures share a common structure: accuracy is high where the label space is small or skewed toward common classes, and degrades wherever the long tail is admitted [25, 31]. That is the empirical signature that motivates the failure analysis below. Figure 6 renders this
1:12
Hoque et al.
signature as a conceptual performance ceiling: reported accuracy rises as the label space narrows and falls as the long tail is admitted, regardless of model family. Study Text2Weak [72], 2024 Pure Self-Attention [86], 2021 Key Term Extraction [71], 2026 VulnBERTa [80], 2024 V2W-BERT [16], 2021 Temporal LR [26], 2024 FixV2W [74], 2025 CVEDrill [1], 2023 ThreatZoom [2], 2020 VulnScopper [4], 2024 VulnBERTa-XAI [81], 2026 Semantic Sim. [29], 2024 Multi-Taxonomy [79], 2026 RoBERTa-125M [52], 2026 LLM-VulnClass [77], 2025 VWC-BERT [15], 2022 NVD-Correct [73], 2024 CVE2CWE [3], 2024 CWE-CVE-CPE KG [70], 2024 LLM-Triage [23], 2025 Mod. (model family): Trn. (training):
Mod. TR TR TR TR TR ML KG SB DL HY TR TR TR TR HY TR OTH ML KG TR
Trn. OTH FT HYB FT FT SUP OTH FT SUP FT FT FT FT FT OTH FT OTH OTH OTH FT
Dataset NVD NVD NVD NVD NVD + MITRE CVE (NVD-derived) NVD NVD Mixed Mixed NVD NVD Mixed Mixed NVD NVD NVD + MITRE NVD + MITRE NVD NVD
Labels 57 11 57 160 124 n.r. 130 100 364 393 160 130 680 205 n.r. 124 130 25 924 n.r.
TR Transformer/LLM
ML Classical ML
KG Knowledge Graph
DL Deep Learning
HY Hybrid
SB SecureBERT
FT Fine-tuning
SUP Supervised
HYB Hybrid
OTH Other Column codes are expanded in the legend below the table; “ n.r.” marks a value not reported in the source. The evaluation details of the same twenty studies are in Table 2b. Table 2a. Part (a) of the corpus table: methodological approach of the 20-paper CVE-to-CWE automation corpus, with model family, training regime, dataset, and label-space size.
Study
Headline result
Test-set characteristics BF-versus-CWE relevance
Text2Weak [72], 2024
Macro-𝐹 1 : 22.33%
Pure Self-Attention [86], 2021
Accuracy: 90.35%
Top@5 acc. 66.38%; 57-class CWE-1003. Top-10 CWE only; weighted-𝐹 1 89.31%.
Parent-child errors expose hierarchy ambiguity. High accuracy depends on drastic label-space narrowing. continued on the next page
1:13 Table 2b (continued) Study
Headline result
Test-set characteristics BF-versus-CWE relevance
Weighted-𝐹 1 : 70.36% Macro-𝐹 1 59.66%; 57-class; Root-cause term extraction +8.88% over full aligns with BF axis description. extraction. VulnBERTa [80], 2024 Accuracy: 88.5% Tiered over 160 classes; Tiering exposes CWE MCC 0.721 at Tier 1. class-imbalance pressure. V2W-BERT [16], 2021 Accuracy: up to 97% Relaxed prediction; Non-disjoint hierarchy and 48-76% on rare, 61% rare classes are central zero-shot. barriers. Temporal LR [26], 2024 Accuracy: 66% 𝐹 1 0.64; skewed CWE Temporal validation does distribution. not remove skew or ambiguity. FixV2W [74], 2025 Exact-match: 52% 93% same-branch within Mapping repair reveals top-10 (CWE-1003); 69% invalid, generic, and top-10 on exploited CVEs. root-cause-poor labels. CVEDrill [1], 2023 𝐹 1 : 84.91-96.08% Hierarchical hit-rate Strong engineering remains 90.77-97.51%. bound to sink-oriented CWE. ThreatZoom [2], 2020 Accuracy: 92% 75% on MITRE fine-grain; Full-hierarchy ambition (NVD) full hierarchy. exposes multi-path weakness difficulty. VulnScopper [4], 2024 Hits@10: 71% (NVD) CVE-CWE linking; +11.7% Context helps; the target lacks causal-chain syntax. over LLMs. VulnBERTa-XAI [81], Accuracy: 88-97.7% Per-tier; prediction-tree Explainability exposes 2026 90.1%. over-generalization and “Other” pressure. Semantic Sim. [29], Accuracy Similarity-based retrieval Better similarity does not 2024 over CWE-1003 view. dissolve overlap in the target. Multi-Taxonomy [79], Accuracy Cross-taxonomy Error propagation worsens 2026 classification; error as taxonomic breadth propagation observed. expands. RoBERTa-125M [52], Top-1 acc.: 87.4% Macro-𝐹 1 60.7%; 205-class. Top-1 accuracy coexists 2026 with weaker macro coverage. LLM-VulnClass [77], Accuracy Full-hierarchy LLM; Generic LLM ability does 2025 ambiguity persists. not eliminate CWE ambiguity. VWC-BERT [15], 2022 Accuracy Cascade: Upstream CWE uncertainty CVE→CWE→CAPEC; contaminates downstream upstream uncertainty mappings. propagates. NVD-Correct [73], Top-K acc. Repair-ranking; targets Label correction becomes 2024 contestable initial NVD necessary because labels. mappings are unstable. CVE2CWE [3], 2024 Top-1 acc.: 69.9% Top-3 acc. 87.5%; 25-class. Accuracy degrades as class count rises. continued on the next page Key Term Extraction [71], 2026
1:14
Hoque et al.
Table 2b (continued) Study
Headline result
Test-set characteristics BF-versus-CWE relevance
CWE-CVE-CPE KG [70], 2024
Accuracy
924-class repository graph; sparsity limits inference. Full-hierarchy LLM; sibling/parent-child confusions persist.
LLM-Triage [23], 2025 Accuracy
Structured relations help but sparsity limits inference. Sibling and parent-child confusions persist in newer LLM workflows.
Headline results are reported beside their test-set conditions because a number is interpretable only against its label-space size, class composition, and metric; “ n.r.” marks a value not reported in the source. The methodological columns of the same studies are in Table 2a. Table 2b. Part (b) of the corpus table: headline performance and BF-versus-CWE relevance of the same twenty studies.
5.2
Four Adaptive Strategies for Thematic Synthesis
Reading the corpus against the operational criteria of Section 3.1 yields four strategies by which systems accommodate a target taxonomy that resists clean single-label prediction. Each strategy demonstrates one or more of the four CWE failures from a different angle, and the resulting acknowledgment pattern is recorded in Table 2b’s “BF-versus-CWE relevance” column. Strategy 1: Restrict the label space. A first group of systems addresses the intractability of CWEs by narrowing the target. Pure Self-Attention [86] reduces to the top ten CWE identifiers. CVE2CWE [3] reduces to a 25-class view. VulnBERTa [80] and its explainability extension VulnBERTa-XAI [81] operate over 160 frequent classes with tiering, RoBERTa-125M [52] over 205 classes, and Key-Term Extraction [71] and Text2Weak [72] over a 57-class view drawn from CWE-1003 [49], the simplifiedmapping view that the NVD itself applies when labeling CVEs [60]. The Temporal LR baseline [26] reinforces the intractability reading: even with temporal validation, the skewed label distribution sets a hard performance ceiling. Strategy 2: Change the prediction form. A second group changes the shape of the prediction, not the size of the target. V2W-BERT [16] predicts hierarchically; CVEDrill, built on SecureBERT [1], predicts top-𝐾 over the CWE hierarchy with explicit hierarchical hit-rates; ThreatZoom [2] predicts across the full CWE hierarchy; Semantic Similarity [29] reframes the problem as embedding retrieval rather than direct classification; the Multi-Taxonomy transformer [79] extends the prediction across taxonomies; and the CWE-CVE-CPE Knowledge Graph [70] replaces flat classification with relation prediction over a structured graph. The errors of this group cluster around parent-child and sibling CWE relationships; consistent with this, VulnScopper [4] reports that 44.3% of CVEs cataloged in two databases carry a different CWE, a divergence it attributes to the hierarchy permitting either a general or a specific weakness for the same defect. Strategy 3: Repair labels after prediction. A third group treats CWE labels as objects requiring repair. FixV2W [74] applies knowledge-graph embeddings to correct invalid (Prohibited) or insufficiently-specific (Discouraged) NVD mappings; for Prohibited mappings it recovers the exact CWE within the top-10 in 65% of cases, with the correct label at the first rank in 52%, and a samebranch candidate in 93%; and NVD-Correct [73] applies a ranking-and-repair pipeline that targets the cases where the initial NVD assignment is contestable. This group is especially important for the BF argument because it demonstrates that the ecosystem already spends effort repairing the
1:15
Strategy 1 Restrict the label space
Strategy 2 Change the prediction form
Pure SelfAttention
V2W-BERT
Strategy 3 Repair labels
Strategy 4 Push to LLMs
FixV2W
Key Term Extraction
NVD-Correct
CVE2CWE
CVEDrill (SecureBERT)
VulnBERTa
ThreatZoom
VulnBERTa-XAI RoBERTa-125M
Semantic Similarity
Text2Weak
Multi-Taxonomy
Temporal LR
VulnScopper
VWC-BERT (cascade)
LLM-VulnClass LLM-Triage
CWE-CVECPE KG
NIST Bugs Framework Orthogonal classes; small structured class set; explicit cause-operation axes; cause-to-consequence chain Each cluster groups papers by the strategy with which they cope with CWE’s structural failures (F1 to F4); BF responds to the same failures by design rather than by adaptation. Placement reflects the Characterization Schema axes; a paper that also contributes to a secondary strategy appears at its primary cluster only. Strategies of Section 5.2. Fig. 7. Schema overview of the 20-paper CVE-to-CWE automation corpus, clustered by the four adaptive strategies.
target representation after the fact: more than half (55%) of NVD CVEs carry invalid or insufficiently detailed CWE mappings [74], a direct empirical signature of structural inconsistency in the target taxonomy. Strategy 4: Push to LLMs at the boundary. A fourth group reaches for LLMs at the boundary where the prior strategies stop helping. Key Term Extraction [71] uses an LLM to extract rootcause terms before mapping into CWE, which aligns naturally with BF’s axis-by-axis extraction model. LLM-VulnClass [77] and LLM-Triage [23] apply LLMs more directly to the classification task. The counterintuitive headline of LLM-VulnClass [77], namely that a simple TF-IDF baseline outperforms LLM embeddings on the same task, provides an empirical anchor for how classical baselines compare to LLMs at this label-space scale. VWC-BERT [15] reinforces this conclusion from a cascade perspective: upstream CWE uncertainty propagates into the downstream CWEto-CAPEC mapping, so an inherited target-space problem at the CVE-CWE stage contaminates everything that depends on it.
1:16
5.3
Hoque et al.
Historical Progression
Chronologically, the corpus traces a clear progression in which more capable systems do not dissolve the failures observed in the earlier ones. They redirect effort to mitigations (narrower targets, hierarchical ranking, label repair, LLM term extraction) that an orthogonal target with explicit cause and chain syntax would not require. Figure 7 renders this progression as a thematic clustering across the five Characterization Schema axes. 6 6.1
Comparative Analysis of the Bugs Framework Against CWE’s Structural Failures Property-by-Failure Mapping
This subsection shows, for each of the four operational structural failures of CWE (see Section 3.1), the specific design property of the BF that addresses it. The mapping is grounded in the BF specification, where each property is named as the specification names it, and each is supported by verbatim text from NIST Special Publication 800-231 or from the Bojanova et al. class papers that document it. We demonstrate a one-to-one correspondence between these properties and failures F1 through F4, and the bottom panel of Figure 8 records the resulting matrix. Figure 10 schematically renders the four resolutions, corresponding to the panel-for-panel failure schematics of Figure 8. 6.1.1 Orthogonal weakness classes address the non-orthogonal hierarchy (F1). The non-orthogonal hierarchy is defined in Section 3.1. BF addresses this failure through the orthogonality of its weakness classes, a property that the specification directly asserts for these classes [6]. The partition is exact at the level of operations: “Orthogonal means that the intersection of the sets of operations of any two BF classes is the empty set” [6]. Because the operation sets are disjoint, a defect represented by one class cannot also belong to another class at a competing level of the same hierarchy. The orthogonal partition removes the freedom, central to the non-orthogonal hierarchy, to select among ancestor, descendant, or sibling labels for one defect. We note a wording difference for accuracy: the specification applies orthogonal to the classes, realized through disjoint operation sets, and multidimensional to the attribute axes that supply the within-class structure [6]. Orthogonality removes overlap by design; it does not by itself guarantee that an automated procedure will select the single correct class for a given CVE, a practical reduction that a CVE-to-BF system would be intended to test. 6.1.2 Structured composition over a small class set addresses the intractable target space (F2). The intractable target space (F2) is stated operationally in Section 3.1. BF addresses this failure through structured composition over a small and complete set of weakness classes [6, 9].2 BF does not enumerate each weakness as a separate entry; it expresses specificity by composition: a weakness is one value per axis drawn from a small per-phase class, and a complex type is described by combining axis values without adding a new entry. The composing base is deliberately small; a single execution phase is covered by a handful of classes, as when “four language-independent, orthogonal classes that cover all possible kinds of memory-related software bugs and weaknesses” span the memory phase [8]. The contrast is stated in the same literature, which observes that the exhaustive CWE list “is prone to having gaps and overlaps in coverage” [8], whereas the CWE catalog has grown beyond 900 entries [40], the long tail of which Table 2a shows the surveyed systems decline. Because specificity is carried by axis combinations rather than by entries, the effective target along each axis is bounded by design, thereby addressing the cardinality and skew that drive 2 BF project site: https://samate.nist.gov/BF/.
1:17
F1
F2 CWE-119 (class) >900 entries, long tail
parent
CVE
CWE-122 (base) both labels defensible
CWE labels by mapped-CVE count
labels co-apply across levels
long tail defeats supervision
F3
F4 root cause
···
not recorded
sink = the CWE label
{ CWE-20
CWE-119 }
an unordered set; which weakness enables which is not recorded
label names the sink, not the cause CWE structural failure
CWE-190
causal order is lost
Bugs Framework (BF) Design Properties Orthogonal classes
Small, structured class set
Cause-toCause and operation axes consequence chain
F1: Non-orthogonal hierarchy F2: Intractable target space F3: Sink-only labeling F4: Absence of a causal chain Legend:
Primary correspondence
Reinforcing correspondence
Not applicable
Fig. 8. Top: minimal schematic of each structural CWE failure, keyed to Table 1; Figure 10 maps onto these panels with the BF property that addresses each failure. Bottom: correspondence between the four structural CWE failures (Section 3.1, rows) and the BF design properties that address them (columns); the legend below the table defines the three marks.
F2 at the per-axis decision level. The reduction is real but partial, and this comparison claims no more than that. At the combination level, the joint space of valid cause, operation, consequence, and attribute values remains large and long-tailed, so a system that predicts a full chain still confronts a skewed target; faceting relocates the intractability to smaller per-axis decisions; it does not remove it outright [20, 25, 31, 88]. Consistent with this reading, the inter-rater study (bottom panel of Figure 10) records almost perfect agreement on the small cause and operation axes (𝜅 = 0.876 and 𝜅 = 0.900) and only fair agreement on the attribute axis (𝜅 = 0.235), which carries the largest and
1:18
Hoque et al.
CVE identifier 𝑣
Anchor: failure 𝐹 and final error 𝐶𝑛 1
Advisory text 𝐸 dsc Q1: which operation? the class menu 𝑇𝐵 [𝑊𝑛 ] is scored
Fix commit, buggy code 𝐸 cod
climb one link
Advisory pages 𝐸 adv Contamination strip (clean policy) Evidence bundle 𝐸
𝐸
exact lookups
Q3: bug or fault? (stopping test, 𝐶𝑠𝑛 ∈ 𝑇𝐶 ) fault
Bug: root found, chain complete
Closed tables 𝑇𝐴 ..𝑇𝐻 OWL orthogonality Constrained oracle Ω
Neuro-Symbolic Verification
bug
Q4: the fault becomes the consequence one link back
Symbolic Store and Oracle
shared vocabulary, shapes, restrictions
Q2: which operand facet is wrong? (name, data, type)
candidate chain
Evidence Acquisition
menus, evidence
Backward Derivation Engine (per-link loop)
Vocabulary check 𝑃voc SHACL structure 𝑃str Orthogonality 𝑃ort Conformant chain C
violation report, regeneration Algorithm 1, notation of Table 5. The engine derives one weakness per link; exact lookups hit the closed tables while inference routes to the oracle Ω, and the verifier gates every chain before release. Fig. 9. Architectural process view of the backward derivation pipeline.
most ambiguous per-axis value set. F2 is therefore mitigated but not eliminated: BF lowers the per-axis target to a finite, small set, while the residual skew at the combination level remains a learnability problem that a CVE-to-BF system is intended to measure. 6.1.3 Explicit cause and operation axes address sink-only labeling (F3). Sink-only labeling (F3) carries the operational test given in Section 3.1. BF expresses cause and operation as first-class axes that are distinct from the consequence [6]. The specification records cause and consequence as separate elements of every weakness: “Bugs and faults are causes of security weaknesses, and errors and final errors are their consequences” [6]. The specification names consequence of sink-only labeling explicitly: “Some [CVEs] list the final error at the sink as the root cause instead of the bug
1:19
or hardware-induced fault that starts the chain. Focusing on the final error helps identify mitigation techniques, but the actual root cause must be known and fixed to resolve the vulnerability” [6]. Sink-only labeling assigns the category of the last weakness alone; BF additionally records the first, so that two CVEs which share a sink but differ in their upstream cause receive different descriptions, as the cause and operation axes each carry an independent value. The information that sink-only labeling discards is thus retained by construction, on axes that are distinct from consequence. 6.1.4 The cause-operation-consequence chain addresses the absence of a causal chain (F4). The absence of a causal chain (F4) is fixed in Section 3.1. CWE supplies curated exceptions through its Named Chains and Composites (CWE-709 [47], CWE-678 [46]), which enumerate specific pre-named chains in place of a general by-construction rule under which every vulnerability is a cause-to-consequence chain [37]. BF records a vulnerability as a cause-to-consequence chain of weaknesses [6]. BF gives the chain a syntactic form: each weakness is one cause-operationconsequence instance, and the consequence of one weakness propagates as the cause of the next, so that “The BF formalism supports a deeper understanding of vulnerabilities as chains of weaknesses that adhere to strict causation, propagation, and composition rules” [6]. The cause-to-consequence link is the mechanism that carries the chain forward: a vulnerability “starts with a bug or hardwareinduced fault, propagates through errors that become faults, and ends with a final error that introduces an exploit vector toward a failure” [6]. The specification represents Heartbleed in exactly this form, as a sequence of BF weakness states and a generated weakness chain [6], which shows that the multi-stage vulnerability named in Section 3.1 and depicted in Figure 14 is losslessly representable as a chain, not reduced to a single CWE. The same representation is provided for the BadAlloc pattern that includes CVE-2021-21834 [6], with the explicit five-stage chain DVR ↷ TCM ↷ MMN ↷ MAD ↷ MUS. A chain representation eliminates the absence of a causal chain by design; reconstructing the full chain for a given CVE from its prose description remains a practical inference task, which a CVE-to-BF system would be intended to test. Across the four failures, each BF property addresses one failure directly, and the correspondence is one-to-one by construction. The bottom panel of Figure 8 records this matrix, with the primary correspondence on the diagonal. The inter-rater study (bottom panel of Figure 10) offers a first, partial confirmation of the orthogonality-to-determinism link. The mapping shows that BF addresses each failure by design; whether each is reduced in measured practice is the question that a CVE-toBF pipeline would answer, and the inter-rater result indicates that the reduction holds first on the cause and operation axes.
1:20
Hoque et al.
F1
F2 DVL DVR
MAD MMN MUS
TCV TCM . . .
class: MUS (exactly one)
operation: Read
(cause, operation) → consequence
disjoint operation sets: one class per operation
small closed class set; values composed from per-axis menus
orthogonal weakness classes
structured composition
F3
F4
cause (the bug)
operation
consequence (the sink)
𝑊1
𝑊2
𝑊3
IEX
both ends recorded: first-class cause and operation
consequence ↷ cause propagation: the order is part of the record
cause is never collapsed into sink
per-instance causal chain
BF Taxonomic Axis
Inter-rater Reliability Metrics Agreement Cohen’s 𝜅 92.3% 92.3% 69.2% 46.2% 75.0%
Cause Operation Consequence Attribute Pooled (all 52 cells) Bands, Landis and Koch [30]: moderate 0.41 to 0.60
slight 0.00 to 0.20 substantial 0.61 to 0.80
0.876 0.900 0.626 0.235 0.732
fair 0.21 to 0.40 almost perfect 0.81 to 1.00
Fig. 10. Top: the BF design property addressing each failure, replicating Figure 8 panel for panel; the grounding specification language per cell is in its bottom matrix. Bottom: inter-rater agreement on the blind BF mappings, over 13 CVEs and 52 axis cells.
7
Case Studies
The four case studies below are stratified, one per BF class type, and selected from the 16-CVE candidate pool in Table 9 by failure density, defined as the number of structural CWE failures a candidate exposes. All four share a pattern: the CWE label moved or was contested, whereas the BF classification did not change under that reassignment, because the disjoint operation sets of the BF classes admit one class and one value per axis (Section 6.1). 7.1
CVE-2021-3156 (Sudo Baron Samedit), BF_INP → BF_MEM
CVE summary. A heap-based buffer overflow in Sudo arises when an argument ending in a single backslash is processed in shell mode, so the parser miscomputes the length and writes past the allocated buffer.
1:21
CVE CVE-2021-3156 (Sudo Baron Samedit)
CVE-2015-0235 (GHOST)
CWE issue BF chain anchor Overlapping bound- DVL → MAD → MUS → ACE ary and overflow labels; label moved across hierarchy levels MMN → MUS Boundary label hides wrong-size allocation upstream
Illustrated resolution Preserves the inputvalidation root and the memory chain leading to code execution.
DVR → TCM → MMN → MAD → MUS
Makes the verify-tocompute-to-allocate-toposition-to-write sequence explicit (BadAlloc pattern).
Over-read sink label DVR → MAD → hides missing verifi- MUS → IEX (with converging MUS cation Clear)
Restores the missingverification cause before the information exposure.
CVE-2021-21834 Integer-overflow (GPAC) and allocation labels name different stages of one chain CVE-2014-0160 (Heartbleed)
Records the allocationstage cause before the outof-bounds memory use.
Table 3. Case study used to contrast CWE labeling with BF chain specification; full chain prose is in Section 7.
Original CWE assignment. Original record characterizes the defect as an off-by-one error (CWE193) [58].3 Current CWE assignment. Primarily assigned to CWE-122 [58], with CWE-193 retained as a secondary label.4 Analysis of the reassignment. The label moved from the root-cause off-by-one to the heap sink as analysis deepened. Because CWE itself chains a root-cause weakness to its sink [47], the reassignment is best read as a choice of which end to surface. Hence, both ends remain accurate descriptions of distinct stages of one chain [58]. CWE failures illustrated. It illustrates non-orthogonal hierarchy (F1) as five CWE entries plausibly apply: CWE-119 at class level, CWE-787 and CWE-120 at base level, CWE-122 as a variant of CWE-787, and CWE-193 as root-cause off-by-one. It illustrates sink-only labeling (F3) because the primary assignment, CWE-122, describes the final heap corruption, while upstream parser bug is not expressible by CWE-122 alone. It exemplifies the absence of a causal chain (F4) because the single per-CVE CWE field collapses the chain to one label and offers no device to express that the off-by-one (CWE-193) causes out-of-bounds write (CWE-787) that manifests as heap overflow (CWE-122). BF classification. Two annotators classified this CVE as cause Erroneous Code, operation Write, consequence Buffer Overflow, and attribute Heap [6]. The upstream stages of the chain map to DVL and MAD before MUS Write sink, and the failure class is ACE, since the heap write yields local privilege escalation to root [6, 67] (see Figure 11). 3 CWE-193:https://cwe.mitre.org/data/definitions/193.html. 4 CWE-122:https://cwe.mitre.org/data/definitions/122.html.
1:22
Hoque et al.
CAUSE
OPERATION
CONSEQUENCE
Validate Check escape characters
Invalid Data Error - propagates
BF class: DVL
Erroneous Code Bug (code defect)
Invalid Data → Wrong Size BF class: MAD
Reposition Move pointer by size
Wrong Size Fault (from W1 error)
Overbound Pointer Error - propagates
Overbound Pointer → Overbound Pointer BF class: MUS
Write Write via pointer
Overbound Pointer Fault (from W2 error)
Buffer Overflow Final Error
Failure: ACE Arbitrary Code Execution LEGEND Cause (Bug or Fault) Consequence (Error, propagates) Propagation between weaknesses
Operation (BF class identifier) Final Error (exploit entry point)
The annotators’ four-axis label (cause Erroneous Code, operation Write, consequence Buffer Overflow, attribute Heap) reads the chain at its sink weakness; both assigned it identically (Section 9). Fig. 11. BF specification of the Sudo Baron Samedit vulnerability (CVE-2021-3156).
BF properties that address each illustrated failure. For F1, the orthogonal weakness classes of Section 6.1 address it. For F3, the cause axis records the erroneous code that the sink-only CWE-122 omits. For F4, the defect is recorded as a chain whose consequence feeds the next cause, which the single CWE cannot encode. 7.2
CVE-2015-0235 (GHOST), BF_MEM
CVE summary. GHOST is a heap-based buffer overflow in the GNU C library’s gethostbyname and gethostbyname2 functions [68]. A specially crafted hostname triggers an undersized allocation in __nss_hostname_digits_dots: the size computation totals three buffer entities, whereas the subsequent pointer layout prepares four, so the width of one alias pointer is omitted from the allocated size, and the strcpy of the hostname then writes past the buffer [68]. Original CWE assignment. The NVD record originally labeled CVE-2015-0235 with the Class-level entry CWE-119 [41], Improper Restriction of Operations within the Bounds of a Memory Buffer.5 Current CWE assignment. The label was later reassigned to CWE-787, Out-of-bounds Write [56],6 with CWE-122 as the variant-level descendant; CWE-787 is a Base and CWE-119 a Class [41, 42, 48]. 5 CWE-119: https://cwe.mitre.org/data/definitions/119.html. 6 CWE-787: https://cwe.mitre.org/data/definitions/787.html.
1:23
CAUSE BF class: MMN
Wrong Size Fault (computed size)
OPERATION
CONSEQUENCE
Allocate Allocate with wrong size
Insufficient Size Error - propagates
Insufficient Size → Insufficient Size BF class: MUS
Write Write past end
Insufficient Size Fault (from W1 error)
LEGEND
Buffer Overflow Final Error
Failure: ACE arbitrary code execution via the heap overflow
Cause (Bug or Fault) Consequence (Error, propagates) Propagation between weaknesses
Operation (BF class identifier) Final Error (exploit entry point)
From a wrong-size allocation to a write past the allocated buffer; the failure class is ACE. Reference stratum: axis values follow the BF specification. Fig. 12. BF specification of the GHOST glibc vulnerability (CVE-2015-0235) as a chain of weaknesses.
Analysis of the reassignment. The available labels record the sink, not the upstream cause. The root cause of the defect is that the allocator allocates an insufficient buffer size; the out-of-bounds write is the manifestation, not the bug. CWE failures illustrated. This CVE shows non-orthogonal hierarchy (F1) through the co-applicability of CWE-119, CWE-787, and CWE-122 to the same defect; it exemplifies sink-only labeling (F3) because the allocation-stage cause, the undersized buffer, is not expressed by any of the CWEs actually assigned to this CVE. It illustrates the absence of a causal chain (F4) because the allocation cause and the use sink are two distinct stages of a chain that the CWE assignment collapses into one label. BF classification. The BF specification represents this case through the memory bugs model [8]. The first BF state is in the MMN class as a ⟨Wrong Size, Allocate⟩ → Insufficient Size weakness. The second BF state is in the MUS class as a ⟨Insufficient Size, Write⟩ → Buffer Overflow weakness, with the attribute on the Heap address state. Figure 12 depicts the chain, with the allocation defect upstream of the write. BF properties that address each illustrated failure. For F1, orthogonal weakness classes address the failure as before. For F3, the MMN Allocate operation and its Wrong Size cause record the allocation-stage bug. For F4, the cause-operation-consequence chain represents the allocate-thenwrite sequence as two linked weaknesses rather than a single sink label. 7.3 CVE-2021-21834 (GPAC), BF_DAT → BF_MEM (BadAlloc pattern) CVE summary. CVE-2021-21834 is an integer overflow that leads to a heap-based buffer overflow in the GPAC multimedia framework. A crafted input causes a size computation to wrap around and produce a value much smaller than the intended buffer size; the write overruns the allocation [13]. The vulnerability follows the BadAlloc pattern [6].
1:24
Hoque et al.
Original CWE assignment. The defect has been associated with CWE-190, and with CWE-787 [44, 48],7 and with the broader CWE-119. Current CWE assignment. CWE-190 names the integer overflow stage; CWE-787 names the out-of-bounds write stage; CWE-119 sits at the class level above both. Analysis of the reassignment. NIST SP 800-231 [6] makes the same point that each label reads one stage of one chain, for the structurally analogous CVE-2018-5907, where the entire chain is CWE-20 → CWE-190 → CWE-119. CAUSE
OPERATION
CONSEQUENCE
Verify size input check
Inconsistent Value Error - propagates
BF class: DVR
Missing/Erroneous Code Bug (absent size check)
Inconsistent Value → Wrong Argument BF class: TCM
Wrong Argument Fault (from W1 error)
Calculate co64 atom size calc
Wrap Around Error - propagates
Wrap Around → Wrong Size BF class: MMN
Wrong Size Fault (from W2 error)
Allocate under-sized buffer
Insufficient Size Error - propagates
Insufficient Size → Insufficient Size BF class: MAD
Reposition pointer past bound
Insufficient Size Fault (from W3 error)
Overbound Pointer Error - propagates
Overbound Pointer → Overbound Pointer BF class: MUS
Write write past end
Overbound Pointer Fault (from W4 error)
Buffer Overflow Final Error
Failure: ACE remote code execution (RCE) LEGEND Cause (Bug or Fault) Consequence (Error, propagates) Propagation between weaknesses
Operation (BF class identifier) Final Error (exploit entry point)
The five-stage BadAlloc pattern: an unverified size input wraps in the co64 size arithmetic, under-allocates the buffer, repositions the pointer past its bound, and the write overflows; the failure class is ACE. Reference stratum: axis values follow the BF specification. Fig. 13. BF specification of the GPAC MPEG-4 vulnerability (CVE-2021-21834) as a BadAlloc-pattern chain. 7 CWE-190: https://cwe.mitre.org/data/definitions/190.html; CWE-787: https://cwe.mitre.org/data/definitions/ 787.html.
1:25
CWE failures illustrated. This CVE exemplifies the absence of a causal chain (F4): the chain spans five BF weakness stages (DVR, TCM, MMN, MAD, MUS) and a failure stage (ACE / RCE), and no single CWE entry can compose them. It also shows the non-orthogonal hierarchy (F1) through the co-applicability of CWE-119, CWE-190, and CWE-787, as well as sink-only labeling (F3). BF classification. NIST SP 800-231 [6] specifies the BadAlloc pattern chain explicitly as DVR ↷ TCM ↷ MMN ↷ MAD ↷ MUS. Reading CVE-2021-21834 against this template, the first weakness is a ⟨Missing/Erroneous Code, Verify⟩ → Inconsistent Value weakness in DVR; the second is a ⟨Wrong Argument, Calculate⟩ → Wrap Around weakness in TCM; the third is a ⟨Wrong Size, Allocate⟩ → Insufficient Size weakness in MMN; the fourth is a ⟨Insufficient Size, Reposition⟩ → Overbound Pointer weakness in MAD; and the fifth is a ⟨Overbound Pointer, Write⟩ → Buffer Overflow weakness in MUS. Figure 13 depicts the five stages. BF properties that address each illustrated failure. For F1, orthogonal weakness classes map each stage to a single class. For F3, the explicit cause axis records the Missing/Erroneous Code bug at the head of the chain. For F4, the cause-consequence chain composes the five weaknesses into one representation: Inconsistent Value ↷ Wrong Argument; Wrap Around ↷ Wrong Size; Insufficient Size ↷ Reposition fault; Overbound Pointer ↷ Write fault. 7.4
CVE-2014-0160 (Heartbleed), BF_INP → BF_MEM, with convergence
CVE summary. Heartbleed (CVE-2014-0160) was a high-severity information-disclosure vulnerability in the TLS and DTLS heartbeat extension of the OpenSSL cryptographic library, with a CVSS v3.1 base score of 7.5 [55]. In OpenSSL 1.0.1 through 1.0.1f, the heartbeat handler omits a bounds check before a memcpy() whose length is taken from an attacker-controlled payload-length field [17, 62]. With each request, the responder discloses up to 64KB of memory through buffer over-reads [6, 12, 55]. Original CWE assignment. The NVD record originally labeled CVE-2014-0160 with the Class-level entry CWE-119 [41], Improper Restriction of Operations within the Bounds of a Memory Buffer.8 Current CWE assignment. The label was later narrowed to CWE-125, Out-of-bounds Read [43], with9 related labels CWE-126 (Buffer Over-read) and CWE-20 (Improper Input Validation) sometimes referenced as upstream causes. Analysis of the reassignment. NIST SP 800-231 records the issue plainly: “CVE-2014-0160 Heartbleed lists the final error at the sink, buffer over-read, as the root cause, while it is missing input verification that leads to pointer reposition over the upper bound and then to buffer over-read” [6]. CWE failures illustrated. This CVE illustrates sink-only labeling (F3): the assigned CWE describes the over-read manifestation, not the missing input verification that produced it. It illustrates the absence of a causal chain (F4) because the defect spans three weaknesses in the main chain and a second converging chain (MUS Clear) that the single CWE-125 cannot represent. It exemplifies non-orthogonal hierarchy (F1) through the co-applicability of CWE-125, CWE-126, and CWE-20 to different stages of the same defect. BF classification. NIST SP 800-231 [6] specifies the BF chain for Heartbleed. The main chain comprises three weaknesses: • ⟨Missing Code, Verify⟩ → Inconsistent Value in DVR; 8 CWE-119: https://cwe.mitre.org/data/definitions/119.html. 9 CWE-125: https://cwe.mitre.org/data/definitions/125.html.
1:26
Hoque et al.
CAUSE
OPERATION
CONSEQUENCE
Verify Check payload length
Inconsistent Value Error - propagates
BF class: DVR
Missing Code Bug (code defect)
Inconsistent Value → Wrong Size BF class: MAD
Reposition Move pointer by size
Wrong Size Fault (from W1 error)
Overbound Pointer Error - propagates
Overbound Pointer → Overbound Pointer BF class: MUS
Overbound Pointer Fault (from W2 error)
Read Read via pointer
Buffer Over-Read Final Error
Failure: IEX Information Exposure LEGEND Cause (Bug or Fault) Consequence (Error, propagates) Propagation between weaknesses
Operation (BF class identifier) Final Error (exploit entry point)
Fig. 14. BF formal specification of the Heartbleed vulnerability (CVE-2014-0160) as a chain of weaknesses.
• ⟨Wrong Size, Reposition⟩ → Overbound Pointer in MAD; • ⟨Overbound Pointer, Read⟩ → Buffer Over-Read in MUS. A converging chain in MUS records the parallel weakness ⟨Missing Code, Clear⟩ → Not Cleared Object. The two chains converge at the IEX failure: “Either the missing Verify bug or the missing Clear bug has to be fixed to avoid this security failure” [6]. BF properties that address each illustrated failure. For F3, the explicit cause axis records the Missing Code bug at the Verify operation. For F4, the cause-consequence chain combines the three mainchain weaknesses and the converging MUS Clear weakness into a single specification, without collapsing them into CWE-125. For F1, the disjoint operation sets across DVR, MAD, MUS, and the IEX failure class prevent freedom of ancestor-or-sibling labels. 8
Empirical Evaluation
Study 1 measures the reproducibility of manual CVE-to-BF mapping. Study 2 measures the determinism of automated CVE-to-BF mapping. Both studies report reproducibility, not human-validated accuracy. 8.1
Study 1: Inter-Rater Reliability of Manual Mapping
Design. Two annotators independently and anonymously mapped 13 CVEs onto the four BF axes of cause, operation, consequence, and attribute. Agreement is quantified per axis with Cohen’s 𝜅 [14] and interpreted against the Landis-Koch bands [30]. The statistic measures chance-corrected reproducibility of the mapping procedure.
1:27
Results. Cause and operation are reproducible at almost-perfect agreement, with 𝜅 = 0.876 and 𝜅 = 0.900 respectively; consequence reaches substantial agreement at 𝜅 = 0.626; attribute reaches fair agreement at 𝜅 = 0.235; and the pooled statistic over all 52 cells is 𝜅 = 0.732. The pattern supports the orthogonality argument on the axes where the specification constrains the choice most tightly, and it locates the principal limitation on the attribute axis. 8.2
Study 2: Automated Determinism of CVE-to-BF Mapping
The study is implemented using an open-source, reproducible FastAPI application for real-time agentic inference [36]. Because a determinism figure measured on one deployment cannot separate what the task contributes from what the model and its serving stack contribute, the study executes the same generation design as 2 combinations: each combination is one model deployment at its own prompt budget. Experiment Setup and Configuration. Two open-weight deployments participate: Gemma 4 31B (the primary deployment) and GPT-OSS 120B, each invoked through the vLLM OpenAI-compatible endpoint. Every model call uses temperature 0, and a single prompt-budget regime completes the design. Under the per-model budget (combinations individual-max-capacity), the budget is derived from each deployment’s own context window, so each deployment reads as much evidence as it can hold and neither is held to the other’s ceiling. Evidence follows a clean-evidence policy: the model receives only the NVD advisory text and the fix-commit patch material. 8.3
The Formal Derivation Procedure
The evidence-based generation of Study 2 instantiates one underlying procedure, which Algorithm 1 states in the notation of Table 5 and Figure 9 renders as an architectural process. The procedure operationalizes backward bug identification as defined in SP 800-231: the walk starts at the observable failure, fills one weakness triple per link, and terminates when a cause resolves to a bug rather than a fault [6]. The closed tables 𝑇𝐴 to 𝑇𝐻 are transcriptions of the vocabulary tables of SP 800-231, that is the class types and member classes, the operations per class, the bug values, the fault and error values with their type dimension, the final-error values with the failure classes, and the operation and operand attribute values; the transcription adds nothing and removes nothing, so every value the algorithm can generate already exists in the standard. A fixed sequence of five lookups then derives each link. The class assignment is the deterministic map 𝜅, because the closed value sets assign each error value to exactly one producing class; the operation, cause, and attribute assignments are semantic-inference selections; and the stopping test at line 12 is a set membership test against the bug table 𝑇𝐶 , because a fault is a good operation over a bad operand and so always points one link further back. In contrast, a bug is the improper operation itself and therefore ends the chain. When the promoted fault belongs to a different class type than the current link, the algorithm records a class-type crossing, which SP 800-231 permits as a meaningful value propagation between class types [6]. The split between the two lookup modes defines the neuro-symbolic boundary of the architecture. A deterministic lookup, such as 𝜅 or the stopping test, is a set or function evaluation over the closed tables and never involves a model. A semantic-inference selection is the oracle call Ω(𝑞, 𝑀 | 𝐸) of Table 5, and the apparatus constrains it three ways: the prompt presents the complete menu 𝑀 of legal values, decoding is constrained by a JSON schema whose value domain is exactly 𝑀, and the temperature is zero. The oracle therefore selects among vocabulary entries and can never introduce a value from outside the closed sets; when the model transport fails, or the constrained decoder rejects the schema, the step raises a typed error, and the run halts rather than degrading to unconstrained generation. The verification phase then re-establishes every property symbolically,
1:28
Hoque et al.
Algorithm 1 : Backward BF Chain Derivation with Neuro-Symbolic Verification Require: CVE 𝑣; evidence 𝐸 = (𝐸 dsc, 𝐸 cod, 𝐸 adv ); closed tables 𝑇𝐴 ..𝑇𝐻 ; layer bound 𝑁 max Ensure: chain C = ⟨𝑊𝑁 ↷ · · · ↷𝑊1 ⟩ → 𝐹 ; audit R 1: {Derivation phase (backward, sink to root)} 2: 𝐹 ← Ω∗ 𝑞 flr, 𝑇𝐴 [_FLR] | 𝐸 dsc {failure classes} 3: 𝐶𝑛 1 ← Ω 𝑞 err, values(𝑇𝐸 ) | 𝐸 dsc {final error} 4: 𝑛 ← 1 5: repeat 6: 𝑊𝑛 ← 𝜅 (𝐶𝑛𝑛 ) {deterministic: the consequence fixes the class} 7: 𝑂𝑝𝑛 ← Ω 𝑞 op, 𝑇𝐵 [𝑊𝑛 ] | 𝐸 dsc, 𝐸 cod 8: {the full operation menu of 𝑊𝑛 is scored; one is kept} 9: 𝐶𝑠𝑛 ← Ω 𝑞 cs, 𝑇𝐶 ∪ 𝑇𝐷 [𝑊𝑛 ] | 𝐸 cod, Obj(𝑂𝑝𝑛 ) 10: 𝐴𝑛 ← Ω-filled attributes over 𝑇𝐺 [𝑊𝑛 ] and 𝑇𝐻 11: R ← R ∪ {record(𝑛, query, result, anchor)} 12: if 𝐶𝑠𝑛 ∈ 𝑇𝐶 then 13: break {a bug is an improper operation: root found} 14: end if 15: 𝐶𝑛𝑛+1 ← 𝐶𝑠𝑛 {the error out of 𝑊𝑛+1 equals the fault into 𝑊𝑛 } 16: if 𝜏 (𝜅 (𝐶𝑛𝑛+1 )) ≠ 𝜏 (𝑊𝑛 ) then 17: R ← R ∪ {crossing(𝑛)} 18: {meaningful value propagation across class types} 19: end if 20: 𝑛 ←𝑛+1 21: until 𝑛 > 𝑁 max 22: {Verification phase (symbolic, deterministic)} 23: 𝑃 voc ← ∀𝑥 ∈ C : 𝑥 ∈ 𝑇𝐴 ..𝑇𝐻 {closed vocabulary} 24: 𝑃 str ← C |= Σ {Σ: SHACL shapes; one cause, operation, consequence per 𝑊 ; one sink} 25: 𝑃 ort ← ∀𝑛 : class(𝑂𝑝𝑛 ) = 𝑊𝑛 {OWL restrictions: the operation fixes the class} 26: 𝑃 lnk ← ∀𝑛 > 1 : 𝐶𝑠𝑛 = 𝐶𝑛𝑛+1 {linking rule} 27: if ¬(𝑃voc ∧ 𝑃 str ∧ 𝑃 ort ∧ 𝑃 lnk ) then 28: return (⊥, R ∪ {violation report}) 29: end if 30: return (C, R)
independently of the oracle. The vocabulary predicate 𝑃 voc confirms closed-set membership. The structural predicate 𝑃 str validates the chain against a set of Shapes Constraint Language (SHACL) shapes [28], which require exactly one cause, one operation, and one consequence per weakness and exactly one sink weakness per chain. The orthogonality predicate 𝑃ort confirms that each selected operation belongs to its link’s class, which the accompanying Web Ontology Language (OWL) ontology encodes as a class restriction per operation [85]. The linking predicate 𝑃 lnk confirms that the cause of each link equals the consequence of the link behind it. A violation returns to the derivation engine as a structured report and the affected steps regenerate, so an accepted chain always satisfies all four predicates; the vocabulary, the ontology, the shapes, and every conformant chain are released with the codebase as machine-readable artifacts, and each accepted chain carries the full audit trail R of its per-step queries, results, and evidence anchors. The generation design. The pipeline is evidence-based and stepwise, where for each CVE, the model derives a BF weakness chain from the CVE’s evidence bundle alone, with one constrained oracle call per derivation step, over repeated independent rounds. Each generated chain passes
1:29
Frame construction
Removed
Remaining 363,394
Baseline corpus (CVE JSON 5.0 records) E1: record not in PUBLISHED state E2: description under 20 words
17,642 27,792
345,752 317,960
E3: no reference link E4: NIST-authored BF specification
0 21
317,960 317,939
Stratum S1: _INP family
Frame size 55,798
Allocated 175
S2: _MEM family
40,803
128
S3: _DAT family
17,013
54
S4a: CWE outside current BF scope
36,245
114
S4b: no usable CWE in the record Total
168,080 317,939
529 1,000
Table 4. Sampling frame and stratified allocation of the evaluation draw (seed 20260706); allocation is proportional with a fixed largest-remainder rule.
through the symbolic verifier of Section 8.3 and every stored chain carries its deployment identifier and budget mode, so a combination is identified from the data rather than folder. CVE selection. The sampling frame is derived from a seeded, stratified probability sample of 1,000 CVEs. The population is every published record of the curated CVE corpus that does not appear in a frozen 26-CVE exclusion set. Stratification uses a two-layer CWE-to-BF-family crosswalk, where an unassigned record moves between strata, thereby costing subgroup precision. One deviation is recorded: the NVD enrichment API was unreachable, so the primary CWE was resolved, which inflates the no-usable-CWE stratum S4b. Each deployment’s evaluation corpus is then resolved from this frame: the 20 CVEs whose NIST-verified BF chains are published are explicitly named, and the frame’s leading records are graded by the confidence of their retrievable evidence until the high- and medium-confidence quotas are filled. Each deployment’s resolved corpus is therefore the verified class plus its own graded selections, and the per-combination totals that carry 3 completed rounds are reported with the results rather than fixed in advance. The verified 20 add a reading that the drawn records are excluded by design: for these CVEs, a derived chain can be compared against a published answer chain, and that comparison is reported case by case, never as a pooled accuracy score. Every combination targets 3 independent rounds per CVE. Cross-combination comparisons are computed on a balanced window of CVEs that carry 3 completed rounds on every combination; within-combination readings are computed per combination over that combination’s own CVEs with 3 completed rounds, capped at 3. Metrics. Chain-level determinism compares every pair among a group’s chains (a group is one CVE’s completed rounds within one combination) after canonicalization to the ordered weakness tuples of cause, operation, consequence, and finality plus the failure multiset: exact determinism 𝐷 is the fraction of identical pairs, structural similarity 𝑆 scores partial agreement on a 0 to 1 scale, a per-axis decomposition over eight independent axes (chain length, failure set, and the root and
1:30
Hoque et al.
sink values of cause, operation, and consequence) locates where variation lives, and each group’s 𝐷 carries a Wilson 95% interval [87]. The weakness class is not a coordinate for comparison. Notation Directory Symbol 𝑣
Description
𝐸
clean evidence bundle (𝐸 dsc, 𝐸 cod, 𝐸 adv ): advisory text, buggy source with its fix commit, and referenced advisory pages
𝑇𝐴
BF class types and their member classes
𝑇𝐵
BF operations per class (one operation name belongs to exactly one class)
𝑇𝐶
bug values (code bugs and specification bugs)
𝑇𝐷
fault and error values with their type dimension
𝑇𝐸
final-error values and security-failure classes
𝑇𝐺
operation-attribute values (Mechanism, Source Code, Execution Space)
𝑇𝐻
operand-attribute values over the eight operand facets
𝑊𝑛
weakness 𝑛, counted backward from the sink: (𝐶𝑠𝑛 , 𝑂𝑝𝑛 , 𝐶𝑛𝑛 , 𝐴𝑛 ), that is cause, operation, consequence, attributes
𝐹 𝜅
the failure classes of the chain, 𝐹 ⊆ 𝑇𝐴 [_FLR]
𝜏
class-type map: 𝜏 (𝑐) is the class type of class 𝑐 in 𝑇𝐴
Ω
constrained semantic-inference oracle: Ω(𝑞, 𝑀 | 𝐸) ∈ 𝑀 returns exactly one element of the closed menu 𝑀 for question 𝑞 under evidence 𝐸; Ω ∗ is its multiselect form, Ω∗ (𝑞, 𝑀 | 𝐸) ⊆ 𝑀
𝑃voc, 𝑃str,
verification predicates: closed-vocabulary membership, structural
𝑃ort, 𝑃lnk
shape conformance, operation-class orthogonality, and link consistency
C, R 𝑁 max
the derived chain and its audit trail layer bound (nine, the number of BF weakness classes)
CVE identifier under analysis
class map: 𝜅 (𝑒) is the unique class whose error set contains the value 𝑒, per the orthogonality of the closed value sets
Table 5. Notation for Algorithm 1; the closed vocabulary tables 𝑇𝐴 to 𝑇𝐻 are transcribed verbatim from NIST SP 800-231 [6].
Justification of the metric suite. The structural similarity is a single normalized agreement count under optimal alignment: for a chain pair (𝑎, 𝑏) the weakness sequences are aligned by dynamic programming with per-field partial credit, and 𝑚(𝑎, 𝑏) + |𝐹𝑎 ∩ 𝐹𝑏 | 𝑆 (𝑎, 𝑏) = , (1) 4 · max(|𝑊𝑎 |, |𝑊𝑏 |) + max(|𝐹𝑎 |, |𝐹𝑏 |) where 𝑚(𝑎, 𝑏) counts matched weakness fields under the optimal alignment, 𝐹 is the failure multiset, 𝑊 the weakness sequence, and four is the number of independently compared fields per weakness (cause, operation, consequence, finality). 𝑆 equals one exactly on identical canonical chains. Two
1:31
readings of a chain are reported below. The whole chain is every compared field of every weakness; the root-to-sink class-chain is the sequence of BF class types alone, read from the root weakness to the sink, so two chains share a class-chain when they pass through the same classes whatever they assign beneath them. The round count per CVE follows from the clustered pair structure: pairs grow as 𝑘2 , while groups remain independent units, so many moderate-𝑘 groups dominate a few large-𝑘 groups. At 3 rounds, 𝐷 is therefore a per-CVE indicator on the admissible grid {0, 31 , 23 , 1} rather than a continuous rate. 8.4
Study 2 Results
Deployment
Prompt budget
Groups 𝑛
𝐷 mean
𝐷 median
𝑆 mean
Fully deterministic
Gemma 4 31B
Own
85
0.890
1.000
0.968
72 of 85
GPT-OSS 120B
Own
85
0.094
0.000
0.577
3 of 85
Bands, Landis and Koch [30]: moderate 0.41 to 0.60
slight 0.00 to 0.20 substantial 0.61 to 0.80
fair 0.21 to 0.40 almost perfect 0.81 to 1.00
Fig. 15. Determinism of the evidence-based derivation for each of the 2 model-budget combinations, over one balanced window of 85 CVEs at 3 rounds.
The balanced cross-combination window. Figure 15 and Figure 16 report determinism per combination over the balanced window of CVEs carrying 3 completed rounds, which admits a window of size 85 at 3 rounds per combination. On that window Gemma 4 31B reproduces identical chains for most CVEs (𝐷 = 0.890, 𝑆 = 0.968), while GPT-OSS 120B reproduces almost none (𝐷 = 0.094 with 𝑆 = 0.577). The GPT-OSS 120B similarity score sits well above zero, and the per-axis block of Figure 16 places its disagreement on the whole-chain, sink, and failure axes rather than on the root. Within-combination determinism. Figure 20 extends the per-axis reading to each combination’s full coverage: 352 CVEs on Gemma 4 31B and 94 on GPT-OSS 120B. Gemma 4 31B holds every axis high; on GPT-OSS 120B the profile sits far lower on every axis. Figure 17 states the two-number rule: the count of CVEs whose class-chain reproduces against the count whose whole chain does. On Gemma 4 31B the class-chain reproduces identically for more CVEs than the whole chain does (312 against 282 of 352), so the residual variation lives below the class-chain in attribute-level and value-level fields. The gap is widest where determinism is lowest: on GPT-OSS 120B the class-chain reproduces for 19 CVEs against 4 whole chains. Determinism by evidence-confidence stratum. Figure 20 compares the fully deterministic share across the evidence-confidence strata. On Gemma 4 31B the high-confidence stratum is fully deterministic for 132 of 145 CVEs (mean 𝐷 0.938) while the medium-confidence stratum is fully deterministic for 132 of 185 (mean 𝐷 0.793). On GPT-OSS 120B both strata are fully deterministic for a small minority (3 of 28 against 0 of 7), on counts small enough that the comparison is descriptive only. The strata are observed rather than assigned, so these are differences across groups, not effects of the group, and Appendix B records the corresponding conditions. The verified-corpus reading. Figure 21 shows which class-chains each deployment builds against the class-chains the verified corpus contains; Figure 18 reports the share of rounds whose class-chain equals the published class-chain; and Figure 19 plots that per-CVE indicator against whole-chain
1:32
Hoque et al.
Deployment Prompt budget CVE groups, 𝑛
Gemma 4 31B
GPT-OSS 120B
Own 85 3 0.890 1.000 0.000 0.847 0.968
Own 85 3 0.094 0.000 0.000 0.035 0.577
0.976 0.992 0.918 0.953 0.969 0.992 0.953 0.984
0.502 0.678 0.333 0.345 0.514 0.729 0.784 0.573
Rounds, 𝑅 Exact determinism 𝐷, mean 𝐷, median 𝐷, minimum Fully deterministic share Structural similarity 𝑆, mean Per-axis agreement chain length root cause root operation root consequence sink cause sink operation sink consequence failure set
Bands, Landis and Koch [30]: slight 0.00 to 0.20 fair 0.21 to 0.40 moderate 0.41 to 0.60 substantial 0.61 to 0.80 almost perfect 0.81 to 1.00 Every combination runs at its deployment’s own context ceiling (Own), so a difference across a row is attributable to the deployment. Configuration rows are not shaded. Every value is measured. Fig. 16. Determinism of the evidence-based derivation by model-budget combination, over one balanced window of 85 CVEs at 3 rounds.
whole chain identical
class-chain identical
Gemma 4 31B (𝑛 = 352) GPT-OSS 120B (𝑛 = 94) 0
0.25 0.50 0.75 share of CVEs with all rounds identical
1.0
Share of CVEs whose repeated rounds produce an identical whole chain (filled) against an identical class-chain (open). Both deployments read to their own context ceiling. Computed per combination over that combination’s own CVEs with 3 completed rounds (per-combination 𝑛 at left). Fig. 17. The two determinism numbers of the evidence-based derivation, one row per combination.
determinism. The three readings agree: the derived class-chains concentrate on short two- and three-class paths where the verified corpus concentrates on DVR → MAD → MUS and its relatives, and a reliably reproduced chain most often reproduces a class-chain that differs from the published
1:33 Gemma 4 31B
GPT-OSS 120B
both at their own context ceiling verified class-chain (measured)
CVE-2013-4934 (2) CVE-2015-5221 (2) CVE-2017-17833 (2) CVE-2006-2362 (3) CVE-2007-1320 (3) CVE-2008-4539 (3) CVE-2014-0160 (3) CVE-2018-14557 (3) CVE-2019-14814 (3) CVE-2007-6429 (4) CVE-2013-4930 (4) CVE-2015-0235 (4) CVE-2021-21834 (5) CVE-2022-34835 (7)
MAD→MMN MMN→MUS MAD→MMN DVR→MAD→MUS DVR→MAD→MUS DVR→MAD→MUS DVR→MAD→MUS DCL→MAD→MUS DVR→MAD→MUS DVR→TCM→MMN→MUS DVR→TCM→TCV→MMN TCM→MMN→MAD→MUS DVR→TCM→MMN→MAD→MUS DCL→TCV→TCM→TCV→TCV→MAD→MUS 0
1/3
2/3
1
share of 3 rounds whose class-chain equals the verified class-chain Over the NIST-verified CVEs with a published verified class-chain (14 out of 20), chain length in parentheses, up to one mark per combination and CVE on the round grid {0, 13 , 32 , 1}; a combination without 3 completed rounds for a CVE carries no mark on that row. The right column prints the published verified class-chain. Fig. 18. Measured per-CVE verified-class-chain agreement over the NIST-verified CVEs, sorted by verified chain length.
one. The mark cluster in the lower right region of Figure 19, high 𝐷 at zero agreement, is therefore the modal outcome: the pipeline is stable on a rendering of the vulnerability that is legal BF and is not the published rendering, a behavior this paper reads against the framework properties recorded in Appendix B. 8.5
Interpretation and Scope
First, determinism under a baseline configuration is a property of the deployment, not of the task, so the task admits deterministic automation but does not guarantee it. Second, where determinism is high, the residual variation is localized below the class-chain, in attribute-level fields, which extends the Study 1 pattern: the axes the specification constrains most tightly reproduce best in both studies. Third, stability and fidelity are distinct properties: the determinism measured here is a claim about repeatability under the baseline configurations and the evidence-based input representation, never a claim of accuracy relative to expert-adjudicated chains. Appendix B records the corresponding threats.
1:34
Hoque et al. Gemma 4 31B (𝑛 = 14) GPT-OSS 120B (𝑛 = 12) both at their own context ceiling 𝐷 = 0.8 verified-class-chain agreement ratio
1
2/3 agreement = 0.5 1/3
0 0
1/3
2/3
1
whole-chain exact determinism 𝐷 ({0, 31 , 23 , 1}) Whole-chain exact determinism 𝐷 against the per-CVE share of rounds whose class-chain equals the published verified class-chain, over the verified CVEs with 3 completed rounds in that combination and a published verified class-chain (per-combination 𝑛 in the legend). Both measures live on the round grid {0, 13 , 32 , 1} and marks carry a small fixed dodge per combination. The 𝑦 series is a per-CVE indicator, not a pooled accuracy score. Fig. 19. Measured determinism against verified-class-chain agreement for the NIST-verified CVEs, one mark per CVE and combination.
9
Scope and Boundary Conditions
Consistent with reporting-guideline practice for systematic reviews [63], the paper’s claims apply within stated boundaries, recorded here so the comparison is judged against what was established, not against a wider construction. BF coverage scope. The BF taxonomy as of NIST SP 800-231, July 2024, defines a finite set of weakness and failure class types: Input/Output Check (INP); Memory (MEM); Data Type (DAT); and Failure (FLR) [6]. BF covers a defined subset of the vulnerability classes that CWE covers, and CWE entries outside this set are out of scope for the comparison. Restriction of comparison. The paper’s comparison applies within BF’s current coverage. CVEs whose root cause lies outside the four BF class types named above are addressed in future work, as BF extends across the forthcoming NIST class-type specifications [6], not by the comparison reported here. Study scope. The paper presents four case studies, one per BF class type, and an inter-rater study computed over 𝑁 = 13 CVEs that form the clean unmapped stratum: CVEs that carry neither a published BF mapping nor prior exposure to the derivation procedure’s reasoning. The unmapped stratum spans three class types, with 9 CVEs in BF INP, 2 in BF DAT, and 2 in BF MEM; BF FLR appears on the consequence axis of multiple CVEs as Information Exposure and Injection final errors rather than as a weakness class type.
1:35 1.0 0.75 Gemma 4 31B (𝑛 = 352)
GPT-OSS 120B (𝑛 = 94)
chain length
0.979
0.514
root cause root operation root consequence
0.994
0.652
0.909
0.351
0.934
0.376
sink cause sink operation
0.951
0.518
0.979
0.734
1.0
sink consequence
0.955
0.798
0.75
failure set
0.989
0.599
0.50
mean pairwise agreement 0 share per axis:
0.50 0.25 0 19 15 verified
145 28 high confidence
185 7 medium confidence
(a) fully deterministic share by evidence-confidence stratum
0.25 0 19 15 verified
330 35 drawn complement
(b) fully deterministic share by corpus source
1
1 is the expected value; higher is better. Both combinations read up to their own deployment’s context ceiling. A cell of 1.000 means every pair of rounds agreed on that axis for every CVE in that combination; 0.000 means no pair agreed. Each column is computed on that combination’s own CVEs with 3 completed rounds, capped at 3. Figure 16 instead uses the balanced window every combination covers, so its 𝑛 is smaller.
Gemma 4 31B GPT-OSS 120B both at their own context ceiling; whiskers: Wilson 95%; 𝑛 under each mark; point when 𝑛 < 5; dash when 𝑛 =0 Left: share of fully deterministic CVEs (𝐷 = 1.0 over 3 rounds) per group and combination, with Wilson 95% whiskers and per-group 𝑛, across (a) the evidence-confidence class of the selected CVEs (the NIST-verified 20 beside the high- and medium-confidence selections) and (b) corpus source (verified against the confidence-selected drawn complement). Computed per combination over that combination’s own CVEs with 3 completed rounds; a group with 𝑛 < 5 renders as a point and is read as descriptive only. Fig. 20. Left: measured determinism of the evidence-based derivation by evidence-confidence stratum, per group and combination. Right: measured per-axis determinism profile of the evidence-based derivation, one column per combination; reading notes and the scale key are inside the panel.
Annotator setup. Two independent human annotators produced the anonymous mappings on an identical template, and the per-axis agreement was computed on their labels. Case studies cite the resolved labels from the inter-rater study where available, and the BF specification for case-study CVEs with published NIST chains, that is, CVE-2014-0160 (Heartbleed) and the BadAlloc pattern, which includes CVE-2021-21834. The remaining case-study mappings take their BF axis values from the inter-rater study (CVE-2021-3156) and the established memory-bugs class specification (CVE-2015-0235), as stated in each case study’s BF classification paragraph, and are not implied to carry an independent agreement statistic.
1:36
Hoque et al.
BF maturity. BF is an evolving framework [6, 53]: NIST has announced companion class-type specifications, and the BF Taxonomy at SAMATE is maintained as a live resource independently of the specification snapshot used here. The four-to-four mapping in Section 6.1 and the case-study labels in Section 7 are grounded in NIST SP 800-231 as the specification of record at the date of writing; a proposed CVE-to-BF system would consume the BF taxonomy as a live resource instead of hard-coding this snapshot. NIST-verified corpus (𝑛 = 14 out of 20) DVR MAD MMN DCL DVR DVR TCM DVR DCL
MAD MUS MMN MUS MAD MUS TCM MMN MUS TCM TCV MMN MMN MAD MUS TCM MMN MAD TCV TCM TCV
class-chain frequency:
MUS TCV
1 to 2
Gemma 4 31B (𝑛 = 352)
MAD
MUS
3 to 9
×5 ×2 ×1 ×1 ×1 ×1 ×1 ×1 ×1
DVL DVL DVL MAD DVL DVR DVL other: 28 class-chains
×81 ×74 ×66 ×23 ×108
GPT-OSS 120B (𝑛 = 94) DVL DVL MUS NRS DVL TCM DVL other: 32 class-chains
×16 ×14 ×8 ×6 ×50
10 or more
Left: the distinct verified class-chains of the NIST-verified corpus (14 out of 20), one row per class-chain, frequency at right and encoded by chip shade on one shared blue ramp. Right: the measured modal derived class-chains per deployment at its own context ceiling, one panel per deployment over its own CVEs with 3 completed rounds (per-panel 𝑛 in the header). Fig. 21. Class-chains, root to sink, in the multi-model evaluation.
Synthesis-mode boundary. Section 5 and the integrated survey tables (see Table 2b) cover all 20 papers in the corpus. Where a methodological field is not reported in the source, the table cell uses the n.r. marker defined in the table legend. The boundary between sourced and schema-derived placement is marked in Table 2a itself through the footnoted cells so that a reader can separate authorial synthesis from verbatim reporting at the point of use. 10
Conclusion
This paper presented a literature-grounded case study of NIST SP 800-231 as a target for automated vulnerability classification. The evidence suggests that current CVE-to-CWE automation faces a structural target-space problem. At the same time, the NIST Bugs Framework offers a more structured and causally meaningful representation through its compositional design. Our empirical studies show that LLM-based BF classification can achieve substantial repeatability under stable deployments, although reproduced chains do not always match published verified chains. Overall, the findings position BF as a promising and more automation-suitable target for future vulnerability classification research, with future work focused on improving per-axis validation, annotation consistency, and alignment with verified BF chains.
1:37
Acknowledgments This research was carried out in the PATENT Lab within the Department of Computer Science at The University of Alabama. The opinions and conclusions expressed in this paper are those of the authors and do not necessarily represent the official policies or positions of their affiliated institutions. References [1] Ehsan Aghaei, Ehab Al-Shaer, and Waseem Shadid. 2023. Automated CVE Analysis for Threat Prioritization and Impact Prediction. arXiv:2309.03040 [cs.CR] doi:10.48550/arXiv.2309.03040 [2] Ehsan Aghaei, Waseem Shadid, and Ehab Al-Shaer. 2020. ThreatZoom: Hierarchical Neural Network for CVEs to CWEs Classification. In Proc. 16th EAI International Conference on Security and Privacy in Communication Networks (SecureComm) (LNICST, Vol. 335). 23–41. doi:10.1007/978-3-030-63086-7_2 [3] Massimiliano Albanese, Oluwaseun Adebiyi, and Frank Onovae. 2024. CVE2CWE: Automated Mapping of Software Vulnerabilities to Weaknesses Based on CVE Descriptions. In Proc. 21st International Conference on Security and Cryptography (SECRYPT). SCITEPRESS, 500–507. [4] Daniel Alfasi, Tal Shapira, and Arik Bar Shalom. 2024. Unveiling Hidden Links Between Unseen Security Entities. In Proc. 3rd International Workshop on Graph Neural Networking (GNNet). ACM. doi:10.1145/ 3694811.3697819 Also available as arXiv:2403.02014. [5] Luca Allodi and Fabio Massacci. 2014. Comparing Vulnerability Severity and Exploits Using CaseControl Studies. ACM Transactions on Information and System Security (TISSEC) 17, 1 (2014), 1:1–1:20. doi:10.1145/2630069 [6] Irena Bojanova. 2024. Bugs Framework (BF): Formalizing Cybersecurity Weaknesses and Vulnerabilities. Technical Report NIST Special Publication 800-231. National Institute of Standards and Technology. doi:10.6028/NIST.SP.800-231 [7] Irena Bojanova, Paul E. Black, Yaacov Yesha, and Yan Wu. 2016. The Bugs Framework (BF): A Structured Approach to Express Bugs. In Proc. IEEE International Conference on Software Quality, Reliability and Security (QRS). 175–182. doi:10.1109/QRS.2016.29 [8] Irena Bojanova and Carlos Eduardo Galhardo. 2021. Classifying memory bugs using bugs framework approach. In 2021 IEEE 45th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, 1157–1164. [9] Irena Bojanova, Carlos E. Galhardo, and Sara Moshtari. 2021. Input/Output Check Bugs Taxonomy: Injection Errors in Spotlight. In Proc. IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW). 111–120. doi:10.1109/ISSREW53611.2021.00052 [10] Sergio Caltagirone, Andrew Pendergast, and Christopher Betz. 2013. The Diamond Model of Intrusion Analysis. Technical Report. Center for Cyber Intelligence Analysis and Threat Research. https://apps. dtic.mil/sti/citations/ADA586960 [11] Trisha Chakraborty, Shaswata Mitra, Sudip Mittal, and Maxwell Young. 2022. AI_Adaptive_POW: An AI assisted Proof of Work (POW) framework for DDoS defense. Software Impacts 13 (2022), 100335. [12] CISA. 2014. OpenSSL Heartbleed Vulnerability (CVE-2014-0160). [Online]. Cybersecurity and Infrastructure Security Agency. Available: https://www.cisa.gov/news-events/alerts/2014/04/08/opensslheartbleed-vulnerability-cve-2014-0160. [13] Cisco Talos Intelligence Group. 2021. TALOS-2021-1297: GPAC Project on Advanced Content MPEG-4 Integer Overflow Vulnerabilities. [Online]. Available: https://talosintelligence.com/vulnerability_reports/ TALOS-2021-1297. CVE-2021-21834 through CVE-2021-21852. [14] Jacob Cohen. 1960. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20, 1 (1960), 37–46. doi:10.1177/001316446002000104 [15] Siddhartha Shankar Das, Mahantesh Halappanavar, Antonino Tumeo, Edoardo Serra, Alex Pothen, and Ehab Al-Shaer. 2022. VWC-BERT: Scaling Vulnerability-Weakness-Exploit Mapping on Modern AI Accelerators. In Proc. IEEE International Conference on Big Data (Big Data). IEEE. doi:10.1109/BigData55660. 2022.10020622
1:38
Hoque et al.
[16] Siddhartha Shankar Das, Edoardo Serra, Mahantesh Halappanavar, Alex Pothen, and Ehab Al-Shaer. 2021. V2W-BERT: A Framework for Effective Hierarchical Multiclass Classification of Software Vulnerabilities. In Proc. IEEE 8th International Conference on Data Science and Advanced Analytics (DSAA). 1–12. doi:10. 1109/DSAA53316.2021.9564227 [17] Zakir Durumeric, James Kasten, David Adrian, J. Alex Halderman, Michael Bailey, Frank Li, Nicolas Weaver, Johanna Amann, Jethro Beekman, Mathias Payer, and Vern Paxson. 2014. The Matter of Heartbleed. In Proc. 2014 Conference on Internet Measurement Conference (IMC). ACM, 475–488. doi:10. 1145/2663716.2663755 [18] Forum of Incident Response and Security Teams. 2019. Common Vulnerability Scoring System v3.1: Specification Document. [Online]. Available: https://www.first.org/cvss/v3.1/specification-document. [19] Carlos E. Galhardo, Paul E. Black, and Irena Bojanova. 2020. Measurements of the Most Significant Software Security Weaknesses. In Proc. Annual Computer Security Applications Conference (ACSAC). 154–164. doi:10.1145/3427228.3427257 [20] Haibo He and Edwardo A. Garcia. 2009. Learning from Imbalanced Data. IEEE Transactions on Knowledge and Data Engineering 21, 9 (2009), 1263–1284. doi:10.1109/TKDE.2008.239 [21] Eric M. Hutchins, Michael J. Cloppert, and Rohan M. Amin. 2011. Intelligence-Driven Computer Network Defense Informed by Analysis of Adversary Campaigns and Intrusion Kill Chains. In Proc. 6th Annual International Conference on Information Warfare and Security (ICIW). Lockheed Martin Corporation, 113–125. [22] Jay Jacobs, Sasha Romanosky, Benjamin Edwards, Idris Adjerid, and Michael Roytman. 2021. Exploit Prediction Scoring System (EPSS). Digital Threats: Research and Practice 2, 3 (2021), 20:1–20:17. doi:10. 1145/3436242 [23] Mohammad Jalili Torkamani, Joey Ng, Nikita Mehrotra, Mahinthan Chandramohan, Padmanabhan Krishnan, and Rahul Purandare. 2025. Streamlining Security Vulnerability Triage with Large Language Models. arXiv:2501.18908 [cs.SE] doi:10.48550/arXiv.2501.18908 [24] Peter E. Kaloroumakis and Michael J. Smith. 2021. Toward a Knowledge Graph of Cybersecurity Countermeasures. Technical Report Case 20-2034. The MITRE Corporation. https://d3fend.mitre.org/resources/ D3FEND.pdf [25] Sabrina Kaniewski, Fabian Schmidt, Markus Enzweiler, Michael Menth, and Tobias Heer. 2026. A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models. ACM Trans. Softw. Eng. Methodol. (May 2026). doi:10.1145/3815425 Just Accepted. [26] Kishan Karthik S, Suraj Singh, S Jeevan, Namruth Reddy, Minal Moharir, and Mohana. 2024. Temporal Analysis and Common Weakness Enumeration (CWE) Code Prediction for Software Vulnerabilities Using Machine Learning. In 2024 8th International Conference on Computational System and Information Technology for Sustainable Solutions (CSITSS). 1–6. doi:10.1109/CSITSS64042.2024.10816791 [27] Barbara Kitchenham. 2004. Procedures for Performing Systematic Reviews. Technical Report TR/SE-0401, NICTA-0400011T.1. Keele University and NICTA. [28] Holger Knublauch and Dimitris Kontokostas. 2017. Shapes Constraint Language (SHACL). W3C Recommendation. https://www.w3.org/TR/shacl/. [29] Kethan Kota, Aiswarya Mahesh, S. Justus Daniel Nesakumar, Saroj Sahay, and Mary Anbarasi. 2024. CWE Prediction Using CVE Description: The Semantic Similarity Approach. Procedia Computer Science 235 (2024), 1167–1178. doi:10.1016/j.procs.2024.04.111 [30] J. Richard Landis and Gary G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data. Biometrics 33, 1 (1977), 159–174. doi:10.2307/2529310 [31] Yikun Li, Ngoc Tan Bui, Ting Zhang, Chengran Yang, Xin Zhou, Martin Weyssow, Jinfeng Jiang, Junkai Chen, Huihui Huang, Huu Hung Nguyen, et al. 2025. Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses? arXiv preprint arXiv:2507.21817 (2025). [32] Jeff Luszcz. 2018. Apache Struts 2: how technical and development gaps caused the Equifax Breach. Network Security 2018, 1 (2018), 5–8. doi:10.1016/S1353-4858(18)30005-9 [33] Peter Mell, Karen Scarfone, and Sasha Romanosky. 2007. A Complete Guide to the Common Vulnerability Scoring System Version 2.0. Technical Report. Forum of Incident Response and Security Teams (FIRST).
1:39 https://www.first.org/cvss/v2/guide [34] Shaswata Mitra, Subash Neupane, Trisha Chakraborty, Sudip Mittal, Aritran Piplai, Manas Gaur, and Shahram Rahimi. 2024. Localintel: Generating organizational threat intelligence from global and local cyber knowledge. In International Symposium on Foundations and Practice of Security. Springer, 63–78. [35] Shaswata Mitra, Subash Neupane, Martin Duclos, Sudip Mittal, Aritran Piplai, Md Rayhanur Rahman, Edward Zieglar, and Shahram Rahimi. 2025. FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection. arXiv preprint arXiv:2508.18684 (2025). [36] Shaswata Mitra, Raj Patel, Sudip Mittal, Md Rayhanur Rahman, and Shahram Rahimi. 2026. Agentic AI for Cyber Defense: A Survey of Failure Modes, Applications, and Research Opportunities. (2026). [37] MITRE Corporation. 2024. CVE → CWE Mapping “Root Cause Mapping” Guidance. [Online]. Available: https://cwe.mitre.org/documents/cwe_usage/guidance.html. Document Version 1.1, released 22 March 2024. [38] MITRE Corporation. 2025. Common Attack Pattern Enumeration and Classification (CAPEC). [Online]. Available: https://capec.mitre.org/. [39] MITRE Corporation. 2025. Common Vulnerabilities and Exposures (CVE). [Online]. Available: https: //cve.mitre.org/. [40] MITRE Corporation. 2025. Common Weakness Enumeration (CWE). [Online]. Available: https://cwe. mitre.org/. [41] MITRE Corporation. 2025. CWE-119: Improper Restriction of Operations within the Bounds of a Memory Buffer. [Online]. Available: https://cwe.mitre.org/data/definitions/119.html. [42] MITRE Corporation. 2025. CWE-122: Heap-based Buffer Overflow. [Online]. Available: https://cwe.mitre. org/data/definitions/122.html. [43] MITRE Corporation. 2025. CWE-125: Out-of-bounds Read. [Online]. Available: https://cwe.mitre.org/ data/definitions/125.html. [44] MITRE Corporation. 2025. CWE-190: Integer Overflow or Wraparound. [Online]. Available: https: //cwe.mitre.org/data/definitions/190.html. [45] MITRE Corporation. 2025. CWE-193: Off-by-one Error. [Online]. Available: https://cwe.mitre.org/data/ definitions/193.html. [46] MITRE Corporation. 2025. CWE-678: Composites. [Online]. Available: https://cwe.mitre.org/data/ definitions/678.html. [47] MITRE Corporation. 2025. CWE-709: Named Chains. [Online]. Available: https://cwe.mitre.org/data/ definitions/709.html. [48] MITRE Corporation. 2025. CWE-787: Out-of-bounds Write. [Online]. Available: https://cwe.mitre.org/ data/definitions/787.html. [49] MITRE Corporation. 2025. CWE View-1003: Weaknesses for Simplified Mapping of Published Vulnerabilities. [Online]. Available: https://cwe.mitre.org/data/definitions/1003.html. [50] MITRE Corporation. 2025. D3FEND: A Knowledge Graph of Cybersecurity Countermeasure Techniques. [Online]. Available: https://d3fend.mitre.org/. Version 1.0 (stable release, January 2025); current ontology version 1.4.0 (March 2026). [51] MITRE Corporation. 2025. MITRE ATT&CK Knowledge Base. [Online]. Available: https://attack.mitre. org/. [52] Nikita Mosievskiy. 2026. Fine-tuning RoBERTa for CVE-to-CWE Classification: A 125M Parameter Model Competitive with LLMs. arXiv:2603.14911 [cs.CR] https://arxiv.org/abs/2603.14911 [53] National Institute of Standards and Technology. 2026. Bugs Framework (BF) Project Site. [Online]. NIST SAMATE. Available: https://samate.nist.gov/BF/. Accessed June 2026. [54] Robert C. Nickerson, Upkar Varshney, and Jan Muntermann. 2013. A Method for Taxonomy Development and Its Application in Information Systems. European Journal of Information Systems 22, 3 (2013), 336–359. doi:10.1057/ejis.2012.26 [55] NIST. 2014. CVE-2014-0160 Detail. [Online]. National Vulnerability Database (NVD). Available: https: //nvd.nist.gov/vuln/detail/CVE-2014-0160. [56] NIST. 2015. CVE-2015-0235 Detail. [Online]. National Vulnerability Database (NVD). Available: https: //nvd.nist.gov/vuln/detail/CVE-2015-0235.
1:40
Hoque et al.
[57] NIST. 2017. CVE-2017-5638 Detail. [Online]. National Vulnerability Database (NVD). Available: https: //nvd.nist.gov/vuln/detail/CVE-2017-5638. [58] NIST. 2021. CVE-2021-3156 Detail. [Online]. National Vulnerability Database (NVD). Available: https: //nvd.nist.gov/vuln/detail/CVE-2021-3156. Published 26 January 2021; last modified 10 November 2025. [59] NIST. 2025. CVE-2018-5907 Detail. [Online]. National Vulnerability Database (NVD). Available: https: //nvd.nist.gov/vuln/detail/CVE-2018-5907. [60] NIST. 2025. National Vulnerability Database (NVD). [Online]. Available: https://nvd.nist.gov/. [61] NIST. 2026. NIST Updates NVD Operations to Address Record CVE Growth. [Online]. Available: https: //www.nist.gov/news-events/news/2026/04/nist-updates-nvd-operations-address-record-cve-growth. [62] OWASP Foundation. 2025. Heartbleed Bug. [Online]. Available: https://owasp.org/www-community/ vulnerabilities/Heartbleed_Bug. [63] Matthew J. Page, Joanne E. McKenzie, Patrick M. Bossuyt, Isabelle Boutron, Tammy C. Hoffmann, Cynthia D. Mulrow, Larissa Shamseer, Jennifer M. Tetzlaff, Elie A. Akl, Sue E. Brennan, Roger Chou, Julie Glanville, Jeremy M. Grimshaw, Asbjørn Hróbjartsson, Manoj M. Lalu, Tianjing Li, Elizabeth W. Loder, Evan Mayo-Wilson, Steve McDonald, Luke A. McGuinness, Lesley A. Stewart, James Thomas, Andrea C. Tricco, Vivian A. Welch, Penny Whiting, and David Moher. 2021. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 372 (2021), n71. doi:10.1136/bmj.n71 [64] Mengyuan Pan, Po Wu, Yiwei Zou, Chong Ruan, and Tao Zhang. 2023. An Automatic Vulnerability Classification Framework Based on BiGRU-TextCNN. Procedia Computer Science 222 (2023), 377–386. doi:10.1016/j.procs.2023.08.176 International Neural Network Society Workshop on Deep Learning Innovations and Applications (INNS DLIA 2023). [65] Raj Patel, Shaswata Mitra, Michele Guida, Stefano Iannucci, Sudip Mittal, and Shahram Rahimi. 2026. Agentra: A Supervisable Multi-Agent Framework for Enterprise Intrusion Response. arXiv preprint arXiv:2606.18325 (2026). [66] Kai Petersen, Sairam Vakkalanka, and Ludwik Kuzniarz. 2015. Guidelines for conducting systematic mapping studies in software engineering: An update. Information and Software Technology 64 (2015), 1–18. doi:10.1016/j.infsof.2015.03.007 [67] Qualys Research Team. 2021. Baron Samedit: Heap-Based Buffer Overflow in Sudo (CVE-2021-3156). [Online]. Available: https://www.qualys.com/2021/01/26/cve-2021-3156/baron-samedit-heap-basedoverflow-sudo.txt. Qualys Security Advisory, 26 January 2021. [68] Qualys Security Advisory. 2015. CVE-2015-0235, GHOST: glibc gethostbyname Buffer Overflow. [Online]. Available: https://www.qualys.com/research/security-advisories/GHOST-CVE-2015-0235.txt. [69] Odette Sangupamba Mwilu, Nicolas Prat, and Isabelle Comyn-Wattiau. 2015. Taxonomy Development for Complex Emerging Technologies — The Case of Business Intelligence and Analytics on the Cloud. In Proc. Pacific Asia Conference on Information Systems (PACIS). Singapore. https://hal.science/hal-01636577 HAL Id: hal-01636577. [70] Zhenpeng Shi, Nikolay Matyunin, Kalman Graffi, and David Starobinski. 2024. Uncovering CWE-CVECPE Relations with Threat Knowledge Graphs. ACM Transactions on Privacy and Security 27, 1 (2024), 13:1–13:26. doi:10.1145/3641819 [71] Stefano Simonetto, Ronan Oostveen, Thijs Van Ede, Peter Bosch, and Willem Jonker. 2026. What Matters Most in Vulnerabilities? Key Term Extraction for CVE-to-CWE Mapping with LLMs. In Cryptology and Network Security (CANS 2025), Yongdae Kim, Atsuko Miyaji, and Mehdi Tibouchi (Eds.). Springer Nature Singapore, Singapore, 467–492. doi:10.1007/978-981-95-4434-9_22 [72] Stefano Simonetto, Thijs Sebastiaan van Ede, Peter Bosch, Willem Jonker, and Ronan Oostveen. 2024. Text2Weak: Mapping CVEs to CWEs Using Description Embeddings Analysis. In Proc. 4th Workshop on Artificial Intelligence-Enabled Cybersecurity Analytics. Conference date: 26-08-2024. [73] Şevval Şimşek, Zhenpeng Shi, Howell Xia, David Sastre Medina, and David Starobinski. 2024. Poster: Analyzing and Correcting Inaccurate CVE-CWE Mappings in the National Vulnerability Database. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS ’24). Association for Computing Machinery, Salt Lake City, UT, USA, 5042–5044. doi:10.1145/3658644.3691375 [74] Şevval Şimşek, Howell Xia, Jonah Gluck, David Sastre Medina, and David Starobinski. 2025. Fixing Invalid CVE-CWE Mappings in Threat Databases. In Proc. IEEE 49th Annual Computers, Software, and Applications
1:41 Conference (COMPSAC). IEEE, 950–960. doi:10.1109/COMPSAC65507.2025.00124 The FixV2W method; an extended preprint is available as arXiv:2604.22176. [75] Jonathan M. Spring, Allen Householder, Eric Hatleback, Art Manion, Madison Oliver, Vijay Sarvepalli, Laurie Tyzenhaus, and Charles Yarbrough. 2019. Prioritizing Vulnerability Response: A Stakeholder-Specific Vulnerability Categorization (SSVC). Technical Report. Carnegie Mellon University, Software Engineering Institute. https://resources.sei.cmu.edu/library/asset-view.cfm?assetid=636379 [76] Blake E. Strom, Andy Applebaum, Doug P. Miller, Kathryn C. Nickels, Adam G. Pennington, and Cody B. Thomas. 2020. MITRE ATT&CK: Design and Philosophy. Technical Report MP180360R1. The MITRE Corporation. https://attack.mitre.org/docs/ATTACK_Design_and_Philosophy_March_2020.pdf Originally published July 2018; revised March 2020. [77] Rustam Talibzade, Idilio Drago, and Francesco Bergadano. 2025. On Using LLMs for Vulnerability Classification. LAMPS’25 (2025), 68. [78] Mahzabin Tamanna, Shaswata Mitra, Md Erfan, Ahmed Ryan, Sudip Mittal, Laurie Williams, and Md Rayhanur Rahman. 2026. What Are Adversaries Doing? Automating Tactics, Techniques, and Procedures Extraction: A Systematic Review. arXiv preprint arXiv:2604.02377 (2026). [79] Franco Terranova, Sana Rekbi, Abdelkader Lahmadi, and Isabelle Chrisment. 2026. Multi-Taxonomy Vulnerability Classification with Hierarchically Finetuned Language Models. In DIMVA 2026 - Conference on Detection of Intrusions and Malware and Vulnerability Assessment. Chania, Crete, Greece. https: //hal.science/hal-05500820 [80] Hannu Turtiainen and Andrei Costin. 2024. VulnBERTa: On Automating CWE Weakness Assignment and Improving the Quality of Cybersecurity CVE Vulnerabilities Through ML/NLP. In Proc. 9th IEEE European Symposium on Security and Privacy Workshops (EuroS&PW). IEEE, 618–625. doi:10.1109/EuroSPW61312. 2024.00075 [81] Hannu Turtiainen, Andrei Costin, and Timo Hämäläinen. 2026. VulnBERTa-XAI: Towards Explainable AI for Automating CWE Weakness Assignment and Improving the Quality of Cybersecurity CVE. In Cyber Security, Martti Lehto and Pekka Neittaanmäki (Eds.). Studies in Big Data, Vol. 183. Springer, Cham, 385–432. doi:10.1007/978-3-032-08890-1_16 Equifax to Pay $575 Million as Part of Settlement [82] U.S. Federal Trade Commission. 2019. with FTC, CFPB, and States Related to 2017 Data Breach. Press release. [Online]. Available: https://www.ftc.gov/news-events/news/press-releases/2019/07/equifax-pay-575-million-partsettlement-ftc-cfpb-states-related-2017-data-breach. [83] U.S. Government Accountability Office. 2018. Data Protection: Actions Taken by Equifax and Federal Agencies in Response to the 2017 Breach. Technical Report GAO-18-559. U.S. Government Accountability Office. [Online]. Available: https://www.gao.gov/products/gao-18-559. [84] U.S. House of Representatives, Committee on Oversight and Government Reform. 2018. The Equifax Data Breach: Majority Staff Report. Technical Report. 115th Congress. [Online]. Available: https: //oversight.house.gov/wp-content/uploads/2018/12/Equifax-Report.pdf. [85] W3C OWL Working Group. 2012. OWL 2 Web Ontology Language Document Overview (Second Edition). W3C Recommendation. https://www.w3.org/TR/owl2-overview/. [86] Tianyi Wang, Shengzhi Qin, and Kam Pui Chow. 2021. Towards Vulnerability Types Classification Using Pure Self-Attention: A Common Weakness Enumeration Based Approach. In 2021 IEEE 24th International Conference on Computational Science and Engineering (CSE). 146–153. doi:10.1109/CSE53436.2021.00030 [87] Edwin B. Wilson. 1927. Probable inference, the law of succession, and statistical inference. J. Amer. Statist. Assoc. 22, 158 (1927), 209–212. doi:10.1080/01621459.1927.10502953 [88] Yifan Zhang, Bingyi Kang, Bryan Hooi, Shuicheng Yan, and Jiashi Feng. 2023. Deep Long-Tailed Learning: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10795–10816. doi:10.1109/TPAMI.2023.3268118
A
List of Acronyms
This appendix collects the acronyms and abbreviations used throughout the paper. Table 6 lists the general technical and standards acronyms. Table 7 tabulates the Bugs Framework taxonomy
1:42
Hoque et al.
abbreviations, ordered to replicate the taxonomy structure of Fig. 4, from the BF class types down to their constituent classes. Acronym Expansion ATT&CK
BERT BF CAPEC CPE CTI
Adversarial Tactics, Techniques, and Common Knowledge Bidirectional Encoder Representations from Transformers Bugs Framework Common Attack Pattern Enumeration and Classification Common Platform Enumeration Cyber Threat Intelligence
CVE
Common Vulnerabilities and Exposures
CVSS
Common Vulnerability Scoring System
CWE
Common Weakness Enumeration Exploit Prediction Scoring System
EPSS JSON
JavaScript Object Notation
KG
Knowledge Graph
LLM
Large Language Model
ML
Machine Learning
NIST
National Institute of Standards and Technology
NVD PRISMA
National Vulnerability Database Preferred Reporting Items for Systematic Reviews and MetaAnalyses
SP
Special Publication
Abbr. Name
Category
Class Category
CT CT CT CT C C C C C C C C C C C C C
W W W F INP INP MEM MEM MEM DAT DAT DAT DAT FLR FLR FLR FLR
INP Input/Output Check MEM Memory DAT Data Type FLR Failure DVL Data Validation DVR Data Verification MAD Memory Addressing MMN Memory Management MUS Memory Use DCL Declaration NRS Name Resolution TCV Type Conversion TCM Type Computation IEX Information Exposure ACE Arbitrary Code Execution DOS Denial of Service TPR Data Tampering Category:
CT BF class type C
Class Category:
BF class
W Weakness class type F
Failure class type
INP MEM DAT FLR parent class type of the class Table 7. Bugs Framework taxonomy abbreviations up to the class level, matching Fig. 4; the Category and Class Category codes are expanded in the legend below the table.
Table 6. General technical and standards acronyms used in the paper.
B
Threats to Validity
Several threats condition the claims presented in this paper. The threats are organized into five categories:
1:43
B.1
Construct Validity
Case studies interpret the reassignment of a CVE’s primary CWE as evidence that the original mapping was never uniquely determined. This is indicative rather than conclusive, since a reassignment can instead reflect a deeper root-cause analysis. In addition, four case studies, representing one per BF class, are not enough to prove that these failures are systemic across the wider CVE corpus, so this generalization remains argued rather than measured. B.2
Internal Validity
The case studies were selected based on failure density and anchor status; the full candidate pool of 16 CVEs is published in Table 9 and in the artifact (bf_example_candidates_full_pool.json). Moreover, annotator bias is bounded by the blind two-annotator design whose per-axis agreement Section 8.1 reports; a family-level encoding mismatch drives the low attribute kappa, so determinism language is scoped to the cause and operation axes. Short name
Method family
Year
In corpus
VWC-MAP
V2W-BERT + T5 pipeline
2022
No
Graph + ML ensemble pipeline
2025
No Yes
ThreatCompass CVEDrill
Fine-tuned prioritization
2023
Text2Weak
Embedding retrieval
2024
Yes
Pure Self-Attention Key Term Extraction
Self-attention classification LLM key-term extraction
2021 2026
Yes Yes
Table 8. Anchor papers used as a known-relevant validation set for search recall, with method family, venue year, and final-corpus membership.
B.3
Limitations of the Bugs Framework as an Automation Target
Treating SP 800-231 as the target of automation surfaces five properties of the framework that condition any CVE-to-BF automation, recorded as specification gaps. First, the attribute layer is under-specified: the specification enumerates attribute values but gives little guidance for selecting among them, and nearly all cross-round variation in the automated study concentrates in attribute assignments while the class-chain holds (Section 8.4). Second, the same vulnerability admits multiple legal renderings: chain length is unbounded, the causation rule’s second mode is enumerated for only four propagations, and on the 42-chain verified corpus the strict value-identity reading of causation holds for 21 of 38 adjacent links, so a conforming tool must either implement an unenumerated rule or fail to reconstruct nearly half of the links. Third, a one-weakness chain whose cause resolves to a bug is legal BF while the verified chains decompose the propagation into several phases, and nothing in the specification ranks two conforming outputs of different depth. Fourth, per-class consequence sets are published for two of the nine classes and the cross-class operation flow is stated to be incomplete, so a tool cannot distinguish not-listed from not-allowed. Fifth, a chain is derivable only to the depth the available artifacts expose, so evidence richness bounds chain depth independently of model capability; the evidence ledger accompanies the artifact.
1:44
C
Hoque et al.
Database Search Strategies
This appendix reproduces the search protocol word-to-word that has been executed against each of the six databases during the Identification stage, so the candidate set can be regenerated independently. Two filters were applied uniformly: a publication-year window of January 2018 to April 2026, and a content-type restriction to journal and conference papers. Four databases (arXiv, IEEE Xplore, the ACM Digital Library, and Springer Link) admit field-tagged Boolean queries and were searched with one common keyword concept structure adapted to each database’s syntax. C.1
Anchor Papers: Known-Relevant Validation Set
The anchor papers in Table 8 form a known-relevant validation set which were selected before search execution to check keyword-search recall. An anchor is not guaranteed a place in the included corpus: VWC-MAP and ThreatCompass are downstream pipelines in which CWE is only an intermediate stage toward CAPEC or ATT&CK, so under the exclusion and safety-valve criteria of Section 2 they are excluded from the corpus while still functioning as recall checkpoints. C.2
Case-Study Candidate Pool
Table 9 lists the 16-CVE candidate pool from which Section 7 drew its four case studies, four candidates per BF class-type family. Each candidate is scored by the number of structural CWE failures it exposes; the selected case study attains the maximum count in its family, with anchor status as the tie-breaker. C.3
Identification Yield, Recall, and Query Protocols
Table 10 reports the candidate records contributed per database; pooled and deduplicated, the Identification stage yielded 660 distinct records. The keyword search retrieved all 20 finally included papers and every anchor in Table 8; no supplementary search method was required. Table 11 reproduces the per-database queries verbatim. Google Scholar weights early tokens heavily and degrades on long Boolean strings, so its concept structure was decomposed into twelve short queries run independently and merged; Semantic Scholar was searched by keyword query with a relevance filter; arXiv was restricted to categories cs.CR, cs.CL, and cs.LG; IEEE Xplore used its Command Search interface; the ACM Advanced Search was applied to “Anywhere”; and Springer Link was searched through its field form, since its advanced search weights title and abstract heavily.
1:45 Family
CVE-2021-3156†∗
Original primary CWE CWE-193
CVE-2018-5907
CWE-20
CVE-2017-5638§
CWE-20
CVE-2014-6271 CVE-2015-0235∗
CWE-78 CWE-119
CVE-2019-0708
CWE-416
CVE-20171000353 CVE-2018-20991
CWE-502
CVE-2021-21834∗
CWE-190
CVE-2016-7524
CWE-190
CVE-2017-6500
CVE-2014-0160∗
CWE-190 / CWE-119 CWE-787 / CWE-119 CWE-119
CVE-2017-5638§
CWE-20
CVE-2021-44228
CWE-502
CVE-2014-3566
CWE-310
CVE
Current primary CWE CWE-122 (was CWE-193) CWE-190
Failures
CWE-755 (was CWE-20) CWE-78 (unchanged) CWE-787
3
CWE-416 (unchanged) CWE-502 (unchanged) CWE-416 (unchanged) CWE-190 (CWE680 refined) CWE-190 (unchanged) Varies ¶
2
Varies ¶
2
CWE-125
3
CWE-755 (was CWE-20) CWE-917
3
CWE-310 (unchanged)
2
3 3
INP
MEM
CWE-416
DAT
CVE-2019-13297
FLR
2 3
2 2 3 2 2
3
Reassignment reason Deepened rootcause analysis Deepened rootcause analysis Deepened rootcause analysis Other or unknown Taxonomy refinement (new CWE) Deepened rootcause analysis Error in initial mapping Deepened rootcause analysis Taxonomy refinement (new CWE) Deepened rootcause analysis Deepened rootcause analysis Deepened rootcause analysis Deepened rootcause analysis Deepened rootcause analysis Deepened rootcause analysis Taxonomy refinement (new CWE)
Reading the table. Failures is the number of the four structural CWE failures the candidate exposes, shaded on a two-step scale, two and three. Within each family the selected CVE attains the maximum count, with anchor status as the tie-breaker. Reassignment reason records why the current primary CWE differs from the original, or states that it is unchanged. CWE and reassignment fields reflect the NVD record on May 27, 2026. ∗ Selected as the case study for its family. † Anchor candidate already referenced in this paper; preferred on the tie-breaker over the two other INP candidates that also reach three. § CVE-2017-5638 appears under both INP and FLR: its input-check root cause places it in INP, its failure-type consequence places it in FLR. It is kept in both family pools and is the selected example for neither. ¶ The current primary CWE varies by source for these two records, which share a root cause with CVE-2016-7518 and CVE-2019-13295 respectively. Table 9. Case-study candidate pool of 16 CVEs, four per BF bug-type family, with the four selected case studies marked.
1:46
Hoque et al.
Database Springer Link
Records 414
Semantic Scholar Google Scholar
171 41
arXiv IEEE Xplore
32 22
ACM Digital Library
11
Pooled, deduplicated
660
Table 10. Candidate records contributed per database at the Identification stage, before deduplication; counts overlap across databases.
Table 11. Verbatim per-database query protocols of the Identification stage. Google Scholar Query 1: CVE CWE classification automated mapping Query 2: CVE CWE prediction transformer BERT Query 3: CVE CWE large language model Query 4: NVD weakness enumeration deep learning Query 5: vulnerability classification CWE hierarchy Query 6: CVE description CWE label assignment Query 7: CWE-1003 classification Query 8: CAPEC ATT&CK CVE pipeline classification Query 9: vulnerability triage CWE neural Query 10: SecureBERT vulnerability classification Query 11: V2W-BERT CWE Query 12: weakness type prediction CVE Semantic Scholar Keyword query (2018-01 to 2026-04): "CVE-to-CWE classifier", "CWE label prediction", "automated weakness enumeration", "CWE hierarchy inference", "vulnerability description classification CWE". arXiv Query 1: abs:"CVE" AND abs:"CWE" AND (abs:"classification" OR abs:"mapping" OR abs:"prediction") Query 2: abs:"vulnerability classification" AND abs:"CWE" Query 3: abs:"NVD" AND abs:"weakness" AND (abs:"BERT" OR abs:"LLM" OR abs:"transformer" OR abs:"large language model") Query 4: ti:"CWE" AND (ti:"CVE" OR ti:"vulnerability") Query 5: abs:"CWE-1003" OR abs:"Top-25 CWE"
1:47 OR abs:"CWE hierarchy" Query 6: abs:"CVE description" AND abs:"CWE" AND (abs:"neural" OR abs:"deep" OR abs:"embedding") Query 7: abs:"SecureBERT" OR abs:"V2W-BERT" Query 8: abs:"CTI pipeline" AND abs:"CWE" IEEE Xplore ("All Metadata":"CVE" AND "All Metadata":"CWE" AND ("All Metadata":"classification" OR "All Metadata":"mapping" OR "All Metadata":"prediction" OR "All Metadata":"label")) AND ("All Metadata":"learn*" OR "All Metadata":"neural" OR "All Metadata":"BERT" OR "All Metadata":"transformer" OR "All Metadata":"large language model" OR "All Metadata":"LLM" OR "All Metadata":"embedding" OR "All Metadata":"retrieval") NOT ("Document Title":"survey" AND NOT "Document Title":"systematic") Second pass (title-restricted): ("Document Title":"CWE" AND "Document Title":"CVE") OR ("Document Title":"vulnerability" AND "Document Title":"CWE") OR ("Document Title":"weakness" AND "Abstract":"CVE") ACM Digital Library (Abstract:(CVE AND CWE) AND Abstract:(classification OR mapping OR prediction OR assignment)) AND (Abstract:(learning OR neural OR transformer OR BERT OR LLM OR "large language model" OR embedding OR retrieval)) Second pass (title-focused): Title:(CWE) AND (Title:(CVE) OR Title:(vulnerability) OR Title:(weakness)) Springer Link with all of the words: CVE CWE with at least one of the words: classification mapping prediction transformer BERT
1:48 LLM neural embedding retrieval without the words: malware-detection source-code-vulnerability-detection Second pass: title contains "CWE" OR "weakness enumeration".
Hoque et al.