ConceptioArchivearXiv CS
arXiv CSopen access

A Structured Cyber Threat Intelligence Dataset Using STIX 2.1 Entities and MITRE ATT&CK Mappings

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

A Structured Cyber Threat Intelligence Dataset Using STIX 2.1 Entities and MITRE ATT&CK Mappings

arXiv:2607.23312v1 [cs.CR] 25 Jul 2026

Dipshikha Das, Arnab Banik, Md. Shariful Islam, and Md Rayhanur Rahman Abstract—Cyber threat intelligence (CTI) reports are typically written in unstructured formats, which complicates the extraction and analysis of important entities and adversarial behaviors. Although existing CTI research provides extraction tools, knowledge-graph frameworks, and MITRE ATT&CKmapped datasets, curated report-level datasets that preserve complex entity relationships and normalized adversarial behaviors remain limited. To address this limitation, this study presents a manually constructed dataset of 150 English-language CTI reports each represented as STIX 2.1 based graphs which includes 4,777 STIX entities, 5,817 STIX relationships in total, and 1,273 STIX attack-pattern entities (adversarial-behaviors) mapped to 269 unique MITRE ATT&CK Enterprise techniques and sub-techniques. Twenty five randomly sampled reports were independently assessed by two cybersecurity researchers, which shows substantial inter-rater agreement. Disagreements were subsequently adjudicated to establish a gold-standard reference dataset. Four locally deployed open-source LLMs were evaluated as automated judges against this adjudicated reference sample. Qwen3.6:27B achieved the strongest overall performance, with a maximum kappa score of 0.803, micro-F1 scores exceeding 92%, and false-positive rates below 5%. The dataset provides a benchmark for CTI information extraction, knowledge-graph construction, incident analysis, and threat attribution. The findings further indicate that locally deployed LLMs can support human reviewers in identifying annotation inconsistencies but expert validation remains essential. Index Terms—Cyber Threat Intelligence, Cyber Security, CTI report, STIX, MITRE ATT&CK, Structured CTI Data, CTI Dataset, ATT&CK Tactics and Techniques, LLM-as-a Judge

I. I NTRODUCTION Cyber threat intelligence (CTI) reports provide critical information about attackers, adversarial behaviors and malicious operations [1]. The reports are published by different vendors and are generally written in natural language in unstructured formats. Differences in terminology, writing style, structure, and formatting makes it difficult to analyze and extract actionable intelligence easily [2]. Structured representations can address this problem by organizing CTI information in a standardized format. Structured Threat Information Expression (STIX) provides a standardized vocabulary for representing threat entities, cyber observables, and semantic relationships [3]. On the other hand, MITRE ATT&CK is a widely used Dipshikha Das, Arnab Banik, and Md. Shariful Islam are with the Institute of Information Technology, University of Dhaka, Dhaka, Bangladesh. E-mail: {bsse1218, bsse1230, shariful}@iit.du.ac.bd. Md Rayhanur Rahman is with the Department of Computer Science, The University of Alabama, Tuscaloosa, AL 35487, USA. E-mail: [email protected].

knowledge base that organizes adversarial behaviors into TTPs - Tactics (attackers’ intentions), Techniques (methods of attack), sub-techniques (highly specific methods of attack) and lastly Procedure (Specific or step by step implementation) [4]. Combining these frameworks provides a consistent format for representing CTI information, which makes information extraction easier, and analysis time shorter. Previous studies have developed methods for extracting structured threat information from CTI reports. These methods individually identify attack actions or entities and relationships, or MITRE ATT&CK techniques. Some approaches extract STIX objects and their relationships [5], [6], while others construct or enrich cybersecurity knowledge graphs [2], [7]–[9]. Many existing studies also primarily provide automated extraction tools, custom graph representations, textspan annotations, or statement-level ATT&CK labels [10]– [17]. These representations and tools may not preserve the broader relationships among the entities. Another challenge is that different vendors use their own terminologies for referring to the same adversarial behaviors. As a result, there remains limited availability of curated datasets that represent each CTI report individually through standardized entities, relationships between these entities, and representing adversarial behavior using a normalized format. This work addresses these gaps by introducing a manually constructed and reviewed dataset of 150 English-language CTI reports, each represented as a directed, heterogeneous graph containing STIX 2.1 Domain Objects (SDO), Cyberobservables (SCO), and STIX Relationship Objects (SRO). SDOs represent a cyber threat concept or entity, SCOs represent observable technical data [3]. These SDOs and SCOs are represented as nodes. SROs define the relationships among these STIX objects/entities [3], and these are represented as edges in the graph structure. The dataset contains 4,777 nodes and 5,817 relationship edges. Among the SDOs, attackpattern refers to the adversarial behaviors that attackers use to achieve their goal. To address the earlier mentioned problem of differing terminologies for the same attack behavior, each attack-pattern node is manually mapped to MITRE ATT&CK enterprise techniques and sub-techniques under specific tactics. This produces 1,273 mappings covering 269 unique ATT&CK identifiers across all 15 Enterprise ATT&CK tactics. Existing datasets present entities, relationships, or technique labels separately, but this proposed dataset shows how these entities are connected and represents the adversarial behaviors described

CTI Report Collection 150 Public CTI Reports (2015 - 2026)

Entity and Relation Extraction

Manual ATT&CK Mapping Attack Patterns SDO -> ATT&CK Technique and Sub techniques

STIX 2.1 SDO STIX 2.1 SCO

Report Level Graph Dataset

Human Annotation

Gold Dataset Construction

LLM-as-a Judge Annotation Evaluation

15 Enterprise Tactic

STIX 2.1 SRO

150 nodes files

150 edges files

Annotator 1

Annotator 2

Fig. 1. Overall Dataset Construction and Annotation Workflow

in each report in a consistent format. Two cybersecurity researchers independently annotated a randomly sampled subset of 25 reports. A gold annotation set was constructed by resolving the disagreements between the human annotators. Four local LLMs annotated the sample dataset and the results of these LLM’s were compared against the gold annotation reference set. The main contributions of this work include: • A curated dataset that represents CTI reports as STIX entities and relationships. • Manual mapping of adversarial behaviors (STIX attack pattern SDO) to MITRE ATT&CK Enterprise techniques and sub-techniques. • Human annotation of the sample dataset with substantial inter-rater agreement between the annotators [18], achieving quadratic weighted Cohen’s κw scores of 0.64, 0.67 and 0.62 for Entity Correctness, ATT&CK Mapping, and Relationship Correctness respectively. • A systematic assessment of local open-source LLMs for automated dataset annotation by following the LLM-asa-judge [19] paradigm. II. R ELATED W ORK Previous CTI research has mainly focused on extracting threat information, constructing knowledge graphs, and mapping attack descriptions to MITRE ATT&CK TTPs. Several studies have worked on cybersecurity knowledge graph construction. AttacKG [2] constructs attack-behavior graphs and identifies associated ATT&CK techniques. Its manual evaluation included only 16 reports, which may not fully represent different CTI writing styles and threat domains. The manually reviewed dataset proposed in this work can support broader evaluation of similar extraction and mapping approaches. ThreatKG [7] and TINKER [8] extract and integrate threat information from multiple sources to construct cybersecurity knowledge graphs. CTINEXUS [9] applies incontext learning to construct knowledge graphs from recent CTI reports and can adapt its output to the STIX framework. These systems focus on automated extraction and knowledge aggregation rather than producing a manually reviewed STIX and ATT&CK-based representation of CTI reports. Other studies have focused specifically on STIX-based information extraction. STIXnet [5] extracts STIX entities and

relationships from unstructured CTI reports. The semanticchunking approach in [10] employs consensus filtering to extract CTI triples, whereas the validation-and-repair framework in [6] identifies and repairs invalid STIX 2.1 triples. Researchers also examined the extraction and annotation of ATT&CK TTPs. EXTRACTOR [11] extracts attack behaviors, while LADDER [12] extracts attack patterns from unstructured reports and maps them to the ATT&CK framework. CTI-HAL [13] provides human-annotated ATT&CK mappings for 81 CTI reports and reports substantial inter-annotator agreement [18]. It mainly focuses on statement-level TTP annotations rather than complete STIX-based report representations. AnnoCTR [14] comprises of 400 CTI reports annotated with general entities, including a subset of 120 reports additionally annotated with cybersecurity entities and ATT&CK concepts. Its annotations similarly focus on individual entities and concepts rather than connected STIX entities and relationships. TRAM [15], TTPHunter [16], and TTPXHunter [17] identify ATT&CK techniques from report text but do not extract the entities and relationships needed to capture the full report context. By combining a graph-structured STIX 2.1 representation with standardized ATT&CK mappings, this work provides a foundational ground-truth dataset that preserves entity relationships, adversarial behaviors, and the broader contextual information described in modern CTI reports. III. M ETHODOLOGY AND DATASET C ONSTRUCTION Figure 1 illustrates the overall dataset construction workflow, with human annotation and LLM annotation evaluation. A. Data Collection The dataset consists of English-language cyber threat intelligence reports collected mainly from references available in the ATT&CK knowledge base [4]. Additional reports were gathered from public sources, including The DFIR Report [20], BleepingComputer [21], and other CTI publishing sites. Only reports containing sufficient contextual and technical information were included in the dataset. B. Entity and Relationship Extraction Each report is converted into a structured STIX 2.1 graph representation [3]. The dataset contains 13 SDOs and 16

SCOs. The extracted SDOs include attack patterns, campaigns, courses of action, identities, indicators, infrastructures, intrusion sets, locations, malware, observed data, threat actors, tools, and vulnerabilities. The extracted SCOs include artifacts, autonomous systems, directories, domain names, email addresses, email messages, files, IPv4 addresses, MAC addresses, network traffic, processes, software, URLs, user accounts, Windows Registry keys, and X.509 certificates. Relationships between entities are represented using SROs. The dataset supports SDO–SDO, SDO–SCO, and SCO–SDO relationships and follows standard STIX 2.1-defined SROs [3]. Entities and relationships were extracted only when the source report explicitly supported them. C. Mapping Attack Patterns to MITRE ATT&CK Techniques Adversarial behaviors described in natural language were identified as attack patterns SDO in the dataset. These attack patterns were manually mapped to ATT&CK version 19.1 enterprise techniques and sub-techniques within the defined tactics [4]. Manual mapping was used to ensure accurate and context-aware ATT&CK technique and sub-technique assignments. As the same adversarial behaviors are described differently across CTI reports, comparison and analysis of these behaviors would be difficult without proper ATT&CK mappings. For each behavior, first the corresponding ATT&CK tactic was identified, followed by identifying the most specific and report-supported technique or sub-technique. Attack behavior was mapped to most suitable ATT&CK sub-technique when sufficient evidence was available. Otherwise, it was mapped to the corresponding parent technique. The ATT&CK Enterprise matrix comprises of 15 tactics, 222 techniques, and 475 sub-techniques [4]. Consequently, manual ATT&CK mapping was the most challenging task, as it required each observed adversarial behavior to be mapped to the most suitable technique or sub-technique supported by evidence in the source report.

TABLE I E XTRACTED E NTITIES IN THE N ODE F ILE ID 01

Type Threat actor

Instance APT39

02

Location

Iran

03

Attack pattern

04

Malware

Phishing: Spearphishing Attachment (T1566.001) Powbat

05

Tool

Description An Iranian state-sponsored cyber-espionage group associated with Chafer that conducts personal-information theft. APT39 is assessed with moderate confidence to originate from Iran and operate in support of Iranian national interests. Tactic: Initial Access (TA0001). APT39 uses spearphishing emails containing malicious attachments.

A custom backdoor variant used by APT39 to establish an initial foothold. Windows Credential A legitimate credential-management tool Editor abused by APT39 for credential harvesting and privilege escalation.

TABLE II E XTRACTED R ELATIONSHIPS IN THE E DGE F ILE Source ID 01 01 03 01 01

Target ID 02 03 04 04 05

Relationship located-at uses delivers uses uses

The graph representation of the given example can be seen in Figure 2, where instance and type represents the nodes and relationship represents the edges.

Iran (Location)

located at

uses APT 39(Threat Actor)

Powbat (Malware)

uses uses

delivers T1566.001 (Attack Pattern)

Windows Credential Editor(Tool)

Fig. 2. Graph structure of the given example

D. Example Data Each CTI report consists of two CSV files: a nodes file and an edges file. Each nodes file contains four attributes: a unique report-level identifier (id), the corresponding STIX object category (type), the specific entity name (instance), and contextual details derived directly from the source document (description). The edges file contains sourceId, targetId, and relationship. Each row represents a directed contextual relationship between two entities in the corresponding nodes file. Together, the two files form a STIXinspired graph representation of each report. The CTI report in [22] describes the cyber-espionage activities of the Iranian threat actor APT39. It initially compromises target systems through spear phishing emails with attachments that deliver a custom variant of POWBAT malware. This group also uses Windows Credential Editor tool during the attack. The extracted entities and relationships from this content are shown in Tables I and II respectively.

E. Human Annotation and Inter-Rater Reliability 25 reports (approximately 17% of the dataset) and their corresponding nodes and edges were randomly sampled to establish a human-annotated reference set. The sample set contains 566 nodes, 677 edges, and 147 attack-pattern nodes mapped to Enterprise ATT&CK techniques and sub-techniques. A graduate level and a post-graduate level cybersecurity researcher independently annotated each node and edge. The annotators followed the rubrics presented in Table III. Each node and edge was annotated on a 3-point ordinal scale. The supported label indicates that the existing node/edge is fully consistent with the source report and their definitions. The partially_supported label indicates that the node/edge that has been extracted from the report is not fully supported by its definitions. The unsupported label indicates that the node/edge is not substantiated by the report according to its definitions.

TABLE III H UMAN -A NNOTATION RUBRICS Rubrics A: Entity type correctness B: Attack-mapping correctness C: Relationship correctness

TABLE VI S UMMARY OF THE P ROPOSED CTI DATASET

Evaluation Question The instance and its description align with the assigned STIX SDO/SCO types and are supported by the report. The assigned ATT&CK technique or sub-technique aligns with its official MITRE ATT&CK definition. The relationship between the source and target is supported by the report and aligns with STIX SRO definitions.

The quadratic weighted Cohen’s Kappa Score κw is 0.64 for Type Correctness, 0.67 for ATT&CK Mapping Correctness, and 0.62 for Relationship Correctness. All 3 Rubrics indicate a substantial agreement between the annotators [18]. Table IV presents the Agreement Matrix where S., P.S. and U. refers to Supported, Partially Supported and Unsupported respectively: TABLE IV H UMAN VS H UMAN AGREEMENT M ATRIX Type ATT&CK Edge S. P.S U. S. P.S U. S. P.S U. S. 492 6 12 113 3 3 511 41 3 P.S. 3 24 3 3 13 2 33 13 30 U. 8 3 15 3 1 6 5 30 11

TABLE V S AMPLED DATASET C ORRECTION S UMMARY Type 566 21 53 13.1 525 20 0

ATT&CK 147 9 25 23.1 132 6 0

Value 150 2015–2026 4,777 & 5,817 32 & 39 269 1,273

most. The least frequent SRO types are uploaded-on, derivedfrom, characterizes, disables, and identifies, each with one occurrence. Table VII presents top 5 most frequent entities and relationships in the dataset. TABLE VII T OP F IVE M OST F REQUENT SDO, SCO, AND SRO T YPES SDO Type Freq. SCO Type Freq. SRO Type Freq. attack-pattern 1,272 file 550 uses 2,697 identity 443 domain-name 318 targets 748 malware 354 ipv4-addr 172 indicates 280 location 293 url 145 mitigates 275 tool 230 software 47 based-on 263

C. MITRE ATT&CK Coverage

Disagreements were discussed and resolved through mutual agreement between the annotators. Feedback from the annotators was used to revise and improve the annotated samples. The summary of changes is presented in Table V.

Old Items Count Total Dropped Items Total Updated Items % of Items Changed Gold Supported Count Gold Partially Supported Count Gold Unsupported Count

Characteristic CTI reports Publication period Total Valid entities & relationships Average nodes & relationships per Report Unique ATT&CK IDs Total ATT&CK occurrences

Edge 677 71 95 24.5 598 8 0

IV. DATASET CHARACTERISTICS A. Dataset Scale and Coverage

The dataset contains 269 unique MITRE ATT&CK technique and sub-technique identifiers, comprising of 123 parent techniques and 146 sub-techniques. In total, the dataset contains 749 parent techniques and 524 sub-techniques. The five most frequently mapped attack patterns are Ingress Tool Transfer (T1105), Data Encrypted for Impact (T1486), Obfuscated Files or Information (T1027), Phishing (T1566), and Exfiltration Over C2 Channel (T1041). The mappings span all 15 ATT&CK tactics. The tactic Stealth contains the highest number of mapped rows and unique techniques, followed by Initial Access, Command and Control, and Execution. Overall, the dataset offers broad structural, and behavioral coverage of each CTI report, which can support diverse CTI analysis, information extraction, automated evaluation tasks. V. A SSESSING LLM- AS - A J UDGE FOR DATASET A NNOTATION A. LLM-as-a-Judge Framework

The proposed dataset contains 150 English language CTI reports published between 2015 and 2026. It includes 4,777 entities. It also has 5,817 best-suited directed relationships. The summary of the dataset is shown in Table VI.

Independently re-annotating the complete dataset by multiple cybersecurity experts is impractical because of its cost and time consumption [23]. Therefore, LLMs are employed as automated annotators to provide a scalable, and complementary solution by following the LLM-as-a-judge paradigm [19].

B. Entity And Relationship Distribution

B. Models Used

The dataset contains 3,446 SDOs and 1,331 SCOs. Attackpattern and file are the most frequent SDO and SCO respectively. In contrast, vulnerability is the least frequent SDO with 50 occurrences, while mac-addr is the least frequent SCO, occurring once. Among the relationships, uses occurs the

4 reasoning models–qwen3.6:27b, qwen3:14b, gpt-oss:20b, and gemma4:31b– were run locally through Ollama with provider set default parameters. All 4 models were using Q4_K_M quantization, and 4-bit KV-cache with a maximum prompt size of 32,768 tokens.

The models were run on an external Nvidia RTX 3090 graphics card with 24GB VRam. Local LLMs were selected over commercial APIs to avoid provider-side model changes [24], and reduce cost on repeated runs as the dataset expands. C. Prompting The LLMs received the same rubrics defined in section III as system prompts, and annotated on the same Ordinal Scale as the human annotators. The LLMs were also provided with official definitions as reference text. Chain-ofThought(CoT) reasoning [25] was used for all four models. Each model received batches of items (a maximum of 5 items for qwen3:14b and gpt-oss:20b, and 8 items for qwen3.6:27b and gemma4:31b), and each batch was a standalone LLM call with only the references of the items included in the batch. The output was a json array of the following schema: { reasoning: string, unsupported_claims: [string], confidence: enum(low, medium, high), verdict: enum(supported, partially_supported, unsupported) } D. LLM Assessment Results All four LLMs evaluated the pre and post corrected versions of the 25 samples described in Section III-E. Table VIII presents the pre-corrected sample results. The performance is measured using micro-F1, macro-precision, macro-recall, macro-F1, and quadratic-weighted Cohen’s Kappa scores against the gold reference set (including unsupported class). Micro-F1 measures overall classification performance across all instances, while the macro metrics give equal importance to all three classes. TABLE VIII LLM A SSESSMENT R ESULTS FOR P RE -C ORRECTED S AMPLES Model Micro-F1 Macro-P Macro-R Macro-F1 κw qwen3:14b 0.913 0.416 0.468 0.437 0.414 gpt-oss:20b 0.924 0.463 0.403 0.418 0.258 Type qwen3.6:27b 0.949 0.611 0.550 0.575 0.675 gemma4:31b 0.920 0.439 0.471 0.454 0.421 qwen3:14b 0.918 0.729 0.707 0.711 0.661 gpt-oss:20b 0.891 0.502 0.538 0.518 0.662 ATT&CK qwen3.6:27b 0.939 0.709 0.765 0.719 0.803 gemma4:31b 0.918 0.603 0.636 0.600 0.703 qwen3:14b 0.790 0.470 0.592 0.475 0.452 gpt-oss:20b 0.809 0.477 0.583 0.497 0.467 Edge qwen3.6:27b 0.838 0.497 0.597 0.469 0.346 gemma4:31b 0.799 0.427 0.523 0.436 0.248 Rubric

Table IX presents the post-corrected sample results. As the post-correction reference labels do not contain the unsupported class, Kappa and macro-averaged metrics are not applicable. Therefore, the performance is reported using micro-F1 and the false-positive rate (FPR) for the unsupported class. E. Analysis of Results In the pre-corrected assessment, qwen3.6:27b achieves the highest macro-F1 scores for all 3 rubrics. It also achieves

TABLE IX P OST-C ORRECTION LLM A SSESSMENT R ESULTS Rubric

Model qwen3:14b gpt-oss:20b Type qwen3.6:27b gemma4:31b qwen3:14b gpt-oss:20b ATT&CK qwen3.6:27b gemma4:31b qwen3:14b gpt-oss Edge qwen3.6:27b gemma4:31b

Micro-F1 0.941 0.934 0.958 0.949 0.928 0.935 0.928 0.906 0.799 0.804 0.929 0.880

Unsupported FPR 0.029 0.031 0.006 0.013 0.036 0.043 0.029 0.072 0.127 0.155 0.031 0.089

the highest κw scores of 0.675, and 0.803 on Type, and ATT&CK rubrics which indicate substantial and almost perfect agreement with the gold standard annotation respectively [18]. For Edges rubric model performance varies across the evaluation metrics, qwen3.6:27b achieves the highest micro-F1, but gpt-oss:20b achieves the highest macro-F1, and κw scores. All Micro-F1 scores signal that the models’ accuracy is high for our single label multiclass gold reference set, and the lower macro average score in comparison properly reflects the class imbalance in the gold set. The lower κw scores on the Edges reflect the reasoning limitations of small to medium sized local LLMs, where the the LLMs try to find verbatim evidence for relationships from a given list of accepted sourceId, targetId, and relationship verb as reference. In the post-corrected results, qwen3.6:27b again achieves the highest micro-F1 , and the lowest false positive rates for Type and ATT&CK rubrics respectively. For Edges, gpt-oss scores the highest micro-F1, and qwen3.6 scores the lowest false positive rate. All models scored above 93% in micro-F1, and qwen3.6 scored below 5% in false positive rates. These results show that the model rarely misclassifies the corrected annotations as unsupported. VI. D ISCUSSION AND L IMITATIONS The proposed dataset represents each CTI report in a structured format that represents entities through STIX 2.1 objects and normalizes adversarial behaviors through ATT&CK mappings. The substantial inter-rater agreement across all three rubrics confirms that the resulting dataset is a reliable and consistent representation of the source reports [18]. The correction rates indicate that ATT&CK mapping and relationship extraction are more challenging than entity extraction. These tasks require substantial contextual interpretation because adversarial behaviors and relationships are often described implicitly or using vendor-specific terminology. The LLM annotation evaluation results reflect similar challenges. High micro-F1 but lower macro-F1 and κw scores indicate class imbalance and weaker performance on minority labels. Qwen3.6:27b achieved the strongest overall performance, while relationship assessment remained comparatively

difficult. So local LLMs may assist human reviewers in identifying inconsistencies but cannot replace expert validation. The complete dataset, source code, evaluation prompts and other materials will be made publicly available upon acceptance. The current dataset presents the following constraints: it is limited to 150 public, English language reports, and it does not model temporal attack sequences described in the CTI reports. Direct human assessment covers approximately 17% of the dataset. This subset was rigorously corrected to ensure the reference set contains zero unsupported instances. VII. C ONCLUSION AND F UTURE W ORK This study introduced a manually constructed dataset of 150 CTI reports each represented as directed heterogeneous STIX 2.1 graph based structure. The dataset contains 4,777 entities, 5,817 relationships, and 1,273 adversarial behavior (attack pattern SDO) occurrences mapped to 269 unique MITRE ATT&CK techniques and sub-techniques across all 15 Enterprise tactics. Human assessment of randomly sampled 25 reports indicates substantial inter-rater agreement across the 3 annotation rubrics. Evaluation against the adjudicated gold reference set further showed that locally deployed open-source LLMs, particularly Qwen3.6, can support automated scalable annotation, although expert validation remains necessary. A dataset of this kind with structured STIX representation is valuable for consistently representing complex threat entities, their relationships, and adversarial behaviors described using different terminology across CTI reports. Its standardized and context-preserving representation can support the development and evaluation of CTI report related automated tasks. Future work may incorporate multilingual and diverse sources, and temporal relationships described in the reports. Additional LLMs and prompting strategies may also be evaluated, together with downstream tasks such as information extraction, incident analysis, CTI report clustering, threat attribution, and knowledge-graph construction. R EFERENCES [1] BlueVoyant, “Cyber threat intelligence: Definition, types, and process.” [Online]. Available: https://www.bluevoyant.com/knowledge-center/cy ber-threat-intelligence-cti-definition-types-process. Accessed: Jul. 4, 2026. [2] Z. Li, J. Zeng, Y. Chen, and Z. Liang, “AttacKG: Constructing technique knowledge graphs from cyber threat intelligence reports,” in Proc. Eur. Symp. Res. Comput. Secur. (ESORICS), 2022, pp. 589–609. [3] OASIS Cyber Threat Intelligence Technical Committee, “Introduction to STIX.” [Online]. Available: https://oasis-open.github.io/cti-documen tation/stix/intro.html. Accessed: Jul. 4, 2026. [4] The MITRE Corporation, “MITRE ATT&CK.” [Online]. Available: ht tps://attack.mitre.org/. Accessed: Jul. 8, 2026. [5] F. Marchiori, M. Conti, and N. V. Verde, “STIXnet: A novel and modular solution for extracting all STIX objects in CTI reports,” in Proc. 18th Int. Conf. Availability, Rel. Secur. (ARES), 2023, pp. 1–11. [6] V. Andrejeus, W. Mitchell, S. Cawthon, J. Sullins, E. Dogdu, and R. Choupani, “Automated validation and repair of knowledge graph triples for cyber threat intelligence,” in Proc. IEEE 5th Int. Conf. AI Cybersecurity (ICAIC), 2026. [7] P. Gao, X. Liu, E. Choi, S. Ma, X. Yang, and D. Song, “ThreatKG: An AI-powered system for automated open-source cyber threat intelligence gathering and management,” in Proc. 1st ACM Workshop Large AI Syst. Models Privacy Saf. Anal., 2023, pp. 1–12.

[8] N. Rastogi, S. Dutta, A. Gittens, M. J. Zaki, and C. Aggarwal, “TINKER: A framework for open-source cyberthreat intelligence,” in Proc. IEEE Int. Conf. Trust, Secur. Privacy Comput. Commun. (TrustCom), 2022, pp. 1569–1574. [9] Y. Cheng, O. Bajaber, S. A. Tsegai, D. Song, and P. Gao, “CTINEXUS: Automatic cyber threat intelligence knowledge graph construction using large language models,” in Proc. IEEE Eur. Symp. Secur. Privacy (EuroS&P), 2025, pp. 923–938. [10] S. Waldrop, E. Dogdu, R. Choupani, W. Mitchell, V. Andrejeus, and S. Cawthon, “Semantic chunking and consensus filtering for structured extraction of cyber threat intelligence,” in Proc. IEEE 5th Int. Conf. AI Cybersecurity (ICAIC), 2026. [11] K. Satvat, R. Gjomemo, and V. N. Venkatakrishnan, “Extractor: Extracting attack behavior from threat reports,” in Proc. IEEE Eur. Symp. Secur. Privacy (EuroS&P), 2021, pp. 598–615. [12] M. T. Alam, D. Bhusal, Y. Park, and N. Rastogi, “Looking beyond IoCs: Automatically extracting attack patterns from external CTI,” in Proc. 26th Int. Symp. Res. Attacks, Intrusions Defenses (RAID), 2023, pp. 92–108. [13] S. Della Penna, R. Natella, V. Orbinato, L. Parracino, and L. Pianese, “CTI-HAL: A human-annotated dataset for cyber threat intelligence analysis,” in Proc. IEEE Eur. Symp. Secur. Privacy Workshops (EuroS&PW), 2025, pp. 69–78. [14] L. Lange, M. Müller, G. H. Torbati, et al., “AnnoCTR: A dataset for detecting and linking entities, tactics, and techniques in cyber threat reports,” in Proc. LREC-COLING, 2024, pp. 1147–1160. [15] Center for Threat-Informed Defense, “Threat Report ATT&CK Mapper (TRAM),” MITRE, Aug. 24, 2023. [Online]. Available: https://ctid.mit re.org/projects/threat-report-attck-mapper-tram/. Accessed: Jul. 4, 2026. [16] N. Rani, B. Saha, V. Maurya, and S. K. Shukla, “TTPHunter: Automated extraction of actionable intelligence as TTPs from narrative threat reports,” in Proc. Australasian Comput. Sci. Week (ACSW), 2023, pp. 126–134. [17] N. Rani, B. Saha, V. Maurya, and S. K. Shukla, “TTPXHunter: Actionable threat intelligence extraction as TTPs from finished cyber threat reports,” Digit. Threats Res. Pract., vol. 5, no. 4, pp. 1–19, 2024. [18] J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,” Biometrics, vol. 33, no. 1, pp. 159–174, Mar. 1977, doi: 10.2307/2529310. [19] L. Zheng et al., “Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, 2023, pp. 46595–46623. [20] The DFIR Report, “Actionable cyber threat intelligence.” [Online]. Available: https://thedfirreport.com/. Accessed: Jul. 4, 2026. [21] BleepingComputer, “Security news.” [Online]. Available: https://www. bleepingcomputer.com/news/security/. Accessed: Jul. 4, 2026. [22] Google Cloud, “APT39: An Iranian Cyber Espionage Group Focused on Personal Information. [Online]. Available: https://cloud.google.com /blog/topics/threat-intelligence/apt39-iranian-cyber-espionage-group-f ocused-on-personal-information/. Accessed: Jul. 15, 2026. [23] Y. Yang, O. Agarwal, C. Tar, B. C. Wallace, and A. Nenkova, “Predicting annotation difficulty to improve task routing and model performance for biomedical information extraction,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Lang. Technol. (NAACLHLT), 2019, pp. 1471–1480. [24] L. Chen, M. Zaharia, and J. Zou, “How Is ChatGPT’s Behavior Changing Over Time?” Harvard Data Science Review, vol. 6, no. 2, 2024, doi: 10.1162/99608f92.5317da47. [25] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Proc. 36th Int. Conf. Neural Inf. Process. Syst. (NeurIPS), 2022, pp. 24824–24837.

Record · ID 405578 · SHA-256 e9cc630e6a19829d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.