ConceptioArchivearXiv CS
arXiv CSopen access

Topical Shifts in the Dark Web: A Longitudinal Analysis of Content from the Cybercrime Ecosystem

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2605.15345v1 [cs.CR] 14 May 2026

Topical Shifts in the Dark Web: A Longitudinal Analysis of Content from the Cybercrime Ecosystem Roy Ricaldi

Maximilian Schäfer

Philipp Zech

Department of Computer Science Eindhoven University of Technology Eindhoven, Netherlands [email protected]

Department of Information Systems University of Liechtenstein Vaduz, Liechtenstein [email protected]

Department of Computer Science University of Innsbruck Innsbruck, Austria [email protected]

Luca Allodi

Raffaela Groner

Irdin Pekaric

Department of Computer Science Eindhoven University of Technology Eindhoven, Netherlands [email protected]

Department of Computer Science Chalmers University of Technology Gothenburg, Sweden [email protected]

Department of Information Systems University of Liechtenstein Vaduz, Liechtenstein [email protected]

Abstract—The dark web hosts a dynamic ecosystem of cybercrime forums and marketplaces that adapt to law enforcement pressure, technological change, and economic incentives. Prior research has extracted cyber threat intelligence from these platforms using static snapshots, with limited attention to how discussions evolve over time. In this study, we conduct a longitudinal analysis of 25,065 websites in the dark web using 11,403,638 HTML snapshots (approximately 1245.38 GB) collected over six years. We develop a longitudinal topic-modeling framework combining domain-specific embeddings, density-based clustering and temporal aggregation to measure topic prevalence and lifecycle at the website level. Our analysis identifies 55 thematic clusters. We find that ≈ 75% of total discussion volume is concentrated in a small set of persistent core topics, while short-lived themes account for ≈ 3% of activity. The median topic lifespan is 75 months, indicating gradual thematic evolution rather than abrupt replacement. Index Terms—Dark Web Analysis, Cybercrime Ecosystems, Longitudinal Topic Analysis, Cyber Threat Intelligence, Natural Language Processing

1. Introduction The dark web constitutes a hidden segment of the internet that enables anonymous communication and hosting of services through specialized routing protocols such as Tor. While it provides privacy protections for journalists, activists, and whistleblowers, it also serves as a major infrastructure for cybercrime. Prior measurements estimate that 57% of dark web content is associated with illicit activity [1], and Tor Metrics data indicate that daily users have doubled since 2022, reaching a peak of 7 million users [2]. This growth highlights the continued importance of the dark web as a space where anonymity, financial motives, and technology intersect. Despite more than two decades of existence, the dark web remains only partially understood, particularly with

regard to how cybercrime-related content evolves over time [3, 4]. Research has largely approached dark web platforms through Cyber Threat Intelligence (CTI), focusing on extracting actionable signals such as indicators of compromise, exploit discussions, or marketplace listings. Early work relied on dictionary-based or fuzzing techniques, often suffering from limited recall or high false positives [5, 6]. More recent studies apply machine learning and natural language processing to improve detection and classification of threat-relevant information [7]. However, most existing studies treat dark web data as static snapshots [8, 9] or focus on a limited set of marketplaces rather than the broader cybercrime ecosystem [10, 11]. While useful for operational intelligence, these approaches provide limited insight into how cybercrimerelated discussions evolve. This is a fundamental limitation in complex threat environments, where security analysis must move beyond static observations and adjust to a continuously evolving attacker ecosystem [12, 13]. This demonstrates the need for similarly changing perspectives in CTI. Dark web platforms are highly dynamic: platforms shut down, rebrand, or migrate; forums fragment or consolidate; and marketplaces adapt to law enforcement pressure, technological developments, and user demand. Yet, how topics and themese discussed and of relevance in these communities and websites emerge, persist, and disappear over time has not been systematically studied longitudinally. Without this perspective, it is difficult to distinguish stable core activities from transient trends. This study addresses this gap by reporting on a longitudinal analysis of dark web forums and marketplaces using automated methods for natural language processing and ML. Relying on repeated HTML snapshots spanning from 2020 to 2026, we analyze how (criminal) content evolves over time, capturing how topics emerge, persist, and decline. C ONTRIBUTIONS . We make the following contributions: We propose a longitudinal measurement framework for analyzing forum and marketplace websites in the dark

web using HTML snapshots. We develop a topic1 discovery pipeline that integrates DARK-BERT embeddings, density-based clustering, and probabilistic topic assignment to model the temporal dynamics of discussion themes. • Using a dataset of over 11 million snapshots from more than 25 thousand dark web websites collected over six years, we provide the first large-scale longitudinal analysis of topic prevalence and turnover in economically motivated cybercrime communities.

signals have motivated the application of machine learning and natural language processing techniques to analyze dark web content.

2.2. Analysis Techniques for the Dark Web The dark web is widely used for illicit activities and has therefore become an important source of CTI [32]. Active measurement approaches deploy controlled systems to observe offender interactions and behavior [33, 34], yet most research focuses on extracting actionable insights from forums and marketplaces through computational content analysis, i.e., deriving conclusions from textual or multimedia data [35]. Within cybersecurity, this approach is commonly applied to mine CTI from unstructured sources[36]. A lot of work treats dark web content as an intelligence source, focusing on extracting structured information such as entities, attack indicators, or threat reports [28, 37, 38]. Reviews further highlight automated extraction of indicators of compromise and behavior patterns [39]. They primarily treat platforms as intelligence feeds rather than evolving socio-technical environments. Another research direction applies NLP and machine learning to detect or predict cyber threats. However, despite the growing use of AI-driven approaches in CTI, their effective adoption in practice remains limited. Recent empirical evidence highlights that challenges such as lack of integration into security workflows, limited trust by analysts, and insufficient robustness and monitoring mechanisms hinder trustworthy deployment [40]. Prior work includes identifying emerging threats in marketplace listings [41], predicting exploit occurrences from forum discussions [42], inferring attacker intent using deep learning [43, 44], and detecting cyberattacks using sentiment signals [45]. Topic modeling is another widely used technique, identifying latent semantic themes within large corpora [46]. Neural approaches such as BERTopic leverage contextual embeddings and clustering to produce interpretable topics [47], and have been applied to uncover discussion themes in dark web data. Finally, prior work also focuses on categorization tasks, including classifying marketplace products and forum posts [48, 49], categorizing onion services [30, 50, 51], and comparing models for detecting illicit activity [52, 53]. Toolkits have also been proposed to identify cybercriminal communities [54], while sentiment and pattern analysis methods are used to characterize content [55].

We release our resources to support software sustainability: https://github.com/irdin-pekaric/WACCO2026.

2. Background and Related Work We provide background information on cybercrime ecosystems and the dark web (§2.1). In addition, we discuss related work on methods that analyze websites on the dark web (§2.2) and pinpoint the research gap (§2.3).

2.1. Cybercrime in the Dark Web For this study, cybercrime refers to illicit activities conducted through digital systems and primarily motivated by economic gain [14, 15]. This includes unauthorized access, misuse, or exchange of digital assets, while politically or ideologically motivated actions are excluded. The primary infrastructure for these activities is located on the dark web via anonymity networks such as Tor and I2P [16]. The ecosystem is structured around two main platform types: marketplaces and forums. Marketplaces enable commoditized exchange of illicit goods using escrow and reputation systems, while forums act as hubs for knowledge exchange, service provision, and network formation [17]. Although central to anonymous transactions, the ecosystem extends beyond the dark web to the clear web and encrypted messaging platforms such as Telegram for recruitment, coordination, and data leaks [15, 18, 19]. This reflects the distributed and adaptive nature of cybercrime communities. The World Wide Web is commonly divided into the surface web, deep web, and dark web [20]. While the surface web is indexed by search engines and the deep web contains restricted-access content, the dark web requires specialized software to access anonymized services [21, 22]. From a structural perspective, dark web networks rely on decentralized routing and hub-like connectivity patterns that enable anonymous communication [23, 24]. While these mechanisms provide privacy protections [25], they are also widely used for illicit trade, with marketplaces facilitating transactions through encryption and cryptocurrency-based payments [26, 27]. Prior work shows that dark web content exhibits distinct linguistic and structural characteristics compared to the clear web, including differences in terminology, named entities, and discourse patterns [28, 29]. Large-scale analyses further indicate that the dark web forms a structured ecosystem of recurring content categories and domainspecific vocabulary [30, 31]. These consistent semantic

2.3. Research Gap Prior research has applied NLP and topic modeling techniques to dark web data, but largely outside the scope of economically motivated cybercrime as defined in this study. Early work focuses on uncovering latent themes or communities without centering on cybercriminal activity [56, 57], while more recent approaches use topic modeling for classification or ontology construction in general threat detection [53, 58]. Other studies apply these techniques in the context of terrorism and extremism rather than market-oriented cybercrime [59], and systematic reviews confirm that most work emphasizes broadly

1. We consider specific content categories (e.g., botnets, ransomware, forged document services) as topics rather than very broad umbrella themes such as drugs or weapons.

2

1. Data Preprocessing

3. Pipeline Evaluation

Llama-3.1 8b-instruct

Label Validation

Topics: 85

Final Labels: 55

Representation: c-TF-IDF and KeyBERT

RQ1 Topic Prevalence

Content extraction

HTML Snapshots (11,403,638) Websites: 25,065 Size: 1245.38 gb

2. Content Analysis

Years: 2020-2026

Define websites (path + title) Organize by temporal unit (website level) Exclusion Criteria (>=4 snapshots)

Data Preparation

Language filtering Text Normalization

Clustering: HDBSCAN Dimensionality Reduction: UMAP

Text Cleaning

Embedding: DARK-BERT

Filtered Snapshots: 7,381,762 Filtered Websites: 22,966 W. with ≥4 snapshots: 15,130

RQ2 Topic Lifecycle Longitudinal Analysis Topic Distribution

Topic Modeling

Figure 1: Overview of methodology from data preprocessing, to analysis pipeline, and evaluation

The dataset consists of 11,337,512 HTML snapshots (approximately 1245.38 GB) collected over six years (2020–20263 ). Each snapshot represents a captured HTML version of a “dark” web webpage at a specific point in time. This longitudinal archive allows us to observe how content on dark web platforms changes over repeated observations of the same website. More details on the dataset are provided in Appendix A.2.

defined security threats [60]. Thus, the application of topic modeling to economically motivated cybercrime communities remains limited. In addition, studies predominantly analyze dark web data as static snapshots or short-term corpora, providing limited insight into how topics evolve over time across platforms, and longer periods. To address this gap, this study adopts a longitudinal perspective using repeated HTML snapshots of dark web websites. Specifically, we investigate the following research questions: • RQ1 Topic Prevalence. Which topics are discussed on dark web forums and marketplaces, and how are they distributed across communities? • RQ2 Topic Lifecycle. How do topics change over time in terms of their relative prevalence, persistence, and duration across the observation period? We conduct a longitudinal analysis of HTML snapshots collected over six years, covering 25,065 dark web websites and 11,403,638 snapshots.

3.2. Data Preprocessing Before analysis, raw HTML snapshots were processed to construct longitudinal website histories. The preprocessing stage consists of website grouping, snapshot filtering, content extraction and text normalization. 3.2.1. Website Grouping and Snapshot Exclusion. Snapshots were grouped based on a combination of file path and page title, which together provide a robust identifier for the same webpage across time. Each group represents a single website observed at multiple points in time and was assigned a unique internal identifier. Duplicate and near-duplicate snapshots were removed during preprocessing. Within each group, snapshots were sorted chronologically using their creation timestamps. Groups containing four or less snapshots were removed, as they do not provide sufficient temporal depth to analyze content evolution. Additionally, metadata entries without corresponding HTML files were excluded to ensure dataset consistency. 3.2.2. Content Extraction and Text Normalization. Raw HTML snapshots were converted into plain text to isolate meaningful webpage content and remove structural noise. HTML parsing and content extraction were performed using the trafilatura library,4 which extracts the main textual content of web pages while excluding scripts, navigation menus and other non-informational elements. To ensure linguistic consistency, extracted text was filtered by language using the langdetect library. Only

3. Methodology We adopt a snapshot-based longitudinal design to analyze how discussions on dark web platforms evolve over time. The methodology consists of three main stages: (1) preprocessing and structuring HTML snapshots into longitudinal website histories, (2) discovering discussion topics using an embedding-based clustering pipeline, and (3) analyzing topic prevalence and lifecycle over time. An overview of the complete workflow is shown in Figure 1 and details on reproducibility are in Appendix A.1.

3.1. Dataset and Study Design The primary unit of our longitudinal analysis is a website, defined as a stable webpage identity observed across multiple HTML snapshots. The data used in this study was collected with the Dark Web Monitor Tool by CFLW Cyber Strategies, an organization specializing in cybersecurity and dark web monitoring. Their services support law enforcement agencies and public institutions2 .

3. Earliest snapshot timestamp: 2020-01-01 14:37:52; Latest snapshot timestamp: 2026-01-30 19:13:34. 4. Trafilatura is specifically designed for web content extraction and reliably removes navigation elements, scripts and boilerplate HTML structures.

2. All data utilized in this study can be obtained by researchers after submitting a research proposal with IRB approval.

3

3.3.4. Topic Representation and Labeling. To interpret the discovered clusters, representative keywords are extracted using class-based TF-IDF, which identifies terms that distinguish each cluster from the rest of the corpus. KeyBERT is additionally used to extract representative keyphrases that improve semantic readability. Because keyword lists can remain ambiguous for large-scale analysis, clusters are assigned semantic topic labels using the large language model (LLM) Llama-3.18B-Instruct. While LLMs enable scalable and interpretable labeling of large text corpora, their use in security-critical contexts must be carefully considered, as they can introduce reduce independent reasoning in human decisionmaking [62]. Thus, the model receives representative keywords and example documents from a cluster to generate a clear and concise human-readable topic name. For example, one cluster produced the keywords product, bought, vendor, shipping, escrow, and bitcoin. These terms reflect transactional discussions commonly found on illicit marketplaces. Based on these keywords and example documents, the language model generated the label Online Shopping. This process converts machinegenerated keyword sets into interpretable thematic labels suitable for longitudinal analysis. All generated labels were manually validated to ensure semantic accuracy and consistency. Labels that were similar, or a subset of another (e.g., label ”Market Transactions” aggregated to ”Online Shopping”) were merged to ensure proper representation. Details of the validation and label merging protocol are provided in Appendix A.4.

English-language snapshots were retained for analysis in order to avoid any noise into the topic modeling process. The extracted text was then normalized through lowercasing and removal of URLs, email addresses, punctuation, special characters, numbers, and repeated characters. Tokens were lemmatized using the WordNet lemmatizer to reduce words to their base forms. Stopwords were removed and tokens shorter than three characters or longer than 25 characters were discarded5 . The resulting corpus represents a standardized and noise-reduced textual dataset suitable for embedding generation and topic modeling. Details on dataset filtering, inclusion criteria, and temporal handling are documented in Appendix A.2. Individual HTML snapshots may contain multiple discussion threads or mixed topics, particularly in forum index pages or aggregated marketplace listings. In our approach, each snapshot is treated as a single document and represented as a probabilistic mixture over topics. Subthread-level segmentation is not performed, which may lead to blended topic representations in pages containing heterogeneous content. This limitation is further discussed in 5.3 and represents an avenue for future work.

3.3. Content Analysis Pipeline After preprocessing, each snapshot is passed through a structured topic discovery pipeline that converts textual content into interpretable thematic categories. The pipeline follows a sequential representation–clustering–labeling workflow illustrated in Figure 1. 3.3.1. Semantic Embedding. Each cleaned snapshot is converted into a semantic vector representation using contextual embeddings. We employ DARK-BERT [61], a domain-adapted transformer model trained on cybersecurity and darknet-related corpora, to generate document embeddings. Domain-specific embedding models have been shown to better capture specialized vocabulary and semantic relationships in technical security discussions. Other embedding models, including general and cybersecurity models, were evaluated during pipeline development. DARK-BERT consistently produced more coherent topic structures and semantically consistent document clusters and after manual verification was selected for the final pipeline. The model selection procedure and comparative evaluation are described in Appendix A.3. 3.3.2. Dimensionality Reduction. Because transformer embeddings are high-dimensional, we use Uniform Manifold Approximation and Projection (UMAP) to project embeddings to lower-dimensional space while preserving semantic neighborhood structure. This improves clustering stability and separates semantically distinct discussions while keeping relationships between related activities. 3.3.3. Topic Clustering. Topics are discovered using density-based clustering (HDBSCAN). Unlike centroidbased clustering methods, HDBSCAN does not require specifying the number of clusters in advance and can identify dense thematic clusters while treating sparse content as noise. Each resulting cluster corresponds to a candidate discussion topic representing a recurring theme.

3.4. Longitudinal Topic Analysis To analyze thematic time dynamics, each snapshot is represented as a probability distribution over discovered topics using BERTopic’s soft assignment mechanism. Rather than assigning a single topic to each snapshot, this approach captures mixtures of topics within a document. These topic distributions are then aggregated across snapshots belonging to the same website and time interval, producing time-indexed topic prevalence measures. This representation allows gradual thematic changes to be measured over time without forcing abrupt topic switches. Two quantitative measures are derived from these temporal topic distributions to address the research questions. Topic Prevalence (RQ1). Topic prevalence is measured as the aggregated topic probability mass across snapshots within each time interval. This metric quantifies how strongly each topic is represented across the broader ecosystem of cybercrime forums and marketplaces. Topic Lifecycle (RQ2). Topic lifecycle is measured using temporal change indicators, including topic lifespan, growth and decay rate. These metrics capture how topics emerge, persist and decline over the observed period.

3.5. Validation and Reliability To ensure the semantic validity and consistency of the discovered topics, a manual validation process was conducted. Two independent annotators reviewed the automatically generated topic labels based on representative keywords and documents. Annotators assessed whether each label accurately reflected the underlying content and

5. Such unusually long strings are typically non-linguistic artifacts (e.g., concatenated words, encoding errors or residual URLs) that do not contribute meaningful semantic information and may negatively impact embedding quality and topic coherence.

4

Transactional

8.54 5.26 2.68 2.18 1.31 1.38

Prod.

This study analyzes textual content from publicly accessible dark web forums and marketplaces. No interaction with users, accounts, or services occurred, and no attempts were made to access restricted areas, in line with established internet-mediated research guidelines [63]. Given the potential presence of illicit or sensitive material, only text necessary for aggregate analysis was processed, and no individuals or accounts were profiled. Results are reported exclusively at the aggregated level, focusing on ecosystem-wide patterns [64]. All analysis was conducted offline on archived snapshots, and no sensitive operational details (e.g., access methods, credentials, or live URLs) are disclosed. The dataset was preprocessed by the data provider to remove personally identifiable information. No direct threats or actionable intelligence targeting specific individuals were identified, and no reporting to authorities was required.

%Corp.

558,658 343,545 175,188 142,684 85,679 90,145

Stolen Bank and Payment Acc. Forged Document Services Counterfeit Money

147,206 146,661 147,332

2.25 2.24 2.25

Infras.

Category Topic Labels

3.6. Ethical Considerations

S. Count

Online Shopping Transaction Protection Online Banking Money Making Opportunities Prepaid Cards Vendor Sales & Shipping

Infrastructure and Hosting VPN Services Databases Browser Hijacking

308,950 138,165 190,900 92,592

4.73 2.11 2.92 1.42

Community

TABLE 1: Top 20 topics labels by corpus share, mapped to category

reworded or refined labels where necessary. In addition, they independently evaluated potential overlap between topics and merged them when substantial semantic similarity was identified. We provide further details on the validation procedure including examples in Appendix A.5.

Torrents and Files Forum Reputation Forum Features Forum Security Vendor Channels and Prom. WikiLeaks Politics and War

1,027,774 904,299 524,838 117,604 76,691 209,978 235,484

15.71 13.83 8.02 1.80 1.17 3.21 3.60

The aggregated distribution shows that community (47.38%) and transactional (20.59%) activities dominate the ecosystem, together accounting for the majority of observed discussion volume, while infrastructure (13.65%) and product-related topics (6.75%) represent smaller portions of the corpus, as shown in Table 1. The most prevalent topics are Torrents and Files (15.71%) and Forum Reputation (13.83%), followed by Online Shopping (8.54%) and Forum Features (8.02%). Beyond these leading topics, the remaining distribution consists of a larger number of lower-frequency topics, each contributing less than 8% of the corpus. The least prevalent topics in the top 20 set include Browser Hijacking (1.42%) and Vendor Sales & Shipping (1.38%), which are still considerably higher when compared to Forum Credentials (0.01%), Digital Payment Card Services (0.03%), Leaked Corporate Files (0.03%), and Banned Accounts (0.03%) which were the least prevalent out of the 55 labels. The complete set of 55 topic labels is shown and defined in Appendix A.6.

4. Results We present the results of the longitudinal topic analysis. The topic discovery pipeline outputted 85 distinct discussion topics extracted from 7,381,762a HTML snapshot. For each label produced by Llama (e.g., Databases, Forum Security, Prepaid Cards), topics were inspected to assess whether they reflected a consistent and interpretable theme and were distinguishable from other topics. After label verification and correction, they were aggregated by merging similar labels, which resulted in 55 final labeled topics. To improve interpretability, we also group the labels into four overarching categories based on their semantic structure and functional roles within the ecosystem (see Table 1). The Transactional category captures activities related to financial exchange, including payment mechanisms, transaction security, and monetization processes. The Products category represents the supply side of the ecosystem, covering the trade of illicit goods and services. The Infrastructure category includes the technical and operational components that enable cybercriminal activity, such as tools and supporting technologies. Finally, the Community category captures the social and informational layer of the ecosystem, including coordination, trust-building and knowledge sharing among participants, primarily within forum environments. The remainder of this section addresses the research questions. First, we analyze how topics are distributed across platforms and how prevalent they are (§4.1), and then examine their emergence and decay through time (§4.2).

4.1.1. Distribution Across Platforms. Topic prevalence differed substantially between forums and marketplaces. Figure 2 compares topic composition across marketplaces, forums, and other site types. Naturally, forum-related topics such as Forum Reputation, Forum Features, and Forum Security are more prominent in forums, while transaction- and product-related topics such as Online Shopping, Stolen Bank and Payment Accounts, and Forged Document Services show higher representation in marketplaces. Several topics, including Infrastructure and Hosting and Online Shopping, appear across multiple site types with varying proportions. Only few topics represent a somewhat even share across platforms, amongst them, Forum Reputation and Transaction Protection. Overall, the distribution indicates differences in topic composition between platform types, with certain topics concentrated more strongly in specific environments.

4.1. Topic Prevalence (RQ1) First, we analyze how discussion topics are distributed across dark web marketplaces and forums, and how their prominence varies across the observed period.

4.1.2. Temporal Prevalence. Topic prevalence evolves over time, wherein the changes are gradual rather than

5

30

Mean = 61.5 Median = 68.0

Forum Reputation Torrents and Files

25

Forum Features

Number of topics

Infrastructure and Hosting Online Shopping

Topic Label

Databases Online Banking Money Making Opportunities WikiLeaks Transaction Protection

20

15

10

5

VPN Services Counterfeit Money

0

Forum Security

20

30

40

50

60

70

Topic lifespan (months)

Forged Document Services

Marketplace Forum Other

Stolen Bank and Payment Accounts 0

5

10

15

20

Figure 4: Topic Lifespan Distribution.

25

Percentage within site type

identified themes are either continuous or recurring, meaning that no final topic cluster appears only once across the observation window. Thus, the dominant topics are not only highly prevalent, but also structurally persistent over time. This indicates that the dark web ecosystem is organized around stable thematic functions.

Figure 2: Distribution of marketplace vs forum vs other of top 20 topic labels in corpus. 100 Forum Reputation Other Topics Online Shopping Torrents and Files WikiLeaks Forum Features Infrastructure and Hosting Online Banking Counterfeit Money Money Making Opportunities Forum Security Databases Transaction Protection VPN Services Politics and War Forged Document Services Vendor Sales & Shipping Prepaid Cards Vendor Channels and Promotions Browser Hijacking Stolen Bank and Payment Accounts

% share of activity

80

60

40

20

0

2020

2021

2022

2023

2024

2025

A NSWER TO RQ1: Topic prevalence is highly concentrated in a small number of dominant themes. Five topics explain more than half of total discussion volume, and thirteen topics explain three quarters of the corpus. At the same time, these dominant topics remain persistent across the observation window, which indicates that the ecosystem is structured around a stable set of recurring functions rather than a broad or rapidly changing distribution of themes.

2026

Time

Figure 3: Temporal Prevalence of Top 20 Topics.

4.2. Topic Lifecycle (RQ2) While prevalence captures how strongly topics are represented within the ecosystem, lifecycle analysis examines how long topics remain active and how their activity levels change over time. 4.2.1. Topic Lifespan Distribution. Figure 4 shows the distribution of topic lifespans measured in months across the observation period. The distribution is strongly skewed toward longer durations, with most topics remaining active for a substantial portion of the six-year timeframe. The median lifespan is approximately 68 months, while the mean lifespan is 61.5 months, indicating that the majority of topics persist over multiple years. Only a small number of topics exhibit shorter lifespans, suggesting that short-lived discussions are relatively uncommon. Instead, the ecosystem is dominated by longstanding topics that remain continuously present, with variation occurring primarily in their level of activity. This pattern indicates that topic turnover is limited and that most thematic structures persist over time. 4.2.2. Temporal Activity and Persistence. Most topics remain active across large portions of the observation period, indicating persistent thematic structures. Variation is visible primarily in relative intensity rather than in complete disappearance. Figure 5 shows that activity is redistributed across topics over time, but the underlying topic set remains largely stable. Figure 6 further illustrates that even the most prevalent topics follow smooth trajectories rather than abrupt discontinuities.

disruptive. As shown in Figure 3, the top 20 topics persist across the observation period, with shifts primarily occurring in their relative prominence (not through new theme introduction). Core topics such as Forum Reputation and Online Shopping remain consistently visible across multiple intervals. Overall, the ecosystem is characterized by stability within a fixed thematic structure. While individual topics rise or decline in importance, the dominant set of topics remains largely unchanged. This suggests that temporal dynamics are driven by a redistribution of attention within an established core and not by continuous thematic turnover. Within this stable structure, we observe notable variations in prominence. Between 2020 and 2022, Forum Reputation dominated the discourse, followed by WikiLeaks, Online Shopping, and Infrastructure and Hosting, with Online Banking gaining relative importance toward the end of the period. 4.1.3. Quantitative prevalence indicators. The top five topics account for 53.15% of total discussion volume, the top ten account for 68.67%, and the top twenty account for 88.36% of the corpus. Only five topics are required to explain 50% of all observed activity, and thirteen topics explain 75%. This shows that topic prevalence is not broadly distributed across many equally important themes, but ordered around a small set of dominant activities. The concentration pattern is reinforced by recurrence behavior. At the level of the final 55 grouped topics, all

6

1.0

Databases

Representative Topic Growth and Decay Trajectories

Money Making Opportunities 0.8

Prepaid Cards

50

Torrents and Files Vendor Sales & Shipping

0.6

0.4

Politics and War Vendor Channels and Promotions VPN Services Transaction Protection

0.2

Stolen Bank and Payment Acc... Browser Hijacking Infrastructure and Hosting

peak

40 % share of activity

Topic category

Forum Features

Relative activity within topic (0–1)

Online Banking Counterfeit Money

Stable: AI-Generated Image and Video Services

Bursting: Databases

30

20

10 peak

0

Forum Security Forum Reputation

Emerging: Forum Features

Forged Document Services

Declining: Forum Reputation peak

50

Online Shopping WikiLeaks

40

0.0

20

% share of activity

20 20 Q1 20 20 Q2 20 20 Q3 20 20 Q4 21 20 Q1 21 20 Q2 21 20 Q3 21 20 Q4 22 20 Q1 22 20 Q2 22 20 Q3 22 20 Q4 23 20 Q1 23 20 Q2 23 20 Q3 23 20 Q4 24 20 Q1 24 20 Q2 24 20 Q3 24 20 Q4 25 20 Q1 25 20 Q2 25 20 Q3 25 20 Q4 26 Q 1

peak

Time (quarters)

Figure 5: Topic Lifecycle.

30

20

10

20 20 Q1 20 20 Q2 20 20 Q3 20 20 Q4 21 20 Q1 21 20 Q2 21 20 Q3 21 20 Q4 22 20 Q1 22 20 Q2 22 20 Q3 22 20 Q4 23 20 Q1 23 20 Q2 23 20 Q3 23 20 Q4 24 20 Q1 24 20 Q2 24 20 Q3 24 20 Q4 25 20 Q1 25 20 Q2 25 20 Q3 25 20 Q4 26 Q 1

20

20 20 20 Q1 20 20 Q2 20 20 Q3 20 20 Q4 21 20 Q1 21 20 Q2 21 20 Q3 21 20 Q4 22 20 Q1 22 20 Q2 22 20 Q3 22 20 Q4 23 20 Q1 23 20 Q2 23 20 Q3 23 20 Q4 24 20 Q1 24 20 Q2 24 20 Q3 24 20 Q4 25 20 Q1 25 20 Q2 25 20 Q3 25 20 Q4 26 Q 1

0

50

Time (quarters)

% share of activity

Time (quarters)

Figure 7: Representative Growth and Decay of Topics.

40

Torrents and Files Forum Reputation Online Shopping Forum Features Infrastructure and Hosting Transaction Protection Forum Security Counterfeit Money Databases Online Banking

30

20

Growth–Decay Dynamics of Topics

Growth–Decay Dynamics (Dense Cluster Focus)

Outliers / Extreme Topics

Banned Accounts Forum Security

0

Bitcoin Transactions

Hacked Social Media Accounts

Infrastructure and Hosting Safety Concerns

0.0

Leaked Data Transaction Protection

WikiLeaks

Dynamic class Bursting Stable Emerging Declining

−1

10 Decay slope after peak

Online Shopping

20

20 20 Q1 20 20 Q2 20 20 Q3 20 20 Q4 21 20 Q1 21 20 Q2 21 20 Q3 21 20 Q4 22 20 Q1 22 20 Q2 22 20 Q3 22 20 Q4 23 20 Q1 23 20 Q2 23 20 Q3 23 20 Q4 24 20 Q1 24 20 Q2 24 20 Q3 24 20 Q4 25 20 Q1 25 20 Q2 25 20 Q3 25 20 Q4 26 Q 1

0

−0.5 −2 Forum Reputation

−1.0 −3

Time (quarters)

Lifespan 8 periods 18 periods 25 periods

Databases

−1.5

−4

Figure 6: Topic Temporal Prevalence (Top ten Topics).

Torrents and Files

−5

−2.0 0.0

The concentration of activity in long-lived topics implies that cybercrime communities are not defined by rapid thematic churn. Instead, the same core activities remain active for longer periods, while their relative prominence changes in response to broader contextual conditions. 4.2.3. Growth and Decay Dynamics. Figure 7 shows representative growth and decay trajectories across topics. Distinct patterns are observed. Databases exhibits a bursting pattern, with a rapid increase to a clear peak followed by a relatively sharp decline. In contrast, AI-Generated Image and Video Services remains stable, maintaining a relatively constant activity level over time. Forum Features shows an emerging pattern, with gradual growth across multiple periods leading to a late peak. Conversely, Forum Reputation displays a declining pattern, with higher early activity followed by a steady decrease. Figure 8 summarizes these dynamics using growth and decay slopes. Most topics cluster in a central region, indicating moderate growth and moderate decay. This includes topics such as Forum Reputation, Online Shopping, and Infrastructure and Hosting, which exhibit balanced and gradual lifecycle dynamics. A small number of topics appear as outliers. Torrents and Files and Forum Features show higher growth slopes, indicating faster increases before peak, combined with stronger post-peak declines. In contrast, more stable topics are located near the origin, reflecting limited variation in both growth and decay. 4.2.4. Recurring vs One-Off Topics. Figure 9 shows a strong concentration of topics along the diagonal (recurrence ≈ lifespan), indicating that a large share of topics are active in nearly all observed periods. These continuous topics reach maximum lifespan (25 periods) with consistently high recurrence counts, suggesting persistent

0.5

1.0

1.5 2.0 Growth slope before peak

2.5

3.0

Forum Features

0

2

4 6 8 10 Growth slope before peak

12

Figure 8: Growth and Decay Dynamics of Topics.

and sustained activity. In contrast, a smaller set of topics is located in the lower-left region of the figure, characterized by low lifespan (< 10 periods) and low recurrence (< 10), indicating short-lived and episodic behavior. The distribution is therefore bimodal, with limited presence of topics in intermediate ranges of lifespan and recurrence. This suggests that topics tend to be either stable and continuously active or short-lived and sporadic, rather than gradually emerging and decaying. The zoomed cluster further shows that even among long-lived topics (lifespan ≈ 25), recurrence varies substantially, reflecting differences in activity intensity rather than survival. 4.2.5. Quantitative lifecycle indicators. The lifecycle results provide a more explicit picture of topic persistence and change. The median topic lifespan is 75 months (25 periods), while the mean lifespan is 68.45 months (22.82 periods). The longest surviving topics include Forum Reputation, Online Shopping, and Forum Features, whereas the shortest-lived topics include CAPTCHA, Banned Accounts, and Question and Answer. These results confirm that most topics remain active for substantial portions of the observation window, while genuinely short-lived grouped themes are comparatively rare. Growth and decline are similarly uneven but bounded. The strongest positive trends are observed for Forum Features, Torrents and Files, and Infrastructure and Hosting, while the steepest negative trends are observed for Forum Reputation, WikiLeaks, and Online Shopping. Among topics that disappear before the final observed period, the mean time from peak activity to last observed activity is 28 months. This indicates that even declining topics generally

7

Recurring vs One-Off Topics Topic volume 2406 docs 38475 docs 227649 docs

25

Location Information

20

Safety Concerns Finance and Operations

Hacked Social Media Accounts AI-Generated Image and Video Services

Recurrence count

decline and redistribution of attention within an already established thematic structure. At the same time, the ecosystem is not static. Several topics show meaningful directional trends. Forum Features, Torrents and Files, and Infrastructure and Hosting exhibit the strongest positive trajectories, while Forum Reputation, WikiLeaks, and Online Shopping show the strongest declines. However, even among topics that disappear before the end of the observation period, the average lag between peak activity and disappearance is 28 months. This indicates that decline is typically progressive.

Rocksolid Newsreader Hacking Torrents and Files Transaction Protection Ransomware Leaked Data Social Media Hacking Remote Access Exposed Website Directories Financial Tracing

Leaked Archive File Tree Listings Personal Data Merchant Transactions

15

Question and Answer

Crypto Wallets

Leaked Corporate Files Access Control and Bypass Forum Credentials Leaked Data Archives

10

Password Security Digital Payment Card Services Banned Accounts CAPTCHA

Recurring type Continuous Recurring

5 Continuous topics (y = x)

One-off line

0 0

5

10

15 Lifespan (periods)

20

25

Zoomed Recurring Cluster Forum Features

Online Shopping

Forum Security

File Downloading

25.1

Forged Document Services

Databases

Money Making Opportunities PayPal Scams Tor Search Engines Politics and War Escrow and Scam Protection Bitcoin Transactions Online Banking

Recurrence count

Vendor Channels and Promotions

VPN Services Doxxing WikiLeaks Forum Reputation

25.0

Infrastructure and Hosting Vendor Sales & Shipping Browser Hijacking Stolen Bank and Payment Accounts

Credit Card Data Prepaid Cards

5.2. Implications

Accounts

Hidden Content and Access Links

24.9

Counterfeit Money

24.9

25.0

25.1

Our findings highlight that cybercrime ecosystems evolve through stable, long-lived topics with gradual shifts in prominence, which has several implications for CTI: Longitudinal monitoring is essential. State-of-the-art approaches miss whether topics are emerging, declining or stable. Tracking topic trajectories over time reveals meaningful dynamics that are otherwise invisible. Resource allocation should be stratified. Persistent core topics (e.g., Online Shopping, Infrastructure and Hosting) need continuous automated monitoring, while emerging topics (e.g., Forum Features, Databases) require periodic expert analysis. Short-lived spikes can be handled via lightweight alerts. This allows more efficient use of analyst effort. Distinguishing signal from noise improves operational intelligence. Persistent topics provide reliable intelligence sources, while sustained growth indicates structural change. In contrast, short-lived bursts often reflect temporary fluctuations and should not be overinterpreted. Longitudinal analysis enables evaluation of enforcement impact. Changes in topic prevalence can indicate whether interventions lead to suppression, displacement or rapid recovery. The observed multi-year topic lifespans provide a baseline for distinguishing temporary disruption from structural change. Threat intelligence should adopt longer planning horizons. Since most topics persist for years, organizations can move from reactive monitoring to multi-year strategic planning. Sustained trends justify long-term investment, whereas transient topics do not.

25.2

Lifespan (periods)

Figure 9: Recurring vs One-Off Topics.

fade over multiple periods and not disappear immediately. A NSWER TO RQ2: Topic lifecycles are dominated by persistence and gradual change. The median topic lifespan is 75 months, and even the shortest-lived grouped topics remain active for multiple years. Changes in thematic prominence occur mainly through gradual growth and decline rather than abrupt emergence or disappearance, indicating that cybercrime ecosystems evolve through slow redistribution of attention within a stable topic structure.

5. Discussion We discuss our key findings (§5.1), outline implications (§5.2) as well as limitations and future work (§5.3).

5.1. Key Findings The dark web ecosystem is both structurally concentrated and thematically persistent. A small number of topics account for most observed activity: the top five topics explain 53.15% of discussion volume, the top ten explain 68.67%, and the top twenty explain 88.36% of the corpus. This concentration indicates that cybercrime ecosystems are not organized around a broad and continuously shifting set of dominant concerns, but instead around a limited set of stable and repeatedly observed functions. This concentration is closely linked to their persistence. At the level of the final grouped topic structure, all identified topics are either recurring or continuous rather than one-off. In other words, the most important themes are not only prevalent but also durable over time. Thus, the ecosystem is shaped less by novelty than by continuity in economically relevant activities such as exchange, coordination, reputation building, infrastructure use and transaction protection. The lifecycle analysis confirms this interpretation. Topics are long-lived, with a median lifespan of 75 months and a mean lifespan of 68.45 months. Even the shortestlived grouped topics remain active for at least two years. This indicates that thematic change in cybercrime ecosystems rarely takes the form of abrupt emergence and disappearance. Instead, changes occur through gradual growth,

5.3. Limitations and Future Work We outline limitations that should be considered when interpreting the results and reflect upon future work. Dataset Constraints As the dataset was collected by an external organization using proprietary crawling strategies, the coverage of the dark web ecosystem cannot be considered exhaustive. Snapshot frequency varies across websites and time periods, which may introduce temporal bias where increased crawling is misinterpreted as increased activity. Website identification based on file path and page title may not fully capture rebranding, domain changes, or migration, which can potentially underestimate continuity. Cross-platform activity (e.g., migration to Telegram) is not captured, so observed declines may reflect relocation rather than disappearance. Topic prevalence reflects content volume only, as no traffic or user engagement data are available. Additionally, the observation window

8

(2020–2026) includes major external events but lacks earlier baseline data, which limits our comparison with longer-term trends. Modeling Assumptions Our focus on English-language content ensures semantic consistency but excludes nonEnglish communities, limiting generalizability to the global cybercrime ecosystem. Topic modeling parameters were fixed after initial evaluation. Although they produce coherent and interpretable topics, alternative configurations may reveal different levels of granularity or uncover niche activities currently treated as noise. Furthermore, topics are modeled as static entities over time, which does not capture semantic drift, splitting or merging of themes as discussions evolve. Temporal Aggregation Aggregating snapshots into quarterly intervals reduces noise and mitigates crawling artifacts, but may overlook short-term dynamics. Thus, we may miss events such as enforcement actions, marketplace shutdowns or sudden activity spikes. Platform Differentiation Forums and marketplaces are modeled jointly to enable ecosystem-level analysis and cross-platform comparison. While this provides a unified view, it may obscure platform-specific practices, vocabulary and structural differences (e.g., in platform types). Temporal Validation Our manual validation focused on topic coherence and label accuracy. While we ensured that observed trends reflect actual content changes rather than crawling artifacts, we did not systematically validate topic trajectories against external event timelines, such as enforcement actions or publicly reported data breaches. Interpretations of temporal patterns are therefore based on post-hoc contextualization.

set of sites. The results show that cybercrime ecosystems evolve primarily through gradual shifts in existing topics, while maintaining a stable underlying structure.

ACKNOWLEDGMENT We thank CFLW Cyber Strategies for data collected with their Dark Web Monitor. Part of this research was funded by Hilti. This publication is part of the INTERSECT project, Grant No. NWA.1162.18.301, funded by NWO and the CATRIN project, Grant No. NWA.1215.18.003.

References [1]

E. Essien, “Relevance of the deep web to academic research,” International Journal of Natural and Applied Sciences, vol. 12, pp. 107–113, 2020. [2] “Tor metrics,” 2025. [Online]. Available: https:// metrics.torproject.org/ [3] S. Sobhan, T. Williams, M. J. H. Faruk, J. Rodriguez, M. Tasnim, E. Mathew, J. Wright, and H. Shahriar, “A review of dark web: Trends and future directions,” in 2022 IEEE 46th Annual COMPSAC, 2022, pp. 1780–1785. [4] S. L. Schroer, N. Canevascini, I. Pekaric, P. Widmer, and P. Laskov, “ The Dark Side of the Web: Towards Understanding Various Data Sources in Cyber Threat Intelligence ,” in 2025 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW). Los Alamitos, CA, USA: IEEE Computer Society, Jul. 2025, pp. 79–89. [5] J. Robertson, A. Diab, E. Marin, E. Nunes, V. Paliath, J. Shakarian, and P. Shakarian, Darkweb cyber threat intelligence mining. Cambridge University Press, 2017. [6] E. Nunes, A. Diab, A. Gunn, E. Marin, V. Mishra, V. Paliath, J. Robertson, J. Shakarian, A. Thart, and P. Shakarian, “Darknet and deepnet mining for proactive cybersecurity threat intelligence,” in 2016 IEEE ISI, 2016, pp. 7–12. [7] M. S. A. Basha, K. V. Kumar, and R. N D, “Classifying dark web–related social media discourse using machine learning, deep learning, and transformer models,” in 2025 6th ICICNIS, 2025. [8] G. P. Paoli, J. Aldridge, R. Nathan, and R. Warnes, “Behind the curtain: The illicit trade of firearms, explosives and ammunition on the dark web,” 2017. [9] M. Ball and R. Broadhurst, “Data capture and analysis of darknet markets,” Available at SSRN 3344936, 2021. [10] R. Hoheisel, T. Meurs, J. Wientjes, M. Junger, A. Abhishta, and M. Paquet-Clouston, “Assessing crime disclosure patterns in a large-scale cybercrime forum,” arXiv preprint arXiv:2603.01624, 2026. [11] D. Décary-Hétu and L. Giommoni, “Do police crackdowns disrupt drug cryptomarkets? a longitudinal analysis of the effects of operation onymous,” Crime, Law and Social Change, 2017. [12] M. Felderer and I. Pekaric, “Research challenges in empowering agile teams with security knowledge based on public and private information sources.” 2017. [13] I. Pekaric, R. Groner, A. Raschke, T. Witte, J. G. Adigun, M. Felderer, and M. Tichy, “Bridging safety and security in complex systems: A model-based approach with saft-gt toolchain,” Journal of Systems and Software, 2026. [14] B. Dupont and C. Whelan, “Enhancing relationships between criminology and cybersecurity,” Journal of Criminology, 2021. [15] R. Ricaldi, T. Marjanov, L. Allodi, and A. Hutchings, “Uncovering the trust signals supporting telegram’s cybercrime economy,” in 2025 eCrime, 2025, pp. 1–17. [16] Y. Wang, B. Arief, and J. Hernandez-Castro, “Analysis of security mechanisms of dark web markets,” in Proceedings of the 2024 EICC. NY, USA: Association for Computing Machinery, 2024. [17] E. R. Leukfeldt, E. R. Kleemans, and W. P. Stol, “Cybercriminal networks, social ties and online forums: Social ties versus digital ties within phishing and malware networks,” The British Journal of Criminology, vol. 57, no. 3, pp. 704–722, 2017. [18] F. Hasanti, M. Z. Osman, M. H. Rahman, M. Z. A. Darus, and N. B. Mohd, “A comprehensive study on emerging trends of dark web marketplaces and forums,” in 2024 IEEE ICOCO, 2024. [19] L. Allodi, R. Ricaldi, J. Wientjes, and A. Radu, “Where is dmitry going? framing ’migratory’ decisions in the criminal underground,” 2024. [Online]. Available: https://arxiv.org/abs/2411.16291

Future Directions We could address these limitations via several directions: (i) controlled crawling with consistent snapshot intervals would eliminate temporal bias from uneven crawling schedules and enable finer-grained analysis of short-term dynamics; (ii) extending the framework to non-English content would provide a more complete view of the global cybercrime ecosystem; (iii) separate topic modeling for forums, marketplaces, and other platform types could reveal specialized practices and vocabularies; (iv) incorporating temporal dependencies directly into the modeling process (e.g., using dynamic topic models or sequential clustering) could better capture how topics evolve, split or merge; (v) integrating clear web sources (e.g., Telegram channels, paste sites) would enable tracking of activity migration across infrastructures; (vi) systematic evaluation of clustering parameters could establish robustness bounds and identify whether key findings hold across alternative configurations; (vii) analyzing cross-topic correlations to capture dependencies between themes; and (viii) building on longitudinal topic trajectories to predict topic persistence, emergence or platform migration patterns could allow for forward-looking CTI threat assessment.

6. Conclusions This study examines how content on the dark web evolves to identify patterns relevant to cybersecurity and CTI. Using NLP and topic modeling on longitudinal snapshots, we captured thematic development across a large and diverse

9

[20] J. Sultana and A. K. Jilani, Exploring and Analysing Surface, Deep, Dark Web and Attacks. Cham: Springer, 2021, pp. 97–108. [21] S. Kaur and S. Randhawa, “Dark Web: A Web of Crimes,” Wireless Personal Communications, vol. 112, no. 4, Jun. 2020. [22] D. Kavallieros, D. Myttas, E. Kermitsis, E. Lissaris, G. Giataganas, and E. Darra, Understanding the Dark Web. Cham: Springer International Publishing, 2021, pp. 3–26. [23] M. Campobasso, R. Rădulescu, S. Brons, and L. Allodi, “You can tell a cybercriminal by the company they keep: A framework to infer the relevance of underground communities to the threat landscape,” arXiv preprint arXiv:2306.05898, 2023. [24] I. Pete, J. Hughes, Y. T. Chua, and M. Bada, “A social network analysis and comparison of six dark web forums,” in 2020 IEEE EuroS&PW, 2020, pp. 484–493. [25] S. Davis and B. Arrigo, “The dark web and anonymizing technologies: legal pitfalls, ethical prospects, and policy directions from radical criminology,” Crime, Law and Social Change, vol. 76, no. 4, pp. 367–386, Nov 2021. [26] D. Georgoulias, J. M. Pedersen, M. Falch, and E. Vasilomanolakis, “A qualitative mapping of darkweb marketplaces,” in 2021 APWG eCrime, 2021, pp. 1–15. [27] E. Kermitsis, D. Kavallieros, D. Myttas, E. Lissaris, and G. Giataganas, Dark Web Markets. Cham: Springer International Publishing, 2021, pp. 85–118. [28] V. Varghese, S. Mahalakshmi, and S. Kb, “Extraction of actionable threat intelligence from dark web data,” in ICCC. IEEE, 2023. [29] L. Choshen, D. Eldad, D. Hershcovich, E. Sulem, and O. Abend, “The language of legal and illegal activity on the darknet,” 01 2019. [30] G. Avarikioti, R. Brunner, A. Kiayias, R. Wattenhofer, and D. Zindros, “Structure and content of the visible darknet,” arXiv preprint arXiv:1811.01348, 2018. [31] Y. Jin, E. Jang, Y. Lee, S. Shin, and J.-W. Chung, “Shedding new light on the language of the dark web,” in Proceedings of the 2022 conference of the north American chapter of the association for computational linguistics: human language technologies, 2022. [32] C. Fachkha and M. Debbabi, “Darknet as a source of cyber intelligence: Survey, taxonomy, and characterization,” IEEE Communications Surveys & Tutorials, vol. 18, no. 2, 2016. [33] R. Ricaldi, Y. Yalamov, M. Campobasso, L. Allodi, H. Kool, A. Moneva, and E. R. Leukfeldt, “An experimental design to investigate attacker actions on an access-as-a-service ‘criminal’ platform,” in 2025 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), 2025, pp. 109–114. [34] P. W. Wang, X. L. Liao, Y. Qin, and X. Wang, “Into the deep web: Understanding e-commercefraud from autonomous chat with cybercriminals,” in Proceedings of the ISOC Network and Distributed System Security Symposium (NDSS), 2020, 2020. [35] Klaus Krippendorff, Content Analysis : An Introduction to Its Methodology, fourth edition ed. Los Angeles: SAGE, Jan. 2019. [36] L. De-Marcos, J.-A. Medina-Merodio, and Z. Stapic, “Methodologies for data collection and analysis of dark web forum content: A systematic literature review,” Electronics, vol. 14, no. 21, p. 4191, 2025. [37] C. Heistracher, S. Schlarb, and F. Ghaffar, “Information extraction from darknet market advertisements and forums,” in Proceedings of the 14th international Conference on emerging security information, systems and Technologies (SECURWARE 2020), 2020. [38] M. Kadoguchi, S. Hayashi, M. Hashimoto, and A. Otsuka, “Exploring the dark web for cyber threat intelligence using machine leaning,” in 2019 IEEE International Conference on Intelligence and Security Informatics (ISI). IEEE, 2019, pp. 200–202. [39] M. R. Rahman, R. Mahdavi-Hezaveh, and L. Williams, “A literature review on mining cyberthreat intelligence from unstructured texts,” in 2020 ICDMW. IEEE, 2020, pp. 516–525. [40] E. Karaosman, A. Rizvani, and I. Pekaric, “Security Barriers to Trustworthy AI-Driven Cyber Threat Intelligence in Finance: Evidence from Practitioners,” in The Sixteenth ACM Conference on Data and Application Security and Privacy (CODASPY), 2026. [41] F. Dong, S. Yuan, H. Ou, and L. Liu, “New cyber threat discovery from darknet marketplaces,” in 2018 IEEE Conference on Big Data and Analytics (ICBDA). IEEE, 2018, pp. 62–67. [42] N. Tavabi, P. Goyal, M. Almukaynizi, P. Shakarian, and K. Lerman, “Darkembed: Exploit prediction with neural language models,” in Proceedings of the AAAI Conference, vol. 32, no. 1, 2018. [43] K. S. Sangher, A. Singh, and H. M. Pandey, “Lstm and bert based transformers models for cyber threat intelligence for intent identification of social media platforms exploitation from darknet forums,” International Journal of Information Technology, 2024.

[44] K. S. Sangher, A. Singh, H. M. Pandey, and V. Kumar, “Towards safe cyber practices: Developing a proactive cyber-threat intelligence system for dark web forum content by identifying cybercrimes,” Information, vol. 14, no. 6, p. 349, 2023. [45] B. Mardassa, A. Beza, A. Al Madhan, and M. Aldwairi, “Sentiment analysis of hacker forums with deep learning to predict potential cyberattacks,” in 2024 15th Annual Undergraduate Research Conference on Applied Computing (URC). IEEE, 2024, pp. 1–6. [46] D. Jiang, C. Zhang, and Y. Song, Topic Models. Singapore: Springer Nature Singapore, 2023, pp. 27–46. [47] M. Grootendorst, “Bertopic: Neural topic modeling with a class-based tf-idf procedure,” 2022. [Online]. Available: https: //arxiv.org/abs/2203.05794 [48] C. Heistracher, F. Mignet, and S. Schlarb, “Machine learning techniques for the classification of product descriptions from darknet marketplaces.” in ICAI, 2020, pp. 128–137. [49] C. A. Murty and P. H. Rughani, “Dark web text classification by learning through svm optimization,” Journal of Advances in Information Technology, vol. 13, no. 6, pp. 624–631, 2022. [50] S. Ghosh, A. Das, P. Porras, V. Yegneswaran, and A. Gehani, “Automated categorization of onion sites for analyzing the darkweb ecosystem,” in Proceedings of the 23rd ACM SIGKDD, 2017. [51] J. Pastor Galindo, H.-Â. Sandlin, F. G. Mármol, G. Bovet, and G. M. Pérez, “A big data architecture for early identification and categorization of dark web sites,” Future Generation Computer Systems, vol. 157, pp. 67–81, 2024. [52] A. Dalvi, A. Shah, P. Desai, R. Chavan, and S. Bhirud, “A comparative analysis of models for dark web data classification,” in International Joint Conference on Advances in Computational Intelligence. Springer, 2022, pp. 245–257. [53] G.-Y. Shin, Y. Jang, D.-W. Kim, S. Park, A.-R. Park, Y. Kim, and M.-M. Han, “Dark side of the web: Dark web classification based on textcnn and topic modeling weight,” IEEE Access, 2023. [54] C. Chen, C. Peersman, M. Edwards, Z. Ursani, and A. Rashid, “Amoc: A multifaceted machine learning-based toolkit for analysing cybercriminal communities on the darknet,” in 2021 IEEE International Conference on Big Data. IEEE, 2021. [55] C. Murty and P. H. Rughani, “Sentiment & pattern analysis for identifying nature of the content hosted in the dark web,” Indian J. Comput Sci Eng, vol. 12, no. 6, 2021. [56] L. Yang, F. Liu, J. M. Kizza, and R. K. Ege, “Discovering topics from dark websites,” in Proceedings of the IEEE Symposium on Computational Intelligence in Cyber Security, 2009. [57] G. L’Huillier, H. Alvarez, S. A. Rı́os, and F. Aguilera, “Topic-based social network analysis for virtual communities of interests in the dark web,” in ACM SIGKDD Explorations Newsletter, 2011. [58] R. Basheer and B. Alkhatib, “Darkonto: An ontology construction approach for dark web community discussions through topic modeling and ontology learning,” Human Behavior and Emerging Technologies, 2024. [59] E. Sönmez and K. Seçkin Codal, “Analyzing a dark web forum page in the context of terrorism: a topic modeling approach,” Security Journal, vol. 37, no. 4, pp. 1360–1381, 2024. [60] S. Nazah, S. Huda, J. Abawajy, and M. M. Hassan, “Evolution of dark web threat analysis and detection: A systematic approach,” IEEE Access, vol. 8, pp. 171 796–171 819, 2020. [61] Y. Jin, E. Jang, J. Cui, J.-W. Chung, Y. Lee, and S. Shin, “Darkbert: A language model for the dark side of the internet,” in Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), 2023, pp. 7515–7533. [62] I. Pekaric, P. Zech, and T. Mattson, “Llms in cybersecurity: Friend or foe in the human decision loop?” arXiv preprint arXiv:2509.06595, 2025. [63] C. Hewson and T. Buchanan, “Ethics guidelines for internetmediated research.” The British Psychological Society, 2013. [64] B. Pickering, S. Roth, and C. Webber, “Ethical approaches to studying cybercrime: considerations, practice and experience in the united kingdom,” in Researching Cybercrimes: Methodologies, Ethics, and Critical Approaches. Springer, 2021, pp. 347–369. [65] M. Pfister, G. Apruzzese, and I. Pekaric, “Department-Specific Security Awareness Campaigns: A Cross-Organizational Study of HR and Accounting,” in 2025 APWG Symposium on Electronic Crime Research (eCrime), 2025. [66] J. Ave, I. Pekaric, M. Frohner, and G. Apruzzese, ““bot lane noob”: Towards Practical Deployment of NLP-based Toxicity Detectors in Video Games,” in European Symposium on Research in Computer Security (ESORICS), 2026.

10

Appendix A.

with our definition of cybercrime as economically motivated activity, the dataset was pre-filtered by the provider to exclude content categories outside this scope. This includes explicit adult material, purely ideological or political content (e.g., extremism), and other non-economic or non-cybercrime-related domains. Additionally, personally identifiable information and sensitive operational details were removed during preprocessing. The resulting dataset focuses on textual content relevant to economically driven cybercriminal activity, suitable for aggregate analysis of ecosystem-level patterns.

A.1. Reproducibility Overview This section summarizes the configuration required to reproduce the content analysis pipeline. All preprocessing and modeling steps were executed offline on archived HTML snapshots using deterministic settings where applicable. The pipeline was implemented in Python using the following libraries: • BERTopic • HuggingFace and sentence-transformers • UMAP-learn and HDBSCAN • scikit-learn • Trafilatura (HTML text extraction) HTML content was extracted using Trafilatura’s default extraction settings and only English-language text was retained. Text preprocessing included: • lowercase normalization • removal of URLs, email addresses, punctuation, numbers, and special characters • WordNet lemmatization • standard English stopword removal • token length filtering (3–25 characters) Websites with fewer than four snapshots were excluded from the dataset. Semantic embeddings were generated using DarkBERT sentence embeddings, selected through comparative evaluation (Appendix A.3). Dimensionality reduction and clustering used the following configuration: • UMAP: n components = 5, metric=cosine, fixed random state • HDBSCAN: metric=euclidean, min cluster size= 80, min samples= 90, prediction data= True Topic representation used class-based TF-IDF (c-TFIDF) with KeyBERT keyphrase refinement. Topic labels were generated using the Llama language model from representative keywords and sample documents, followed by manual validation (Appendix A.5). Snapshots were assigned probabilistic topic distributions using BERTopic soft assignment. For temporal analysis, snapshots were ordered chronologically at the website level and topic probabilities were aggregated within time intervals to compute the longitudinal prevalence and topic turnover measures used in the main analysis.

A.2.2. Website Selection. The dataset consists of archived HTML snapshots collected over multiple years. A website was defined as a stable page identity observed repeatedly across time. To construct longitudinal histories: • Snapshots were grouped by identical file path and page title • Each group was assigned a unique internal website identifier Websites with fewer than four snapshots were excluded, as they do not provide sufficient temporal coverage for trend analysis. The number of snapshots per website varies substantially (mean = 321.42, median = 10), reflecting differences in crawling frequency and site availability. As continuous uptime information is not available, website persistence is approximated through repeated observations across time. A.2.3. Snapshot Validity Criteria. Individual snapshots were retained only if they satisfied all of the following conditions: • A valid timestamp was available • HTML content could be successfully parsed • Extracted text length exceeded 50 characters Snapshots failing any condition were discarded to prevent noise from incomplete crawls or placeholder pages. A.2.4. Content-Based Filtering. After text extraction, additional filtering steps were applied: Language Restriction. Only English-language content was retained. Non-English snapshots were excluded to ensure semantic consistency in embedding and topic modeling. Duplicate Removal. Near-duplicate snapshots within the same website and time interval were removed. Two snapshots were considered duplicates if their cleaned textual content matched after normalization. Non-Informational Pages. Pages containing only navigation elements, login prompts, error pages, or empty marketplace listings were excluded. These pages typically lack meaningful discussion content and can bias topic distributions.

A.2. Dataset Filtering and Temporal Handling This section documents the criteria used to determine which websites, snapshots, and textual content were included in the analysis. The goal of the filtering process was to ensure longitudinal consistency while minimizing noise unrelated to cybercrime discussions.

A.2.5. Temporal Consistency Handling. Snapshots were ordered chronologically using timestamps. When multiple snapshots existed within the same time interval, they were aggregated during analysis rather than removed, ensuring that content updates were preserved without overweighting high-frequency crawls. The filtering criteria were designed to balance coverage and reliability:

A.2.1. Details on the Dataset. The dataset consists of archived HTML snapshots of dark web forums and marketplaces collected between 2020 and 2026 by CFLW Cyber Strategies, a third-party provider specializing in cyber threat intelligence. It includes textual content from publicly accessible pages such as forum discussions, marketplace listings, and related informational pages. In line

11

Minimum snapshot requirement ensures longitudinal validity • Language filtering ensures semantic comparability • Content filtering removes structural website noise These steps produce a dataset suitable for measuring thematic prevalence and turnover over time, while reducing artifacts caused by crawling behavior or nondiscussion pages.

documents as outliers, indicating more selective clustering. This behavior suggests improved embedding geometry where unrelated discussions are rejected rather than weakly merged.

TABLE 2: Model Comparison Across Parameters (Best result in bold) Model AttackBERT AttackBERT AttackBERT AttackBERT AttackBERT AttackBERT AttackBERT AttackBERT

A.3. Model Comparison Experiments This section documents the experimental comparison of candidate embedding models and clustering configurations used in the topic modeling pipeline. The objective was to select the configuration that maximized semantic coherence, minimized outlier assignments, and produced interpretable topic clusters suitable for longitudinal analysis.

DarkBERT DarkBERT DarkBERT DarkBERT DarkBERT DarkBERT DarkBERT DarkBERT

A.3.1. Candidate Embedding Models. The following embedding models were evaluated: • MiniLM (all-MiniLM-L6-v2) • DarkBERT • AttackBERT Each model was integrated into the BERTopic framework using identical dimensionality reduction and clustering procedures to ensure comparability.

MiniLM MiniLM MiniLM MiniLM MiniLM MiniLM MiniLM MiniLM

A.3.2. Evaluation Metrics. Model performance was evaluated using a combination of quantitative clustering metrics and qualitative interpretability assessment. • Topic Coherence (c v) — measures semantic similarity among top words in each topic. Higher values indicate more interpretable topics. • Number of Topics — constrained to remain below 100 to maintain analytical usability while preserving granularity. • Outlier Count — number of documents not assigned to any cluster. Higher values indicate stricter noise filtering but risk excluding meaningful content. • Cluster Stability — consistency of topic structure across repeated runs with different random seeds. • Cluster Size Distribution — minimum and average cluster size used to evaluate fragmentation vs overgeneralization. • Keyword Interpretability — manual inspection of representative keywords and documents.

Parameters Topics AttackBERT 500 50 2392 700 70 1994 900 90 1774 1200 120 1532 1500 150 1274 2000 200 820 2500 250 659 3000 300 568 DarkBERT 500 50 2323 700 70 2021 900 90 1791 1200 120 1596 1500 150 1407 2000 200 837 2500 250 695 3000 300 607 MiniLM 500 50 2272 2006 700 70 900 90 1785 1200 120 1602 1500 150 1468 2000 200 921 2500 250 707 3000 300 629

Outliers

Min

1021396 975141 901760 975868 963499 1235456 1391916 1507640

500 701 900 1200 1500 2001 2510 3004

1377521 975145 845592 855891 838806 1173336 1286506 1375889

501 700 900 1203 1500 2008 2503 3000

1062998 1101910 1006363 884327 937896 1141357 1377228 1526143

500 700 900 1200 1504 2000 2506 3000

A.3.4. Results. Qualitative inspection of resulting topics further confirmed that domain-specific embeddings produced more meaningful clusters. DarkBERT grouped infrastructure-related terms (e.g., tor, onion, server), marketplace terminology (e.g., product, paypal, western union), and community interaction terms (e.g., message, member, buyer protection) into coherent themes, whereas general-purpose embeddings produced clusters dominated by generic web vocabulary. These observations supported the final model selection.

A.4. Labelling Topics with Llama Llama was used for semantic structuring and annotation consistency rather than primary classification.

A.3.3. Embedding Model Comparison. We evaluated three embedding models (MiniLM, DarkBERT, and AttackBERT) under identical BERTopic configurations across three clustering strategies: • min cluster size = min samples • min samples = min cluster size - 10 • min samples = min cluster size + 10 Evaluation considered topic granularity (75–90 topics target), outlier filtering, and minimum cluster population. Table 2 report clustering outcomes under progressively stricter density constraints. The equal-density setting provides a baseline comparison, while decreasing min samples relaxes cluster acceptance and increasing it enforces stronger semantic separation. Across configurations, DarkBERT consistently produced topic counts within the desired analytical range and assigned more

A.4.1. Prompt Structure. All prompts followed a structured format consisting of a task definition, labeling rules, and an explicit output constraint. The model was instructed to generate concise topic labels directly from representative keywords. The exact prompt template used is shown below: Task: Generate a short topic label. Rules: - Output ONLY the label - Maximum 4 words - No punctuation - No explanation Topic words: <top words> Label:

12

Here, <top words> represents the list of representative keywords extracted for each topic using c-TF-IDF and KeyBERT.

Distribution of Topic Labels

Torrents and Files Forum Reputation Online Shopping Forum Features Infrastructure and Hosting

A.4.2. Generation Settings. Topic labels were generated using controlled sampling to balance consistency and flexibility. The generation parameters were: • Temperature: 0.1 • Sampling: enabled (do_sample=True) • Maximum new tokens: 12 • End-of-sequence token: default model EOS token A low temperature was used to ensure stable and reproducible outputs while still allowing minor variation when generating short labels. The combination of constrained prompting and limited token generation ensured that outputs remained concise and consistent across topics.

Transaction Protection Forum Security Counterfeit Money Topic Label

Databases Online Banking Politics and War Money Making Opportunities Vendor Channels and Promotions VPN Services WikiLeaks Forged Document Services Prepaid Cards Stolen Bank and Payment Accounts Browser Hijacking Vendor Sales & Shipping

0

2

4

6

8 10 Percentage of Corpus

12

14

16

Figure 10: Distribution of topic labels in corpus.

A.5. Human Validation of Topic Labels Annotators corrected these to better align with the actual topic content. Topic Merging. Multiple clusters were found to represent highly similar or identical concepts. Annotators identified such overlaps and merged them into a single, unified topic to avoid redundancy and improve analytical clarity.

To ensure that generated topic labels accurately reflected the underlying keyword clusters, a manual validation step was performed. Topic labels were initially generated using the large language model Llama-3.1-8B-Instruct, which produced candidate topic names based on representative keywords extracted for each cluster. All generated labels were then manually reviewed. For each topic, the Llama-generated label was evaluated against the corresponding keyword set to determine whether it accurately represented the underlying theme of the cluster. Labels that were judged to be representative were retained without modification. When a label did not adequately describe the keyword set or produced an incorrect interpretation of the topic, a corrected label was assigned manually based on the keywords and their semantic context.

A.5.3. Examples of Label Corrections and Topic Merges. Table 3 presents representative examples of how topic labels were refined during manual validation. A.5.4. Discussion of Refinements. The examples in Table 3 illustrate common issues in automatically generated topic labels. First, prompt artifacts and non-descriptive instructions were frequently present in the original labels, requiring cleaning and reinterpretation. Second, several labels were overly broad or ambiguous, necessitating semantic correction based on the associated keywords. Finally, a substantial number of topics exhibited high overlap, particularly in domains such as file sharing, financial transactions, and leaked data, and were therefore merged to reduce redundancy.

A.5.1. Validation Procedure. Two independent annotators reviewed all automatically generated topic labels produced by the topic modeling pipeline. Annotators assessed whether each label accurately represented the underlying documents and associated keywords, and performed corrections where necessary. The validation process consisted of three main actions: (1) cleaning and rewording labels to improve clarity and remove non-descriptive or extraneous text, (2) aligning labels with domain-relevant terminology, and (3) identifying and merging semantically overlapping topics. Annotators performed these steps independently before meeting to resolve disagreements and produce a final, consensusbased set of topic labels. The inter-code reliability score was calculated following similar procedure presented in works by Pfister et al. [65] and Ave et al. [66]: Cohen κ=0.96, which indicated strong agreeability.

A.5.5. Impact on Final Topic Set. The human validation process reduced noise, improved label clarity, and ensured that each topic corresponded to a distinct and meaningful concept. The resulting refined topic set forms the basis for all subsequent analyses of topic prevalence (RQ1) and lifecycle dynamics (RQ2).

A.6. Detailed Topic Definitions This appendix provides an overview of the topics identified in our analysis. Topics are grouped into higherlevel categories reflecting their functional role within the ecosystem (e.g., transactional, product-related, infrastructure, and community-oriented). For reference and clarity, each topic is assigned a unique identifier (e.g., T1, P3, I2, C4). Topics marked with ∗ correspond to the ten most prevalent topics in the corpus, while those marked with † fall within the top twenty. The distribution of the top 20 topic labels in the corpus is shown in Figure 10.

A.5.2. Types of Corrections. The manual validation resulted in three primary types of refinements: Label Cleaning and Rewording. Many automatically generated labels contained extraneous instructions, formatting artifacts, or overly verbose descriptions (e.g., “Here is the topic label: . . . ”). These were simplified into concise and interpretable labels. Semantic Correction. In some cases, labels did not accurately reflect the underlying keywords or documents.

13

TABLE 3: Examples of topic label corrections and merges after human validation Topic ID(s)

Llama Original Label(s)

Human Cleaned / Interpreted Label

Human Final Label

0 1, 31, 32, 38, 45, 52, 56, 57, 60, 62, 65, 70, 80 19 15, 22, 25

Membership Status Torrents and File Sizes; Torrent Files; Directories; Mobile Phone Torrents

Shop/Forum Membership Torrent File Listings / File Structures

Forum Reputation Torrents and Files

Transaction Support Transaction Protection; Vendor Protection; Cloned Credit Card Protection Hacking and Password Security Counterfeit Currency and Surveillance; Counterfeit Money Transfer; Counterfeit Money Exposed Web Directories; File Listings; Directory Identifiers Bitcoin; Bitcoin Transactions; Bitcoin Transaction

Scam Warnings Financial Protection Mechanisms

Escrow and Scam Protection Transaction Protection

Social Media Account Access Counterfeit Financial Activity

Social Media Hacking Counterfeit Money

Exposed File Systems

Exposed Website Directories

Bitcoin Payments

Bitcoin Transactions

27 12, 24, 82 43, 44, 48, 49, 50, 59, 69, 71, 75 46, 53, 67

TABLE 4: Transactional Topics

ID T1 T2 T3 T4 T5 T6 T7 T8 T9 T10 T11 T12

Topic ∗

Online Shopping Transaction Protection† Online Banking† Money Making Opportunities† Prepaid Cards† Vendor Sales & Shipping† Escrow and Scam Protection† PayPal Scams† Bitcoin Transactions Crypto Wallets Merchant Transactions Financial Tracing

Count

%

558,270 243,077 175,989 156,609 121,707 90,047 89,166 74,718 22,191 6,301 5,977 4,778

8.54 3.71 2.69 2.39 1.86 1.38 1.36 1.14 0.34 0.10 0.09 0.07

Description Buying and selling goods or services. Methods to secure or guarantee transactions. Banking access, transfers, and fraud. Methods for generating illicit income. Use and trade of prepaid payment cards. Logistics and delivery of goods. Third-party services to reduce fraud risk. Fraud schemes involving PayPal. Payments using Bitcoin. Management of cryptocurrency wallets. Payment processing for vendors. Tracking financial flows.

TABLE 5: Product-Related Topics

ID P1 P2 P3 P4 P5 P6 P7 P8 P9 P10

Topic ∗

Counterfeit Money Forged Document Services† Stolen Bank and Pmt. Accs.† Credit Card Data Social Media Hacking Hacked Social Media Accounts Ransomware Hacking Personal Data Accounts

Count

%

197,892 122,344 120,963 87,250 43,383 23,559 16,101 15,216 6,686 6,270

3.02 1.87 1.85 1.33 0.66 0.36 0.25 0.23 0.10 0.10

Description Production and sale of fake currency. Creation and sale of fake documents. Compromised financial accounts. Trade of stolen credit card data. Compromising social media accounts. Sale of hacked social accounts. Malware used for extortion. General unauthorized access techniques. Trade of personal information. Sale of compromised accounts.

TABLE 6: Infrastructure Topics

ID

Topic

Count

%

I1 I2 I3 I4 I5 I6 I7 I8 I9 I10 I11

Infrastructure and Hosting∗ Databases∗ VPN Services† Browser Hijacking† Exposed Website Directories Hidden Cont. and Access Links Rocksolid Newsreader Tor Search Engines Remote Access CAPTCHA Access Control and Bypass

461,151 192,833 145,070 92,553 66,696 38,594 38,475 33,629 5,018 4,363 2,015

7.05 2.95 2.22 1.42 1.02 0.59 0.59 0.51 0.08 0.07 0.03

14

Description Hosting services and server infrastructure. Storage and access to data collections. Privacy and anonymity via VPNs. Redirecting or controlling browsers. Listings of accessible or vulnerable sites. Access to restricted or hidden resources. Tool for accessing content feeds. Search tools for dark web content. Tools for remote system control. CAPTCHA solving or bypass methods. Circumventing access restrictions.

TABLE 7: Community Topics

ID C1 C2 C3 C4 C5 C6 C7 C8 C9 C10 C11 C12 C13 C14 C15 C16 C17 C18 C19 C20 C21

Topic ∗

Torrents and Files Forum Reputation∗ Forum Features∗ Forum Security∗ Vendor Channels and Promos.† WikiLeaks† Politics and War† Doxxing Location Information File Downloading Leaked Data Leaked Arch. File Tree Listings Leaked Data Archives Safety Concerns Question and Answer Banned Accounts Leaked Corporate Files Forum Credentials Password Security Finance and Operations Digital Payment Card Services

Count

%

1,026,851 903,883 524,089 204,509 150,399 127,343 159,594 49,577 25,867 23,972 19,650 9,357 7,766 4,630 2,594 2,281 1,732 977 3,014 16,096 1,605

15.71 13.83 8.02 3.13 2.30 1.95 2.44 0.76 0.40 0.37 0.30 0.14 0.12 0.07 0.04 0.03 0.03 0.01 0.05 0.25 0.02

15

Description File sharing via torrent systems. User trust, ratings, and credibility. Platform functionality and usability. Security practices within forums. Vendor advertising and communication. Sharing and discussion of leaked information. Discussions on geopolitical events. Publishing personal information. Sharing geographic details. General downloading practices. Discussion of leaked datasets. Structured listings of leaked files. Collections of leaked data. Risk and precaution discussions. General help and information exchange. Account restrictions and recovery. Corporate data breach discussions. Access credentials for forum accounts. Password protection and cracking. Operational coordination and planning. Services related to card payments.

Record · ID 196417 · SHA-256 29089c5f6d14c2e1
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.