Analyzing Concentration, Temporal Routines and Targeting in Public Ransomware Leak Site Data
arXiv:2605.24559v1 [cs.CR] 23 May 2026
Lea Müller[0009-0009-6538-8046] and York Yannikos[0009-0001-2751-5253] Fraunhofer Institute for Secure Information Technology (SIT) National Research Center for Applied Cybersecurity (ATHENE), Rheinstr. 75, 64295 Darmstadt, Germany {lea.mueller,york.yannikos}@sit.fraunhofer.de
Abstract. Ransomware has grown to become one of the most damaging types of cybercrime, affecting private and public organizations in any sector. While early types of ransomware targeted many victims via automated attacks, ransomware groups have started to specifically target organizations and companies in the expectation of receiving larger ransoms. To increase the pressure on victims, most groups host so-called data leak sites, where information about their victims is made public. The shift towards ‘human-operated’ ransomware together with easily accessible behavioral traces available from data leak sites makes research investigating operational regularities of ransomware groups of interest. Using leak site posts as behavioral traces of ransomware groups, we created a dataset consisting of over 27,000 posts from 325 groups. Based on this dataset, we analyzed victim concentration, temporal routines and targeting regularities. Our findings suggest that groups do not behave entirely random. Instead, the observable traces found on leak sites show concentration of activity, temporal routines and selective patterns. Keywords: Ransomware · Data leak site · Operational regularities.
1
Introduction
Ransomware has grown to become one of the most financially damaging types of cybercrime over the past decade, with ransom payments amounting to USD 820 million in 2025 alone [4]; a figure that does not even take into account the damage incurred through loss of business or reputation. Affecting private and public organizations in any sector [19], ransomware is described as a key threat by both industry [26] and public authorities [5,10]. For example, according to the European Union Agency for Cybersecurity (ENISA) [5], ransomware attacks and data breaches account for over 95% of cybercriminal activities affecting EU organizations, with the latter directly resulting from the former. Europol describes ransomware as a key threat in the European Union and observes a steady increase in attacks [6]. While the majority of ransomware attacks affect small and medium-sized enterprises, critical infrastructure and governments have also suffered significant incidents in the past [29].
2
L. Müller, Y. Yannikos
Ransomware is a type of malware that encrypts or otherwise blocks access to systems or data [2,19] in order to extort payment from a victim [19]. In the past, ransomware targeted many victims via automated attacks [18], demanding relatively low ransom amounts. With the rise of ‘big game hunting’ and Ransomware-as-a-Service (RaaS), ransomware groups started to more specifically target high-value organizations. With RaaS, cybercriminals can obtain ransomware from operators to execute attacks, eliminating the need to create their own malware. This enables criminals with minimal programming skills to orchestrate ransomware attacks and earn money through extortion [15]. RaaS structures often involve a core group developing the malware and coordinating operations, and so-called affiliates performing attacks and sharing profit with the core group [29]. ‘Big game hunting’ describes targeted ransomware attacks, whereby groups target large companies and organizations rather than individuals in the expectation of receiving larger payments [2,14,15,19]. Since targeting specific, high-prospect organizations has increased the resources cybercriminals invest, a shift in extortion tactics has been observed. From 2019 onward, ransomware groups have started to host so-called data leak sites to increase pressure on their victims [8,22]. In 2020, this procedure became a trend, with more and more ransomware groups using data leak sites as an additional lever of extortion [8]. Data leak sites are dark web platforms operated by ransomware groups. These sites are used to announce new victims and often include a deadline for them to pay a ransom to prevent their data from being published. While early forms of ransomware simply encrypted a victim’s data and demanded a ransom for decryption, extortion now focuses on coercing victims to pay for their data to not be released publicly [6]. The trend of targeting specific organizations has led perpetrators to manually perform attacks instead of using autonomous malware [19,25]. This shift towards ‘human-operated’ ransomware – in contrast to earlier ransomware families spreading autonomously [16] – has made research investigating operational regularities of ransomware groups of interest. If attacks are performed by humans, routines or targeting preferences may be observed by analyzing ransomware groups’ behavior. While the proliferation of big game hunting has shifted ransomware groups’ focus on companies and organizations, little is known about other factors that may influence temporal routines or targeting preferences, such as working conditions or victim characteristics. Most research in the field of ransomware focuses on its evolution (e.g., [18]) or approaches to detection and defense (e.g., [2,19,25]); however, ransomware groups’ operational behavior, like temporal routines or targeting, is rarely analyzed. Existing research often has a narrow focus, with analyses concentrating on either a select set of nations or a select set of ransomware groups. Whelan et al. [30] analyze ransomware attacks in Australia, Canada, New Zealand and UK from 2020 to 2022. Their dataset includes 865 ransomware incidents across the four nations. Data was obtained from a security company that collects information from ransomware groups’ leak sites. They find that different ransomware groups target different countries with varying prevalence, suggesting targeting
Analyzing Ransomware Leak Site Data
3
preferences. They also observe that some sectors are targeted more often, with the sector Industrials being the most prevalently attacked while other sectors, such as education or energy, experience fewer attacks. Kim et al. [12] analyze 20 ransomware groups active in the Arab world, based on 226 incidents disclosed by these groups. Their analysis focuses on groups with a high level of activity in 22 Arab countries, examining incidents reported between 2020 and 2023. In terms of temporal routines, the authors find that incidents do not appear to be connected to specific times or seasons. They also find that the main sectors targeted by ransomware groups are commercial facilities and critical manufacturing. Phipps and Nurse [21] analyze the three ransomware groups Conti, LockBit and BlackCat/ALPHV, based on articles obtained through web search. Their analysis focuses on groups’ ‘origins, structure, organisation, dynamics and nature’, arguing that these areas are key characteristics of ransomware groups. In terms of global targeting, they observe opportunistic behavior, with no specific focus on or exclusion of certain sectors or nations; with an exception of nations linked to the former Soviet Union. Prior studies investigating the targeting behavior of ransomware groups often have a narrow scope, either by concentrating on a select set of ransomware groups (e.g., [21]) or a select set of nations (e.g., [12,30]). Additionally, sectors or nations are often analyzed based on raw incident counts (e.g., [1,12,30]) which neglects that incident counts may be dependent on factors such as a sector’s or nation’s size or economic scale. For example, the United States are consistently named as the nation most affected by ransomware attacks, but are also the nation with the highest economic scale worldwide [32]. To address the gaps identified in prior research, we investigate a large dataset of ransomware leak posts without deliberately excluding certain ransomware groups or nations. Since most ransomware groups operate data leak sites, which are inherently intended as broadcast channels to reach a wide audience and increase pressure on victims, data about ransomware groups’ posting behavior can easily be collected. Treating leak site posts as behavioral traces, this study aims to analyze whether ransomware groups exhibit behavioral regularities across organizational, temporal, geographic and sectoral dimensions. To this end, we create a dataset consisting of over 27,000 leak site posts from 325 groups by aggregating information of two open-source ransomware monitoring tools, RansomLook [23] and Ransomware.live [24]. We then provide an exploratory empirical characterization of ransomware groups’ leak site posting behavior using this dataset. We analyze victim concentration across groups, temporal routines as well as target selection. The objective of this article is to understand whether ransomware groups employ opportunistic behavior, similar to early-day ransomware using automatic targeting, or whether groups show non-random characteristics in their operations. Our analyses are structured around four types of findings: 1) Ecosystem structure, analyzing the concentration of activity across ransomware groups. 2) Temporal routines, analyzing whether temporal patterns can be observed in leak post behavior. 3) Geographic target-selection regularities, analyzing whether
4
L. Müller, Y. Yannikos
certain nations are over-/underrepresented relative to their economic exposure. 4) Sectoral target-selection regularities, analyzing whether certain sectors are over-/underrepresented relative to a market-sector baseline. We make the following contributions: – Analysis of a curated dataset compiled from two open-source ransomware monitoring tools. Our dataset includes 27,629 leak posts from 325 ransomware groups. – Based on analysis of victim concentration, we offer evidence for strong concentration of observable activity on only a few highly active groups. – Based on analysis of temporal routines, we offer evidence for concentration of posting activity on weekdays. Analysis also indicates a seasonal pattern with higher activity in Q4 and lower activity in January. – We compare geographic and sectoral incident distributions with a baseline of economic scale to identify nations and sectors that are over- or underrepresented in ransomware attacks. Overall, while ransomware groups’ behavior is often observed to be unstable [6], our findings suggest that groups do not behave entirely random. Instead, the observable traces found on data leak sites indicate concentration of activity, temporal routines and selective patterns. The rest of this paper is structured as follows: Section 2 presents the methodology, including data collection, curation of our leak post dataset, and the analytical approaches employed. Section 3 presents the results, structured around activity concentration, temporal routines, and target-selection regularities. Section 4 discusses our observations and the work’s limitations.
2
Methods
2.1
Data Collection and Data Curation
Data was collected from two open-source ransomware monitoring tools, RansomLook [23] and Ransomware.live [24]. Both services track ransomware groups’ posts and activities by monitoring leak sites, providing aggregated information about various active and inactive groups. Additionally, some of the entries on Ransomware.live list the organization’s nationality and sector. The first victims listed by RansomLook are dated in 2020. The first victims listed by Ransomware.live are dated in 2013. From both sources, we retrieved data on all active and inactive groups with at least one leak site post. At the time of data collection (March 4, 2026) and after removal of duplicate entries, RansomLook listed 25,761 leak posts from 257 groups and Ransomware.live listed 25,656 leak posts from 298 groups. For each post, the collected data includes: timestamp, name of the ransomware group, post title (usually the name of the organization), nationality of the organization where available, sector of the organization where available. In a first step, for the purpose of merging the data from the two aforementioned services, ransomware groups’ names were brought into a standardized
Analyzing Ransomware Leak Site Data
5
format. As some groups were listed with different names by the two services, names were compared and standardized to merge the data in the next step. After standardizing group names, 325 unique groups remained from both services. In a next step, the two datasets were merged. To be able to merge the two datasets of over 25,000 entries each, we compared the tuples of organization’s name and ransomware group’s name in both datasets automatically. Since slight variations in organizations’ names were observed between the two datasets, we normalized names in a first step and computed similarity between two names in a second step. The similarity of two names was computed as an equation of 2m t , with m being the number of matches and t being the total number of elements in both sequences. This results in a similarity score between 0 (= no similarity) and 1 (= identical sequences). The normalization procedure was improved iteratively by inspecting entries manually and identifying causes of false positives (i.e., two entries that should not be merged receiving a high similarity score) and false negatives (i.e., two entries that should be merged receiving a low similarity score). For example, some ransomware groups add prefixes like ‘full data leak of’ to a leak post, which can have significant but irrelevant impact on the similarity score. Iterative refinement of the normalization procedure allowed us to catch such instances and thus reduce the number of errors. Other normalization steps included turning all characters to lowercase, stripping white spaces, removing prefixes and suffixes, or unescaping HTML entities. The two datasets were then merged based on a comparison of normalized organization names. Two entries were treated as a match if the similarity score was ≥ 0.8. Entries with a similarity score < 0.8 were treated as disparate entries. Matches with a similarity score ≥ 0.8 and < 1.0 were checked manually and only if the entries were identical, the match was accepted. Matches with a similarity score of 1.0 were accepted automatically and no manual check was done for these instances. Merging the two datasets resulted in a total of 28,127 entries from 325 groups. In a last step, the entire dataset was checked manually to remove entries not containing incident information. For example, some entries were announcements to inform site visitors about new channels or collaborations. 498 entries were removed, resulting in a total of 27,629 leak posts. Additionally, we observed 791 instances where victims’ names had been anonymized. In most cases, these entries did not provide enough information to identify an organization’s nationality or sector. For the purpose of analyzing geographic preferences in targeting, we employed three rule-based steps to identify organizations’ nationality. In the first step, where available, we copied the information about an organization’s nationality as provided by Ransomware.live. In the second step, we identified the nationality from the top-level domain (TLD) for the entries for which a domain was given. This was only possible for leak posts where the ransomware group provided the victim’s domain. In the third step, we implemented a keyword-based approach to identify leak posts that contained the organization’s nationality in their title. Employing these three steps, nationality was identified for 18,041 victims (65.30% of victims). It should be noted that classification of nationality
6
L. Müller, Y. Yannikos
was not achieved for all victims. Firstly, researching over 9,000 organizations manually is not feasible for the purpose of this study. Secondly, an organization may be international and classification into one nationality not possible. Additionally, 791 of the incidents were anonymized, with the majority of anonymized posts not providing enough information to be able to identify the organization or information about its nationality. For the purpose of analyzing sectoral preferences in targeting, we first defined a number of sectors to be part of our analysis. We used the 11 Global Industry Classification Standard (GICS) sectors, expanded to include government and education as additional sectors. The GICS is an industry classification framework developed by MSCI [17] and S&P Dow Jones Indices [27] and classifies organizations into 11 sectors. Government and education institutions, however, are not covered by this classification. After having defined a set of sectors, we employed two rule-based steps to identify the sector of victim organizations. In the first step, where available, we mapped the sector information provided by Ransomware.live to the 13 sectors selected for our analysis. For the remaining entries, classification was realized via a keyword-based approach. Keywords were defined for the 13 sectors. If the name of an organization contained one of the keywords, it was classified into the respective sector. Sector information was identified for 16,868 victims (61.05% of victims). It should be noted that classification of the sector was not achieved for all victims. Firstly, researching over 10,000 organizations manually is not feasible for the purpose of this study. Secondly, over 791 organizations’ names were anonymized, with the majority of anonymized posts not providing enough information to be able to identify the organization or the sector it operates in. 2.2
Analyses
Based on the dataset of leak site posts, we employed explorative data analyses to understand activity concentration across groups, temporal behavior, and targetselection regularities. Victim Concentration across Groups To analyze victim concentration across groups, we calculated the number of leak site posts per group. To assess concentration, we calculated the concentration ratio CRk for k = 0.01n, 0.05n, 0.10n, 0.20n, with n = 325 being the number of ransomware groups in our dataset. I.e., we calculated the concentration ratio for the largest 1%, 5%, 10%, and 20% of groups. Temporal Behavior We analyzed the distribution of leak site posts over weekdays by calculating weekend-weekday ratio based on raw leak post counts per weekday. We also calculated daily averages of z-standardized ransomware incident counts to identify days with activities above or below average. Z-standardized values were used to control for long-term trends. It should be noted that we only
Analyzing Ransomware Leak Site Data
7
considered incidents between January 2018 and 04 March 2026. Earlier incidents were excluded due to rare occurrence of leak site posts. We analyzed the distribution of leak site posts over months using annual z-standardized values to control for long-term trends. Due to rare occurrence of leak site posts prior to 2019, we only considered posts between January 2019 and February 2026. Geographic Regularities To identify geographic regularities, we first calculated each nation’s relative number of ransomware attacks. It should be noted that relative numbers were calculated based on the total number of organizations for which we were able to identify the nationality (see Section 2.1). This means that we set the figure of 18,041 incidents as 100% in order to calculate each nation’s share, so as to avoid artificially generating lower relative numbers. Since raw percentages fail to capture nations’ economic scale, we calculated deviations of observed incident distribution from an economic baseline. Deviations were calculated for all nations that account for a cumulative 95% of all ransomware incidents (52 nations). As a baseline, we used gross domestic product (GDP) as a proxy for nations’ economic scale. Nations’ GDP in 2024 was sourced from [32]. Since Taiwan is not listed in this database, GDP for Taiwan in 2024 was sourced from [9]. Deviations from the baseline indicate whether a nation is over- or underrepresented relative to its economic weight. Sectoral Regularities To identify sectoral regularities, we first calculated each sector’s relative number of ransomware attacks. Similar to the analysis of geographic regularities, we set the figure of 16,868 incidents as 100% (see Section 2.1). Since raw sectoral distributions may lack informative value, we calculated the deviation of ransomware incident distribution from a baseline representing each sector’s economic scale. The sectoral baseline was derived from data about value added to GDP per industry, as published by the U.S. Bureau of Economic Analysis (BEA) [28]. The industries listed by BEA were mapped to the 13 sectors (eleven GICS sectors, expanded to include government and education) and value added was calculated as a sum of the mapped industries. As the sectoral baseline, we used relative value added by each sector as a proxy for economic scale. We then calculated the deviation of each sector’s share of ransomware incidents from the distribution in the sectoral baseline.
3
Results
3.1
Victim Concentration across Groups
The number of leak site posts ranged from 1 for the smallest groups to 3,038 for the biggest group, with a mean of 85.01 leak posts per group. Figure 1 shows the Lorenz curve of the victim concentration across ransomware groups, indicating a substantial concentration in the distribution of incidents. We report the concentration ratio for the largest 1%, 5%, 10%, and 20% of groups.
8
L. Müller, Y. Yannikos
CR0.01n = CR3 = 21.29%; CR0.05n = CR16 = 51.84%; CR0.10n = CR33 = 67.45%; CR0.20n = CR65 = 82.66%. Additionally, we report the share of groups needed to explain 50%, 75%, 90%, and 95% of attacks. The top 15 groups (4.62% of groups) account for over 50% of attacks; the top 46 groups (14.15%) account for 75% of attacks; the top 97 groups (29.85%) account for 90% of attacks; and the top 135 groups (41.54%) account for 95% of attacks. This illustrates that the bottom 190 groups (58.46%) account for less than 5% of all observed ransomware incidents.