Identifying Security Platform Product Abuse with Machine Learning Shaefer Drew∗ , Michael Brautbar∗ , Paul Knight† , Edward Raff∗ , Lana Peric-McDermott‡ , Simran Sarin∗ , Nickolas Machado§ , Hanna Albright∗ , Vitaly Zaytsev∗ ∗ USA
† Canada
‡ Ireland
§ Brazil
arXiv:2609.21303v1 [cs.CR] 18 Sep 2026
CrowdStrike {shaefer.drew, michael.brautbar, paul.knight, edward.raff, lana.pericmcdermott, simran.sarin, nickolas.machado, hanna.albright, vitaly.zaytsev}@crowdstrike.com telemetry and provide suspicious abuse events with more context. Furthermore, it adds explainability to the framework to guide investigators and support their investigation. We contribute a study of the system’s design process, results, and validation, to help bridge the gap between industry needs and data mining and storage systems. Our goal is to help advance the science of security and the evolving needs in data systems to support this kind of work, which they were often not originally designed for. The process begins by engineering features from multiple data sources, including Browser Fingerprint (BFP) data, interesting events (which we will define), IP context data, firmographic data, and security tickets. These features then pass through an anomaly-detection algorithm to both filter out normal instances and calculate an anomaly score for downstream risk scoring. Afterward, features from each data source are ranked and ingested into a data-source risk-scoring equation to calculate each data source’s risk score. These, along with the anomaly score, are inputs to a final risk-scoring equation that calculates I. I NTRODUCTION the platform events’ final risk. This final risk score has its Some threat actors have exploited enterprise security plat- weights tuned adaptively using Bayesian optimization [2], [3]. forms as attack vectors. An example is a threat actor gaining Next, a threshold is determined, above which abuse leads are Single Sign-On (SSO) access to a customer account via a sent to the product security investigators. In addition to the phishing campaign, then using it to log in to the Software-as-a- final score, four (4) levels of explainability are implemented Service (SaaS) user interface (UI) and disable specific security to point investigators in the right direction. This involves detections. This poses a significant risk to the security platform sending data source risk scores and attributing risk factors and its customers, particularly since the platform is designed to predictions, using SHapley Additive exPlanations (SHAP) to protect them. To protect customers and the product, it is to explain local feature contributions for anomaly predictions important to detect all forms of product abuse. [4], and highlighting the overlapping interesting events for the Most known existing solutions to security product abuse rely user session. on rules-based alerts. Examples could be "Platform Access Finally, the model is evaluated based on known product from a non-standard browser". In our customer base, these rules abuse incident coverage and alert volume. Our machine learning haven’t generated many high-efficacy leads, often producing framework can increase product abuse coverage by 35% while too many false positives with minimal coverage of real product reducing alert volume by 30%. From these results, we conclude abuse. This leads to alert fatigue among investigators assigned that this framework is a superior solution to baseline rule-based to these leads [1]. These alert rules are often siloed to single methods, as it detects more product abuse while reducing data sources and lack statistical rigor. analyst fatigue. As a solution, CrowdStrike has developed a machine-learning II. R ELATED W ORK framework to detect product abuse on security platforms and generate leads for product-abuse investigators. This solution The tools needed for system administrators and investigators uses anomaly detection and adaptive risk scoring, and is fit to to perform cybersecurity defensive operations are almost always multiple data sources to take full advantage of security platform “dual use”, in that they can be leveraged for both good
Abstract—Product abuse is an individually rare, but growing, problem across the SaaS industry. Highly sophisticated threat actors can misuse security platforms within customer environments or conduct bypass experiments on the product itself. Threat actors can leverage living-off-the-land (LOTL) attacks to avoid using cumbersome, frequently detected malware. Remediating this threat requires collecting multiple data modalities across different types of databases, addressing a cold-start problem in the intrinsic rarity of such sophisticated but dangerous events, and designing within the constraints of real-world deployment (e.g., cost, user behavior, performance, etc). To wit, we provide the first study of such a whole-system defense, especially with respect to a deployed and operational capability. Our results show an increase in product abuse coverage by 35%, a 30% reduction in monthly alerts, and adaptability to changes in malicious actors’ behavior. We review both the constraints we considered in designing the system to meet operational requirements and a retrospective evaluation of the value of explainable features and counterfactual performance on previously identified attacks. Index Terms—product abuse, anomaly detection, risk scoring, machine learning, cybersecurity, explainability
Data Source Description (defense) and bad (cyber attacks) purposes [5]–[7]. In this work, we describe, in as much detail as possible without Browser Fingerprint (BFP) Browser Device Fingerprints and attributes Interesting Events Platform Events that have potential for abuse revealing sensitive security details, how we developed a system IP Context IP behaviors and associations for detecting threat actors attempting to abuse CrowdStrike Firmographic Customer Account Data products in malicious ways (e.g., searching for vulnerabilities Overlapping threat hunting, MDR, Security Tickets and incident response tickets in CrowdStrike systems, or testing intrusion techniques to TABLE I confirm operational success). All TPs found through these I NPUT DATA S OURCES AND F EATURES . E ACH DATA SOURCE IS USED TO methods have been remediated. There is minimal academic CALCULATE FEATURES USED AS INPUT TO ANOMALY DETECTION AND RISK literature on this topic at all, let alone from real-world industrial SCORING . N OTABLY, EACH COMES FROM A DIFFERENT KIND OF DATA STORE AND EVOLUTION PERIOD , PROVIDING LONG - SCALE MACRO ( E . G ., deployment. FIRMOGRAPHIC ) AND SHORT- SCALE ( E . G ., IP CONTEXT ) INFORMATION . There are some different solutions out there for product abuse in adjacent but unrelated fields, such as Google Apigee abuse detection [8], which detects API abuse using machine learning but doesn’t address the complexities and data sources platform interactions. Each platform interaction collects an involved in typical EDR product abuse. Our detection has a event and a browser fingerprint for that device. To reduce the broader scope, covering the whole EDR platform (UI, API, event space, interactions are filtered to only those with the and Real-Time Response) using a wider range of data sources potential for product abuse. For example, dismissing endpoint and techniques tailored to EDR-specific abuse patterns that the detections as an event could enable product abuse, as threat API-focused Apigee solution cannot adequately handle. actors could use it to evade defenses. These interesting events Non-industry work has been published on sub-components are further enriched with IP context, such as VPN information, of this problem, but does not consider the aforementioned geography, etc. Customer firmographic information is also whole-product scope. Living-off-the-land (LOTL) work often joined, detailing high-level information about the customer focuses on limited and synthetic data due to the intrinsic such as employee count. Finally, our internal ticketing system, difficulty of “acquiring” the scripts used by attackers [9]–[12]. consisting of threat hunting leads, managed services, and Another avenue to produce abuse (though not the only method) response tickets, are joined in a way to indicate overlap with the is “account takeover”. This occurs when a malicious actor event itself for that user. For example, if [email protected] manages to phish or otherwise obtain control of a pre-existing is currently under investigation with an open ticket and she legitimate account. The methods of how that is technically is seen dismissing endpoint detections, this will increase the achieved have also been studied in many domains like finance, suspicious signal. All of these data sources are combined to e-commerce, and more [13]–[19]. Because defense-in-depth provide valuable context for the user and their interactions on is necessary in modern security systems, our study focuses the security platform. on the post-acquisition stage of malicious account acquisition, where multiple systems’ data must be integrated to produce A. Features For each data source, features and risk indicators were an effective, viable solution. The first approach attempted in real-world cybersecurity is engineered to best discover anomalies and malicious product still having experts write rule-based detection systems. While abuse signals. Because this system is live and used to stop real basic rule-based detection systems can identify simple abuse attacks, we do not specify all feature detail to avoid enabling patterns, they fail to detect sophisticated attacks that mimic le- attackers to circumvent it. For Browser Fingerprint data, this concerned creating gitimate administrative actions, and have long been recognized for producing too many false positives [20]. The rule-based customer and user baselines based on historical activity to foundations are still widely used even in modern ML systems, differentiate normal from unusual behavior. Inspired by prior be it using ML to help find new rule candidates [21], [22], user authentication research [28], [29], historical device data building reports for investigators [23], [24], or as a component was used to estimate, for each device attribute and user, the in a larger ML system [25]–[27]. Our method addresses the probability that the attribute belongs to that device, producing need for an intelligent system that can learn normal behavioral a value between 0 and 1: closer to 0 indicates the attribute is patterns across multiple dimensions, identify subtle deviations uncommon for that user, closer to 1 indicates it is frequently that indicate abuse without relying on predefined rules or observed. For example, a user who has only used Chrome will static thresholds, and explain those deviations to guide analyst score high on a Chrome event, but low the first time they use Firefox. investigation. Baselines are created for users by calculating the probability III. DATA of a device attribute belonging to that user P (A = x | u) = N (A=x,u)+α , where N (A = x, u) is the number of Browser In order to detect platform product abuse, we looked towards N (u)+β the telemetry of security data available on CrowdStrike’s Falcon FingerPrint (BFP) events with attribute A = x for user u, and platform. For product abuse detection, 5 different data sources N (u) is the total number of BFP events for user u. A for user u) were chosen to capture user signal and event context. Table To avoid zero probabilities, α = (#new smooths (#BFP events) I describes these input data sources. The process starts with based on how often entirely new attributes appear for a user,
while β = 1 adds a constant to bound the probability between (blue) for each of these input data sources. In step 3, an anomaly 0 and 1 and prevent infinite values. detection model (Isolation Forest [30]) (yellow) ingests these Beyond BFP, interesting events were another key data source features and flags the top anomalous events according to a for the product abuse modeling pipeline. Not only does it filter defined contamination ratio threshold. A contamination ratio down the important events, but it also adds many valuable is the proportion of outliers suspected to "contaminate" the features. As discussed before, interesting events represent data set. For example, a contamination ratio of 0.05 will mark events with potential for product abuse if used by malicious 5% of the data set as anomalies. In step 4 (red and green), actors. Occurrences are converted into count vectors as features. these filtered proportions of anomalous events (red) are sent to IP context features mainly consist of categorical features, step 5, while the normal events (green) are dropped. In step 5, such as the VPN in use. These features provide a granular data source risk scores (dark blue) are computed using features interpretation of the IP and its associations collected from the from their respective data sources (dotted lines connecting risk open internet. Features are categorical (VPN name), continuous, scores to their sources). The anomaly score (gray) is also stored and binary. These features help collect signal towards the IP’s in step 5 from the anomaly detector (yellow). Both the anomaly reputation. Firmographic features concern customer account score and the data source scores are used as input to the step 6 information. These fields create features themselves but are risk score function (brown). The risk score function calculates also combined with other data sources to engineer new features. a weighted risk score from all scores in step 5 and outputs Security ticket features involve features indicating whether or a final risk score in step 7 (orange). The highest risk scores not there is an overlapping ticket for that user and event as above a threshold are sent to product security investigators. well as specific attributions, such as the threat actor seen on the ticket. Joining all of these sources enabled substantial context and adjacent signaling to be added to the product abuse model. B. Labels Due to the rarity of product abuse, labels for the historical corpus were sourced using multiple methods. Below outlines the labeling sources used to mark malicious product abuse events: 1) Existing Rules-Based True Positives - Using Product Security’s alert API, true positive product abuse events that the existing rules caught were labeled as malicious and joined into the dataset. The high and critical alerts were used to form a baseline to measure against. 1. Modeling Diagram. Outlines the ML-based framework of how anomaly 2) Threat Hunting Product Abuse Tickets - There also Fig. detection and risk scoring are combined to generate product abuse leads exist threat hunting tickets specifically for historical product abuse. We can label certain users as malicious based on these tickets. 3) Emulations - The Product Security team ran a handful A. Anomaly Detection of red teaming exercises on internal accounts, conducting The first modeling portion of the framework involves numerous types of product abuse. These samples were anomaly detection, which acts as both a pre-filter and an labeled and joined into the dataset. input into the final risk score. We use an Isolation Forest, an 4) Manual Labels - Manual labels came from manual unsupervised algorithm that randomly splits trees on features investigations that didn’t start as an automated alert, recursively until leaf purity or maximum path length is reached. where Product Security investigators identified product Anomaly predictions are those with fewer splits, reflecting a abuse against the platform. These instances were used to statistical difference from the expected number of splits under label more malicious product abuse events in the dataset. a random null model [30]. Even with the labels collected using the methods above, For this model, a contamination ratio r was pre-selected there was still an extreme imbalance of actual product abuse. based on downstream classification results on a random sample Therefore, unsupervised learning methods were preferred. of the data. Recall was plotted in Figure 2 for different IV. M ETHOD values of r, showing the trade-off between the proportion Figure 1 shows the modeling diagram for the product abuse of events to label as anomalous vs. the proportion of malicious model that generates the final risk score. The process starts with events captured. All anomaly predictions were filtered through platform events for step 1 (purple). Tied to these platform events the pipeline; all normal predictions were filtered out. The are Browser Fingerprint data, potentially abusive interesting anomaly score for all anomalous instances was stored for platform events, IP context data, firmographic data, and further usage. This significantly filtered down the event space correlated security tickets. Features are engineered in step 2 and had downstream effects on the final risk score.
Recall (%)
100
TABLE II R ISK FACTORS AND R ANKINGS . E XAMPLES OF RISK FACTOR FEATURES AND THEIR RANKINGS PROVIDED BY SUBJECT MATTER EXPERTS . T HESE ARE USED AS INPUT TO DATA SOURCE RISK SCORES . V ERY /E XTREMELY BAD RANKS CONTRIBUTE TO HIGHER MARGINAL RISK INCREASES THAN SLIGHTLY BAD RANKS .
80 60 40 0
0.1 0.15 5 · 10−2 Contamination Ratio (r)
0.2
Risk Factor Suspicious VPN Event 12 (creating a new user) Threat Actor Association
Data Source IP Context Interesting Events Security Tickets
Rank Very Bad Slightly Bad Extremely Bad
1) Strict Hierarchy: Maintains EB >VB >B >SB for bad factors, G >SG for good factors 2) Dominance Principle: Highest present tier dominates the calculation, with lower tiers providing marginal contributions. 3) Bounded Output: Final score always between 0 and 1 4) Monotonically Increasing with respect to negative B. Adaptive Risk Scoring factors Monotonically Decreasing with respect to good factors 5) The next part of the pipeline involved adaptive risk scoring. 6) Concave Growth Curve: Each factor demonstrates Anomaly detection alone struggles to separate malicious from diminishing returns through exponential decay benign. For example, a rare enterprise VPN may raise the Asymptotic Approach: Factor contributions approach 7) anomaly score despite being a benign signal. Since this problem their maximum limits as counts increase lacks a large labeled corpus, risk scoring combines anomalous 8) Tunable Sensitivity: Alpha parameter adjusts how signals with suspicious signals, beginning with data source quickly factors reach maximum effect risk scores that reflect weighted risk factors. These are then 9) Diminishing Returns: First instance of each factor has ingested by a final risk function weighting the anomaly score the highest impact, with decreasing marginal impact for and data-source scores, with scores above a threshold sent as additional instances leads for analyst investigation. 10) Base Risk Adjustment: Starts with a configurable 1) Data Source Risk Scores: The following equation shows baseline that can be tuned to application context how data-source-level risk scores are calculated using a Equation 1 shows how risk scores combine contributions weighted, tiered dominance approach with exponential decay across risk factor tiers. Since the lack of data makes a pure and are bounded between 0 and 1 using a clipping function machine learning approach difficult, we instead work with (property 3). The risk equation takes a tiered dominance domain experts to define informed ordinal categories with approach (property 2), adding domain-tier and lower-tier weights set by hyper-parameter tuning. This proved effective contributions and subtracting good contributions from the base when merged this with the more data-heavy unsupervised score. A clipping function is used rather than an alternative bounded transformation such as sigmoid to preserve linear relaapproach. To create the input for the data source risk scores, Product tionships between risk factors and the final score, maintaining Security investigators ranked the importance of risk features interpretability and the additive nature of the contributions. The (risk factors) on a scale from Somewhat Good (SG) to baseline is configurable; however, it is kept constant for this Extremely Bad (EB). These ordinal rankings are based on research (property 10). domain expert guidance and reflect their expertise. "Bad" R = max(0, min(1, b + D + L − G)) (1) factors increase the data source risk score while "Good" factors decrease the data source risk score. An example could be The ordinal factor counts, in decreasing severity, are: EB "Suspicious VPN" falling under "Very Bad". This represents (Extremely Bad), VB (Very Bad), B (Bad), SB (Somewhat feature categorization via analyst annotation. It is completely Bad), G (Good), and SG (Somewhat Good). Each tier has separate from the target variable labels. a corresponding dominant-tier weight W(·) (e.g., WEB ) and, L Table II shows how certain features can be ranked as risk where applicable, a lower-tier weight W(·) (e.g., WVLB ). The factors. Each factor is one-hot boolean encoded with a 0 or a remaining parameters are α (sensitivity/decay rate) and b ≥ 0 1 depending on if the factor was hit for a given event. (base score). Given these ranked factors, an equation was needed to Equation 2 represents the dominant tier contribution and combine risk scores in a meaningful way, satisfying specific Equation 3 the lower-tier marginal contributions. Let the bad properties to be viable for deployment and interpretability by tiers be ordered t1 > t2 > t3 > t4 corresponding to (EB, VB, maintainers and investigators. Below are the properties used B, SB), and let k ∗ = min{k : tk > 0} denote the index of the to derive the data source risk equation: highest-severity tier present. Each tier’s contribution follows an Fig. 2. Recall vs Contamination Ratio (what fraction of the dataset we choose to mark as anomalous, as ranked by our model). Based on historical data we can acheive 100% recall considering only 20% of potential alerts, and still highly effective 40% at ≤ 2.5%, giving us an effective means to balance analyst availability and “hunt” for such rare but dangerous events.
exponential decay (1−e−α·tk ) that is monotonically increasing with diminishing returns (properties 4,6,8,9), bounded at 1 as tk → ∞. The dominant tier receives weight Wtk∗ while all lower tiers receive marginal weights WtLk , maintaining strict hierarchy (property 1) and dominance (property 2). The sensitivity parameter α was kept constant across all datasets. D = Wtk∗ · (1 − e−α·tk∗ )
L=
X
(2)
WtLk · (1 − e−α·tk )
(3)
k>k∗
Equation 4 is a weighted equation that combines all the "good" risk factors, in the "Good" tier. It is monotonically increasing with a strict hierarchy of G >SG (properties 1, 5). G = WG · (1 − e−αG ) + WSG · (1 − e−αSG )
(4)
The overall risk score equation 1 combines the Dominant (D), Lower (L), and Good (G) contributions with the base score (b) to create a data source risk score between 0 and 1. Figure 3 shows how risk factors are weighted differently and use exponential decay functions. The first factor has the greatest influence on the risk score, with each successive factor contributing a progressively smaller marginal gain or loss. See Appendix A and figure 5 for the complete charts, which follow the same trend.
Risk Score
0.8 0.6 0.4 0.2
Very Bad Bad
0
1
2 3 (V)B Value (Integer)
4
5
Fig. 3. Visualizes the impact of a risk factor on the initial risk score, based on the number of factors that hit at each rank. Very Bad and Bad are represented in the figure.
2) Final Risk Score: Once the data source risk scores and the anomaly risk score are calculated and stored, these scores are combined into the final risk score. Intuitively, this risk score is supposed to represent the ratio of malicious to normal. Using a conceptual application inspired by Likelihood Ratio Testing, this is shown by combining the risk scores to create a sense of maliciousness in the numerator while having 1− anomaly score in the denominator representing "normal" [31]. This creates an amplification effect as the anomaly score approaches 1. The following equation shows how the final product abuse lead risk score is computed by combining all the different data source risk scores and the anomaly score.
Equation 5 shows the weighted combination of all the data source risk scores and anomaly score. scoresi = W1 · ARSi + W2 · IPRSi + W3 · STRSi + W4 · IERSi + W5 · BFPRSi + W6 · FRSi
(5)
where each feature is continuous ∈ [0, 1]: ARSi : Anomaly Risk Score for alert i IPRSi : IP Risk Score for alert i STRSi : Security Ticket Risk Score for alert i IERSi : Interesting Events Risk Score for alert i BFPRSi : BFP Risk Score for alert i FRSi : Firmographic Risk Score for alert i Equation 6 then normalizes this score to be between 0 and 1. Ni =
scorei − minj scorej maxz scorez − minj scorej
(6)
Ni Ri = 1+W0 (1−ARS is how the final risk score is calcui) lated, dividing the normalized numerator by the weighted (1−anomaly score), adding 1 to smooth the score and constrain it to between 0 and 1. These risk weights are tuned with Bayesian Optimization using Gaussian processes [2], [3]. A k-fold bootstrapped approach is taken where a subset of data is randomly sampled, predicted on using a set of weights, and average recall is recorded across each iteration when constrained to a predefined flag rate. Bayesian weight tuning allows the risk score to be dynamic, in that we can re-tune the weights with an updated corpus, assigning more weight to data source risk scores that indicate a stronger signal based on the labels. Since the historical dataset has very few malicious labels, the weighttuning dataset is drawn from the same dataset as the evaluations. This decision was made because of the lack of labeled data, which required a bootstrapped approach, and because of the unsupervised nature of this method. While this may present leakage concerns, the model was still evaluated against live production data that wasn’t seen in any of the weight tuning process. Even against the unseen data, the model significantly outperformed baselines. Once the final risk scores are calculated, a threshold is created to select only the top leads to show investigators. This threshold is calculated to satisfy volume constraints, setting the threshold to flag only the top leads based on the sample set. This threshold ensures that fewer average monthly leads are flagged vs existing methods on the historical sample.
C. Explainability The final risk score alone doesn’t meet all the requirements of a lead generator. As such, it was required to be user-friendly for the security investigators. Due to the high dimensionality of input data, simply sending a risk score by itself is insufficient. There is a "cold start" problem that arises when multiple data sources produce these leads, and investigators have no clue where to begin their investigation. Should they start with the browser fingerprint, the event itself, or the IP? A lead generation model needs to tell investigators not just "what" to look at,
but "why" and "where" to look first. To speed up investigation time and solve for this cold start problem, explainability was built in as a major component of this model. 1) Data Source Risk Attribution: Starting with the data source and anomaly risk scores, investigators can identify which sources had worse risk factors, giving them a starting point for investigation. subsubsection IV-C1 shows example risk lead predictions. Product security investigators see the final score and they can also see the data source risk scores that went into the calculation, giving them an idea of where to begin investigating. We can also use this risk score setup to explain why events were flagged as malicious. We do this by simply sending a list of which risk factors hit for each event and their corresponding values. TABLE III R ISK S CORE A NALYSIS BY E VENT. S HOWS EXAMPLE DATA SOURCE RISK SCORES , THE ANOMALY SCORE , AND FINAL RISK SCORE FOR 2 EVENTS ,
In the example local explanation, among other indicators, it showed the BFP hash had a low probability of belonging to that IP, it was using a suspicious browser, and performed interesting platform event 13.
f(x) = 0.416 +0.034
0.05 = bfp_p_hash_given_ip 1 = bfp_suspicious_browser 1 = event_13 1 = event_16 3 = event_70 1 = ip_suspicious_proxy 2 = event_21 0.03 = bfp_p_browser_given_customer 1 = security_has_threat_actor 1 = ip_suspicious_vpn
+0.034 +0.033 +0.031 +0.028 +0.025 +0.023 +0.022 +0.021 +0.020
0.000 0.005 0.010 0.015 0.020 0.025 0.030 0.035 0.040
E[f(X)] = 0.145
ALLOWING INVESTIGATORS TO UNDERSTAND WHICH DATA SOURCES CONTRIBUTED THE MOST TO RISK .
Event a b
Anomaly
IP
Security
Events
Firmographic
Total
0.25 0.30
0.45 0.78
0.39 0.44
0.40 0.34
0.20 0.60
0.86 0.89
Fig. 4. Local Explanation of Top 10 SHAP Features for Isolation Forest, signifying the top feature contributions to the anomaly score. These were shared with investigators (who could resolve events to specific types via a database) who used them to identify discrepancies between initial modeling and how attacks tend to work, and successfully iterate to our current solution.
Table IV provides an example of explaining leads by 3) Interesting Events Explainability: Lastly, interesting highlighting the risk factors and their values that were hit events themselves can be flagged as an explanation. Since that for a single prediction. This user performed interesting event was a pre-filter for flagging leads in the first place, investigators 15 exactly 3 times, has a threat actor associated with an are always going to look into the events themselves. So, overlapping security ticket, and was using a suspicious browser. providing the event name helps show them exactly what the In the actual queue, interesting events show the description, not user is doing to abuse the product. For example, they may be just the mapping. Interesting event 15 may represent something creating a new admin for the purpose of privilege escalation. like a defense evasion or privilege escalation indicator. Explainability also served as a means of model improvement and debugging. When running a trial version, investigators noticed trends in the explained features leading to false TABLE IV R ISK FACTORS E XPLANATION E XAMPLE . T HIS DEMONSTRATES THE RISK positives, prompting changes to feature engineering and risk EXPLANATION FOR A SINGLE EXAMPLE PRODUCT ABUSE LEAD , WHERE factor tiers, with some factors lowered, removed, or added THE FACTORS CONTRIBUTING TO THE RISK SCORE INCLUDED THE USER based on their prevalence in false positives. PERFORMING INTERESTING EVENT 15 THREE TIMES , HAVING AN ASSOCIATED THREAT ACTOR IN AN OVERLAPPING SECURITY TICKET, AND USING A SUSPICIOUS BROWSER TO PERFORM THESE ACTIONS . T HESE EXPLANATIONS HELP POINT INVESTIGATORS IN THE RIGHT DIRECTION .
Risk Factor Interesting Event 15 Threat Actor Association Suspicious Browser
Hit Count 3 1 1
2) SHAP Anomaly Risk Attribution: Next, using Isolation Forest’s compatibility with SHAP, we flag the top 10 local feature contributions based on Shapley values, a game-theoretic approach connecting optimal credit allocation to local explanations [4]. SHAP scores for an Isolation Forest analyze marginal feature contributions to path length, with shorter paths yielding higher SHAP values, helping investigators understand why leads were flagged as anomalous. Figure 4 shows the local SHAP explanations for the Isolation Forest model for a single anomaly prediction. This visual uses anomaly score instead of path length. This helps investigators understand why this event may have been flagged as an anomaly.
V. E VALUATION AND R ESULTS The model was evaluated using two methods: a counterfactual replay against a historical corpus of known product abuse examples (i.e., testing what the model would have caught had it been live during past incidents), and a live beta release against never-before-seen data. Both evaluations showed significant efficacy gains and reduced alert volume vs. existing rules-based alerts. A. Evaluation Against Historical Corpus Using the labels discussed earlier, a historical labeled corpus of 28 true product abuse labels was constructed. The model impact was evaluated for two different purposes against the training corpus: (1) Improve product abuse coverage compared to existing methods, protecting customers more with highfidelity detections. (2) Reduce the number of leads sent to investigators compared to existing methods, minimizing alert fatigue for investigators.
Stat
Model
Rules-Based
To evaluate the model’s performance across these 2 areas, Precision 13.3% [9.0, 19.2] 3.7% [2.9, 4.6] the chosen metrics were recall and average monthly leads. Avg Monthly Lead Reduction 57.6% [54.3, 60.7] NA Recall, also known as True Positive Rate, is the percentage of TABLE VI known product abuse incidents in the historical sample dataset O UR ML SOLUTION WAS DEPLOYED AGAINST LIVE DATA IN A 5- MONTH “ BETA” RELEASE , SHOWING FAR HIGHER PRECISION THAN THE that the model captures (Recall = T P/(T P + F N )). Average RULE - BASED APPROACH . B RACKETED VALUES ARE 95% CI S (W ILSON monthly leads is simply the number of leads sent divided by SCORE INTERVAL [33] FOR PRECISION ; EXACT P OISSON RATE - RATIO [32] the number of months in the sample dataset. Investigators can FOR LEAD REDUCTION ). G IVEN THE HIGH - RISK AND LOW- ENOUGH - TO - STAFF LOAD OF ALERTS , THE MODEL IS CONTINUING only handle so many leads before becoming fatigued; therefore, EXPANDED ROLL OUT DUE TO ITS SUCCESS . this was controlled ahead of time with a conservative flag rate for sending leads. Fewer monthly leads along with higher recall represent a successful model in both facets. Table V shows how the model catches more true product of these originated from specific risk factors being hit when abuse events while flagging fewer leads than existing rules- they shouldn’t have been, which led to modifying that field based methods (baseline). The product abuse recall increased upstream. Similar trends were spotted, and modifications were 35% vs the baseline (95% CI [15, 56] pts, paired Wald interval), made to feature engineering and risk factor bucketing. The model is also adaptable to new labels, as we collect and average monthly leads decreased 30% (95% CI [22%, 37%], exact Poisson rate-ratio method [32]), exceeding goals for both these new labels over time and can build an updated corpus with them. The risk score weights can be automatically reefficacy and alert reduction. tuned based on this labeled corpus, additional risk factors may Stat Model Rules-Based be added, problematic ones may be dropped, etc. Recall 39% [24, 58] 4% [1, 18] Due to the transparency and explainability of each detection, Avg Monthly Lead Reduction 30% [22, 37] NA it is often apparent to investigators why it is a true positive TABLE V U SING A HISTORICAL CORPUS OF 28 PRODUCT ABUSE CASES (3- MONTH or false positive and whether that indicates a one-time failure WINDOW ), THE NEW APPROACH CATCHES 9.8× MORE INSTANCES . or a symptom of the model that could be adjusted. Overall, B RACKETED VALUES ARE 95% CI S (W ILSON SCORE INTERVAL [33] FOR explainability is not only a tool for investigators to understand RECALL ; EXACT P OISSON RATE - RATIO [32] FOR LEAD REDUCTION ). T HIS IS A SIGNIFICANT OPERATIONAL ADVANTAGE , MITIGATING WHAT WAS where to start their investigation, but it is also a tool for PREVIOUSLY A REACTIVE - ONLY ISSUE . filtering out false positives and incorporating feedback into model adjustments down the line. VI. C ONCLUSION By combining multiple data streams covering different time As an additional evaluation and learning step, the model was scales and update frequencies, we build an effective strategy released in “beta mode”, where it predicted against live, neverfor identifying product abuse that far exceeds rules-based before-seen data. During the beta release, investigators labeled approaches. A mix of unsupervised anomaly detection is predictions in a limited capacity vs existing queues. Since this combined with domain-expert ordinal coding to devise a final wasn’t reliant on a single labeled corpus but on two queues with risk score that enables finding product abuse within the data very different proportions of alerts reviewed (the new model 10× more effectively. Incorporating explainability enabled rapid vs the rule-based method), efficacy was measured by precision iteration and information retrieval by investigators to improve and alert volume. Investigators reviewed 1,925 rules-based the design. alerts vs 173 ML alerts in the time period of 5 months, with ML alerts showing a significant precision increase and ACKNOWLEDGMENT alert volume decrease vs existing rules alerts. Efficacy was The authors used Claude [34] throughout the preparation of measured by precision (Precision = T P/(T P + F P )) and this manuscript to assist with language editing, phrasing, LaTeX alert volume. formatting, and revisions. All technical content, experimental The average monthly leads metric was used again to calculate design, results, and conclusions are the authors’ own. the average monthly leads. For this evaluation, the beta model R EFERENCES ran for 5 months against live data. Table VI shows how the beta model, despite having far fewer [1] B. Gelman, S. Taoufiq, T. Vörös, and K. Berlin, “That Escalated Quickly: An ML Framework for Alert Prioritization,” Feb. 2023, arXiv:2302.06648 alerts reviewed, exceeded rules precision by 9.6 pts (95% CI [cs]. [Online]. Available: http://arxiv.org/abs/2302.06648 [4.5, 14.7]) and reduced alerts fired by 57.6% (95% CI [54.3%, [2] T. Head, M. Kumar, H. Nahrstaedt, G. Louppe, and I. Shcherbatyi, 60.7%]) compared to rules-based alerts. In total, 173 ML alerts “scikit-optimize/scikit-optimize: v0.9.0,” Oct. 2021, bSD-3-Clause license. [Online]. Available: https://doi.org/10.5281/zenodo.5565057 were reviewed, resulting in 23 TPs (19 of which were novel, [3] B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, and N. de Freitas, missed by the high/critical rule queue); 1,925 rules-based alerts “Taking the human out of the loop: A review of Bayesian optimization,” were reviewed, with 71 TPs discovered. Proceedings of the IEEE, vol. 104, no. 1, pp. 148–175, Jan. 2016. [4] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model The beta model was also very beneficial for model debugging predictions,” arXiv preprint arXiv:1705.07874, 2017, 31st Conference and improvements. Due to the explainability aspects, trends on Neural Information Processing Systems (NIPS 2017), Long Beach, were identified that tended to produce false positives. Some CA, USA. B. Beta Mode Evaluation
[5] A. S. M. Irwin, “Double-Edged Sword: Dual-Purpose Cyber Security Methods,” in Cyber Weaponry: Issues and Implications of Digital Arms, H. Prunckun, Ed. Cham: Springer International Publishing, 2018, pp. 101–112. [Online]. Available: https://doi.org/10.1007/ 978-3-319-74107-9_8 [6] E. Raff and C. Nicholas, “A Survey of Machine Learning Methods and Challenges for Windows Malware Classification,” in NeurIPS 2020 Workshop: ML Retrospectives, Surveys & Meta-Analyses (ML-RSA), 2020, arXiv: 2006.09271. [Online]. Available: http://arxiv.org/abs/2006.09271 [7] E. Raff, M. Ashkenazi, S. Samtani, D. J. Elkind, and S. Krasser, “Cybersecurity is the True Frontier for Generative AI Success or Failure,” in 2026 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), Jul. 2026, pp. 242–252, iSSN: 2768-0657. [Online]. Available: https://ieeexplore.ieee.org/document/11632077/authors [8] Google, “Abuse detection | apigee | google cloud,” Google, 2025. [Online]. Available: https://cloud.google.com/apigee/docs/api-security/ abuse-detection [9] F. Barr-Smith, X. Ugarte-Pedrero, M. Graziano, R. Spolaor, and I. Martinovic, “Survivalism: Systematic Analysis of Windows Malware Living-Off-The-Land,” in 2021 IEEE Symposium on Security and Privacy (SP), May 2021, pp. 1557–1574, iSSN: 2375-1207. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9519480 [10] R. Ning, W. Bu, J. Yang, and S. Duan, “A Survey of Detection Methods Research on Living-Off-The-Land Techniques,” in 2023 IEEE International Conference on Sensors, Electronics and Computer Engineering (ICSECE), Aug. 2023, pp. 159–164. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10263445 [11] D. Trizna, L. Demetrio, B. Biggio, and F. Roli, “Robust Synthetic Data-Driven Detection of Living-Off-the-Land Reverse Shells,” Dec. 2024, arXiv:2402.18329 [cs]. [Online]. Available: http://arxiv.org/abs/ 2402.18329 [12] R. Stamp, “Living-off-the-Land Abuse Detection Using Natural Language Processing and Supervised Learning,” Aug. 2022, arXiv:2208.12836 [cs]. [Online]. Available: http://arxiv.org/abs/2208.12836 [13] V. Haupert, D. Maier, and T. Müller, “Paying the Price for Disruption: How a FinTech Allowed Account Takeover,” in Proceedings of the 1st Reversing and Offensive-oriented Trends Symposium, ser. ROOTS. New York, NY, USA: Association for Computing Machinery, Nov. 2017, pp. 1–10. [Online]. Available: https://doi.org/10.1145/3150376.3150383 [14] G. Milka, “Anatomy of Account Takeover,” 2018. [Online]. Available: https://www.usenix.org/conference/enigma2018/presentation/milka/ [15] P. Doerfler, K. Thomas, M. Marincenko, J. Ranieri, Y. Jiang, A. Moscicki, and D. McCoy, “Evaluating Login Challenges as aDefense Against Account Takeover,” in The World Wide Web Conference, ser. WWW ’19. New York, NY, USA: Association for Computing Machinery, May 2019, pp. 372–382. [Online]. Available: https://doi.org/10.1145/3308558.3313481 [16] R. Santoso and A. A. S. Gunawan, “Detecting Account Takeover (ATO) in Fintech Companies Using Machine Learning,” in 2024 6th International Conference on Cybernetics and Intelligent System (ICORIS), Nov. 2024, pp. 1–6. [Online]. Available: https://ieeexplore. ieee.org/abstract/document/10903690 [17] R. Kawase, F. Diana, M. Czeladka, M. Schüler, and M. Faust, “Internet Fraud: The Case of Account Takeover in Online Marketplace,” in Proceedings of the 30th ACM Conference on Hypertext and Social Media, ser. HT ’19. New York, NY, USA: Association for Computing Machinery, Sep. 2019, pp. 181–190. [Online]. Available: https://doi.org/10.1145/3342220.3343651 [18] J. Tao, H. Wang, and T. Xiong, “Selective Graph Attention Networks for Account Takeover Detection,” in 2018 IEEE International Conference on Data Mining Workshops (ICDMW), Nov. 2018, pp. 49–54, iSSN: 23759259. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/ 8637408 [19] M. Gao, “Account Takeover Detection on E-Commerce Platforms,” in 2022 IEEE International Conference on Smart Computing (SMARTCOMP), Jun. 2022, pp. 196–197, iSSN: 2693-8340. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9821104 [20] R. Bace and P. Mell, “NIST Special Publication on Intrusion Detection Systems,” NIST, Tech. Rep. OMB No. 074-0188, Jan. 2001. [Online]. Available: https://apps.dtic.mil/sti/html/tr/ADA393326/ [21] I. J. King, R. Ramirez, B. Bowman, and H. H. Huang, “Trail: A Knowledge Graph-Based Approach for Attributing Advanced Persistent Threats,” in 2025 IEEE 41st International Conference on Data
Engineering (ICDE), May 2025, pp. 1207–1220, iSSN: 2375-026X. [Online]. Available: https://ieeexplore.ieee.org/document/11113100 [22] S. Gupta, F. Lu, A. Barlow, E. Raff, F. Ferraro, C. Matuszek, C. Nicholas, and J. Holt, “Living off the Analyst: Harvesting Features from Yara Rules for Malware Detection,” in 2024 IEEE International Conference on Big Data (BigData), Dec. 2024, pp. 2624–2634, iSSN: 2573-2978. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10825735 [23] P. Gao, X. Liu, E. Choi, B. Soman, C. Mishra, K. Farris, and D. Song, “A System for Automated Open-Source Threat Intelligence Gathering and Management,” in Proceedings of the 2021 International Conference on Management of Data, ser. SIGMOD ’21. New York, NY, USA: Association for Computing Machinery, Jun. 2021, pp. 2716–2720. [Online]. Available: https://dl.acm.org/doi/10.1145/3448016.3452745 [24] P. Gao, X. Xiao, Z. Li, K. Jee, F. Xu, S. R. Kulkarni, and P. Mittal, “A query system for efficiently investigating complex attack behaviors for enterprise security,” Proc. VLDB Endow., vol. 12, no. 12, pp. 1802–1805, Aug. 2019. [Online]. Available: https://doi.org/10.14778/3352063.3352070 [25] M. Saqib, B. C. Fung, P. Charland, and A. Walenstein, “GAGE: Genetic Algorithm-Based Graph Explainer for Malware Analysis,” in 2024 IEEE 40th International Conference on Data Engineering (ICDE), May 2024, pp. 2258–2270, iSSN: 2375-026X. [Online]. Available: https://ieeexplore.ieee.org/document/10598144 [26] T. Song, M. Organokov, L. Gulikers, G. Grassi, G. Carofiglio, and M. Meo, “Advancing Cloud-Native Cyber Threat Detection with Graph-Based Feature Engineering,” in 2025 IEEE 41st International Conference on Data Engineering (ICDE). Los Alamitos, CA, USA: IEEE Computer Society, May 2025, pp. 4291–4297. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/ICDE65448.2025.00321 [27] Y. Sui, X. Wang, T. Cui, T. Xiao, C. He, S. Zhang, Y. Zhang, X. Yang, Y. Sun, and D. Pei, “Bridging the Gap: LLM-Powered Transfer Learning for Log Anomaly Detection in New Software Systems,” in 2025 IEEE 41st International Conference on Data Engineering (ICDE), May 2025, pp. 4414–4427, iSSN: 2375-026X. [Online]. Available: https://ieeexplore.ieee.org/document/11113176 [28] Clarence Chio and David Freeman, Machine Learning and Security. Sebastopol, CA: O"Reilly, 2018, vol. 1. [29] D. Freeman, S. Jain, M. Dürmuth, B. Biggio, and G. Giacinto, “Who are you? a statistical approach to measuring user authenticity,” in Proceedings 2016 network and distributed system security symposium, 2016. [30] F. T. Liu, K. M. Ting, and Z.-H. Zhou, “Isolation forest,” in 2008 Eighth IEEE International Conference on Data Mining. IEEE, 2008, pp. 413–422. [31] J. Grana, D. Wolpert, J. Neil, D. Xie, T. Bhattacharya, and R. Bent, “A likelihood ratio detector for identifying within-perimeter computer network attacks.” Sep. 2016, arXiv:1609.00104 [cs.CR]. [Online]. Available: https://arxiv.org/pdf/1609.00104 [32] J. Przyborowski and H. Wilenski, “Homogeneity of results in testing samples from poisson series: With an application to testing clover seed for dodder,” Biometrika, vol. 31, no. 3/4, pp. 313–323, 1940. [Online]. Available: http://www.jstor.org/stable/2332612 [33] E. B. Wilson, “Probable inference, the law of succession, and statistical inference,” Journal of the American Statistical Association, vol. 22, no. 158, pp. 209–212, 1927. [Online]. Available: http: //www.jstor.org/stable/2276774 [34] Anthropic, “Claude,” Large language model, 2026. [Online]. Available: https://claude.ai
A PPENDIX
Fig. 5. Risk Score Diagram