The Anatomy of Scam Scenarios: Large-Scale Characterization and Conversation-Aware Detection Shang Ma¶ , Chen Yanai‡ , Avichai Ben‡ , Zichen Liu† , Yanfang Ye¶ , Xusheng Xiao† † Arizona State University, ‡ Charm Security, ¶ University of Notre Dame
arXiv:2606.16052v1 [cs.CR] 14 Jun 2026
[email protected], {zliu396,xusheng.xiao}@asu.edu, {sma5,yye7}@nd.edu
Abstract—Online scams have become a pervasive global threat, causing substantial financial, psychological, and operational harm. Scammers embed psychological techniques (PTs) within reusable operational schemes to scale scam campaigns with minimal adaptation. However, existing studies often analyze PTs as isolated features, overlooking the recurring scam scenarios in which they are systematically deployed. To address this gap, we first conduct a large-scale empirical study to jointly characterize scam scenarios and their associated PTs. Specifically, we develop a data-driven pipeline to derive a hierarchical taxonomy of scam scenarios, consisting of 18 finegrained scenarios grouped into 6 high-level tactics based on their PT profiles. Furthermore, to transfer this scenario-level knowledge to practical defense, we design a conversation-aware scam scenario detection approach for financial-institution customer interactions, enabling timely warning and intervention. Our study on 102,054 real-world scam incident reports, spanning 2024-02-01 to 2025-10-31, reveals that PT usage is significantly associated with scam scenarios. We further show that scammers organize scenarios around different operational goals, such as broad victim exposure, high victim conversion, and high-value extraction, and reuse infrastructure, including IP addresses, domains, email addresses, and phone numbers, to launch coordinated campaigns at scale. Evaluation on 1,115 real-world customer-service conversations from an industry partner shows that our approach achieves 84.41% tactic-level classification accuracy (84.14% F1) and ranks the correct scam scenario within its top three predictions in up to 91.04% of cases. These results demonstrate that scenario knowledge distilled from large-scale scam incident reports can effectively support early detection and characterization of scam-related customer-service conversations.
1. Introduction Online scams have become a global epidemic, inflicting severe financial and psychological harm on individuals and organizations. Their rapid proliferation is fueled by lowcost, fast, and often anonymous channels such as SMS and email. According to the U.S. Federal Trade Commission (FTC) 2025 data, consumers lost $12.5 billion to fraud, a 25% increase from the prior year, with investment fraud leading at $5.7 billion in reported losses [1]. Identity theft cases exceeded 1.15 million, driven by sophisticated AI-
assisted tactics including job offers, romance scams, and SMS-based fraud [2], [3], [4]. Globally, Europe experienced billions in losses, with Germany and France alone reporting C10.6 billion and C7.6 billion [5], [6], [7], while largescale operations from East and Southeast Asia, such as crypto fraud and pig-butchering schemes, caused massive financial damage [8], [9], [10]. AI-enabled impersonation scams accounted for over $17 billion, and Southeast Asia’s scam ecosystem has evolved into a $70B+ illicit marketplace often tied to forced-labor scam centers. Beyond victim losses, scams impose significant operational, financial, and regulatory burdens on financial institutions. Banks must analyze vast volumes of communications (call transcripts, chat logs, emails) to detect potential fraud, often requiring manual review, inter-institution coordination, and evidence collection. Institutions also face liability for reimbursing victims of authorized push payment scams, where customers are tricked into transferring funds [11], [12]. Regulatory bodies have introduced stricter monitoring and reimbursement obligations, intensifying compliance pressure. These challenges underscore the urgent need for automated systems that identify scam signals in customer communications, enabling earlier intervention and reducing both investigative workload and financial losses. The effectiveness and scalability of scams largely stem from the systematic use of psychological techniques (PT) and social engineering techniques, such as authority, urgency, reciprocity, and social proof, to manipulate victims’ decisions and actions [13], [14]. These techniques are embedded within carefully crafted scam narratives that appear legitimate and contextually relevant [15], [16]. For example, a smishing campaign may impersonate a highway toll service, claiming an unpaid $12.51 fee that must be paid immediately via a provided link to avoid fines or license suspension [17], [18]. Such messages combine multiple PTs: authority by impersonating official agencies, fear through immediate penalties, and credibility via links resembling official sites [19]. Once such a narrative proves effective, scammers can reuse it across victims and geographic regions with minor contextual adaptations, allowing them to scale operations while minimizing the cost of crafting new scams [20]. Although prior work has studied PTs for scam detection and generation [15], [16], [21], [22], [23], these studies often treat PTs as isolated persuasive signals and pay less
attention to the recurring narrative structures in which they are deployed. In practice, scammers do not simply stack PTs arbitrarily. They embed them into reusable operational schemes, such as unpaid toll notifications, fake job offers, investment opportunities, or account-security alerts. These schemes define the scam’s premise, victim role, communication flow, and monetization path, while PTs shape how the victim is persuaded within that scheme. Our preliminary study shows that scam narratives naturally group into recurring scam schemes, each characterized by a distinct combination of PTs. We refer to such a recurring operational scheme as a scam scenario. To better understand scam scenarios and inform the development of effective defenses against online scams, we first conduct a large-scale empirical study based on realworld scam incident reports. We adopt a data-driven pipeline to discover a hierarchical taxonomy of scam scenarios, where fine-grained scenarios are organized into higher-level tactics based on their associations with PTs. We then conduct a detailed analytical study to examine the distribution of these scenarios, assess differences in their risk profiles, and characterize how scammers adapt their tactics and operations across scenarios. Building on the taxonomy and these empirical insights, we further develop an approach for detecting scam scenarios in customer-service conversations within financial institutions, where early recognition of the underlying scam scenario can support timely intervention [24], [25]. Empirical Study of Scam Scenarios. Existing studies are often limited to small datasets, specific scam scenarios, or narrow time frames [26], [27], [28], [29], [30]. Consequently, there is still no systematic, large-scale analysis that jointly characterizes scam scenarios and the PTs used by scammers across scenarios. To bridge this gap, we curate a dataset of 102,054 scam incident reports from BBB Scam Tracker [31], covering the period from 2024-02-01 to 202510-31 (see Figure 2 for an example scam incident report). We first leverage topic modeling to cluster scam reports based on recurring textual patterns, then apply LLM-assisted summarization, manual inspection, and expert-aligned refinement to consolidate noisy topics into 18 distinct scam scenarios. We then analyze PT usage within each scenario and find a significant association between scam scenarios and PTs (χ2 = 28052, p < 0.001) revealing that PTs are not uniformly distributed across scenarios but instead reflect scenario-specific persuasion strategies. Based on this observation, we further organize the 18 scenarios into 6 higher-level tactics according to their shared PT profiles, as shown in Table 1. Our analysis reveals several critical findings that fundamentally advance the understanding of how scammers structure, adapt, and scale their operations across scenarios. First, scammer weaponize PTs in scenarios by stacking them to exploit victims: 75.55% of scam incidents contain at least two PTs, and each additional PT is associated with a 26.27% increase in expected monetary loss. Second, different scenarios optimize for different operational goals, such as broad victim exposure, high victim conversion, or high-
value extraction. Third, infrastructure analysis shows that scammers reuse operational resources to launch coordinated campaigns at scale: over 30% of scam incidents are linked to distinct scam campaigns, and the largest campaign alone contains 3,809 incidents, spans all 18 scenarios. Finally, our case studies reveal that scammers adapt scenarios to seasonal events, cause harms beyond direct monetary loss, and operate international, cross-platform campaigns. Conversation-Aware Scam Scenario Detection. The preceding analysis studies scam scenarios from complete, retrospective reports. In practice, however, scam defense often happens during partial and unfolding interactions. In particular, when a suspicious transaction is flagged, the defender contacts the customer and asks about the incident. The customer then reveals information incrementally, such as how the interaction began, who the scammer claimed to be, and what payment was requested. As illustrated in Figure 1, the defender must infer the underlying scam scenario from incomplete and evolving customer utterances to enable timely intervention [24], [25]. This setting motivates a key question: can scenario-level knowledge learned from largescale scam reports support early scam scenario detection in customer-service conversations? This problem can be formalized as a sequential text classification task, where the model updates its prediction as more conversation turns become available [32], [33], [34]. However, applying this formulation to our setting faces three challenges. (1) Incomplete Context. Customers typically describe incidents from memory, resulting in partial, fragmented, or uncertain context. (2) Uncertain Evidence Onset. The evidence needed to identify the underlying scam scenario may emerge only after several turns of a customerservice conversation. Early turns are often noisy, containing irrelevant information, such as greetings or identity verification. (3) Diverse Narratives. Real-world customer narratives vary widely in communication style, emotional state, and level of detail, even when they describe the same scam scenario. Building on these insights, we present EARS (Early Anti-scam Recognition System), which detects the underlying scam scenario from unfolding customer-service conversations. To handle incomplete context, EARS first pretrains a scenario classifier on the large-scale incident-report corpus. Although incident reports differ from conversations in format, they describe the same underlying scam scenarios from the victim’s perspective, enabling the model to learn scenario-level semantics before adapting to conversation data. To handle uncertain evidence onset, EARS then finetunes the classifier on customer-service conversations under turn-level supervision: at each turn, it predicts from customer narrative observed so far and conditions the prediction on the turn’s discourse phase, withholding judgment while the exchange is benign and committing to a scenario as soon as diagnostic evidence appears. To stay robust across the stylistic and structural variation of real customer narratives, EARS further augments the training conversations with an LLM-based, label-preserving rewriting pipeline. Together, these components turn report-level scenario knowledge into
c1 : Hello? a1 : Hello, my name is David Okonkwo, and I’m calling from Bank of America Fraud Prevention. I’m calling regarding a recent transaction on your account that our system flagged as potentially unusual. Could I please confirm your full name and the last four digits of your account number? c2 : Yes, this is Bradley Myers. My account number ends in 5486. a2 : Thank you. We’re looking at a payment of $500 that was made recently. Do you recall making a $500 payment using a gift card for a service? c3 : Yes, I did make a payment of about $500 recently, and it was with gift cards. It was for my utility bill, which was quite urgent. a3 : We flagged this because using gift cards to pay a utility bill is quite unusual. Could you tell me why you used gift cards for this payment instead of a more traditional method? c4 : The utility service said my power would be cut off very soon if I didn’t pay immediately. They insisted on gift cards for a “rapid processing payment” to avoid the shutoff. a4 : Which utility service did you speak with, and how did they instruct you to make this payment? c5 : It was Idaho Power. They said it was the only way to process the payment quickly to avoid immediate power disconnection.
a5 : Legitimate utility companies, including Idaho Power, would never ask for payment using gift cards. They typically accept payments through their official website, by mail, or over the phone with a credit/debit card or bank transfer. c6 : Oh dear. So, if they don’t ask for gift cards, then. . . was that not really Idaho Power? I was told it was the only way to prevent my power from being cut off immediately. a6 : Based on what you’ve described, it sounds like you may have been targeted by an authority imposter scam. Scammers often pretend to be public agencies such as utility companies and demand immediate payment via gift cards to create a sense of urgency. c7 : It was a scam, then. I. . . I did have a feeling, a slight unease, but they sounded so terribly official, and quite urgent. What does this mean for my account, exactly? a7 : I’m very sorry to hear that. The $500 payment you made with gift cards was likely sent to a scammer. We will mark this transaction as fraudulent and proceed with a fraud investigation, including steps to secure your account and explore recovery options.
[ Agent initiates transfer to the specialized fraud department; call closes with a commitment to a callback within 5–10 minutes. ]
Figure 1: Example customer-service conversation in which the banking agent (a) investigates a flagged $500 payment with the customer (c). In this case, the scammer impersonates the public utility service and threatens a power shutoff to pressure the victim into paying with gift cards. an early, deployable detector for live conversations. We evaluate EARS on a test set of 1,115 customerservice conversations comprising 14,406 turns and covering all tactics and scenarios in our taxonomy. EARS achieves up to 84.41% tactic-level accuracy (84.14 F1) and ranks the correct scam scenario among its top-3 predictions in over 90% of conversation turns. Relative to a strong incidentreport-pretrained encoder, our two conversation-level adaptations, LLM-based data augmentation and discourse-phase injection, jointly improve scenario macro-F1 from 64.87 to 71.39. EARS also outperforms LLM-based scam scenario detection: in scenario macro-F1, it exceeds the few-shot GPT-5.4 baseline, one of the strongest commercial LLMs, by 25.62 points. These results show that the scenario taxonomy derived from real-world incident reports transfers to customer-service conversations and supports earlier recognition of scam-related customer cases.
2. Background 2.1. Human Vulnerabilities and Psychological Exploits Human vulnerabilities are intrinsic cognitive, emotional, and social tendencies that influence how people perceive risk, trust others, and make decisions [35]. Human-centric cyber attacks exploit these vulnerabilities through psychological rather than purely technical means [15], [16]. Prior work has characterized these manipulation strategies as Psychological Techniques (PTs) [26] and Human Vulnerability Exploits (HVEs) [36], both describing recurring methods used by scammers to influence victims. Building on these frameworks, we refine and extend the PT taxonomy using
the HVE framework, yielding nine PT categories: Authority, Phantom Riches, Fear and Intimidation, Liking, Urgency and Scarcity, Pretext and Trust, Evoking Social Norms, Consistency, and Social Proof. Definitions are provided in Appendix Table 13.
2.2. Classification of Scam Types
A scam taxonomy is essential for understanding, measuring, and mitigating online scams, providing researchers with a common vocabulary for analysis and practitioners with a framework for education, triage, and defense [36]. Existing resources such as the BBB Scam Tracker glossary define 32 scam types [37], but these labels are often flat, coarse, and semantically overlapping; for example, the category phishing may encompass distinct scenarios such as delivery impersonation, account-security alerts, fake subscriptions, and credential-theft campaigns (Figure 2). More systematic efforts, such as the Stanford Center on Longevity and FINRA fraud taxonomy, organize fraud into hierarchical dimensions including claimed identity, victim-facing pretext, promised benefit or threatened loss, and requested victim action [38]. However, manually curated taxonomies are costly to maintain and struggle to keep pace with evolving scam ecosystems, which increasingly incorporate new technologies, platforms, and monetization models such as cryptocurrency scams, data-breach extortion, and remote-work employment fraud. These challenges motivate a data-driven, continuously updateable taxonomy derived from real-world scam reports.
https://www.bbb.org/scamtracker/lookupscam/1291795 Description. I was on the phone with the USPS and the USPS online trying to locate a package and it came up to fill out a form to be able to track a package, it then asks for a $5 refundable deposit, USPS advised me that was not them. I called and assumed this was resolved. A charge was taken out of my account for $65 for a service I did not ask for. When I called to see, it was not even on my phone number, but a work cell number, they said they cx the “subscription” I didn’t know I had but would not refund my money. Scam Type: Phishing Scammer: (800) 556-3410 Date: May 15, 2026 Email: [email protected]
Scam ID: 1291795 Victim Loc.: OK, 73065 Dollars Lost: $65 URL: justanswer
Figure 2: Example of a scam incident report.
3. Empirical Study of Scam Scenarios In our empirical study, we first employ a data-driven pipeline to extract scam scenarios from a large corpus of real-world scam incident reports. Building on these extracted scenarios, we next examine their relationships with PTs and systematically characterize them along multiple dimensions, including risk profiles, scammers’ operational strategies, and infrastructure usage.
3.1. Datasets We collect scam incident reports from BBB Scam Tracker [31], a public platform where victims and targets of scams voluntarily submit reports containing rich textual descriptions and structured metadata. Starting from 177,989 raw reports, we apply a two-stage preprocessing pipeline. First, we remove business-fraud cases outside the scope of this study (e.g., undelivered online purchases) using a lightweight LLM-based classifier (Qwen3-1.7B); manual evaluation of 100 sampled reports found no misclassified irrelevant cases. Second, we remove duplicate submissions and filter out overly short reports (≤27 words, corresponding to the 20th percentile in length) that are unlikely to contain useful information. After preprocessing, the final dataset contains 102,054 high-quality scam incident reports spanning 2024-02-01 to 2025-10-31.
3.2. Data-Driven Scam Scenario Discovery Topic Modeling. Scammers often reuse effective scenarios across victims and geographic regions with only minor contextual adaptations. To capture such recurring patterns, we apply topic modeling to partition a large corpus of scam incident reports into clusters, where each cluster serves as a candidate scam scenario. Specifically, we employ BERTopic [39], which first embeds each report into a dense semantic space using a pretrained language model, then clusters semantically similar reports into preliminary topics, and finally extracts representative topic keywords via class-based TF-IDF. We set the number of topics to 100, intentionally overestimating the expected number of true scam scenarios. This over-segmentation reduces cluster
heterogeneity and more effectively isolates outliers, i.e., reports that do not clearly align with any coherent topic. Overall, the model assigns 64,963 of 102,054 reports to 99 non-outlier topics, while the remaining 36.34% are marked as outliers. This relatively high outlier rate is expected, given the noisy and heterogeneous nature of real-world scam incident reports, and it helps prevent ambiguous cases from being forced into poorly formed clusters. Table 14 in Appendix gives examples of the BERTopic outputs. The clustering metrics show that the raw topics are locally coherent but globally overlapping. Specifically, the average centroid similarity is 0.6769, indicating that reports within the same topic are semantically close, and the keyword diversity score is 0.8677, suggesting that each topic captures distinct recurring narratives. However, the cosine silhouette score is only 0.0395, meaning that many topics are weakly separated from nearby topics in the embedding space. Thus, the predefined 100-topic setting yields finegrained clusters that often split the same scam scenario into multiple semantically related topics. To consolidate these raw topic clusters into a smaller set of distinct scam scenarios, we apply a two-step LLM-human collaborative annotation pipeline. Topic Merging and Scenario Derivation. Based on the top50 keywords associated with each topic, we prompt an LLM to generate a concise one- to three-sentence summary (see prompt template in Figure 17), with representative examples provided in Table 14. BERTopic further provides a relevance score between each report and its assigned topic, enabling identification of the most representative reports per topic. We then perform a manual inspection of topic keywords, LLMgenerated summaries, and the top-5 representative reports, and merge topics that correspond to the same underlying scam scenario but differ due to lexical variations or reporting styles. This consolidation reduces the initial 99 topics to 18 distinct scam scenarios. At the same time, we create an initial name and definition of each scenario during this annotation process, and align them with feedback from security experts at Charm Security and existing scam classification systems [37], [38]. This step ensures that each scenario is semantically accurate, distinct from the others, grounded in the data, and consistent with practitioner-facing terminology. Scenario Refinement. Topic modeling provides only rough scenario clusters. For example, a scam incident report may be assigned to an incorrect cluster because it shares similar wording with reports from another scenario. Thus, topic modeling alone is insufficient for producing high-quality scenario labels for subsequent analysis. However, manually annotating the entire dataset is expensive. To address this, we manually annotate reports from each scenario cluster, and refine the definition of each scenario. This process is part of our active learning pipeline, which will be elaborated in Subsection 4.2.
3.3. Taxonomy of Scam Scenarios In total, our scenario discovery process yields 32,746 scam incident reports annotated with 18 scam scenarios,
TABLE 1: Hierarchical taxonomy of scam scenarios grouped by major psychological technique drives Scenario
Description
Tactic
Retail Insurance & Warranty E-Commerce
Impersonating a retail brand, selling counterfeit goods, or fake ads for non-existent deals. Consumer & Service Scams
Financial Services
Impersonating a bank/credit card company, alleging account problems, or offering fraudu- PT Drive: Credibility lent debt relief.
Impersonating a provider, alleging policy issues, or offering fake health/auto/home policies. Scams that exploit trust in legitimate businesses, Impersonating e-commerce platforms or delivery services, or purchase scams where items brands, or service providers by impersonating them or offering fraudulent products/services. are never delivered.
Tech & Online Service Impersonating a tech company, alleging subscription/virus issues, fake tech support, or recovery scams. Impersonating the government, such as the IRS, SSA, or other agencies to demand fake Authority & Compliance Scams Scams that leverage the perceived authority to payments or steal personal information. Impersonating the law enforcer such as court, lawyer, or the FBI to allege warrants or coerce victims through fear and intimidation. criminal involvement, demanding payment to “resolve” it. PT Drive: Authority, Fear and Intimidation
Government Legal
Lottery & Sweepstakes Victims are offered lottery, prize, or products (e.g., donations, heritage, fake social media Windfall Scams giveaways) but must pay a “fee” or “tax” to claim winnings. Scams that lure victims with promises of unexpected Fake grants, scholarships, or aid programs requiring upfront fees or personal information. financial gains, prizes, or exceptional deals.
Funds, Grants & Aid Investment & Trading Good Deals
Promising high-return, low-risk investments via fraudulent platforms, fake stock tips, etc. PT Drive: Phantom Riches Fake check/overpayment scams, fraudulent travel packages, or ads for non-existent deals. Threatening to have compromising photos or videos of the victim and threatens to release Extortion Scams Scams that coerce victims into payment via them unless a payment (often crypto) is made. Threatening to have hacked the victim’s device, encrypted their files (ransomware), or extortions. stolen their data, demanding payment for its return. PT Drive: Fear and Intimidation
Sextortion Hack & Data Breach
Fake applications to steal personally identifiable information, or requiring payment for Employment Scams Scams that exploit job seekers by offering fake “training”/“equipment” for non-existent jobs. Tricking victims into performing “trial” work with a promise of payment that never comes, employment. such as package delivery, data labeling. PT Drive: Phantom Riches, Consistency
Fake Job Offer Unpaid Labor
Building a fake relationship to manipulate the victim into sending money or investing in Relationship & Trust Scams Scams that build or exploit personal and emotional fraudulent platforms. connections to manipulate victims Soliciting donations for fake charitable causes, such as disasters, wars.
Pig Butchering Charity Friends & Relatives
Claiming to be a friend/relative member and needing financial help due to an emergency. PT Drive: Evoking Social Norms
Retail
79
22
20
26
14
39
Insurance & Warranty
75
34
48
68
4
6
E-commerce
73
52
45
48
2
5
7
6
8
3 3
1
Financial Services
67
32
18
43
23
16
1
3
Tech & Online Service
72
53
58
41
11
11
3
2
Government
57
77
70
59
3
4
11
4 3
60 64
74
87
53
3
3
75
38
10
27
68
13
19
2
2 5
5
Funds, Grants & Aid
60
28
10
43
62
5
2
2
1
3
Investment & Trading
48
28
24
18
71
34
6
4
13
14
Good Deals
62
13
10
40
63
28
8
10
4
7
Sextortion
60
30
72
43
14
14
13
14
1
18
Hack & Data Breach
53
12
72
10
11
16
14
5
5
Fake Job Offer
89
52
2
32
59
61
5
30
10
Unpaid Labor
76
24
10
11
72
61
11
6
8
Pig Butchering
75
41
20
20
75
50
15
30
23
6
Charity
51
30
16
11
10
8
40
2
3
7
Friends & Relatives
65
36
27
27
21
24
36
7
6
40
Prevalence (%)
Legal Lottery & Sweepstakes
d Cre
80
2
20
lity ibi
ity
Au
r& ncy y& tom Fea tion rgenc rcity Phan ches nsiste Ri Co ida U Sca tim
r tho
In
l cia So s E. Norm
ing Lik
l cia So oof Pr
6
t tex Pre rust T &
0
Figure 3: PT distribution across scam scenarios (Finding 1). with their distribution reported in Appendix Figure 13. To further understand how different scenarios exploit victims, we next analyze their associated PTs. Specifically, we build a PT classifier following prior work [26], apply it to our dataset, and examine the prevalence of PTs across scenarios. Finding 1: Scam scenarios exhibit distinct but recurring psychological technique profiles. Our chi-square test shows a significant association between scam scenarios and PTs (χ2 = 28052, df = 153, p < 0.001), with a moderate effect size measured by Cramer’s V = 0.195. This indicates that PTs are not uniformly distributed across scenarios. As
shown in Figure 3, Credibility is widely used across nearly all scenarios, which is intuitive because most scams leverage impersonation to gain victims’ trust. Beyond this common pattern, however, scenarios show distinctive PT profiles. For example, Government and Legal scams rely heavily on Authority and Fear & Intimidation, while Charity and Friends & Relatives scams show stronger use of Evoking Social Norms. These patterns suggest that scammers craft different victim-facing narratives while exploiting similar underlying psychological levers. To capture this tactic-level regularity, we introduce an additional categorization layer above the 18 scenarios. This layer groups scenarios that share similar PT drives and scammer tactics. The resulting hierarchical taxonomy is shown in Table 1, where each high-level category groups tactically related scenarios under a major PT drive.
3.4. Analysis As shown in Figure 2, each scam incident report contains rich telemetry about the scammer and the incident, enabling us to examine how these signals characterize different scam scenarios. In this part, we primarily analyze scenarios at the tactic level, as it combines scenario semantics with PTs and provides a cleaner basis for visualization. Fine-grained scenario-level results are provided in Appendix B.
Median loss
All reports Reports w/ $loss
15000 12500
0.8
10000
0.6
7500 0.4
5000
0.2
2500
0.0
0 0
1
2
3
4
5
6
Number of PTs per Scam Incident
Figure 4: PT combinations and dollar lost (Finding 2). We further validate this trend statistically. Using 6,050 reports with valid loss values, we find a significant positive Spearman correlation between the number of PTs and monetary loss (ρ = 0.398, p = 1.07 × 10−228 ). The association remains significant after controlling for scam scenario and report length using a regression model with heteroscedasticity-robust standard errors: each additional PT is associated with a 26.27% increase in expected monetary loss (p < 0.001, 95% CI: 10.29%–44.56%). These results suggest that PT combinations are linked to scam severity, as multiple psychological levers can simultaneously build trust, create pressure, and drive harmful actions. Median Loss
# Reports 1k
$1k
5k
15k
$500 $0 0%
10%
20%
30%
40%
50%
60%
Loss Rate
Figure 5: Incident report volume, monetary loss rate, and median loss across scenario tactics (Finding 3). Finding 3: Scammers tune scenarios and tactics for different strategic goals: mass exposure, high victim conversion, or high-value extraction. Figure 5 characterizes each scam scenario tactic along three dimensions: prevalence, measured by the number of reports; conversion rate, measured by the fraction of reports with positive monetary loss; and severity, measured by median dollar loss. We observe a diversity of tactic profiles among these dimensions. For example, Consumer & Services and Authority & Compliance scams
account for the most incident reports, suggesting broad exposure but their median losses are relatively low, indicating limited per-victim financial harm. In contrast, Relationship & Trust scams have the highest conversion rate, with nearly half of reports involving monetary loss, consistent with their personalized nature: scammers either gradually build emotional trust (Pig Butchering) or impersonate existing social relationships (Friends & Relatives). Employment and Extortion scams show a different profile: they are lower in volume and conversion, but produce the highest typical losses. Employment scams often exploit desperate and financially vulnerable job seekers who are willing to pay more fees for a potential job opportunity, while Extortion scams rely on threats or ransom-like pressure that could lead to severe reputation harm (Sexortion) or personal loss (Hack & Data Breach). Overall, these results indicate that scammers optimize tactics with different goals: exposure, victim conversion, and profit gain. Resilient Domain
Mean log(1 + loss)
1.0
Number of Reports
Normalized Loss Metric
Finding 2: Scammers weaponize PTs by stacking them to manipulate victims: each additional PT coincides 26.27% increase in expected monetary loss. As shown in Figure 4, PT stacking is common in real-world scams: 75.55% of incident reports contain at least two PTs. Moreover, reports with more PTs tend to incur higher losses, suggesting that psychological complexity is associated with scam severity. However, effective PT stacking is scenario-dependent rather than simply additive. As shown in Figure 3, Government and Legal scams rely heavily on Authority and Fear and Intimidation (both > 70%) while rarely using Liking, Social Proof, or Evoking Social Norms, which are difficult to reconcile with threat-based narratives. These findings suggest that scammers combine multiple PTs only when they are coherent with the underlying scam playbook.
# Domains
30%
100 800 2.4k
20%
10% 5%
10%
15%
20%
25%
30%
Disposable Domain
Figure 6: Resilient and disposable domain rates across tactics (Finding 4). Finding 4: Scammers exploit infrastructure for different strategic goals across scenarios and tactics: disposable domains for mass exposure, and resilient domains for sustained engagement. We further analyze scam campaign infrastructure using website domains extracted from incident reports. Among 37,197 reports, 8,595 contain valid domains, yielding 5,486 unique domains. For each domain, we collect WHOIS, DNS, ASN/hosting, CDN/Cloudflare usage, and HTTP liveness signals, and classify domains as disposable (e.g., newly registered, short-lived within one year, or unreachable) or resilient (e.g., persistently live and CDN/Cloudflare-backed). As shown in Figure 6, tactics exhibit distinct infrastructure patterns. Authority & Compliance scams have the highest disposable-domain rate and the lowest resilient-domain rate, reflecting rapid deployment and abandonment consistent with broad-exposure campaigns. In contrast, Relationship & Trust scams show the highest resilient-domain rate and lower disposable usage, consistent with sustained infrastructure needed for prolonged trustbuilding and repeated interactions. Overall, scam infrastructure is not uniform but adapted to the operational goals of each tactic. Finding 5: Scammers reuse infrastructure to coordinate campaigns: over 30% of scam incidents are linked to distinct campaigns, and the largest campaign alone accounts for more than 10% of all incidents. To explore whether the network infrastructure is reused across incidents, we build a graph to connect scam incidents potentially operated by the
C1 C2 C3 C4 C5
#Tactics #Scenarios
3,809 559,666 494 / 209 / 665 / 1,116 698 6,010 0 / 0 / 16 / 21 350 1,643 1 / 0 / 7 / 2 307 6,176 5 / 4 / 5 / 11 287 44,084 211 / 7 / 183 / 78
6 5 4 6 5
18 12 13 9 6
same scammers. Specifically, each node is a scam incident, and two incidents are connected if they share an IP address, domain, email address, or phone number. Each connected component of the graph therefore represents a potential scam campaign. Table 2 summarizes the five largest components by number of reports, total loss, shared infrastructure, and covered tactics. The largest component, C1 , contains 3,809 reports, causes $559K in reported losses, and spans all tactics and scenarios, accounting for more than 10% of all incidents in our dataset. This data reveals complex cross-tactic infrastructures that allow scammers to scale and sustain their operations.
3.5. Case studies Case Study 1: Seasonal scam events. We examine whether scam narratives align with real-world seasonal events, focusing on the U.S. tax season and the back-to-school season. For tax-related scams, we study IRS impersonation scams, which belong to the “Authority & Compliance– Government” scenario; for back-to-school scams, we study student loan scams, which belong to the “Windfall–Funds, Grants & Aid” scenario. We identify both cases using keywords over the description field of scam incident reports. As shown in Figure 7, IRS impersonation scams rise sharply during Jan–Apr 2025 and peak around the tax filing deadline, while student loan scams increase during Jul–Sep 2025, aligning with the back-to-school period. These patterns suggest that scammers adapt their narratives to timely real-world contexts, exploiting periods when victims are more attentive to specific institutions, deadlines, or financial needs. Tax season (Jan–Apr 2025)
120 80 40 0
9 # Report
# Reports
160
Dec 2024
Feb 2025
Apr Jun 2025 2025 Month
Aug 2025
(a) IRS impersonation
Oct 2025
Back-to-school season (Jul–Sep 2025)
6
40 100 89 27.4% 80 30 60 62 18.2% 20 40 33 10 20 4.5% 0 0 onlybor only Both o f n i a l itive Free Sens
% with monetary loss
Cluster #Reports Loss ($) Infra. (D/I/E/P)
and government IDs) and unpaid labor (e.g., package reshipping, fake reviews, and mystery shopping). As shown in Figure 8, scammers frequently obtain identity- and financecritical information, with bank accounts, home addresses, SSNs, and government IDs among the most common disclosures. These harms can expose victims to long-term identity theft. In contrast, free-labor scams are also more likely to involve monetary loss than sensitive-info-only scams (27.4% vs. 4.5%), because victims are often asked to pay upfront costs such as shipping, gift cards, or training fees.
Number of reports
TABLE 2: Top scam campaign clusters identified in our dataset. Infra. (D/I/E/P) denotes the number of shared domains, IPs, emails, and phone numbers within the cluster (Finding 5).
(a) Different scam outcomes.
Bank account Home address SSN Other Name Government ID Phone number Driver's license Email Other PII
25 25 21 18 18 15 14 12
0
44
30
20
40
Number of reports
(b) Sensitive information leak.
Figure 8: Case studies on employment scams. Case Study 3: An international and cross-platform counterfeit scam. We conduct a case study on cluster C5 in Table 2. As shown in Figure 9, the cluster is anchored by three dense subclusters: two Shopifyresolved IPs (23.227.38.65 and 23.227.38.32) and one Alibaba Cloud subcluster centered on three co-hosting IPs (47.76.127.217, 47.91.170.222, and 8.218.208.240). These subclusters are linked via a shared identity, where the domain vivistylee.com (Alibaba Cloud subcluster) connects through the email [email protected] to storefronts in the Shopify subcluster. Manual inspection of incident reports further reveals a consistent cross-platform scam flow (see [40], [41] for the corersponding scam incident reports): scammers register domains and emails (e.g., vivistylee.com, [email protected]) on Alibaba Cloud infrastructure, then run Facebook and Instagram ads impersonating brands such as Sundance Catalogue to promote heavily discounted products. Victims who click these ads are redirected to seemingly legitimate storefronts, place orders, and ultimately receive suspicious or counterfeit shipments from China. Overall, this case illustrates how scammers integrate cloud infrastructure, e-commerce storefronts, reusable email identities, and social media advertising to execute an international counterfeitshopping campaign.
3 0
Dec 2024
Feb 2025
Apr Jun 2025 2025 Month
Aug 2025
Oct 2025
4. Conversation-Aware Scam Scenario Detection
(b) Student loan scams
Figure 7: Case studies on seasonal scam trends. Case Study 2: Victims lose more than money. We manually annotated 500 randomly sampled Employment-scam incident reports and recorded two non-monetary outcomes: sensitive information disclosure (e.g., SSN, bank accounts,
Building on scenario-level knowledge learned from large-scale scam reports, we present EARS, a conversationaware framework for early scam scenario detection in customer-service conversations. As illustrated in Figure 10, EARS consists of a training phase, including Incident Report Pretraining and Transfer Learning for Customer-Service
IP Domain Email Phone Leaf
Figure 9: Visualization of the C5 scam campaign cluster. Nodes with degree greater than one are colored by infrastructure type, while degree-one nodes are shown as gray leaves. The top-degree IP nodes and subcluster connection nodes are labeled. Conversations, and an inference phase. We next describe the threat model and then detail each phase of the framework.
4.1. Threat Model We consider online scams in which an external attacker crafts deceptive narratives and applies PTs to induce harmful victim actions (e.g., monetary transfer). The attacker communicates via channels such as SMS, email, or social media, often impersonating trusted entities (e.g., agencies, employers, or service providers). We assume a social-engineering setting without technical compromise: the attacker does not gain access to victim devices, banking infrastructure, or communication platforms, but instead relies on scalable template reuse and contextual adaptation to enhance credibility. We further assume no access to the victim’s personal data, no direct interaction with the attacker, and exclude data poisoning, infrastructure compromise, automated transaction blocking, and white-box adversaries optimizing against the deployed model. The defender is a financial institution or anti-scam service that observes customer-provided narratives or customerservice conversations during fraud investigations. Given partial and evolving information of the incident, the defender aims to infer the underlying scam scenario as early as possible. Figure 1 illustrates an example conversation for the “Government” scam scenario.
4.2. Incident Report Pretraining We pretrain a scenario classifier on large-scale incident reports to learn transferable scenario-level semantics for conversations. Because scenario labels are not directly available and require manual annotation of clustering results (as discussed Subsection 3.2), we construct the dataset using a two-stage pipeline that combines LLM-based scam content extraction and human verification under active learning. LLM-Based Scam Content Extraction. Incident reports are noisy, often containing emotional complaints, informal language, and irrelevant details, making manual annotation
expensive. Prior work shows that LLMs effectively extract scam-relevant content from such data [26]. We therefore use an LLM to generate structured draft annotations, prompting it to extract scam evidence, predict the scam scenario, and provide a brief rationale. Although direct LLM-based classification achieves only about 60% accuracy, it reliably surfaces useful evidence for human verification. Thus, we use LLM outputs as annotation assistance rather than final labels. The prompt is provided in Appendix Figure 15. Multi-Round Active Learning. Starting from LLM-assisted annotations, we perform multi-round active learning. In each round, we first uniformly sample an equal number of reports per scenario for human labeling to maintain a balanced dataset. We then expand the labels through similarity-based propagation, using cosine similarity between incident report text embeddings with a threshold of 0.7, which substantially increases coverage while preserving label quality based on spot checks. Next, we train a scenario classifier on the expanded set and apply it to the remaining unlabeled data. We select the least confident predictions for human review in the next round, prioritizing informative samples over random selection. After each round, we evaluate performance on the labeled set and adjust the annotation budget based on observed gains. Data Quality Control. The human annotation process involved three PhD-level annotators in computer science. We craft a detailed guideline document defining the taxonomy and labeling protocol, and conduct a calibration stage on a trial set spanning all scenarios to align annotator understanding. Each report is independently labeled by two annotators, and disagreements are resolved by the third annotator, with the final decision determined by majority agreement. The detailed human review protocol is provided in Appendix A.
4.3. Transfer Learning for Customer-Service Conversations We next describe how we conduct the transfer learning that adapts the classifier to the conversational setting. Formal Definition. We represent a customer-service conversation as an ordered sequence of agent–customer utterance pairs, C =
(a1 , c1 ), (a2 , c2 ), . . . , (aT , cT ) ,
(1)
where at and ct denote the agent and customer utterances at turn t, respectively, and T is the total number of turns. The customer narrative through turn t is the concatenation of customer-side utterances, Nt = c1 ⊕ c2 ⊕ · · · ⊕ ct .
(2)
We further define two label spaces, i.e., Yscn = {L EGITIMATE} ∪ S over the scam scenarios S and Ytac = {L EGITIMATE} ∪ T over the tactics T . Phase-Aware Modeling. In reality, these customer-service conversations usually unfold in three broad phases: U NRE LATED phase contains routine or irrelevant content, such as greetings. NARRATIVE phase contains general background
Inference gist:
Incident Report Pretraining Incident Report Pretraining LLM-Based Scam Content Extraction
Pretrain
+ Multi-Round Active learning
Pretrained Classifier Encoder
turns:
pred:
greet
Legit
ID check "IRS arrest threat"
Legit Gov
Scam Incident Reports
Gov
Transfer Learning for Conversations LLM-Based Data Augmentation
Pretrained Encoder
Narrative Prefix with Phase and Scenario Labels
Phase Info Builder Customer Service Conversation
Early Alert
Gov
Customer Utterances Customer Utterances with Phase Labels Customer Utterances with Phase Labels with Phase Labels
Turn-Pos. Emb. Phase-State Feature Turn Offset
Phase-State Proj. Offset Emb.
Latest turn
T Scenario Head
T Loss - PATL
T Tactic Head
T
Narrative Prefix
Phase Classifier
T Accumulate from t=0 to T
T Scenario / Tactic Classifier
Scenario / Tactic Classifier
Predicted Onset : Legitimate :
Phase-State Turn Offset Feature
T
Phase Classifier
T
Phase Info Builder
Updated in Training
Figure 10: Overview of EARS. or administrative information, such as identity verification or transaction confirmation. S CAM E VIDENCE phase contains information useful for identifying the scam scenario, such as the claimed identity of the scammer, the pretext, or the requested payment action. For example in Figure 1, c1 is a greeting (U NRELATED), c2 provides routine administrative operations (NARRATIVE), and diagnostic evidence appears only from c3 onward (S CAM E VIDENCE). To capture this progression, we train a per-turn phase classifier |P|
gψ : ct 7→ ∆ , P = {U NRELATED, NARRATIVE, S CAM E VIDENCE}, (3) which predicts a phase distribution over P . Specifically, we employ a Phase Info Builder to construct two phase-aware signals. First, it converts the predicted phase information into a phase-state feature vector ϕt ∈ R11 , encoding both the current phase and its progression across turns. Second, to model the temporal proximity to the emergence of scam indicators, it encodes the turn offset dt = t − te , where te is the onset turn, defined as the first turn labeled as S CAM E VIDENCE. Details of the phase-state feature vector are provided in Appendix Table 10. LLM-Based Data Augmentation. To better mimic the variability of real customers, we use LLM to mutate each conversation with two complementary augmentation modes. The first mode performs fine-grained paraphrasing. The LLM rewrites each turn in place, adding spoken style and non-structural lexical variation. The second mode performs coarse-grained conversation-level rewriting. The LLM rewrites the full conversation with varied persona, open style, and narrative structure, while preserving the underlying scenario. An independent LLM judge then checks whether the rewritten conversation preserves the original scenario. Mutations that fail this check are rejected. Prompt templates are given in Appendix Figure 16. The final training pool is a mixture of source conversations and accepted mutations from both modes. Model Architecture. Our goal is to train a scenario classifier fθ that predicts the scam scenario and tactic at each
turn: tac fθ : (Nt , ϕt , dt ) 7→ pscn , t , pt
(4)
where pscn and ptac are distributions over scenario and tactic t t labels given the narrative prefix Nt , phase-state features ϕt , and turn offset dt . We initialize fθ from the incident-report classifier (subsection 4.2) and reuse its transformer encoder to encode the narrative prefix: hinc t = einc (Nt ).
(5)
We make each prefix encoding position-aware with a learnable turn-position embedding epos (t), giving the base conversation representation h̄t = hinc t + epos (t),
(6)
which is also used by the no-phase baseline. EARS then injects the two phase signals into the same hidden space through independent parameters: ht = h̄t + Wϕ ϕt + eoff (dt ),
(7)
where Wϕ is a learnable projection of the phase-state vector ϕt and eoff (·) is a learnable embedding of the evidence offset dt . The final hidden representation ht is fed into two parallel heads that predict the scam scenario and tactic: pscn = softmax(Wscn ht ), t
ptac = softmax(Wtac ht ). t (8) This design transfers scenario-level semantics learned from incident reports while adapting to the turn-by-turn structure of customer-service conversations. Training. As discussed above, customer-service conversations progress through three phases. During the U NRE LATED and NARRATIVE phases, EARS outputs L EGITI MATE ; once the conversation enters the S CAM E VIDENCE phase, EARS predicts the corresponding scam scenario and tactic labels. To model this setting, we design a phaseaware turn loss (PATL) for training the scenario and tactic classifier. Let pt be the predicted distribution at turn t, yt ∈ Y = {L EGITIMATE} ∪ Yscn be the classification label, and V be the set of supervised turns. The loss is defined as P wt αy CE(pt , yt ) LPATL (p, y) = t∈V P t , (9) t∈V wt
(10)
encouraging the model to focus on diagnostic evidence while reducing bias toward frequent classes. Inference. During live customer-service conversations, EARS processes turns incrementally. At current turn t, EARS first applies the phase classifier gψ to the current customer utterance ct to determine the phase, P = gψ (ct ). If P = S CAM E VIDENCE, the current turn is marked as the S CAM E VIDENCE onset turn, t̂. For every subsequent turn t > t̂, the utterance ct is passed to the scenario classifier. EARS then employs the Phase Info Builder to construct the phase-state feature vector ϕt from ct and the narrative prefix Nt (i.e., c1 ⊕ c2 ⊕ · · · ⊕ ct ), and predicts the scenario and and ptac tactic distributions, pscn t , respectively. EARS then t emits the scenario and tactic prediction: ŷtscn = arg max pscn t (y), y
ŷttac = arg max ptac t (y). y
(11) The system alerts the customer-service agent with the predicted scenario and tactic to support timely intervention, such as at a6 in Figure 1 and c3 in Figure 10.
5. Evaluation on Scam Scenario Detection 5.1. Evaluation Setup Dataset. We evaluate our approach on a customer-service conversation dataset provided by Charm Security, a leading security company focused on fraud and cybercrime. Because real customer conversations are subject to strict privacy and data-sharing restrictions, the partner generated synthetic conversations based on real fraud-investigation calls. This choice reflects institutional privacy constraints rather than a deployment limitation, as the detector can be applied directly to institution-internal conversations without exposing customer data externally. The dataset contains 1,115 customer-service conversations, with 948 for training, and 167 for testing. Every turn carries both a tactic label and a scenario label. The per-scenario conversation counts and turn statistics are reported in Table 11. Our data augmentation further expands the original 948 training set into 2,277 conversations. Specifically, the finegrained mode produces 7,061 accepted variants and 523 rejected variants, yielding a 93.1% acceptance rate; the coarse-grained mode produces 4,590 accepted variants and 2,994 rejected variants, yielding a 60.5% acceptance rate. Implementation. We implement EARS with three stateof-the-art transformer encoders, RoBERTa [42], ELECTRA [43], and DeBERTa [44], and report results for each
Round 1 Round 2 Round 3 Human-annotated reports Similarity-expanded reports 91.28
87.32
87 85
LLM
83.98
DeBERTa
83
R1
R2
R3
Active Learning Rounds
(a) Tactic
2,202 36,527
91.67
91 89
1,523 28,710
2,413 37,569 83.62
Accuracy (%)
L = LPATL (pscn , y scn ) + LPATL (ptac , y tac ),
TABLE 3: Dataset size across active learning rounds.
Accuracy (%)
where wt is a turn weight and αyt an inverse-frequency class weight (balancing power 0.5) that supports long-tail classes. We up-weight the decisive turns: wt = 2.5 at the first S CAM E VIDENCE onset turn, wt = 1.5 for other S CAM E VIDENCE and NARRATIVE turns, and wt = 1.0 for U NRELATED turns. The scenario and tactic heads are optimized jointly with the combined objective:
82 78 74 70 66
84.38
79.81
LLM
67.50
R1
DeBERTa
R2
R3
Active Learning Rounds
(b) Scenario
Figure 11: Pretraining performance across active learning rounds. Dashed horizontal lines denote direct LLM classification. backbone. Owing to the limited size of the conversation training data and the cost of full fine-tuning, we adopt the base-size variant of each encoder (hidden size 768). Each fold follows the same two-stage recipe: incident-report pretraining (subsection 4.2) initializes the backbone, and conversation fine-tuning then trains the turn-position embedding, phase-conditioned projection, and two task heads. Fine-tuning uses AdamW with weight decay 0.01, encoder learning rate 1×10−5 , head/injection learning rate 8×10−5 , batch size 1 with 4 gradient-accumulation steps (effective batch 4), and 10 epochs; the checkpoint is selected by the validation multi-task score. Inputs are truncated to max turns = 64 turns and a 512-token prefix budget. The phase classifier is trained with weighted cross entropy and AdamW, using a learning rate 2 × 10−5 , maximum length 256, batch size 16, and 5 epochs. Evaluation Metrics. We report the accuracy and macroF1 to measure the classification performance under 5-fold cross-validation. Research Questions. We aim to answer the following research questions: • RQ1: How effective is active learning for incident report pretraining? • RQ2: How effective is EARS in detecting scam scenarios in customer-service conversations? • RQ3: How does each component contribute to EARS’s overall performance?
5.2. RQ1: Incident Reports Pretraining EARS employs multi-round active learning to improve pretraining performance. Table 3 summarizes dataset growth across three active-learning rounds, and Figure 11 reports the corresponding pretraining performance. The number of human-annotated reports increases from 1,523 to 2,413, while similarity-based expansion enlarges the training set from 28,710 to 37,569 reports. As the training corpus grows, tactic accuracy improves from 87.32% to 89.33%
TABLE 4: Incident reports pretraining performance. Model
Tactic Acc. Scenario Acc. Tactic F1 Scenario F1
RoBERTa 90.93±2.13 81.69±4.59 90.99±2.33 79.21±6.74 ELECTRA 91.80±1.72 81.82±3.10 92.01±1.80 81.18±4.13 DeBERTa 91.67±1.34 84.38±1.47 92.05±1.52 85.23±1.43
TABLE 5: Tactic classification results on customer-service conversations. Method
Variants
Tactic F1
Tactic Acc
38.33 51.84 53.51 64.87
46.48 58.03 59.16 67.12
Llama-3.1-8B GPT-OSS-20B Qwen3-32B GPT-5.4
62.11±5.49 64.51±0.79 66.97±2.36 70.94±2.64
63.07±4.53 66.97±0.71 68.70±1.92 72.13±2.60
Transformer
RoBERTa ELECTRA DeBERTa
81.32±1.62 80.67±2.45 80.92±1.40
82.09±1.45 81.21±2.23 81.34±1.24
EARS
RoBERTa ELECTRA DeBERTa
83.35±1.20 83.80±1.11 84.14±1.57 84.41±1.60 83.80±0.54 83.97±0.66
Llama-3.1-8B GPT-OSS-20B LLM Zero-shot Qwen3-32B GPT-5.4
LLM Few-shot
and scenario accuracy from 79.81% to 82.74%, substantially outperforming direct LLM classification (83.98% and 67.50%, respectively). Performance gains largely saturate after Round 2, with tactic accuracy increasing by only 0.05 percentage points in Round 3 (89.28% → 89.33%). We therefore stop the active-learning process after the third round. Table 4 shows the performance of models trained on the final 37,569 incident reports. All models achieve strong tactic-classification performance, exceeding 90% accuracy and F1, indicating that high-level scam tactics are readily distinguishable from incident reports. In contrast, scenario classification is more challenging because of the greater diversity among scenarios; nevertheless, the model still achieves strong performance. DeBERTa achieves the best results, with 84.38% scenario accuracy, 85.23% scenario F1, and 92.05% tactic F1. These results validate the effectiveness of our active learning pipeline and establish a strong foundation for transferring scenario-level knowledge to conversational scam detection. Detailed results across all scenarios and tactics are reported in Appendix Table 12.
5.3. RQ2: Customer-Service Conversations Baselines. We compare EARS against three groups of baselines. First, LLM Zero-shot prompts an LLM with the scenario taxonomy and the customer-service conversation, and asks it to predict the underlying tactic and scenario without task-specific examples. Second, LLM Few-shot uses the same prompt setting but includes five labeled examples with classification rationales. Third, Transformer fine-tunes off-the-shelf transformer encoders and a phase classifier on the conversation dataset without incorporating EARS’s data augmentation. To ensure a fair comparison, we adopt the same inference procedure as EARS (subsection 4.3): the transformer encoder is invoked only after the phase classifier
TABLE 6: Scenario classification results on customerservice conversations. Acc@3 counts a turn as correct if the ground-truth scenario appears among the model’s top-3 predicted scenarios. Method
Variants
Llama-3.1-8B GPT-OSS-20B LLM Zero-shot Qwen3-32B GPT-5.4
Scn F1
Scn Acc
Acc@3
15.56 29.44 33.15 42.52
32.15 44.11 45.88 52.61
47.52 57.48 68.12 78.35
41.52±2.75 47.29±1.25 50.84±2.01 54.23±1.61
65.23±4.19 68.62±1.90 73.07±3.38 81.90±1.01
LLM Few-shot
Llama-3.1-8B GPT-OSS-20B Qwen3-32B GPT-5.4
30.74±3.80 34.59±1.18 42.14±3.02 45.77±1.89
Transformer
RoBERTa ELECTRA DeBERTa
66.15±2.35 71.39±2.26 88.88±1.87 64.89±2.70 70.13±2.36 87.71±1.57 65.39±1.38 71.21±1.85 91.64±1.42
EARS
RoBERTa ELECTRA DeBERTa
70.33±1.68 73.72±1.37 89.81±1.07 71.39±1.60 74.96±2.04 88.61±1.16 69.83±2.01 75.17±1.25 91.04±0.94
identifies an onset turn; otherwise, all preceding turns are treated as legitimate. Overall Effectiveness. Table 5 and Table 6 show the overall performance of EARS in tactic and scenario classification. EARS achieves the best macro-F1 and accuracy on both tasks. For tactic classification, EARS achieves 84.14% macro-F1 and 84.41% accuracy, improving over the best Transformer baseline by 2.82 and 2.32 points, and over the best few-shot LLM by 13.20 and 12.28 points. For scenario classification, EARS obtains up to 71.39% macro-F1 and 75.17% accuracy, outperforming the best Transformer baseline by 5.24 and 3.78 percentage points, respectively. The improvement is more substantial over LLM baselines: compared with the best few-shot LLM, EARS improves scenario macro-F1 by 25.62 points and accuracy by 20.94 points. Beyond higher means, EARS is also more stable across folds: its standard deviations are lower than the Transformer baseline’s on nearly all metrics. These results show that directly prompting LLMs remains insufficient for reliable conversation-level scam understanding, while EARS effectively transfers scenario knowledge to customer-service conversations through augmentation and phase-aware modeling. For Acc@3, EARS achieves performance comparable to the Transformer baselines, while maintaining a slight overall advantage. Although DeBERTa attains a marginally higher mean Acc@3, it exhibits greater variance across folds. This is likely because each tactic is associated with roughly three scenarios on average; once the correct tactic is identified, the correct scenario is often included among the top three predictions, making Acc@3 less sensitive to fine-grained scenario ranking differences. The roughly ten-point gap between EARS’s tactic and scenario classification performance is itself informative. Based on our manual inspection of the classification errors, nearly half still preserve the correct tactic, as the model often confuses closely related scenarios within the same tactic (e.g., fake job offer vs. unpaid labor under employment scam). These near-misses are well covered by the top-3 predictions, which contain the correct scenario for nearly 90% of these sibling confusions and reach up to 91.04%
Per-turn
ELECTRA
Scenario accuracy (%)
RoBERTa 77.89
80 75
76.57
68.62
70
72.46
DeBERTa 78.95
77.82
77.09
74.49
73.09
TABLE 7: Ablation study of context-aware scenario detection. Results report scenario-classification F1; values in parentheses indicate the improvement over the previous row.
Overall
78.42
78.06
76.95
73.97 77.01 74.27
73.33
74.27
67.66
60
0
56.17
56.05 1
2
3
4
5
6
0
1
2
3
4
5
6
0
1
2
3
4
5
ELECTRA
RoBERTa
DeBERTa
Baseline (no aug., no phase) 64.87 65.82 65.22 + Augmentation 2× 67.70 (+2.84) 70.00 (+4.18) 67.39 (+2.17) + Phase-aware modeling (EARS) 71.39 (+3.68) 70.33 (+0.34) 69.83 (+2.44)
65
55 57.84
Added factor
6
Offset from scam evidence onset
Figure 12: Per-turn scenario classification accuracy of EARS. Acc@3. Therefore, surfacing the top-3 candidate scenarios provides useful triage support and better reflects the model’s ability to handle the taxonomy’s ambiguity. Late vs. Early Correctness. Our goal is to detect scam scenarios as early as possible, but early prediction naturally trades off with accuracy: as the conversation progresses, customers reveal more scam evidence, allowing the model to make more confident and accurate predictions. To quantify this tradeoff, we evaluate EARS at different turns and measure how scenario-classification accuracy changes over time. Specifically, we define the turn offset as d = t − te , where d = 0 denotes the first turn in the S CAM E VIDENCE phase. Figure 12 reports scenario classification accuracy with varying offset. Overall, accuracy increases steadily as the offset grows. At the onset turn (d = 0), EARS is less accurate (∼56– 58%), which is expected because limited evidence is available at this point to enable accurate classification. Nevertheless, a single additional turn is decisive: at d = 1, scenario accuracy improves by 11–16 percentage points and already reaches the overall accuracy. Subsequent gains are smaller but consistent, reaching around 78–79% by d = 6. These results show that EARS typically identifies the correct scenario within one to two turns after scam evidence first appears, supporting early intervention while still benefiting from additional evidence as the conversation unfolds.
5.4. RQ3: Ablations In this part, we ablate two techniques used for transferring the incident-report knowledge to customer-service conversations: LLM-based data augmentation and phase-aware modeling. Table 7 shows that both techniques consistently improve scenario-classification F1 across all three backbone models. Starting from the baseline without augmentation or phase modeling, adding LLM-based augmentation improves F1 by 2.84%, 4.18%, and 2.17% for ELECTRA, RoBERTa, and DeBERTa, respectively. This suggests that augmented conversations help bridge the distribution gap between incident reports and customer-service dialogues by exposing the model to more diverse conversational expressions of the same scam scenarios. Phase-aware modeling further improves performance on top of augmentation, with additional gains of 3.68%, 0.34%, and 2.44% for ELECTRA,
TABLE 8: Augmentation-ratio sweep results without phaseaware modeling. Values are scenario macro-F1 in percentage points. Backbone
0×
1×
2×
4×
8×
ELECTRA 64.87 66.52 67.70 66.50 68.13 RoBERTa 65.82 65.97 70.00 69.77 69.01 DeBERTa 65.22 66.45 67.39 69.08 66.46
RoBERTa, and DeBERTa, respectively. The improvement indicates that explicitly modeling conversation phases helps the classifier focus on scenario-relevant turns, rather than treating all turns equally. Data Augmentation Ratio. Table 8 reports the performance of EARS under different augmentation ratios without phase conditioning. Overall, data augmentation improves scenario macro-F1 for all backbones, with gains over the no-augmentation baseline ranging from 3.26% to 4.18% at their best settings. The optimal ratio varies across backbones: ELECTRA achieves its best F1 at 8× (68.13%), RoBERTa peaks at 2× (70.00%), and DeBERTa performs best at 4× (69.08%). However, larger augmentation ratios do not always yield further gains, as performance drops for RoBERTa at 4× and for DeBERTa at 8×. This suggests that moderate augmentation is generally effective, while excessive augmentation may introduce noise or redundant conversational patterns. Phase Input at Inference Time. Table 9 replace EARS’s phase classifier with an oracle setting that uses the groundtruth phase label at inference time. We can see that oracle phase labels consistently improve performance across all backbones, with moderate gains: scenario F1 increases by 2.69%, 2.97%, and 2.27% for ELECTRA, RoBERTa, and DeBERTa, respectively. Similarly, scenario accuracy improves by 4.29%, 4.55%, and 3.67%. These gaps indicate that the learned phase classifier already supplies useful phase information, while also showing that more accurate phase prediction could further improve scenario detection.
6. Discussion Extendable Scenario Taxonomy. Our taxonomy captures the major scam scenarios observed in our dataset, but it is not exhaustive. In practice, rarer scenarios may appear infrequently, and new scenarios may emerge over time. To handle reports that are unidentifiable or too ambiguous during annotation and model training, we include an “Other” scenario during human annotation. Future work can study the temporal evolution of scam scenarios and update the taxonomy as new patterns emerge, which can be well supported by following our empirical study methodology.
TABLE 9: Inference-time phase-aware supervision. “Oracle” uses the ground-truth phase label instead of the phase classifier’s output. Backbone
Setting
Scenario F1 Scenario Acc
ELECTRA
Phase classifier (EARS) Oracle
71.39±1.60 74.08±1.57
74.96±2.04 79.25±1.86
RoBERTa
Phase classifier (EARS) Oracle
70.33±1.68 73.30±1.83
73.72±1.37 78.27±1.37
DeBERTa
Phase classifier (EARS) Oracle
69.83±2.01 72.10±2.14
75.17±1.25 78.84±1.27
Unbalanced Scenario Distribution. As shown in Figure 13, scam scenarios follow a skewed, long-tailed distribution, which is expected in real-world scam reports but can affect detection performance. Although we use similaritybased expansion during active learning to look for more samples, some scenarios remain inherently rare in BBB Scam Tracker. Future work can mitigate this limitation by incorporating additional scam-reporting platforms [45], [46] or by using synthetic data augmentation to improve coverage of low-frequency scenarios. Data Augmentation. Our two-mode LLM mutation pipeline improves robustness without requiring new labels, and the independent judge keeps label drift low. It nonetheless has inherent limits. Because mutation is label-preserving by construction, it broadens lexical and stylistic coverage within existing scenarios but cannot introduce genuinely new scenarios or improve coverage of the long tail. Our ablation shows diminishing returns beyond a 2× ratio (Table 8). Future work could explore controllable or personaconditioned generation, augmentation targeted at rare scenarios, and mixing in additional real conversations under the same privacy constraints. Conversation Context Window. Our detector inherits the 512-token context limit of base-size transformer encoders, and we cap inputs at max turns = 64 with a 512-token prefix budget. In our corpus this is usually sufficient (mean 12.9 turns and ≈ 350 sub-word tokens per conversation), but the longest conversations are truncated, which drops earlier context that may carry diagnostic evidence, and real customer-service conversations can be longer still. Several directions could relax this constraint: (i) long-context encoders that widen the window directly; (ii) hierarchical or streaming encoding that encodes each turn once and aggregates across turns, which both avoids re-truncating long prefixes and removes the cost of re-encoding the entire prefix at every turn during live, turn-by-turn deployment; and (iii) compress or summarize historical turns by following research in LLM agent memory [47].
7. Related Work Empirical Studies of Online Scams. Most prior efforts focus on specific forms of online scams. Park et al. [27] studied targeted Nigerian scams on Craigslist using honeypot advertisements and automated scammer interactions. Miramirkhani et al. [28] and Acharya et al. [29] conducted
large-scale studies of technical support scams, revealing how scammers combine deceptive webpages, malvertising, phone calls, remote-access tools, and payment channels to monetize victims. Acharya and Holz [30] examined pigbutchering scams, highlighting their long-term grooming process and severe financial harm. Kharraz et al. [48] and Ma et al. [49] studied web- and mobile-ad-driven scam campaigns, showing that such campaigns not only monetize user engagement but also expose users to broader security threats, including malware and identity theft. These studies provide valuable insights into individual scam types, but lack a comprehensive measurement of the broader scam ecosystem. Early Classification in Harmful Conversations. Prior work has studied early classification as a sequential decision problem, in which a model must balance prediction accuracy with the cost of waiting for more evidence [50], [51], [52]. Dialogue classification work further shows that multi-turn context is essential for understanding intent, risk, and user state, because critical evidence may be distributed across turns rather than appearing in a single utterance [32], [33], [34]. Similar ideas have been applied to harmful conversation detection, such as identifying risky online interactions or abusive dialogues from partial conversation prefixes [24], [25]. Building on these insights, our work studies early scam scenario detection in customer-service conversations, where the defender must classify a scam scenario from a progressively revealed customer narrative. Conversation-based Scam Detection. The most relevant research direction to our setting is conversation-based scam detection. Traditional methods detect vishing or fraud from conversation transcripts using handcrafted signals, such as clustering scam signatures derived from speech acts [53] or applying decision trees over linguistic features [54]. Recent efforts further explore LLMs for real-time scam detection from ongoing conversations [22], [23]. Our setting differs because the defender observes a second-hand customer narrative elicited by a bank representative, rather than the original attacker-victim conversation.
8. Conclusion In this paper, we conduct a large-scale study of scam scenarios using 102,054 real-world incident reports and construct a hierarchical taxonomy of 18 scenarios grouped into 6 high-level tactics based on associated PTs. Our analysis reveals that scammers reuse PTs, scenarios, and infrastructure across diverse contexts to scale their operations. Building on these insights, we develop a conversationaware scam scenario detection method for customer-service conversations that enables early detection from partial and evolving dialogues. Experiments on 1,115 conversations synthesized from real customer-service calls demonstrate that our approach effectively identifies scam scenarios in practical settings and that scenario-level knowledge provides a practical and interpretable foundation for both analysis and defense.
References [1]
[2]
[3]
FTC, “Uncovering hidden fraud trends in 2025: The rise of job scams and data exploitation,” 2025. [Online]. Available: https: //www.moodys.com/web/en/us/kyc/resources/insights/uncovering-hid den-fraud-trends-the-rise-of-job-scams-and-data-exploitation.html ——, “How To Recognize and Avoid Phishing Scams,” 2025. [Online]. Available: https://consumer.ftc.gov/articles/how-recognize -avoid-phishing-scams B. P. Department, “Benicia Police Department’s Post,” 2025. [Online]. Available: https://www.facebook.com/beniciapd/posts/-in-2 025-identity-theft-cases-exceeded-115-million-in-the-us-in-2025-f inancial-f/1197685249179145/
[4]
S. News, “Love shouldn’t cost a thing, but romance scammers bilk $1 billion from Americans each year,” 2025. [Online]. Available: https://spectrumlocalnews.com/nc/triad/politics/2026/02/13/romance -scams-are-costing-americans--1-billion-a-year
[5]
I. Euro, “How online financial fraud is taking Europeans for billions,” 2025. [Online]. Available: https://www.investigate-europe. eu/themes/investigations/-scam-europe
[6]
GASA, “Germans lose C10.6 billion to scams in 12 months,” 2025. [Online]. Available: https://www.gasa.org/post/germans-lose-10-6-b illion-to-scams-in-12-months
[7]
BioCatch, “French lose C7.6 billion to scams in 12 months,” 2025. [Online]. Available: https://www.biocatch.com/press-release/french-l ose-billions-to-scams-in-12-months
[8]
GASA, “New Study Reveals 63% of Southeast Asians Experienced Scams in Past Year, Prompting Urgent Call for Stronger Anti-Scam Measures,” 2025. [Online]. Available: https://www.gasa.org/post/ne w-study-reveals-63-of-southeast-asians-experienced-scams-in-pas t-year
[9]
Chainalysis, “Record $17 Billion Estimated Stolen in Crypto Scams and Fraud in 2025 as Impersonation Tactics and AI Enablement Surge,” 2026. [Online]. Available: https://www.chainalysis.com/blog /crypto-scams-2026/
[10] A. Shimbun, “Japan crime rate surges as fraud losses hit record 324.1 billion yen,” 2026. [Online]. Available: https: //www.asahi.com/ajw/articles/16351757 [11] Payment Systems Regulator, “Authorised push payment (app) scams performance report,” 2024. [Online]. Available: https: //www.psr.org.uk/ [12] UK Finance, “Annual fraud report 2024,” 2024. [Online]. Available: https://www.ukfinance.org.uk/policy-and-guidance/reports-and-publi cations/annual-fraud-report-2024 [13] L. Bustio-Martı́nez, V. Herrera-Semenets, J. L. Garcı́a-Mendoza, M. Á. Álvarez-Carmona, J. Á. González-Ordiano, L. Zúñiga-Morales, J. E. Quiróz-Ibarra, P. A. Santander-Molina, and J. van den Berg, “Uncovering phishing attacks using principles of persuasion analysis,” Journal of Network and Computer Applications (JNCA), vol. 230, p. 103964, 2024. [14] T. T. Longtchi, R. M. Rodriguez, K. Gwartney, E. Ear, D. P. Azari, C. P. Kelley, and S. Xu, “Quantifying psychological sophistication of malicious emails,” IEEE Access, vol. 12, pp. 187 512–187 535, 2024. [15] R. Montanẽz Rodriguez and S. Xu, “Cyber social engineering kill chain,” in Proceedings of the International Conference on Science of Cyber Security (SciSec). Springer, 2022, pp. 487–504. [16] T. T. Longtchi, R. M. Rodriguez, L. Al-Shawaf, A. Atyabi, and S. Xu, “Internet-based social engineering psychology, attacks, and defenses: a survey,” Proceedings of the IEEE, vol. 112, no. 3, pp. 210–246, 2024. [17] FBI Internet Crime Complaint Center (IC3), “Smishing scam regarding debt for road toll services,” https://www.ic3.gov/PSA/2024/PSA 240412, 2024, public Service Announcement.
[18] A. Vakulov, “Fake toll messages are flooding phones in a nationwide scam,” Forbes, 2025. [Online]. Available: https: //www.forbes.com/sites/alexvakulov/2025/04/06/fake-toll-message s-are-flooding-phones-in-a-nationwide-scam/ [19] Cyber Defense Magazine, “High speed smishing: The psychology behind toll road scams,” https://www.cyberdefensemagazine.com/hig h-speed-smishing-the-psychology-behind-toll-road-scams//, 2024. [20] K. Krombholz, H. Hobel, M. Huber, and E. Weippl, “Advanced social engineering attacks,” Journal of Information Security and applications (JISA), vol. 22, pp. 113–122, 2015. [21] M. Schmitt and I. Fléchais, “Digital deception: Generative artificial intelligence in social engineering and phishing,” Artificial Intelligence Review, vol. 57, no. 324, 2024. [22] Z. Shen, S. Yan, Y. Zhang, X. Luo, G. Ngai, and E. Y. Fu, “” it warned me just at the right moment”: Exploring llm-based real-time detection of phone scams,” in Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2025, pp. 1–7. [23] W. Sun, S. Ma, Y. Li, T. Ma, Z. Wang, C. Nelson, X. Xiao, and Y. Ye, “Prescam: A benchmark for predicting scam progression from early conversations,” arXiv preprint arXiv:2605.12243, 2026. [24] J. An, S. Ryu, H. Do, Y. Kim, J. Ok, and G. Lee, “Revisiting early detection of sexual predators via turn-level optimization,” in Proceedings of the Conference of Americas Chapter of the Association for Computational Linguistics (NAACL), 2025, pp. 4713–4724. [25] K. Chehbouni, M. De Cock, G. Caporossi, A. Taik, R. Rabbany, and G. Farnadi, “Enhancing privacy in the early detection of sexual predators through federated learning and differential privacy,” in Proceedings of the AAAI conference on artificial intelligence (AAAI), vol. 39, no. 27, 2025, pp. 27 887–27 895. [26] S. Ma, T. Ma, J. Liu, W. Song, Z. Liang, X. Xiao, and Y. Ye, “Psyscam: A benchmark for psychological techniques in real-world scams,” in Findings of the Association for Computational Linguistics: EMNLP 2025, 2025, pp. 12 623–12 637. [27] Y. Park, J. Jones, D. McCoy, E. Shi, and M. Jakobsson, “Scambaiter: Understanding targeted nigerian scams on craigslist,” system (SYSTEM), vol. 1, p. 2, 2014. [28] N. Miramirkhani, O. Starov, and N. Nikiforakis, “Dial one for scam: A large-scale analysis of technical support scams,” Proceedings of the Network and Distributed System Security Symposium (NDSS), 2017. [29] B. Acharya, M. Saad, A. E. Cinà, L. Schönherr, H. Dai Nguyen, A. Oest, P. Vadrevu, and T. Holz, “Conning the crypto conman: Endto-end analysis of cryptocurrency-based technical support scams,” in 2024 IEEE Symposium on Security and Privacy (SP), 2024, pp. 17– 35. [30] B. Acharya and T. Holz, “An explorative study of pig butchering scams,” arXiv preprint arXiv:2412.15423 (APA), 2024. [31] Better Business Bureau, “BBB Scam Tracker,” Accessed: March, 2026. [Online]. Available: https://www.bbb.org/scamtracker/lookupsc am [32] I. Serban, A. Sordoni, R. Lowe, L. Charlin, J. Pineau, A. Courville, and Y. Bengio, “A hierarchical latent variable encoder-decoder model for generating dialogues,” in Proceedings of the AAAI conference on artificial intelligence (AAAI), vol. 31, no. 1, 2017. [33] V. Raheja and J. Tetreault, “Dialogue act classification with contextaware self-attention,” in Proceedings of the Conference of North American Chapter of the Association for Computational Linguistics (NAACL), 2019, pp. 3727–3733. [34] C. Qu, L. Yang, W. B. Croft, Y. Zhang, J. R. Trippas, and M. Qiu, “User intent prediction in information-seeking conversations,” in Proceedings of the Conference on Human Information Interaction and Retrieval (CHIIR), 2019, pp. 25–33.
[35] T. Sharma and M. Bashir, “An analysis of phishing emails and how the human vulnerabilities are exploited,” in Proceedings of the International Conference on Applied Human Factors and Ergonomics (AHFE), 2020, pp. 49–55. [36] A. Ben, T. Rahav, D. Illaev, A. Nahon, and A. Grushka, “The human vulnerabilities & exploits (hve) framework,” 2026. [37] Better Business Bureau Scam Tracker, “Scam type glossary,” Accessed: April, 2026. [Online]. Available: https://scamsurvivaltool kit.bbbmarketplacetrust.org/scam-type-glossary/ [38] M. Beals, M. DeLiema, and M. Deevy, “Framework for a taxonomy of fraud,” Financial Fraud Research Center, 2015. [39] M. Grootendorst, “Bertopic: Neural topic modeling with a class-based tf-idf procedure,” arXiv preprint arXiv:2203.05794 (APA), 2022. [40] Better Business Bureau Scam Tracker, “Scam Report of Sundance Catalogue Impersonation via Facebook Ads,” Accessed: May, 2026. [Online]. Available: https://www.bbb.org/scamtracker/lookupscam/1 080362 [41] ——, “Scam Report of Sundance Catalogue Impersonation via Instagram Ads,” Accessed: May, 2026. [Online]. Available: https: //www.bbb.org/scamtracker/lookupscam/1080362 [42] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A robustly optimized bert pretraining approach,” arXiv preprint arXiv:1907.11692 (APA), 2019. [43] K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning, “Electra: Pretraining text encoders as discriminators rather than generators,” in Proceedings of the International Conference on Learning Representations (ICLR), 2020. [44] P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding-enhanced bert with disentangled attention,” in Proceedings of the International Conference on Learning Representations (ICLR), 2021. [45] Chainabuse, “Chainabuse: Report and Combat Crypto Scams,” 2026. [Online]. Available: https://chainabuse.com/ [46] ScamSearch, “Global Scam Database,” 2026. [Online]. Available: https://scamsearch.io/ [47] W. Xu, Z. Liang, K. Mei, H. Gao, J. Tan, and Y. Zhang, “A-mem: Agentic memory for llm agents,” in Advances in Neural Information Processing Systems (NIPS), 2025, pp. 17 577–17 604. [48] A. Kharraz, W. Robertson, and E. Kirda, “Surveylance: Automatically detecting online survey scams,” in Proceedings of the IEEE Symposium on Security and Privacy (SP), 2018, pp. 70–86. [49] S. Ma, C. Chen, S. Yang, S. Hou, T. J.-J. Li, X. Xiao, T. Xie, and Y. Ye, “Careful about what app promotion ads recommend! detecting and explaining malware promotion via app promotion graph,” Proceedings of the Network and Distributed System Security Symposium (NDSS), 2024. [50] G. Dulac-Arnold, L. Denoyer, and P. Gallinari, “Text classification: A sequential reading approach,” in Proceedings of the European conference on information retrieval (ECIR), 2011, pp. 411–423. [51] S. G. Burdisso, M. Errecalde, and M. Montes-y Gómez, “A text classification framework for simple and effective early depression detection over social media streams,” Expert Systems with Applications, vol. 133, pp. 182–197, 2019. [52] ——, “τ -ss3: A text classifier with dynamic n-grams for early risk detection over text streams,” Pattern Recognition Letters, vol. 138, pp. 130–137, 2020. [53] A. Derakhshan, I. G. Harris, and M. Behzadi, “Detecting telephonebased social engineering attacks using scam signatures,” in Proceedings of the ACM workshop on security and privacy analytics (IWSPA), 2021, pp. 67–73. [54] N. Bajaj, T. G. Constance, M. Rajwadi, J. Wall, M. Moniri, C. Glackin, N. Cannings, C. Woodruff, and J. Laird, “Fraud detection in telephone conversations for financial services using linguistic features,” arXiv preprint arXiv:1912.04748, 2019.
Appendix A. Human Review Protocol Each annotation instance corresponds to one scam incident report. Annotators are asked to assign one dominant scenario label from our taxonomy. When a report contains multiple scam elements, annotators label the scenario that best captures the scammer’s primary pretext, victim-facing narrative, and requested action. If no scenario applies, annotators assign the label “Other.” If the report does not contain enough information to make a confident decision, annotators mark it as “need review” for later adjudication. Guidelines and Decision Rules. Annotators follow a twostep decision process. First, they identify the high-level tactic by examining the prominent PTs used in scammer’s pretext. Second, within the identified tactic, annotators examine the report content to select the most appropriate scenario. Agreement and Adjudication. To assess the reliability of the human annotation process, we measure inter-annotator agreement on the double-labeled reports using Cohen’s κ, which accounts for agreement that may occur by chance. The two primary annotators achieved substantial agreement (κ = 0.75), indicating that the scenario taxonomy and annotation guideline were consistently understood and applied. Cases with disagreement were further reviewed by a third annotator. If two annotators agreed after adjudication, the majority label was used as the final label; otherwise, the case was discussed jointly until consensus was reached. This process helps ensure that the final scenario labels are both reliable and grounded in the report narratives.
Appendix B. Additional Figure and Tables TABLE 10: The two phase-aware signals constructed by the Phase Info Builder. # Feature
Description
1 cur p unrelated Indicator that the current turn is U NRELATED. 2 cur p narrative Indicator that the current turn is NARRATIVE. 3 cur p evidence Indicator that the current turn is S CAM E VIDENCE. 4 pre mean p narr Running mean of the narrative channel over the prefix, i.e., the share of NARRATIVE turns so far. 5 pre mean p evid Running mean of the evidence channel over the prefix. 6 pre max p evid Running maximum of the evidence channel, i.e., whether an evidence turn has appeared in the prefix. 7 related frac Fraction of prefix turns labeled scam-related, i.e., NARRATIVE or S CAM E VIDENCE. 8 evidence frac Fraction of prefix turns labeled S CAM E VIDENCE. 9 first evid norm Normalized position of the first evidence turn within the prefix; 1 if no evidence turn has appeared yet. 10 since evid norm Normalized number of turns elapsed since the first evidence turn; 0 before evidence appears. 11 narr2evid frac Fraction of adjacent prefix turn pairs that switch from NARRATIVE to S CAM E VIDENCE, capturing escalation into evidence. 12 evid offset
Signed offset dt = t − te to the S CAM E VIDENCE onset turn. Mapped to the hidden space by the learnable embedding eoff (dt ).
Tactic / Scenario
#All Turns
Authority & Compliance Government Legal Other Authority Windfall Good Deals Investment & Trading Fund, Grants & Aid Lottery & Sweepstakes Consumer & Services Financial Services Tech & Online Service E-commerce Retail Insurance & Warranty Employment Fake Job Offer Unpaid Labor Relationship & Trust Pig Butchering Friend & Relative Charity Other Relationship Extortion Sextortion Hack & Data Breach Total
te
109 65 36
10.7 3.3 11.8 3.2 13.5 4.2
71 45 23 20
11.7 14.4 11.8 11.1
4.2 3.8 2.7 3.0
70 58 47 43 24
11.9 13.4 12.1 12.3 10.2
3.0 3.7 3.7 3.8 3.6
135 87
12.7 3.1 14.9 4.8
45 41 29 51
16.6 12.8 12.2 14.4
68 48
14.1 3.9 15.0 4.1
1,115
12.9 3.8
5.3 3.9 4.4 3.6
7 347
6,000
8,000
# of Reports
Figure 13: Distribution of scam scenarios. Resilient Domain
TABLE 11: Customer-service conversation dataset statistics. #All is the total conversation count. Turns is the mean number of turns per conversation. te the mean annotated onset turn of S CAM E VIDENCE.
Financial Services 3 989 Fake Job Offer Legal 3 153 2 789 E-commerce Lottery & Sweepstakes 2 537 Insurance & Warranty 2 412 2 398 Government 2 283 Retail Funds, Grants & Aid 1 775 1 602 Tech & Online Service Unpaid Labor 771 594 Good Deals Friends & Relatives 289 Pig Butchering 202 Investment & Trading 191 Sextortion 167 Hack & Data Breach 150 Charity 97 0 2,000 4,000
40%
# Domains
30%
100 500 1k
20% 10% 0%
10%
20%
30%
40%
Disposable Domain 1. Retail 10. Funds, Grants and Aid 2. Fake Job Offer 11. Legal 3. Government 12. Pig Butchering 4. Unpaid Labor 13. Investment and Trading 5. E-commerce 14. Insurance and Warranty 6. Lottery and Sweepstakes 15. Friends and Relatives 7. Financial Services 16. Hack and Data Breach 8. Tech and Online Service 17. Sextortion 9. Good Deals 18. Charity
Figure 14: Resilient and disposable domain rates across scam scenarios (Finding 4).
TABLE 12: F1 score of scam incident report classification (%). Tactic / Scenario
Sup. RoB. ELEC. DeB.
Authority & Compliance 61 91.63 Government 29 91.92 Legal 32 89.47 Consumer & Services 137 88.10 E-commerce 24 86.65 Financial Services 34 72.40 Insurance & Warranty 27 91.14 Retail 26 74.60 Tech & Online Service 26 81.33 Employment 90 96.92 Fake Job Offer 50 89.41 Unpaid Labor 40 84.48 Extortion 24 91.74 Hack & Data Breach 9 23.04 Sextortion 15 87.94 Relationship & Trust 42 87.50 Charity 14 74.66 Friend and Relative 14 71.99 Pig Butchering 14 77.24 Windfall 107 90.05 Fund, Grants and Aid 26 87.70 Good Deals 20 81.32 Investment and Trading 16 74.10 Lottery and Sweepstakes 45 86.36
91.16 88.57 87.88 89.50 86.29 67.74 92.07 69.19 84.24 95.50 87.43 84.50 94.84 45.16 87.77 88.90 83.92 74.33 79.56 92.15 91.63 83.12 79.54 88.34
91.41 90.83 91.64 89.24 85.73 75.07 90.13 68.29 82.48 95.80 88.68 83.69 93.81 70.71 94.84 91.06 92.18 84.49 85.75 91.01 88.93 87.58 86.24 86.90
461 90.99 461 79.21
92.01 81.18
92.05 85.23
Overall Tactic Overall Scenario
S YSTEM P ROMPT You are an expert at analyzing scam incident reports. A scam scenario is the recurring operational scheme that captures the core strategy used by the scammer, including the claimed pretext, the victim-facing narrative, and the requested action or payment method. Given a scam incident report, your task is to identify the scam scenario that best describes the underlying scheme. The complete taxonomy of considered scenarios is provided below. {scenario taxonomy}
U SER P ROMPT Scam Incident Report: {description} Analysis Steps: 1) Identify the dominant psychological technique that makes this scam succeed, and use it to determine the underlying tactic. 2) Based on the victim-facing narrative, select the most appropriate scenario from the taxonomy. Output Requirements: Provide a concise summary of your reasoning, and return the result strictly as a JSON object: { "tactic": {tactic name} | "Other", "scenario": {scenario name} | "Other", "summary": {brief
classification rationale}, "reference": {scam-relevant content} }
{few-shot examples}
Figure 15: Prompt template used by the LLM annotator to produce draft scenario labels for scam incident reports.
TABLE 13: Definition of the psychological techniques used in this paper. Psychological Technique
Description
Authority
Leveraging the tendency of people to obey authority figures. Scammers impersonate roles such as government officials, executives, police officers, lawyers, doctors, or religious leaders to gain compliance. Exploiting desire and reward-seeking by presenting promises of large gains, easy money, or unusually attractive opportunities that can override rational judgment. Triggering fear, panic, or anxiety to pressure victims into immediate compliance and suppress careful reasoning. Persuading victims by appearing friendly, charming, supportive, or flattering, thereby increasing rapport and reducing suspicion. Creating a sense of time pressure or limited availability so that victims act quickly without sufficient verification or reflection. Fabricating a credible story, scenario, or identity to establish legitimacy and gain the victim’s trust. Exploiting socially desirable norms such as kindness, helpfulness, reciprocity, and obligation, e.g., making victims feel they should return a favor or cooperate. Exploiting the tendency of people to behave consistently with their prior commitments, statements, or actions. Influencing victims by referencing the behavior, approval, or participation of others, suggesting that the scam is legitimate because many people are doing it.
Phantom Riches Fear and Intimidation Liking Urgency and Scarcity Credibility Evoking Social Norms Consistency Social Proof
TABLE 14: Raw outputs of the BERTopic model. The raw topics provide candidate clusters for scenario construction rather than final scenario labels. Some topics exhibit semantic overlap, such as loan-related Topics 3 and 6, and job-related Topics 5, 14, and 16. This motivates the subsequent manual inspection, merging, and refinement steps used to construct the final scam scenario taxonomy. Topic
Reports
Top-10 keywords
LLM summary based on top-50 keywords
0
8,524
card, bank, check, told, flight, credit, credit card, gift, cards, won
1
5,978
2
5,407
3
4,408
4
4,395
5
3,880
6
2,495
7
2,443
8
2,388
9
1,645
10
1,358
11
1,296
12
1,287
13
1,260
14
1,011
15
938
16
904
17
832
18
789
19
771
warranty, letter, coverage, home, home warranty, notice, property, vehicle, mortgage, final leave, days, earn, merchants, days paid, paid, paid annual, parttime, annual leave, paid annual leave loan, calls, numbers, different, calling, approved, times, different numbers, applied, list court, case, office, documents, legal, place, calling, served, voicemail, filing packages, job, interview, package, shipping, position, month, offer, linkedin, check press, loan, approval, underwriting, connect, opt, personal loan, just need, quick, longer refund, return, product, ordered, china, item, dress, items, shoes, order withdraw, trading, funds, crypto, investment, platform, withdrawal, deposit, invest, wallet invoice, norton, subscription, mcafee, geek, renewal, payment, geek squad, squad, consulting facebook, site, instagram, fake, people, page, scammer, fb, real, using medicare, insurance, medical, health, health insurance, doctor, patient, knee, hospital, plan door, garage, car, driveway, garage door, parking, repair, told, contractor, price form, beneficial, ownership, letter, filing, annual, records service, mandatory, reporting, ein interview, position, teams, microsoft teams, microsoft, hiring, manager, entry, data entry, data paypal, invoice, transaction, paypal account, usd, received email, invoice number, order, purchase, bitcoin trial period, trial, period, hr, global streaming, wbd global, wbd, streaming, 3day paid, 3day passport, government, renewal, site, application, renew, form, government website, global, personal information toll, avoid, unpaid, late, unpaid toll, balance, settle, link, late fees, outstanding tax, filings, tax debt, end month, settled, taxes, wage, wage garnishment, garnishment, overdue
Mixed payment and consumer scams involving credit cards, bank checks, prizes, travel reservations, gift cards, or cancellation-related charges. Warranty-renewal or coverage notice scams that pressure victims through mail about home, vehicle, mortgage, or property protection plans. Remote part-time job scams promising easy earnings, flexible schedules, paid leave, and short daily work for merchant visibility or similar tasks. Loan-related robocall or telemarketing scams that repeatedly contact victims from different numbers with claims of loan approval. Legal threat scams that claim victims have court cases, legal documents, filings, or pending service of process to induce fear and response. Fake job and package-reshipping scams where victims are recruited for shippingrelated positions and asked to receive or forward packages. Personal loan approval scams using scripted messages that ask victims to confirm information, connect with agents, or proceed with funding. Online shopping scams involving low-quality, counterfeit, or misrepresented products, often followed by failed refund or return requests. Crypto, trading, and investment scams where victims are shown profits but are blocked from withdrawing funds without additional deposits or fees. Fake subscription renewal and tech-support invoice scams impersonating Norton, McAfee, Geek Squad, or similar service providers. Social media impersonation scams using fake Facebook or Instagram profiles, pages, ads, photos, or hacked accounts to lure victims. Medicare, health insurance, and medical billing scams involving insurance plans, doctors, medical supplies, labs, or patient information. Local service scams involving garage doors, car repairs, towing, contractors, locks, or home repair work with inflated prices or incomplete jobs. Business compliance filing scams that send official-looking letters about beneficial ownership, annual filings, records services, or penalties. Fake hiring scams for data-entry or remote positions, often conducted through Microsoft Teams, Signal, or fake HR/hiring manager personas. PayPal invoice scams that send fake transaction or purchase notices, often involving unauthorized charges, Bitcoin, or seller payment claims. Fake streaming or media-company job scams advertising paid trial periods, daily tasks, promotions, and customer service roles. Fake government service websites for passport renewal or applications that collect personal information, fees, and credit card details. Toll-payment phishing scams that send messages about unpaid tolls, late fees, or outstanding balances and direct victims to malicious links. Tax debt relief scams claiming to reduce overdue taxes, stop wage garnishment, or settle tax filings through relief programs.
P ROMPT U SER P ROMPT (M ODE 1: F INE -G RAINED PARAPHRASING ) Task: Rewrite {num turns} customer turns to sound like a real phone-call transcript. Target Label: {binary label} / {tactic} / {scenario} Rules: • Mode: {mutation mode} (filler, hesitation, self-repair, ASR-like noise, paraphrase). • Hard rules: No scam evidence before turn {first evidence turn}; label ({binary label} / {tactic} / {scenario}) stays exact; do not weaken the core scam mechanism. • Forbidden: {forbidden drifts} Source Turns: {source turns table} Output: JSON array of {num turns} strings.
S YSTEM P ROMPT You are an expert in analyzing online scam reports. Given the top representative keywords of a topic produced by topic modeling over scam incident reports, characterize the underlying scam scenario. A coherent topic typically corresponds to a recurring scheme with a common pretext, victim-facing narrative, and requested action or payment method; surface this scheme rather than restating keywords.
P ROMPT U SER P ROMPT (M ODE 2: C OARSE R EWRITING ) Task: Rewrite the {num turns} customer turns into a noticeably different but label-faithful conversation. Target Label: {binary label} / {tactic} / {scenario} Rules: • Keep the same number of turns and the first evidence turn at {first evidence turn}. • Make tactic and scenario cues clear; do not drift to another tactic or scenario. • Freely change wording, local details, and style; do not optimize for lexical overlap. Source Turns: {source turns table} Output: JSON array of {num turns} strings.
P ROMPT I NDEPENDENT LLM J UDGE (S YSTEM + U SER P ROMPT ) System: You are a strict label-integrity judge, independent of the generator. Decide only whether the mutated sample still matches the ground-truth label. Return only JSON. Context: Ground-truth {binary label} / {tactic} / {scenario}; evidence onset {first evidence turn}; taxonomy: tactics {all tactic options}, scenarios for {tactic}: {tactic scenario options}. Inputs: {source turns table}, {mutated turns table} Output: "accept": {bool}
Figure 16: Prompt templates for LLM-based data augmentation.
U SER P ROMPT Top-K Keywords: {keyword list} Analysis Steps: 1) Jointly infer the likely pretext, victim-facing narrative, and requested action or payment method. 2) Judge whether the keywords cohesively support a single scenario or mix unrelated schemes. 3) Summarize the scenario in 1–3 sentences without verbatim keyword repetition. Output (JSON only): { "summary": {1--3 sentence description}, "coherence": "high" | "medium" | "low", "rationale": {brief justification} }
Figure 17: Prompt template for generating scenario summaries for each BERTopic cluster.