Engineered Persuasion: Evaluating Personalized Pretexts in LLM-Generated Spear Phishing
arXiv:2609.04410v1 [cs.CR] 3 Sep 2026
Jerson Francia, Derek Hansen, Benjamin Schooley, and Shydra Valynn Murray Department of Electrical and Computer Engineering Brigham Young University Provo, UT 84602, USA [email protected] (J.F.); [email protected] (D.H.); ben [email protected] (B.S.); [email protected] (S.V.M.) Abstract—Large language models can insert workplace details into phishing pretexts at low cost, but those details may either support or undermine a message’s credibility. We recruited 180 U.S. working adults to evaluate simulated, AI-generated phishing emails in a disclosed survey. The emails used four cumulative levels of information: workplace (Level 1); recipient name and job title; job responsibilities; and coworker/sharedproject context (Level 4). Participants rated each message’s convincingness from 0 to 100, chose one stated action (open the link, investigate, delete, or report), and explained why their highest- and lowest-rated messages stood out. Across 1,436 valid evaluations, convincingness increased by 2.40 points per personalization level in a sensitivity analysis, while the odds of expressing click intention increased by 28% per level. Among participants who did not express an intention to click, investigation remained common, reporting declined, and deletion increased. A post-hoc descriptive analysis found higher ratings and click intention for messages from a named person who referenced a supplied coworker than for messages from a department or entity. Qualitative coding showed why added detail could help or hurt: details that matched participants’ roles and routines supported credibility, while incorrect, vague, or channel-inappropriate details raised suspicion. Together, the results highlight that personalization is not simply a matter of adding more details: it depends on whether the pretext fits the recipient’s work context. We discuss how this distinction can inform workplace cybersecurity training. Index Terms—phishing, spear phishing, large language models, social engineering, personalization, cybersecurity training
1. Introduction Phishing remains one of the main ways attackers gain initial access, even as organizations improve their technical defenses. The 2026 Verizon Data Breach Investigations Report shows that although cyber breaches are increasingly influenced by vulnerability exploitation, ransomware, and AI-supported activity, attacks targeting people remain a major concern [1]. Reports from the Federal Bureau of Investigation (FBI) Internet Crime Complaint Center (IC3) also
show that phishing and spoofing make up a large share of complaints, with business email compromise (BEC) leading to the highest median losses [2]. The Anti-Phishing Working Group (APWG) recorded 1,003,924 phishing attacks in the first quarter of 2025 alone, showing that phishing also remains a high-volume threat [3]. Although malicious actors succeed with both simple and highly sophisticated phishing attacks, users should not be framed as ”the weakest link.” Susceptibility to phishing is often shaped by workload, time pressure, workplace norms, and imperfect defenses that exist in the organization [4]. The continued scale of phishing suggests that attackers continue to find value in making malicious requests appear routine, relevant, and socially appropriate. Many phishing messages work by presenting a pretext: a fabricated scenario or backstory that gives the recipient a reason to treat the message as legitimate [5], [6]. Personal details can strengthen a pretext when the message uses details that match the recipient’s workplace, job, routine, or coworkers [7], [8]. However, those same details can also make a message less convincing when they are wrong, vague, or inconsistent with normal workplace procedures [8]. Pretexts also depend on who the message claims to be from; a message from a department/entity can carry different authority and familiarity cues compared to a person who references someone you know [9], [10]. Because pretexts depend on selecting and arranging contextual details, tools that can quickly draft many plausible variants may change the scale at which attackers can tailor them. Large language models (LLMs) can reduce the time, expertise, and cost needed to draft convincing phishing content [11]–[13]. Recent peer-reviewed work therefore treats generative AI as part of a broader social-engineering risk, where models can support message drafting, scaling, and repeated adaptation [14], [15]. However, lower generation cost does not mean that every generated phishing message will be as convincing. An LLM-generated pretext still depends on whether the message fits the recipient’s work context, uses plausible relationships, and avoids details that expose the scenario as wrong or unusual. For that reason, we study which added workplace details make LLM-generated phishing messages seem more believable, and which details
make them seem suspicious. We therefore examine how the amount and type of workplace detail used in AI-generated spear phishing pretexts shape perceived convincingness and stated action in a controlled survey. Participants knowingly evaluated simulated phishing messages; the study was not designed to test detection in a mixed inbox or estimate operational phishing success. This disclosed design allowed repeated comparisons and collection of participants’ reasoning without delivering personalized deception to real inboxes. We use personalization level as our measure of personalized pretext depth, meaning how much work-context information was made available for the message. We also analyze coded impersonation style (sender/pretext configuration) as a posthoc exploratory component describing the apparent sender and whether the message invoked a supplied coworker. Building on mixed findings about LLM phishing persuasiveness [11], [13], we compare how participants judged messages generated with four cumulative levels of workplace information. Participants rated convincingness and selected one stated response: open the link, investigate, delete, or report. We use click intention to mean selection of the open-link option in this hypothetical survey; it is not observed clicking. The planned comparison concerns personalization depth; associations with sender/pretext configuration are exploratory because the model, rather than the experimental design, selected the sender role. This paper makes three contributions. First, it provides a within-participant comparison of four cumulative workplace-information tiers rather than treating personalization as simply present or absent. Second, it shows how convincingness and forced-choice stated actions changed across those tiers under the same disclosed survey context, without treating those responses as operational click-through. Third, participant explanations show that personalization depends on contextual fit: accurate role, relationship, and project details can support credibility, while incorrect or channelinappropriate details can expose the pretext. We also report a clearly secondary analysis of the sender/pretext configurations produced by the model.
2. Review of Related Literature We review research on how people judge spear phishing messages, how personalization and persuasion shape those judgments, and how LLMs facilitate the creation of targeted phishing pretexts. Recent complaint and incident data show that spear phishing attacks remain a persistent problem despite advances in technical safeguards. IC3 reports show that phishing/spoofing account for the largest share of complaints, while business email compromise (BEC) drives the highest median losses [2]. APWG reports also point to historically high phishing volumes in early 2025 [3]. Because different industry reports measure phishing in different ways (e.g. complaints reported to the FBI vs confirmed security incidents and data breaches), we cannot reliably determine that phishing attacks are becoming more or less targeted
over time. Instead, we view targeting as a choice attackers make by balancing the cost of tailoring a message and its expected return [1], [16]. How effective this targeting is ultimately depends on how targets interpret and respond to the message. People’s susceptibility to phishing is influenced by the cues present in the message, the recipient’s own knowledge and habits, and the setting in which they receive it [17]–[20]. In those messages, easily spoofed surface cues, like padlocks and URL warnings, can create a false sense of legitimacy [21]; however, training tools can improve awareness to these cues and help users detect phishing [22]. Organizationoriented phishing research likewise emphasizes that susceptibility and response depend on workplace context, organizational controls, and human factors together [23]. However, recent research suggest that people can still fall for both crude and polished attacks, so phishing risk is not limited to advanced spear phishing messages [8], [24]. Thus we focus on one part of that larger picture: how personalized details help build a believable phishing pretext. We define a phishing pretext as the story behind the message: the scenario, role, or request that makes the message appear legitimate. Personalization is when we include specific details that connect this story to the target (e.g., name, job title, responsibilities or coworker), often gathered from open-source intelligence (OSINT) or internal directories. We define personalization depth as the level of the recipient-specific information made available to the message source. This approach builds on prior work comparing phishing messages created with different amounts of target information, although our four levels are specific to the present study [7]. Personalization depth is distinct from persuasion strategy: it describes the information available to create the pretext, whereas persuasion strategies describe how the message encourages action. Strategies such as authority, urgency, social proof, and reward—sometimes called weapons of influence—may appear alone or in combination at any personalization level [6], [18], [25]. Whether these strategies seem plausible also depends on workplace norms. The same request may seem routine when it comes from an expected person in an appropriate tone, but suspicious when the sender or wording does not fit the relationship [26]. Sender identity is another part of a pretext, but it must be interpreted together with the message body and the workplace relationship. An unfamiliar sender may seem plausible when the role and request fit. Conversely, a familiar name may seem suspicious when the request or communication channel does not. Prior research finds that authority, role, and premise alignment shape responses in workplace phishing tests [9], [10]. Law-enforcement materials on BEC similarly describe executive and vendor impersonation as common tactics that exploit organizational roles and familiar relationships [27], [28]. Prior lab and field studies generally find that including personal details can raise susceptibility to phishing, although these effects vary by population, setting, and message [7], [29], [30]. Much prior research treats personalization as either present or absent, or compares broad low- and high-
information conditions [7], [19], [31]. In the SpearSim study, message creators who received more information about a target deceived more recipients than creators with minimal information, and produced more contextually meaningful narratives [7]. These findings motivate a closer examination of information depth, but also raise a broader question: when does additional target information make a phishing pretext seem more convincing, and when does it instead undermine credibility? LLMs add a new reason to revisit this question. They can lower the time, cost, and writing skill needed to create convincing phishing pretexts [11]–[13]. A NeurIPS workshop demonstration showed that automated messages based on target information could support social engineering campaigns at scale [32]. More recent work places generative AI within a broader deception workflow: models can help with drafting messages, generating scenarios, and iterating automatically, while safeguards and proposed mitigations remain insufficient [14], [15]. At the same time, research on LLM-assisted lateral phishing shows that the effectiveness of these capabilities depends on the organizational context in which they are deployed [33], [34]. The closest operational comparison to this paper is Czybik et al.’s large-scale field experiment. It used information from web searches to generate LLM spear phishing emails, and measured clicks after delivering them to around 7,700 recipients [13]. That study addresses operational effectiveness and scalability. Our study asks a different question: in a controlled setting, how do people judge messages generated with cumulative tiers of workplace information, what action do they say they would take, and which contextual matches (or errors) explain those judgments? This setting may not estimate inbox behavior, but its repeated ratings and open-ended explanations provide evidence about perceived convincingness, stated action, and context fit across information-depth conditions. Laboratory and survey measures can provide useful comparisons when researchers clearly separate them from field behavior. For example, the Phishing Email Suspicion Test compares laboratory suspicion ratings with field-email outcomes but treats them as distinct measures [35]. We are not aware of prior humansubject work that deals with cumulative information tiers in LLM-generated messages with both quantitative and qualitative analyses. Together, the literature suggests that perceived legitimacy of spear phishing messages depends on several factors: the information in a pretext, its persuasion strategies, the apparent sender and relationships, and whether the request fits workplace norms. Our analysis focuses on cumulative personalization depth. We separately examine coded impersonation style (sender/pretext configuration) as an exploratory factor, and use participant explanations to identify the characteristics they regarded as convincing or suspicious.
3. Research Questions We posit two research questions and one exploratory question: RQ1. How does personalization level relate to self-reported convincingness and stated action in AI-generated spear phishing pretexts? EQ1. What associations appear between impersonation style (sender/pretext configuration) and these outcomes? RQ2. Which personalized pretext cues do participants report as making spear phishing messages persuasive or suspicious?
4. Methodology This section describes the project’s rationale, design, and methods for data collection and analysis. We recruited working participants, collected information about their roles and workplaces, and used an LLM to generate spear phishing emails for participants to evaluate. Our main outcomes were convincingness, stated response (open the link, investigate, delete, or report), and open-ended explanations. We treat the emails as simulated spear phishing messages, with the amount of personalized work-context detail included in pretexts varying by condition. We later added a post-hoc exploratory analysis of impersonation style, meaning how the LLM chose to portray the sender. Figure 1 summarizes the study workflow.
4.1. Design Rationale Personalization level refers to the amount of recipientspecific workplace information included in a phishing pretext. We use it to measure personalized pretext depth. The study used four levels, from minimal tailoring to detailed workplace context: • Basic (1) – workplace referenced; generic greeting • Low (2) – basic level + participant’s name and job title • Moderate (3) – low level + job responsibilities • High (4) – moderate level + coworker’s name and shared-project details Table 1 uses fictional names to show which fields are added at each level. The scenario and sender remain fixed so the differences are easy to see. Participants did not see these examples. In the study, GPT-4o generated each message separately; the available profile fields were controlled by level, but scenario, tone, persuasion strategy, and sender configuration could vary across outputs. Four altered and deidentified historical outputs from one participant appear in Appendix D; eight additional historical examples, two per level, are included in the reproducibility artifact. The generation prompt in Appendix C followed this cumulative field structure. These cumulative levels represent the amount of workplace information available to the model, not increasing levels of persuasion. We did not assign authority, urgency, social proof, reward, or politeness to
Figure 1. High-level workflow of the study: recruitment, collection of participant context, LLM-based message generation, and participant assessment. TABLE 1. R ESEARCHER - CONSTRUCTED ILLUSTRATION OF CUMULATIVE PERSONALIZATION DEPTH . A LL ENTITIES ARE FICTIONAL ; TEXT SHOWN WAS NOT PRESENTED TO PARTICIPANTS .
Level
Information available
Illustrative addition to a fixed scheduling-workspace pretext
1 2 3 4
Organization Level 1 + recipient name and job title Level 2 + responsibilities Level 3 + coworker and project
“Northbridge Civic Services is moving the weekly scheduling workspace...” “Hello Alex, as an Operations Coordinator at Northbridge Civic Services...” “Because you coordinate weekly staffing schedules...” “Taylor mentioned that the Atlas records-migration pilot may affect the weekly staffing schedules...”
particular levels. Because GPT-4o generated each message independently, these strategies and features could appear and vary at any level. We also left tone and sender role open so the model could construct the rest of the pretext. This choice allowed the LLM to select the sender persona (i.e., impersonation style), which we later analyzed in Subsection 4.5. A field experiment involving names, coworkers, and shared projects would require sending highly personalized deceptive messages to real workplaces. We instead used a disclosed survey so that participants could evaluate these messages without putting themselves, their coworkers, or their organizations at risk. This approach allowed us to compare responses across personalization levels, although it cannot show how people would behave in a real inbox. The study was designed to compare levels of personalization within simulated phishing pretexts, not to measure whether participants could distinguish phishing from legitimate workplace communication. We discuss this tradeoff along with other limitations further in Section 8.
4.2. Participants Before recruitment, a conventional four-group analysis of variance (ANOVA) calculation (α = .05, 1−β = .80, f = 0.25) suggested a target of approximately N = 180 participants [36]. We used this calculation only as a recruitmentplanning benchmark; it was not a power analysis tailored to the repeated-measures mixed models used in the final analysis. Participants were a self-selected convenience sample of working adults recruited through Prolific. Eligibility required current employment outside the home, U.S. residence, and no prior participation in a similar spear phishing message evaluation study. Eligible participants were directed to an online Qualtrics survey. Of the 257 people who started the study, 67 timed out or did not finish. Of the remaining 190, six failed the attention check and were terminated without completing the survey. We excluded two responses completed in under 10 minutes and two with unusable workplace fields, leaving a final ana-
lytic sample of N = 180. Copy/paste was disabled to deter scripted answers. Although the survey was designed to take approximately 20 minutes, the observed median completion time was approximately 30 minutes and 33 seconds. Participants had a mean age of 37.0 years (SD = 11.7, range 19–69). Eighty-nine (49.4%) identified as female and 91 (50.6%) as male. In total, 127 (70.6%) reported fulltime employment and 53 (29.4%) part-time employment. Prior phishing-message exposure ranged from never (2.2%) to daily (8.3%), with the largest groups reporting exposure several times a week or about once a month (28.9% each). Most participants were very (48.3%) or somewhat (45.0%) confident in detecting phishing.
4.3. Ethics and consent All procedures received IRB approval (see Appendix A for details). The Qualtrics survey began with a standalone consent page that participants had to accept to participate. The consent page described the study as voluntary, stated that participants could withdraw without penalty, instructed participants not to provide confidential or sensitive information, and disclosed that simulated spear phishing messages would be fictional and AI-generated. Only consenting U.S. working adults (18+) could proceed. Participant-provided identifiers and workplace-context fields were stored in restricted Qualtrics and university Box environments accessible only to the research team. The appendix documents data handling, retention, deidentification, and third-party processing in detail.
4.4. Survey Procedure After consenting, each participant completed a brief questionnaire covering prior exposure to spear phishing, self-reported confidence in identifying phishing messages, and demographics (age and gender). Participants then provided several fields: full name, workplace, job title, responsibilities, the first name of a known coworker, and a short description of a shared, non-confidential project. These fields were inserted into specific prompts via client-side Qualtrics JavaScript and transmitted to GPT-4o through four API requests. Each request asked for two messages at one personalization level. The survey was designed to produce eight messages per participant. During data collection, four requested outputs (two Level 3 and two Level 4) were not returned. The generation code did not retry failed requests, so we excluded the corresponding four evaluations, leaving 1,436 valid message assessments across 180 participants. The resulting sample counts for each level are 360, 360, 358, and 358. The messages were presented in random order within the survey. For each simulated email message, participants rated its convincingness on a 0–100 visual-analog scale. Participants were then required to select one stated action: open the link, investigate further, delete, or report as phishing. The item did not include an ignore option or allow multiple selections. Participants also rated their confidence in the
selected action on a 0–100 scale. These were hypothetical survey responses; no links were active and no messages were delivered to real inboxes. After evaluating all messages, participants wrote open-ended explanations for the messages they found most and least convincing. These responses form the qualitative dataset.
4.5. Impersonation style We use impersonation style as shorthand for a nominal sender/pretext configuration. The code combines the persona of the apparent sender with whether the message refers to a known coworker. Two researchers coded each generated message after data collection. We define each style as follows: • Style 1 – a department or entity is the apparent sender • Style 2 – a named person not supplied as the coworker is the apparent sender • Style 3 – a named person not supplied as the coworker references the supplied coworker in the message body • Style 4 – the supplied coworker is the apparent sender Impersonation style does not show whether the sender would actually be familiar, authoritative, or legitimate in the participant’s workplace. To decide whether a message used the participant-provided coworker, coders traced the name to the supplied coworker field. This factor was not part of the initial design focus or independently randomized. The historical messages in Appendix D provide observed examples. In the dataset, one Level 2 message happened to use the same common first name as a supplied coworker. We coded it as Style 2 because the coworker field was not available to the model at Level 2. No valid study message was coded as Style 4.
4.6. Data Handling and Analysis We analyzed convincingness with a linear mixed-effects model and click intention with a frequentist logistic mixedeffects model fitted by maximum likelihood. Both models included a participant random intercept to account for repeated evaluations. The primary model specifications were: convincingness ∼ C(level) + (1 | participant), logit[Pr(open-link)] ∼ C(level) + (1 | participant). Level 1 was the reference condition. The three contrasts with Level 1 use unadjusted confidence intervals and pvalues; the separately reported 16 exploratory demographic tests use Holm adjustment. The primary models treated personalization level as a categorical within-participant factor; an ordinal linear-trend specification served as a sensitivity analysis. We summarized the post-hoc sender/pretext styles descriptively. We separately used an auxiliary multinomial logistic model with participant-clustered standard errors to examine investigate, report, and delete selections among responses that did not select open-link, using delete as the
reference response. Demographic variables were examined as exploratory covariates and moderators. The following definitions are used throughout the paper. convincingness Self-reported rating of how convincing each message seemed immediately after exposure, used as a survey measure of perceived credibility (integer 0– 100; higher = more convincing). action Self-reported stated action: open link (click), investigate further, delete, or report as phishing. click intention Selection of the open-link option in the survey. This is a stated response, not an observed click. personalization level Ordinal factor {1, 2, 3, 4} defined in Section 4.1, representing how much workplace detail was used in the pretext. impersonation style The post-hoc category {1, 2, 3, 4} defined in Subsection 4.5, representing the sender/pretext configuration generated by the model.
4.7. Qualitative Data Analysis We used inductive thematic analysis to study participants’ free-text responses [37]. Each participant provided four responses: (i) why the top-ranked (most convincing) message was convincing, (ii) why it was not convincing, (iii) why the bottom-ranked (least convincing) message was convincing, and (iv) why it was not convincing. With four responses per participant, this yielded 720 open-ended explanations. All responses received at least one code. We used the non-thematic Q0 label to responses that did not provide an interpretable explanation, which included unrelated feedback or other non-substantive text. Nineteen responses received only the Q0 label and were excluded from thematic interpretation, leaving 701 responses in the thematic summaries. Four researchers participated in the qualitative analysis workflow, with three of them serving as primary coders. The team first used open coding on a simple random sample of 103 responses (14.3% of the corpus) to draft initial codes and rules. Coders marked every theme present in a response and assigned positive (+) or negative (-) valence. Codes were not mutually exclusive, so a response could receive several codes. Pairwise Cohen’s κ was calculated for each parent code on this independently coded subset, ignoring subcodes and valence. Disagreements involved whether a code applied, where related codes differed, which subcode fit, and whether several codes applied. The coders resolved these issues through discussion, codebook revision, and recoding, with input from the principal investigator when needed. This was done iteratively until the team reached a mean κ of at least .70. The unweighted mean across seven estimable parent codes in the final iteration was κ = .753 (per-code range .678–.951). Neither coder assigned Communication Medium in this subset, so its chance-corrected κ could not
TABLE 2. D ESCRIPTIVE STATISTICS BY PERSONALIZATION LEVEL
Level
Convincingness M ±SD
Click intention (%)
Del:Rep
Basic (1) Low (2) Moderate (3) High (4)
56.6 ± 31.8 59.5 ± 28.7 60.7 ± 29.9 64.3 ± 27.3
31.9 36.4 36.3 43.0
0.37 0.43 0.69 0.72
Note. Level sample sizes are 360, 360, 358, and 358. Del:Rep is the deleteto-report ratio among responses that did not select open-link.
be calculated.One of the primary coders then coded the rest of the responses, with oversight from the rest of the research team. The final codebook contains eight main themes, with definitions and examples described in Section 5.
5. Quantitative Results We present the quantitative analyses across the 1,436 valid evaluations of AI-generated spear phishing messages. Participants rated convincingness at a mean of 60.2 (SD = 29.6) on the 0–100 scale and selected the open-link option in 36.9% of responses. We report the percentages descriptively and use the models to compare outcomes across the tested pretexts.
5.1. Effect of Personalization Level Table 2 summarizes results by personalization level. Compared with Level 1, convincingness was 2.90 points higher at Level 2 (95% confidence interval (CI) [−0.60, 6.39], p = .104), 4.09 points higher at Level 3 (95% CI [0.59, 7.59], p = .022), and 7.61 points higher at Level 4 (95% CI [4.11, 11.11], p < .001) in the linear mixed-effects model. A sensitivity analysis treated the levels as a linear trend and estimated a 2.40-point increase for each additional level (95% CI [1.29, 3.51], p < .001).
Figure 2. Convincingness and click intention by personalization level.
Click intention followed a similar pattern, with the clearest difference at Level 4. Compared with Level 1, the odds
of selecting open-link were not clearly different at Level 2 (odds ratio (OR) = 1.41, 95% CI [0.94, 2.11], p = .099) or Level 3 (OR = 1.39, 95% CI [0.92, 2.08], p = .114), but were higher at Level 4 (OR = 2.29, 95% CI [1.52, 3.43], p < .001). The ordinal sensitivity model estimated 28% higher odds per level (OR = 1.28, 95% CI [1.13, 1.46], p < .001). Figure 3 shows the other three choices among responses that did not select open-link. Investigation remained near half of selections (49–52%). Reporting declined from 37% at Level 1 to 28% at Level 4, while deletion increased from 14% to 20%. An auxiliary clustered multinomial model compared these choices. For every one-level increase in personalization, the relative risk of selecting report rather than delete decreased by 21.7% (relative risk ratio (RRR) = 0.78, 95% CI [0.67, 0.92], p = .002). The estimate comparing investigation with deletion was less precise (RRR = 0.88, 95% CI [0.77, 1.01], p = .068).
TABLE 3. D ESCRIPTIVE STATISTICS BY IMPERSONATION STYLE
Style
Convincingness M ±SD
Click intention (%)
Department/entity (1) Named person (2) Coworker reference (3)
55.2 ± 32.9 59.9 ± 29.2 65.7 ± 26.8
31.0 35.1 47.3
Note. Style sample sizes are 258, 878, and 300. Style 4 was unobserved.
Figure 4. Convincingness and click intention by coded sender/pretext style. These descriptive differences do not isolate an effect of style.
We therefore present the style differences only descriptively and do not treat them as independent effects of sender style.
Figure 3. Distribution of investigate, delete, and report selections among responses that did not select open-link, by personalization level. Percentages are conditioned on non-click selections and may differ slightly from 100% because of rounding.
5.2. Impersonation Style As an exploratory analysis, we examined the sender/pretext styles that GPT-4o produced, described in Table 3. Department/entity messages had the lowest convincingness and click intention (55.2; 31.0%; n = 258). Named-person messages came next (59.9; 35.1%; n = 878). Messages from a named person who referred to the participant-provided coworker had the highest values (65.7; 47.3%; n = 300). No valid message was coded as Style 4. This style analysis was post-hoc and not independently randomized, and styles were strongly imbalanced across personalization levels. All 300 Style 3 messages occurred at Level 4; no Style 3 messages occurred at Levels 1–3. We therefore could not estimate a full categorical Style × Personalization interaction. Figure 5 shows this imbalance. Because only Level 4 supplied the coworker field, any person names at lower levels came from the model itself.
Figure 5. Message counts across personalization levels and coded sender/pretext styles. Empty Style 3 cells make the full categorical interaction non-estimable.
5.3. Other Factors We explored whether age, gender, prior phishingmessage exposure, or confidence in detecting phishing predicted the outcomes or changed the relationship with personalization. None of the 16 overall tests for convincingness or click intention remained significant after Holm correction (all adjusted p ≥ .795). This does not prove that the subgroups responded identically.
TABLE 4. Q UALITATIVE CODES AND DESCRIPTIONS
Name
Description
Personal
Message feels relevant/mismatched to the participant’s current work context Message tone, clarity, formality, and formatting The sender identity and authority cues Link/file interactions, routine requested actions, or sensitive/high-risk actions Pressure/threats to act quickly
Style Source Requested Action Urgency Benefit/Favor/ Reward Medium General
Benefits, opportunities, favors, incentives, gifts, or ego appeals The message channel matches/mismatches expectations Unsupported overall credibility judgments when no substantive theme applies
Note. Personalization is abbreviated as Personal.
6. Qualitative Results This section explains why participants found messages convincing or suspicious. Our thematic coding produced a codebook with eight main themes. As mentioned in Section 4, nineteen responses received only the non-thematic Q0 label because they were non-substantive or uninterpretable, and thus were excluded. For each theme below, we provide its working definition, common indicators, and some illustrative quotes. Participant identifiers (e.g., R1) indicate which participant said each quote. Names of real people and workplaces are redacted. We then report how often each theme appeared and whether comments about it were more often positive or negative. Table 4 gives a short description of each code. Personalization / Relevance. Participants linked credibility of a message to how closely it fit their daily work. Among 329 valenced Personalization/Relevance mentions, 220 (66.9%) were positive and 109 (33.1%) were negative. This total is four higher than the 325 responses containing the theme because four responses included both positive and negative personalization subcodes. Messages that combined personal details—full name, job title, employer, coworker, or current project—were often mentioned as convincing. Messages were described as convincing because they referenced “something I would be doing often in my job” (R11) or “my company’s name and also . . . my job role . . . ” (R164). References to specific work (e.g., “licensing system project” (R164), “informative session” (R16), “ongoing campaign strategies” (R36)) and routine tasks like scheduling—“one of the most integral parts of my job and that part got to me most” (R134)—were also deemed more convincing. Negative mentions described pretexts that did not fit the participant’s real work. Often this occurred because details were simply wrong: one message “said for me to grab my lesson plans, [but] I don’t do lesson plans” (R15), and another participant said, “[I] don’t have an events coor-
dination team” (R59). Other messages were less convincing because the personalized context was too vague: one message “lack[ed] specific details about the. . . purpose or context” (R115), and another noted a “lack of specific detail about the conference” (R61). Overall, accurate details that aligned with participants’ work were more persuasive, while incorrect or vague details created mismatches. Message Style. Tone and styling shaped credibility in several ways. Among valenced Message Style mentions, 46.6% were positive and 53.4% were negative. Convincing messages used a “casual tone” (R87) and felt “naturally friendly. . . not over the top” (R83). Messages seemed genuine when their style matched local norms in their workplace. One participant noted that “it was casual, which is typical of our office” (R31), while others said it “looked professional and relevant to my role,” “used a familiar tone” (R53), or resembled “typical corporate communication” (R76). Negative Message Style mentions described messages as “overly flattering” (R88), “a bit too friendly” (R73), or sounding “cliché,” “boilerplate” (R98), or “mechanical” (R91). Stock phrases like “I hope this finds you well” or “best regards” made one message “more than likely a scam” (R34). Greetings that were too formal or too informal, such as “Cheers” (R13), or “way too white collar for the people I am interacting with” (R14) were also deemed suspicious. Overall, when style cues matched the tone of the workplace, either casual or professional, the messages were more persuasive; when they were mismatched or looked like stock phrases, they were flagged as suspicious. Source. The source of a message, or the perceived authority of its sender, also mattered. Among valenced Source mentions, 42.1% were positive and 57.9% were negative. One participant noted that “the sender is from the [company name] security team. This gives the air of authority and that I should listen to her and follow her instructions” (R20). Those who were convinced explained it was in part due to the titles and roles used in messages such as “a safety compliance officer” (R22), “Investment Analyst” (R64), and senders from clients or internal departments (“human resource department” (R85); “a trusted IT team member connected to my department” (R103)). Messages were less convincing when they arrived “without providing a means of confirming its legitimacy, such as mentioning an internal portal or HR contact, and the sender’s name” (R114), or from an “unfamiliar sender” or “vague source” (R157) with “no prior communication” (R148). Overall, clear and plausible authority cues raised persuasiveness, while the lack thereof raised suspicion. Requested Action. Participants also judged what the message asked them to do, including opening links or files, completing routine tasks, and taking sensitive or risky actions. Among valenced Requested Action mentions, 23.6% were positive and 76.4% were negative. One message felt like “a serious announcement made for a thorough reason” (R90); a breach notice was “very harmful to [their] career and the company’s well being,” making it “very convincing
to prompt me . . . to secure my account” (R46). Thus, a plausible context could make some requested actions appear credible. More often, participants treated the requested action as an indicator of a fake message. Requests for passwords, sensitive documents, or credential checks were considered red flags: “any message to verify credentials is probably . . . a phishing attack” (R106); “we are not sent emails to update the accounts ourselves” (R123). Routine requests also reduced credibility when they did not fit the participant’s role or normal workflow. Overall, requested actions were often suspect, but a compelling context could make them believable. Urgency / Pressure. Urgency and pressure cut both ways. Among valenced Urgency/Pressure mentions, 52.4% were positive and 47.6% were negative. For some, a deadline added realism. One participant said, “failure to do so in 48 hours will lead to delays” “got me thinking” (R146). Another participant highlighted an email from “the ‘IT team at [company name]’” because of its “very professional appearance” and “sense of urgency” (R138). Others saw pressure as a warning: “Once i spot pressure, i become much aware of the possibilities of something fishy. Why must this be urgent? why the pressure?” (R54). Another wrote that “the urgency with which it encourages action to be taken is quite off for me” (R110). Overall, urgency could signal importance or harmful intent, depending on the reader. Benefit / Favor / Reward. Participants also noticed benefits, opportunities, favors, rewards, and appeals to ego. Among valenced Benefit/Favor/Reward mentions, 75.8% were positive and 24.2% were negative. These cues could make a message seem legitimate and motivate action: “The incentive of a $100 gift card added a sense of legitimacy and motivation” (R53), and “[the] request for photos also sounded like a genuine favor from a colleague” (R87). They could also raise suspicion because “unsolicited offers of gifts are common in phishing and spam” (R148), or because a “reward element felt a bit too promotional for a real internal survey, which raised mild suspicion” (R53). Overall, participants more often described these cues as increasing credibility or motivation than as raising suspicion. Communication Medium. Communication channel was mainly a negative cue. Among 62 valenced Communication Medium mentions, 58 (93.5%) were negative and four (6.5%) were positive. One participant would “normally go look in [Microsoft] Teams rather than click on a link to the document” (R13). Another said, “I would likely receive a call from [company name] if this was the case” (R20). Personal invites were expected via call or text (“catch up with me for coffee” (R22)); networking events were “not via email” (R96). Overall, channel mismatches generally made messages less convincing and more suspicious. General. A small group of responses gave an overall judgment without a specific reason, such as “Honestly, nothing at all” (R83), “it doesn’t convince me” (R73), and “There was no aspect that made it less convincing” (R134). We placed these responses in the residual General category. Among 22
TABLE 5. C OVERAGE OF QUALITATIVE CODES ACROSS N =701 RETAINED RESPONSES : OVERALL (% OF RESPONSES WITH CODE PRESENT ) AND WITHIN C ONVINCING /L ESS - CONVINCING AND T OP /B OTTOM STRATA . Code Personal Style Action Source Benefit Urgency Medium General
n responses
% overall
Conv %
Less %
Bottom %
Top %
325 325 161 145 95 82 62 22
46.4 46.4 23.0 20.7 13.6 11.7 8.8 3.1
61.4 45.5 14.2 16.8 19.6 11.6 2.8 2.3
31.2 47.3 31.8 24.6 7.4 11.7 14.9 4.0
44.9 47.7 22.3 19.1 14.3 9.7 8.3 3.7
47.9 45.0 23.6 22.2 12.8 13.7 9.4 2.6
Note. All percentages are shares of retained responses. Convincing (n = 352), Less-convincing (n = 349), Bottom (n = 350), and Top (n = 351) columns report coverage within those strata.
valenced General mentions, 12 (54.5%) were positive and 10 (45.5%) were negative.
6.1. Descriptive Coverage of Qualitative Codes This subsection summarizes how often each category appears in the open-ended responses and how those codes are distributed across Convincing vs. Less convincing explanations and Top vs. Bottom ranked messages. Unless stated otherwise, each percentage is the share of distinct retained responses in which a parent code appeared at least once. Table 5 details the coverage of the codes across the dataset and within the defined strata. Overall coverage. Among the 701 responses, Personalization/Relevance and Message Style appeared most often (325 responses each; 46.4%). Requested Action appeared in 161 responses (23.0%) and Source Cues in 145 (20.7%). Each remaining theme appeared in 3.1–13.6% of responses. Because responses could receive more than one code, these percentages do not sum to 100%. Convincing vs. Less convincing. Explanations for Convincing messages emphasized Personalization more than explanations for Less convincing messages (61.4% vs. 31.2%; ∆=30.2 percentage points) (Figure 6). This pattern complements the quantitative finding by showing that personal details were discussed more often in explanations of convincing messages. However, the presence of Personalization in 31.2% of Less convincing explanations also shows that personalization was not uniformly helpful. Some participants treated mismatched or incomplete personal details as reasons for suspicion. Benefit/Favor/Reward also appeared more often in Convincing explanations (19.6% vs. 7.4%). In contrast, Communication Medium and Requested Action were more common in Less convincing explanations (Table 5). Even a personalized message could lose credibility when its channel or requested action did not fit expectations. Net persuasion index. To summarize whether a theme was described as helping or hurting persuasion, we calculated a net persuasion index from the positive and negative mentions of each code. A response could contribute both a positive and a negative mention. A positive score means the theme
Figure 6. Coverage by Convincing vs. Less-convincing explanation (% within each response stratum). TABLE 6. N ET PERSUASION BY CODE ( VALENCED MENTIONS ONLY ). Code
+
-
Total
Net
%Net
Benefit Personal General Urgency Style Source Action Medium
72 220 12 43 152 61 38 4
23 109 10 39 174 84 123 58
95 329 22 82 326 145 161 62
49 111 2 4 -22 -23 -85 -54
51.6 33.7 9.1 4.9 -6.7 -15.9 -52.8 -87.1
more often made a message convincing, while a negative score means it more often raised suspicion. Formally, %N et = 100 ·
#(+) − #(−) . #(+) + #(−)
Table 6 and Figure 7 detail the net persuasion index for each code. Benefit (+51.6%) and Personalization (+33.7%) skew positive when participants commented on them. The Personalization score should be read together with its negative mentions: personalization was usually helpful when it fit the recipient’s work context, but it also raised suspicion when details were wrong, vague, or implausible. Urgency is near neutral (+4.9%), while Message Style leans slightly negative ( −6.7%). Requested Action (−52.8%) is strongly negative, and Communication Medium (−87.1%) was the strongest negative cue. General is residual-only and should not be interpreted as comparable to the other themes.
7. Discussion This section discusses how participants judged the simulated spear-phishing messages and what they intended to do. We consider how more personalized details can make a message seem more believable but can also raise suspicion when those details do not fit the workplace context. We also discuss coworker references, individual differences, and implications for verification and reporting. Throughout, we interpret the results as relative differences in survey responses across participants. Personalized pretexts increased convincingness. Convincingness rose as personalization increased from Level 1 to
Figure 7. Net persuasion index by code.
Level 4. The percentage expressing click intention also rose descriptively, from 31.9% to 43.0%. The categorical models indicate that evidence was strongest at Level 4 rather than showing a clear difference at every intermediate level. This pattern agrees with laboratory and field findings that role, coworker, and project references can make phishing emails seem better matched to a recipient [7], [13], [29]. Among responses that did not select open-link, investigation remained common at every level. This suggests uncertainty rather than simple acceptance or rejection. Higher personalization was also associated with lower reporting relative to deletion in our model. Whether this pattern would reduce organizational awareness in a real deployment remains a question for future research. Coworker references worked without direct impersonation. Our prompt allowed the LLM to choose the sender persona. Instead of presenting a supplied coworker as the sender, the model usually used that coworker as supporting context. A named person or organizational entity referred to the coworker to support the pretext. No valid message directly impersonated the coworker. One Level 2 message happened to use the same common first name, but the coworker field was not available to the model in that condition, so the message was therefore coded as Style 2. One possible explanation for this pattern is that model safety guardrails may have discouraged direct impersonation, but our data cannot establish that explanation. Another is that the prompt did not explicitly state the coworker to be the intended sender, so it did not consider the scenario. Messages from a named person who referenced a coworker (Style 3) had a higher percentage expressing click intention than messages from a department or entity (47.3% vs. 31.0%), but all Style 3 messages occurred at Level 4, so this descriptive relationship is mixed with personalization level and other message features. These findings are consistent with evidence that authority and role cues influence clicking and reporting in workplace phishing [9], [10] and with guidance describing common BEC tactics that exploit organizational roles and familiar relationships [27], [28]. Impersonation style may therefore be a separate part of a pretext alongside personalization level. Future work should manipulate sender identity
directly in a more balanced design. Direct impersonation may be more persuasive when accurate, but more suspicious when the impersonation includes incorrect details. Individual differences were limited. We also examined whether age, gender, phishing exposure, or self-reported confidence changed how participants responded to the messages. We did not find reliable main or moderation effects after correcting for multiple comparisons (all adjusted p ≥ .795). These null results do not prove that participant groups responded in the same way. The study may not have had enough power to detect smaller subgroup differences. Future work with larger and more diverse samples could better test whether personalization and impersonation affect particular populations or experience levels differently. Context fit shaped persuasion and suspicion. The openended responses help explain why richer personalization was associated with higher convincingness and more click intention. The most frequent themes were Personalization/Relevance and Message Style (46.4% each), with Personalization appearing more often in Convincing explanations than in Less-convincing ones (61.4% vs. 31.2%). Participants described messages as more convincing when details aligned with their role, responsibilities, or active projects, and became suspicious when mismatches existed. This reasoning is consistent with prior work showing that recipients rely on contextual signals under uncertainty and that attackers can exploit those signals to be more persuasive [17], [18], [21]. Message Style (tone, clarity, and formatting) helped when it matched the participants’ workplace norms, while boilerplate or overly flattering phrasing did the opposite. Participants also referenced Source cues, where job titles, departments, or implied authority affected their willingness to comply, aligning with findings that role and authority framing influence workplace phishing responses [9], [10]. The qualitative data also show that personalized cues were not always persuasive. Among valenced Personalization/Relevance mentions, 33.1% were negative, usually because the pretext failed to match the participant’s real work context. Incorrect details, vague references, or details that were too hard to verify made the message easier to reject. Thus, personalization worked through fit , not simply through the number of details. Personalized details helped more often than they hurt, but mismatches remained an important failure mode. The net persuasion results reinforce this pattern. Benefit and Personalization skewed positive, Communication Medium and Requested Action skewed negative, and Message Style, Source Cues, and Urgency depended more heavily on context. Requested actions (e.g., credential checks, document requests, or account verification) were frequently treated as suspicious even when the message looked professional. This agrees with classic findings that people use both surface cues and context to judge legitimacy [21]. Communication Medium was similarly important. Participants noted that legitimate requests would normally arrive through a different channel, such as Teams, phone, or an internal portal, and
that email links felt inappropriate for certain actions. This points to a limitation of LLM-generated pretexts: even when a model can insert plausible personal details, it may not reliably infer which channel or request format is normal in a particular workplace. A request that seems reasonable in Teams, by phone, or through an internal portal may become suspicious when it arrives as an email link. Training should target verification and reporting. Our results point to the need to both reduce clicking and increase reporting. Higher personalization increased convincingness and click intention, and among non-click responses it was associated more clearly with declining reporting than with declining investigation. Training should emphasize that personal details do not prove that a message is legitimate. It should also teach simple verification habits, such as using known portals or known contact methods [22]. Recent work also shows that matching phishing training to a user’s proficiency can improve training outcomes, although our study did not test a training intervention [38]. Future work should test whether training examples matched to employees’ roles and communication channels help them recognize both accurate pretexts and mismatches. Reporting should also be quick and easy so that employees report suspicious messages even when they do not open the link. A practical next step is to compare role-tailored training with generic training. Survey studies can measure stated responses, whereas field studies can measure observed clicking and reporting. LLMs may help generate tailored training examples efficiently, but researchers still need to review them to avoid unsafe or overly invasive content.
8. Limitations and Future Work We enumerate the limitations of this study, as well as suggestions for future research. First, this study is best suited to comparing the four personalization conditions. Each participant rated messages from every level in the same disclosed survey. This design allowed us to compare responses as the model received more workplace context, but it cannot estimate real-world click-through rates. It also allowed participants to assess highly personalized messages without exposing them, their coworkers, or their organizations to deceptive messages in a real workplace. Second, our measures rely on self-report, which limits how directly they represent real behavior. Convincingness captures perceived credibility, while the open-link response captures click intention. Prior phishing research has used laboratory and survey measures to study detection, suspicion, and stated responses under controlled conditions [7], [35]. We follow this approach and treat our outcomes as comparative survey measures rather than real-world behavior. Participants might respond differently in a workplace, especially because they knew they were evaluating simulated spear phishing. The phishing-awareness questions presented before the message evaluations may also have encouraged closer scrutiny than an everyday inbox encounter. As explained in Section 4.1, this is a tradeoff of the disclosed
survey design. These survey cues may have affected the overall responses, even though every personalization condition was evaluated using the same procedure. Third, the survey setting has limited ecological validity. Messages appeared in Qualtrics rather than in participants’ inboxes, and the survey did not reproduce real links, attachments, warning banners, email-client interfaces, or organizational reporting policies. It also did not include legitimate workplace messages, so the study cannot show how well participants distinguish phishing from normal communication. Creating legitimate control messages would require either authentic workplace communications, which could create additional privacy concerns, or separately validated synthetic messages. Generating them from the same profile information would not establish that they reflected normal communication in each participant’s workplace. The forcedchoice response also omitted an ignore option and did not allow sequential or combined actions. Together, these omissions limit how closely the study represents inbox behavior. Future work should test these findings in live or longitudinal settings that measure observed clicking. Fourth, our participants were a self-selected convenience sample of U.S. working adults recruited through Prolific. They may not represent workers in other countries, cultures, or organizational settings. Future studies should test whether the findings generalize to these populations and settings. Fifth, we studied email only. Future work should examine whether personalization operates differently in SMS, workplace messaging platforms, and attacks that span multiple communication channels. Sixth, impersonation style was examined in a post-hoc exploratory analysis. Our coding describes how the generated message presented its sender and pretext, but it cannot establish whether that sender would actually seem familiar or legitimate in the participant’s workplace. Impersonation styles were also not balanced across personalization levels. We therefore cannot separate the influence of impersonation style from personalization and other message features. Future work should use a fully crossed design that independently controls sender identity and personalization level. Finally, the full qualitative corpus was not independently double-coded. The reported agreement applies only to the 103-response codebook-development subset. Future work should independently double-code the full corpus or a larger subset using the finalized codebook.
9. Conclusion This study examined how participants responded to LLM-generated spear phishing pretexts containing different levels of workplace personalization. In our disclosed survey, the strongest differences appeared at Level 4: participants rated the most personalized messages as more convincing and had higher odds of expressing click intention than at Level 1. Personalization was not always persuasive. Details that fit a participant’s role, routines, and workplace context made messages seem more credible, while incorrect, vague,
or inappropriate details raised suspicion. Messages that referenced a coworker also received higher convincingness ratings and a higher click intention. However, all of these messages appeared at Level 4, so a balanced experiment is needed to separate the influence of coworker references from personalization level. Among responses without click intention, investigation remained common, while reporting became less likely relative to deletion as personalization increased. Although these results do not estimate real-world clicking or compromise, they show why personalized details should be treated as claims to verify rather than proof of legitimacy. Security training should reinforce this habit, and organizations should make suspicious messages easy to report.
Appendix A. Ethical Considerations Oversight, consent, and research subject population.. All procedures were conducted under institutional review board (IRB) approval (IRB Number: IRB2025-128). Research subjects provided implied consent via a standalone Qualtrics consent page. The consent page stated that participation was voluntary, subjects could withdraw at any time without penalty, and subjects should not reveal confidential or sensitive information. Recruitment targeted U.S.-based working adults (18+) (through Prolific), excluding minors and other vulnerable populations by design. Subjects recruited via Prolific were compensated $5. Consent materials.. Because the full form contains operational contact details and vendor-policy language, we summarize the consent content here.The consent page disclosed that subjects would provide profile and workplacecontext information, evaluate simulated spear phishing messages, and assess message convincingness. It also stated that the messages were fictional and AI-generated, that GPT4o would be used through the vendor API policy, and that the study involved privacy risk from collection of digital records. The consent materials described secure storage in Qualtrics and a restricted university Box environment and three-year data retention. Stakeholders.. Primary stakeholders include: (i) subjects, whose safety and privacy we must protect, (ii) organizations and third parties referenced in subjectprovided data (e.g. coworkers), who could face risk if identifiable information were disclosed, and (iii) service providers used for recruitment, survey delivery, and LLM message generation. A secondary stakeholder is potential attackers, insofar as our findings about persuasive cues could be misused. Data minimization and third-party references.. To generate simulated spear phishing messages, subjects provided limited work-context fields: name, workplace/employer, job title, job responsibilities, the first name of a coworker, and a short description of a shared nonconfidential project. We explicitly instructed subjects not to enter sensitive information and requested that any coworker
reference be first name only, along with a non-confidential project context. We treated coworker names as third-party personal information requiring additional protection. Risks and mitigations.. The study was classified as minimal risk. Foreseeable risks included (1) privacy risk from collection of identifiers and work context, (2) possible social confusion if subjects later discussed fictional messages that used a coworker name, and (3) psychological confusion or misremembering a simulated message as real. We mitigated these risks by: (i) containing all messages within the survey (no emails were sent to subjects), (ii) presenting clear consent language that messages were fictional and AI-generated, and (iii) removing identifiers and workplace-context fields from the public research artifact and retaining identifiable source records only in restricted research storage under the approved retention schedule. Digital records were stored securely in Qualtrics and a restricted university Box environment with research teamonly access. Data collection ended on May 31, 2025. Restricted study records held by the research team that are not part of the deidentified public artifact are scheduled for deletion by May 31, 2028, consistent with the approved three-year retention period. Third-party processing and model use.. Thirdparty services included Prolific (for recruitment/payment), Qualtrics (for survey delivery), and an OpenAI model used during the survey to generate simulated spear phishing messages from the provided subject information. During data collection, client-side Qualtrics code transmitted the levelspecific workplace fields directly to the OpenAI API. The consent materials disclosed GPT-4o use and represented the model as accessed through the vendor’s API policy. The research team did not opt in to OpenAI API data use for model training, and the historical API key has since been revoked. We report this conservatively: OpenAI’s API data controls state that API inputs and outputs are not used to train models unless the customer opts in, but default abuse-monitoring logs may retain content for up to 30 days unless an organization has modified abuse monitoring or zero data retention [39]. We therefore do not claim that prompts and outputs were never retained by the provider. Generated outputs were stored with the survey records and inspected informally after data collection; participants were not protected by a documented real-time content-screening step. Deception and debriefing.. Subjects were told that the messages were AI-generated and simulated. Since messages were not delivered to real inboxes, we reduce the risk of deception-related harm. Due to these conditions, no poststudy debrief was needed. Ethics decision and dual-use considerations.. Phishing experiments require balancing methodological value against deception, consent, legal, and dual-use concerns
[40]. Because our findings describe what makes phishing pretexts more convincing in a survey setting, they could be misused as well as used defensively. To reduce this risk, we emphasize defensive takeaways (e.g., verification and reporting practices) and we do not release materials that would make it easier to run real attacks, such as subject data, raw message text tied to real organizations, or prompts that are ready-to-use. For transparency and reproducibility, we provide the placeholder-based study prompt in Appendix C. We do not release participant data, raw messages tied to real workplaces, or tools for delivering operational phishing campaigns. We believe the remaining risk is justified by the benefit of helping defenders understand how personalized pretext cues affect perceived convincingness and reporting, which can improve cybersecurity measures.
Appendix B. Open Science We will release the analysis code, deidentified quantitative message-evaluation data needed for the primary personalization and descriptive sender/pretext analyses, sanitized qualitative code and valence labels for aggregate summaries, and deidentified study materials such as the survey instrument, codebook, and consent form with direct contact details removed. The public artifacts contain no source identifiers, participant or coworker names, workplaces, project descriptions, raw generated messages, open-ended response text, demographics, or phishing-background variables, and no reidentification key is retained. Analyses that require restricted demographics or raw qualitative text cannot be reproduced from the public data. A repository link will be provided upon publication.
Appendix C. LLM Prompt Below is the prompt used to generate spear phishing email examples. This version corresponds to personalization Level 4. Level 3 omitted COWORKER and PROJECT. Level 2 also omitted JOBRESP1 and JOBRESP2 but retained NAME and JOB TITLE. Level 1 retained only WORKPLACE. The survey made four API calls per participant and requested two outputs in each call. The historical generation code did not retry when an output was not returned; four fallback evaluations were therefore excluded from analysis. Requests used the OpenAI Chat Completions API with gpt-4o, temperature 0.8, top-p 1, and a maximum of 1,000 output tokens. The level-specific prompt was sent as the system message; settings not specified in the deployed code used API defaults. You are a cybersecurity professor educating your students by simulating a spear phishing attack. Your task is to generate two spear phishing emails directed to the fictional user described in their profile below.
Format each message like a short email. Make the two messages as distinct as possible from one another. Include a link in each message that appears as '[LINK]'. Never use square brackets anywhere else. If you reference people other than those in the profile, use believable names that would plausibly exist in the workplace. If you reference another entity or department, it must either appear in the profile or be a widely recognized, real organization. Do not invent titles, departments, or institutions. If you are unsure whether one exists, exclude it. Each message must include the workplace, the user's name, job title, one job responsibility, the coworker, and the project. Be creative in how you weave this information in. Use the first job responsibility in the first message, and the second job responsibility in the second message. If the message is not official communication, use a more casual tone. Keep each message under 150 words, and make them engaging. Here is the user's workplace: WORKPLACE Here is the user's name: NAME Here is the user's job title: JOB_TITLE Here is the user's first job responsibility: JOBRESP1 Here is the user's second job responsibility: JOBRESP2 Here is a coworker's name: COWORKER Here is something they are working on: PROJECT Enumerate the messages as 1) and 2).
Appendix D. Historical Personalization Examples The following are altered and deidentified historical outputs shown to one study participant. Identifying profile details, including names, organization, and workplace context, have been replaced with fictional values. The replacements preserve the pretext structure relevant to the comparison but are not verbatim. GPT-4o generated each message independently, so scenario, sender, tone, and request were not held constant across levels. The observed convincingness rating
and open-link selection are included only to illustrate that greater information availability did not produce a monotonic response for this participant. Eight additional deidentified historical examples from other participants appear in the reproducibility artifact.
Level 1 (Style 1; rating 100; selected open-link) Available context: workplace. Subject: Important: Meridian Mobility Employee Benefits Portal Update Hello Team Member, Exciting news! Meridian Mobility is rolling out new enhancements to our Employee Benefits Portal, designed to provide you with a more seamless experience. To explore these new features and confirm your current benefit selections, please access the portal through the link below: [LINK] We encourage you to take advantage of the latest updates tailored to meet your needs. If you have any questions, feel free to reach out to our Benefits Support Team. Thank you for your continued dedication to Meridian Mobility. Warm regards, Benefits Administration Meridian Mobility
Level 2 (Style 2; rating 90; selected open-link) Available context: workplace, name, and job title. Subject: Important Change in Production Protocols Hi Alex, I hope this message finds you well. As the Production Operations Manager, it is important for you to be aware of recent updates to our production protocols. We have made several changes to improve efficiency and safety standards. To review these updates, please follow this link: [LINK]. Please familiarize yourself with the new procedures by the end of this week to ensure a smooth transition. Best regards, Michael Thompson Quality Assurance Team, Meridian Mobility
Level 3 (Style 2; rating 50; did not select open-link) Available context: workplace, name, job title, and job responsibilities. Subject: Important Update Required—Quota Tracking System Hi Alex, I hope this message finds you well. We have recently upgraded our quota tracking system to help you manage your team of 1,200 production staff and meet its targets more efficiently. As the Production Operations Manager at Meridian Mobility, your role is crucial in ensuring smooth operations, and this new tool is designed to support that. Please review the new features by clicking the following link: [LINK]. Your feedback is invaluable to us,
and we believe this update will improve your team’s productivity. Best regards, Jessica Thompson Meridian Mobility IT Support
Level 4 (Style 3; rating 70; did not select open-link) Available context: workplace, name, job title, job responsibilities, coworker, and project. Subject: Development on Automated Line Integration Hey Alex, I wanted to keep you updated on the progress Jordan and I are making with the automated line loaders. With your experience keeping automated assembly equipment running smoothly, we would value your input on integrating the new systems. We are considering several options to optimize the process. Please review our implementation plan and provide feedback here: [LINK]. Let us know what you think. Cheers, Morgan
References [1]
Verizon, “2026 data breach investigations report,” 2026, accessed 2026-05-20. [Online]. Available: https://www.verizon.com/business/ reReferences/downloaded/reports/dbir/
[2]
Federal Bureau of Investigation, Internet Crime Complaint Center (IC3), “Internet crime report 2024,” Federal Bureau of Investigation (FBI), Washington, D.C., Tech. Rep., Apr. 2025. [Online]. Available: https://www.ic3.gov/AnnualReport/Reports/2024 IC3Report.pdf
[10] M. P. Steves, K. K. Greene, M. F. Theofanos, N. A. Stanton, and S. Furman, “Categorizing human phishing detection difficulty: A phish scale,” Journal of Cybersecurity, vol. 6, no. 1, p. tyaa009, 2020. [11] F. Heiding, B. Schneier, A. Vishwanath, J. Bernstein, and P. S. Park, “Devising and detecting phishing emails using large language models,” IEEE Access, vol. 12, pp. 42 131–42 146, 2024. [Online]. Available: https://doi.org/10.1109/ACCESS.2024.3375882 [12] H. Khan, M. Alam, S. Al-Kuwari, and Y. Faheem, “Offensive AI: Unification of Email Generation Through GPT-2 Model With a Game-Theoretic Approach for Spear-Phishing Attacks,” Proceedings of the 2021 International Conference on Security and Emerging Trends in Cyber Technologies, pp. 178–184, January 2021, publisher: IET Digital Library. [Online]. Available: https: //digital-library.theiet.org/content/conferences/10.1049/icp.2021.2422 [13] S. Czybik, A. J. Kouam, P. Heubl, J. M. Nold, and K. Rieck, “A largescale study of personalized phishing using large language models,” in 35th USENIX Security Symposium (USENIX Security 26). Baltimore, MD: USENIX Association, Aug. 2026. [Online]. Available: https: //www.usenix.org/conference/usenixsecurity26/presentation/czybik [14] M. Schmitt and I. Flechais, “Digital deception: Generative artificial intelligence in social engineering and phishing,” Artificial Intelligence Review, vol. 57, no. 12, 2024. [Online]. Available: https://doi.org/10.1007/s10462-024-10973-2 [15] S. S. Roy, P. Thota, K. V. Naragam, and S. Nilizadeh, “From chatbots to phishbots?: Phishing scam generation in commercial large language models,” in 2024 IEEE Symposium on Security and Privacy (SP). IEEE, 2024, pp. 36–54. [Online]. Available: https://doi.org/10.1109/SP54263.2024.00182 [16] D. Desai, R. Hegde, and D. Shtil, “ThreatLabz 2025 Phishing Report,” Zscaler, Inc., San Jose, CA, Technical Report, Apr. 2025, based on analysis of over two billion blocked phishing transactions in 2024. [Online]. Available: https://www.zscaler.com/reReferences/ downloaded/industry-reports/threatlabz-phishing-report-2025.pdf [17] Z. Benenson, F. Gassmann, and R. Landwirth, “Unpacking Spear Phishing Susceptibility,” in Financial Cryptography and Data Security, M. Brenner, K. Rohloff, J. Bonneau, A. Miller, P. Y. Ryan, V. Teague, A. Bracciali, M. Sala, F. Pintore, and M. Jakobsson, Eds. Cham: Springer International Publishing, 2017, pp. 610–627.
[3]
APWG, “Phishing activity trends report q1 2025,” Anti-Phishing Working Group, Tech. Rep., 2025. [Online]. Available: https: //docs.apwg.org/reports/apwg trends report q1 2025.pdf
[4]
V. Zimmermann and K. Renaud, “Moving from a ”human-asproblem” to a ”human-as-solution” cybersecurity mindset,” International Journal of Human-Computer Studies, vol. 131, pp. 169–187, 2019.
[5]
M. Workman, “Wisecrackers: A theory-grounded investigation of phishing and pretext social engineering threats to information security,” Journal of the American Society for Information Science and Technology, vol. 59, no. 4, pp. 662–674, Feb. 2008. [Online]. Available: https://doi.org/10.1002/asi.20779
[19] R. A. Alsharida, B. A. S. Al-rimy, M. Al-Emran, and A. Zainal, “A systematic review of multi perspectives on human cybersecurity behavior,” Technology in Society, vol. 73, p. 102258, May 2023. [Online]. Available: https://www.sciencedirect.com/science/article/ pii/S0160791X23000635
[6]
P. Rajivan and C. Gonzalez, “Creative Persuasion: A Study on Adversarial Behaviors and Strategies in Phishing Attacks,” Frontiers in Psychology, vol. 9, p. 135, February 2018. [Online]. Available: https://www.frontiersin.org/articles/10.3389/fpsyg.2018.00135/full
[7]
T. Xu, K. Singh, and P. Rajivan, “Personalized persuasion: Quantifying susceptibility to information exploitation in spearphishing attacks,” Applied Ergonomics, vol. 108, p. 103908, Apr. 2023. [Online]. Available: https://www.sciencedirect.com/science/ article/pii/S0003687022002319
[20] T. Lin, D. E. Capecci, D. M. Ellis, H. A. Rocha, S. Dommaraju, D. S. Oliveira, and N. C. Ebner, “Susceptibility to SpearPhishing Emails: Effects of Internet User Demographics and Email Content,” ACM Transactions on Computer-Human Interaction, vol. 26, no. 5, pp. 32:1–32:28, Jul. 2019. [Online]. Available: https://dl.acm.org/doi/10.1145/3336141
[8]
[9]
V. Distler, “The influence of context on response to spear-phishing attacks: an in-situ deception study,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, ser. CHI ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3544548.3581170 E. J. Williams, J. Hinds, and A. N. Joinson, “Exploring susceptibility to phishing in the workplace,” International Journal of HumanComputer Studies, vol. 120, pp. 1–13, 2018.
[18] D. S. Oliveira, T. Lin, H. Rocha, D. Ellis, S. Dommaraju, H. Yang, D. Weir, S. Marin, and N. C. Ebner, “Empirical analysis of weapons of influence, life domains, and demographic-targeting in modern spam: an age-comparative perspective,” Crime Science, vol. 8, no. 1, p. 3, Dec. 2019. [Online]. Available: https://crimesciencejournal. biomedcentral.com/articles/10.1186/s40163-019-0098-8
[21] R. Dhamija, J. D. Tygar, and M. Hearst, “Why phishing works,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI). Montréal, Québec, Canada: ACM, 2006, pp. 581–590. [Online]. Available: https://people.eecs.berkeley.edu/ ∼tygar/papers/Phishing/why phishing works.pdf [22] G. Canova, M. Volkamer, C. Bergmann, and B. Reinheimer, “Nophish app evaluation: Lab and retention study,” in 2015 APWG Symposium on Electronic Crime Research, 2015, pp. 143–157. [23] K. Althobaiti and N. Alsufyani, “A review of organization-oriented phishing research,” PeerJ Computer Science, vol. 10, p. e2487, 2024. [Online]. Available: https://peerj.com/articles/cs-2487/
[24] P. Burda, L. Allodi, and N. Zannone, “Cognition in social engineering empirical research: A systematic literature review,” ACM Transactions on Computer-Human Interaction, vol. 31, no. 2, pp. 1–55, 2024. [25] A. Ferreira, L. Coventry, and G. Lenzini, “Principles of persuasion in social engineering and their use in phishing,” in Human Aspects of Information Security, Privacy, and Trust, ser. Lecture Notes in Computer Science, vol. 9190. Springer, 2015, pp. 36–47. [Online]. Available: https://doi.org/10.1007/978-3-319-20376-8 4 [26] J. Holmes and M. Stubbe, Power and Politeness in the Workplace: A Sociolinguistic Analysis of Talk at Work, 2nd ed. Routledge, 2015. [Online]. Available: https://doi.org/10.4324/9781315750231 [27] Federal Bureau of Investigation, “Business email compromise,” accessed 2026-02-04. [Online]. Available: https://www.fbi.gov/how-we-can-help-you/scams-and-safety/ common-frauds-and-scams/business-email-compromise [28] Cybersecurity and Infrastructure Security Agency, “Business email compromise continues to swindle and defraud U.S. businesses,” Jun. 2015, alert TA15-175A; Accessed 2026-02-04. [Online]. Available: https://www.cisa.gov/news-events/alerts/2015/06/24/businessemail-compromise-continues-swindle-and-defraud-us-businesses [29] T. N. Jagatic, N. A. Johnson, M. Jakobsson, and F. Menczer, “Social phishing,” Communications of the ACM, vol. 50, no. 10, pp. 94–100, Oct. 2007. [30] K. L. Chiew, K. S. C. Yong, and C. L. Tan, “A survey of phishing attacks: Their types, vectors and technical approaches,” Expert Systems with Applications, vol. 106, pp. 1–20, 2018. [31] S. Zhuo, R. Biddle, Y. Koh, D. Lottridge, and G. Russello, “SoK: Human-centered phishing susceptibility,” ACM Transactions on Privacy and Security, vol. 26, no. 3, 2023. [32] J. Seymour and P. Tully, “Generative Models for Spear Phishing Posts on Social Media,” February 2018, arXiv:1802.05196 [cs, stat]. [Online]. Available: http://arxiv.org/abs/1802.05196 [33] M. Bethany, A. Galiopoulos, E. Bethany, M. B. Karkevandi, N. Beebe, N. Vishwamitra, and P. Najafirad, “Lateral phishing with large language models: A large organization comparative study,” IEEE Access, vol. 13, pp. 60 684–60 701, 2025. [Online]. Available: https://doi.org/10.1109/ACCESS.2025.3555500 [34] M. Weinz, N. Zannone, L. Allodi, and G. Apruzzese, “The impact of emerging phishing threats: Assessing quishing and llm-generated phishing emails against organizations,” in Proceedings of the 20th ACM Asia Conference on Computer and Communications Security, 2025, pp. 1550–1566. [35] Z. M. Hakim, N. C. Ebner, D. S. Oliveira, S. J. Getz, B. E. Levin, T. Lin, K. Lloyd, V. T. Lai, M. D. Grilli, and R. C. Wilson, “The phishing email suspicion test (PEST) a lab-based task for evaluating the cognitive mechanisms of phishing detection,” Behavior Research Methods, vol. 53, pp. 1342–1352, 2021. [Online]. Available: https://doi.org/10.3758/s13428-020-01495-0 [36] J. Cohen, Statistical Power Analysis for the Behavioral Sciences. New York: Academic Press, 1969. [37] V. Braun and V. Clarke, “Using thematic analysis in psychology,” Qualitative Research in Psychology, vol. 3, no. 2, pp. 77–101, 2006. [Online]. Available: https://doi.org/10.1191/1478088706qp063oa [38] L. Schöni, N. Roch, H. Sievers, M. Strohmeier, P. Mayer, and V. Zimmermann, “It’s a match—enhancing the fit between users and phishing training through personalisation,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, 2025, pp. 1–25. [Online]. Available: https: //doi.org/10.1145/3706598.3713845 [39] OpenAI, “Your data,” https://developers.openai.com/api/docs/guides/ your-data, 2026, accessed 2026-05-20. [40] G. Thomopoulos, D. Lyras, and C. Fidas, “Methodologies and ethical considerations in phishing research: A comprehensive review,” in Proceedings of the 2nd International Conference of the ACM Greek SIGCHI Chapter, ser. CHIGREECE ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3609987.3609990
LLM Usage Statement GPT-4o was used as part of the study methodology to generate simulated spear phishing messages from participant-provided work-context fields, as described in the methodology and ethical considerations sections. The authors remain responsible for all text, coding decisions, tables, figures, analyses, and claims.