LLM Spear Phishing Personalization
Abstract
Large language models can insert workplace details into phishing pretexts at low cost, but those details may either support or undermine a messages credibility. We recruited 180 U.S. working adults to evaluate simulated, AI-generated phishing emails in a disclosed survey. The emails used four cumulative levels of information: workplace (Level 1); recipient name and job title; job responsibilities; and coworker/sharedproject context (Level 4). Participants rated each messages convincingness from 0 to 100, chose one stated action (open the link, investigate, delete, or report), and explained why their highest- and lowest-rated messages stood out. Across 1,436 valid evaluations, convincingness increased by 2.40 points per personalization level in a sensitivity analysis, while the odds of expressing click intention increased by 28% per level. Among participants who did not express an intention to click, investigation remained common, reporting declined, and deletion increased. A post-hoc descriptive analysis found higher ratings and click intention for messages from a named person who referenced a supplied coworker than for messages from a department or entity. Qualitative coding showed why added detail could help or hurt: details that matched participants roles and routines supported credibility, while incorrect, vague, or channel-inappropriate details raised suspicion. Together, the results highlight that personalization is not simply a matter of adding more details: it depends on whether the pretext fits the recipients work context. We discuss how this distinction can inform workplace cybersecurity training.