Artificial intelligence is making spear phishing scams more effective, and most people cannot tell the difference between an AI-written lure and a human-written one. That is the core finding of new research from Brigham Young University, published in the peer-reviewed Journal of Cybersecurity and Privacy and highlighted in recent coverage by KSL.com and Phys.org.
"Our research is just the latest to show how sophisticated the current state of AI is," said Jerson Francia, a BYU cybersecurity researcher involved in the work, according to BYU's news release.
What the Researchers Did
The study, authored by a BYU team and available as a preprint on arXiv, set up a controlled comparison between personalized phishing text messages generated by GPT-4 and messages written by human novices — students writing under time constraints — for the same 25 targets.
Participants each received a set of 12 personalized messages crafted specifically for them. Some were AI-generated, some human-written. The researchers used a method they call TRAPD — the Threshold Ranking Approach for Personalized Deception — under which participants ranked the messages, indicated the point at which they would intend to click, explained their reasoning, and judged whether each message had been written by a person or by GPT-4.
The design matters because most earlier phishing studies measured outcomes in the abstract — generic messages sent to generic recipients. This one measured responses to messages personalized for each individual participant, which is precisely how real spear phishing works. The focus on SMS — often called smishing — is also deliberate: text messages carry an implicit trust that email does not, arrive on the device people check most often, and offer far less metadata for filters to work with.
More Clicks, No Better Detection
The headline numbers: GPT-4-generated messages elicited an intention to click more often than student-authored messages — 28% versus 21% — although the researchers note that the difference carries statistical uncertainty given the pilot study's size.
The more unsettling result is what participants could not do. When asked whether a message was written by a human or generated by AI, recipients identified the source no more accurately than chance. In other words, the extra wariness people believe they apply to machine-written text did not show up in the results — study participants were essentially coin-flipping on authorship.
The study also found that context beats content: messages framed around job-related matters were significantly more likely to trigger an intention to click than messages tied to hobbies or social media. For attackers, that is a force multiplier, because employment details are among the easiest personal information to find online.
Why Job-Related Lures Work Best
The finding on work-related messages deserves particular attention from defenders. Spear phishing succeeds through personalization — a message that references your employer, your role, or your current projects feels legitimate precisely because it is specific.
The BYU results suggest that large language models compress that personalization workflow dramatically. What once required hours of manual research and careful drafting by a skilled social engineer can now be approximated from simple prompts at near-zero marginal cost, the researchers write. The barrier to running a large, personalized campaign is no longer skill. It is access to a chatbot and a target list.
What It Means for Defenders
The practical implications are sobering for security teams. Awareness training often teaches employees to spot awkward phrasing, odd tone, or unnatural language — tells that historically exposed phishing attempts. A detector that relies on spotting "robotic" writing is now measurably unreliable, since recipients in this study could not distinguish AI-written from human-written messages at all.
The BYU team frames the findings as evidence that accessible AI-assisted personalization may increase the practical scale of social-engineering threats. Defenses that assume cheap, generic phishing will remain the norm are built on an assumption the data no longer supports. The economics of the attack have flipped: when personalized messages cost nearly nothing to produce, the number of targets an attacker can afford to pursue stops being the constraint, and volume shifts from a signal of fraud to background noise.
The researchers also stress the study's limits. It was a 25-target pilot using SMS messages, with student authors rather than professional phishers — a comparison that, if anything, understates the gap between AI output and average real-world scam quality. The team notes that while human judges could not tell the sources apart, the two message sets remained computationally distinguishable, leaving room for automated detection approaches.
For organizations, the takeaway is procedural rather than technical: verify requests through separate channels, treat urgency and job-related framing as risk signals regardless of how polished the message reads, and assume that well-written no longer means human-written.
Stay Ahead of AI
Tracking how AI is reshaping security and society? Read more AI news on AI Buzz Wire — independent AI coverage, every day.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →