Gut feeling does nothing against AI spear phishing texts

A banker at a credit union sat down at a table with a dozen printed text messages, all of them written for that banker personally, and put them in order from the one most likely to get a click down to the one least likely. One of them stopped the sorting. It looked like something the bank sends out: “alert literally looks like the alert we get [at work] when there’s a fraud.”

AI spear phishing

Half the pile came from GPT-4. The banker was not told which half, and when asked to guess, did about as well as flipping a coin. So did almost everyone else.

That was the setup for a pilot study run at Brigham Young University with 25 volunteers. Each one filled out a survey first, handing over their job, workplace, hobbies, city, and something they had recently posted about online. Those details went into a short prompt template. GPT-4 got the template and produced six messages per person. A pool of undergraduates in a deception course got the same template and produced the rest, working under a fifteen-minute clock for up to four messages apiece.

Then each volunteer sat down with all twelve and sorted them.

The AI messages did slightly better, and the gap means less than it looks

Participants drew a line in their sorted pile: above this, I would have clicked. GPT-4’s messages landed above that line 28% of the time. The student-written ones managed 21.3%.

That is a gap of 6.7 percentage points, and it is tempting to read it as GPT-4 winning. The confidence interval says otherwise. It runs from 2.9 points in favor of the students to 16.3 points in favor of the model, which is another way of saying the study cannot tell you which side is ahead. Twenty-five people is not many. The result establishes neither that GPT-4 was better nor that the two were equal.

The thing worth noticing is what produced the gap. The AI side was one short prompt, filled in from a survey, run once per person. The human side was students who had sat through phishing instruction, whose output was then screened by a review team including two cybersecurity professors, with roughly a third of the submissions thrown out for being incomplete or unusable. A one-line prompt landed in the same neighborhood as that, and it can be run again, and again, for the cost of an API call.

Anything work-related was far more dangerous

Every participant got three flavors of message: one built on their job, one on a hobby, one on something they had posted to social media.

Work won, and not by a little. Job-themed messages cleared the click line 38% of the time. Hobbies managed 19%, social media 17%. This is the one comparison in the study that survives statistical adjustment, so if you take a single thing away from the numbers, take this one.

Personalization is also where the messages fell apart, and the failures are instructive. One participant read a message referencing a coworker and said, “there’s no Mike at work.” Another knew the instructional design team at their university and knew there was no Sarah on it. A message pitching a dance opportunity went straight to the bottom for someone who had quit dancing. Getting personal details into a message is easy now. Getting them right still requires knowing the target, and wrong details are worse than none, because they hand the recipient a reason to be suspicious.

Everyone had a theory about which ones were AI, and none of the theories worked

After the sorting was done, participants were told that one or more of the messages had been written by an AI and were asked to mark them. Across all 300 judgments, they got 52% right. Chance is 50%.

They were not short on reasoning. Some said AI writes too formally. Others said AI writes too generically. Perfect grammar was read as a machine tell, and typos as proof of a human. One participant zeroed in on punctuation: “why are there so many exclamation points all over?” Twelve of the 25 admitted somewhere in the interview that they were mostly going on instinct.

Emojis are an example of how badly the instincts were calibrated. They really were a tell: 66% of the GPT-4 messages had them, against 2% of the human ones. Only five participants mentioned emojis at all, and of those, two took them as a sign of AI while three took them as a sign of a person, on the theory that a chatbot probably could not do emojis.

Software found what the people missed

The researchers converted all 300 messages into embeddings and trained a classifier to sort AI from human. It hit 88.7% balanced accuracy under the strictest test conditions, which meant standardizing the URLs, stripping the emojis, flattening the case and digits and punctuation, and trimming each matched pair of messages down to the length of the shorter one. It was also tested only on people whose messages it had never seen during training, so it was not memorizing individual targets.

People could not find the pattern. A logistic regression found it easily.

Before anyone treats that as a detector, note the fine print. It was trained and tested on one message set, from one model, with one prompt design, against one pool of student writers. It has no proven ability to generalize anywhere else. And anyone who wants to beat this kind of classifier can: research cited in the paper shows that paraphrasing AI text with a detector in the loop knocks several of these tools down hard.

What 25 people on a Tuesday cannot tell you

The messages were printed on cards. Nobody’s phone buzzed, no sender number showed up, no link went anywhere. What was measured is what people said they would click, which is a well-used proxy in phishing research and still just a proxy.

The human comparison was novice students, not professional social engineers, so this is not AI against the best humans available. To detect a difference the size of the one observed with any confidence, the study would need about 100 completed targets rather than 25.

There is also a gap in the paperwork. The exact GPT-4 snapshot and the API logs were never recorded, so while the messages themselves survive and the analysis reproduces, the generation run that produced them cannot be repeated.

The practical advice at the end is short and does not depend on any of the uncertainty above. Check the sender, the channel, the link, and the request against what you would expect to receive. Do not try to decide whether it sounds like a robot. That is the one thing the study shows people cannot do.

Download: The ultimate guide to network operations management

Don't miss