Bizarus.
HomeResearch › A field trial sent 19,500 hospital staff ten phishing lures over eight months. Training changed almost nothing.
Evidence Review

A field trial sent 19,500 hospital staff ten phishing lures over eight months. Training changed almost nothing.

Nearly every organisation trains staff against phishing. An eight-month randomised trial of more than 19,500 hospital employees, published at IEEE Security and Privacy, found annual awareness training had no significant effect on who fell for a simulated lure, and the post-click "you clicked" training reduced click likelihood by about two percentage points. What the email pretended to be mattered far more than whether the reader had been trained. We read the study, its limits, and a converging field experiment, and ask what the evidence actually supports.

19 August 2026 Bizarus

The question

Almost every large organisation now does two things about phishing. It makes staff sit an annual security awareness course, and it runs unannounced phishing simulations, showing a short "you just fell for it" training page to anyone who clicks. Both are close to universal, both consume real time, and until recently almost nobody had measured whether either one actually changes how often employees click. A team from UC San Diego set out to test exactly that, under conditions strong enough to support a causal claim rather than a correlation.

Why it matters

Phishing is not a marginal problem that training is over-engineered against. The authors note a 2023 IBM analysis identifying phishing as the single largest source of successful breaches, roughly 16 percent of the total, and the healthcare setting they studied is among the most heavily targeted: in 2023 the U.S. Department of Health and Human Services logged more than 725 large breach events covering over 133 million records. The intervention under test is also one most readers have personally been subjected to, which is part of what makes the result worth sitting with. It is rare to get a rigorous measurement of something so widely assumed to work.

What the researchers did

The study was an eight-month randomised controlled experiment run across UC San Diego Health, covering more than 19,500 employees. The team sent ten distinct simulated phishing campaigns over the period and randomised which lures employees received and whether, and in what form, they were shown embedded training after clicking. The authors describe it as the largest study of anti-phishing training effectiveness to date, and one of only two to use a randomised design rather than observing whatever an organisation happened to do. It was published at the 46th IEEE Symposium on Security and Privacy in May 2025 (DOI 10.1109/SP61157.2025.00076).

A note on how this piece judged the source. In computer security the peer-reviewed venue of record is the top conference, not a journal, and IEEE Security and Privacy ("Oakland") is one of the field's most selective. SCImago journal quartiles, the usual screen here, do not cleanly apply to a conference, so this piece assessed the paper directly: a large real-world field experiment with a randomised design, an author group with a strong track record in enterprise phishing research, claims kept carefully inside what the design supports, and an independent field study reaching the same conclusion (below). That is the paper-level standard this platform uses when a field's structure makes quartile the wrong instrument.

What they found

On the two interventions the study was built to test, the results were close to flat.

Two further findings sharpen the picture. Over the eight months and ten campaigns, the cumulative share of employees who clicked at least one lure climbed from around 10 percent in the first month to more than half by the eighth: enough exposure and most people eventually slip once. And the content of the lure mattered enormously. A message posing as an Outlook password update was clicked by 1.82 percent of recipients, while one posing as a change to the organisation's vacation policy was clicked by 30.8 percent. The gap between those two numbers is far larger than any gap the training produced.

The authors' own summary is measured:

Taken together, our results suggest that anti-phishing training programs, in their current and commonly deployed forms, are unlikely to offer significant practical value in reducing phishing risks.

What the finding does not establish

The careful reading matters here, because the headline invites overreach in both directions.

This is one organisation, in healthcare, in the United States. Susceptibility and the effect of training could differ elsewhere, and the study cannot settle that. "Failing a simulation" means clicking a link the researchers controlled; it is a proxy for real compromise, not the same event, and the step from click to breach involves further defences. The randomisation covered which lures people received and the embedded training; the comparison on annual training was necessarily observational within that population, since employees take the mandatory course on the organisation's schedule rather than by coin flip, so that particular result is an association inside a strong study rather than a randomised contrast.

Most importantly, the finding is not that training can never work. It tests training as it is actually deployed today: generic, annual, and delivered at the worst possible moment, right after someone has clicked and wants the interruption gone. A two percentage point effect is small, and the authors judge it poor value for the cost, but it is not zero, and better-designed or better-timed approaches were outside the study's scope. The claim is about the common forms, and it is disciplined enough to say so.

Reading it in context

A single null result is easy to wave away as a quirk of one site. This one does not stand alone. A separate 2025 field study at a large European university hospital, published at the ACM Conference on Computer and Communications Security (DOI 10.1145/3719027.3765164), tested a range of common in-situ anti-phishing interventions and likewise found them largely ineffective at changing risky behaviour. Two large field experiments, different continents, different teams, pointing the same way. Convergence like that is the real signal, and it is a good instance of a habit this platform keeps returning to: weigh the body of evidence, not the most quotable study.

There is a more general lesson in it for anyone who builds or evaluates interventions. Awareness training is intuitively obvious. The mechanism is legible, the cost is visible, the annual ritual is reassuring, and it can survive for years on plausibility alone, because plausibility feels like evidence until someone runs the experiment. When the experiment finally runs, the effect turns out to live somewhere other than expected: not in the trained person but in the situation. A lure dressed as an internal HR notice beat a security-flavoured one by more than an order of magnitude, which suggests susceptibility is largely situational rather than a fixed trait that a course tops up.

That is why the authors' recommendation points at systems rather than people: phishing-resistant two-factor authentication, and password managers that autofill only on the correct domain. Both remove the human decision rather than trying to improve it. It is the same move that works in most places where reliability matters, which is to change the environment so the safe outcome does not depend on every person being vigilant every time.

What remains open

Whether training that is timed to the moment of risk, or designed very differently from a compliance module, can do better than two points. Whether the pattern holds outside healthcare and outside the United States. How behaviour evolves over years rather than months as employees habituate to being tested. And where, precisely, the cost-effectiveness line sits, given that even a small real effect can be worth having if it is cheap enough. The study does not answer these. What it does, well, is remove the comfortable assumption that the current approach is obviously working, and hand the burden of proof back to anyone who wants to keep spending on it.

CybersecurityEvidence & Evaluationphishingsecurity awareness trainingrandomized controlled trialfield experimenthuman factorstwo-factor authentication
← All Research
© 2026 Bizarus AI