The half that gets cut
Training a user to report creates a queue. Most programmes buy the training and never staff the queue, which teaches people that reporting achieves nothing.
A user reads a suspicious email, hesitates, and presses the button an awareness programme taught them to press. That press is where the programme either holds or quietly stops working, and it sits outside the line item that paid for the training. Teaching someone to report is a skill you buy by the seat. Receiving what they report is a rota, an owner and a reply, and the question before you sign is who has all three.
One control, two halves
An awareness programme is sold as training and measured as training, and it does not stop at the user. A trained user does not produce safety. A trained user produces a message. Every hour of training you buy raises the number of messages arriving at the address the report button points at, which makes that address part of the control.
The balance between the two halves has also moved. Awareness curricula spent twenty years teaching surface tells, the broken grammar, the odd greeting, the sentence that read like a machine translation. Generated text removes those tells in every language at once. Detection keeps its job and loses cues. Reporting gains work, because what the user is left holding is a feeling that something was off, and a feeling is only useful if it has somewhere to go.
The UK National Cyber Security Centre puts that requirement into its guidance on defending your organisation against phishing, asking whether the reporting process is clear, simple and quick to use, and telling organisations to give feedback quickly on what action was taken so people see their reports make a difference. That is not a courtesy. It is the input a user needs to keep doing what you paid to teach them.
Lain, Kostiainen and Capkun measured it at ETH Zurich across 15 months and more than 14,000 employees, published at IEEE Security and Privacy in 2022. Employees who received positive feedback after a report went on to report more than those who received none, a difference the authors record as statistically significant, and across 15 months they found no reporting fatigue so long as reporting stayed easy. A user who presses the button twice and hears nothing has learned something too, and it is the opposite of the lesson on the invoice.
Why the second half gets cut
The two halves have different shapes on a purchase order, and that decides which survives the budget round. Training is a per-seat licence, a content library, and a completion export for the auditor who asks for evidence that training happened. The queue is a named owner, a response time, and fifteen minutes of someone's Tuesday morning. Procurement can raise a purchase order against a licence. It has no instrument for a rota, which becomes something an existing team absorbs.
The workload argument does not survive the data. In the ETH Zurich study the report button reached over 21,000 employees and collected 7,191 reports of non-simulated email. Automated triage disposed of the bulk, and around 1.5 emails a day needed a decision from a human administrator. That study puts average reporting accuracy at 68 percent, so around a third of what arrives is not a threat. A shared mailbox with no filter in front of it is why the queue gets abandoned.
A falling click rate is not evidence
A click rate is a function of the population and the lure, and a dashboard varies only one of them on purpose. Make the next simulation slightly more obvious and the line falls, and the chart you take to the board looks like one a genuine improvement would produce. Generated lures sharpen that problem. When each campaign is written rather than drawn from a template library, two consecutive campaigns are not the same instrument, and the gap between their click rates carries the difficulty of the writing as much as the state of your users.
The lure effect has been measured. An 8 month randomised controlled trial at UC San Diego Health, ten simulated campaigns sent to over 19,500 employees and published at IEEE Security and Privacy in 2025, found that plain lures drew clicks from 1 to 2 percent of users while other lures produced failure rates upwards of 30 percent, which the authors describe as far outstripping the benefit attributable to training. Their instruction to anyone holding a before and after chart is to read it against an untrained control group.
Two measures survive both objections, and both belong to the second half. Report rate is the share of recipients who pressed the button, and it rises only when users believe pressing it causes something. Time to first report is how long a campaign sat in inboxes before anyone raised it, and it decides whether the remaining copies can be pulled in time. In the ETH Zurich study around 10 percent of reports arrived within five minutes, 20 percent within fifteen, and 30 to 40 percent within half an hour. The NCSC asks for the same shift, counting reports alongside clicks. If nobody raises your own simulation until the next afternoon, that is a statement about the button and the queue, not about your users.
A simulation only measures in the language the user reads
There is another way to produce a flattering click rate, and it is easier to do by accident. Send the simulation in a language the office does not work in.
Hasegawa, Yamashita, Akiyama and Mori surveyed 862 non-native English speakers in Germany, South Korea and Japan for SOUPS 2021 and found that participants, particularly those less confident in English, ignored English emails without careful inspection more readily than emails in their own language. Not clicking because the mail was skipped and not clicking because the user judged it are the same event on a dashboard and different events in the building.
An attacker has no such problem. A campaign aimed at a Turkish-speaking office arrives in Turkish, and it now arrives in fluent Turkish. A mail that read like a machine translation used to be one of the few cues that worked without any training, and it carried more weight outside English than inside it. Generated text takes that cue away and leaves the reader holding a suspicion with no evidence. Run an English-only simulation there and you have measured reading comprehension and filed it under security awareness.
Four things have to exist in the language of the desk, and a platform can localise some and leave the rest. The simulation templates, the training module assigned after a failure, the wording on the button, and the reply the user gets back.
Five questions that separate the halves
Ask these of whichever platform is in front of you.
- Where does a report land, who owns that destination by name, and what response time is in their rota?
- What runs automatically before a human sees a report, and what happens to the harmless share?
- Does the reporting user get an answer, and does the platform generate it?
- When one report identifies a live campaign, can the same message be pulled from every other mailbox, and from which console?
- Can the platform break report rate and time to first report down by department, or does it only count clicks?
KnowBe4, Proofpoint and Hoxhunt all ship a button that lets a user report an email. The questions are about what it is wired to.
Where this sits in the portfolio
The human layer area of the portfolio holds one product, BeamSec. PhishPro sends the simulation, Academy assigns training where the failure was, and PhishTrace is the second half, the queue a report lands in with an owner and a reply attached. Its training and simulation content is published in English and Turkish, which sets where the language test above can be passed today.