The training completion report says 100%. The phishing simulation says a third of the same staff clicked.
Every MSP selling security awareness training lives between those two numbers, and in 2025 that gap finally got measured properly, by a hospital system that let researchers run a real experiment on its own employees.
What they found changes the job: what you should sell, what you should measure, and what you should put in front of a client at renewal. No vendor shortlist, no completion-rate theater.
TL;DR
- Report rate beats completion rate. Effective security awareness training moves the share of staff who report a phish, not the share who finish a module.
- The headline datasets disagree. KnowBe4 measured an 86% drop in simulated clicks; a UC San Diego trial measured about two percentage points.
- Risk is concentrated. Elevate Security and Cyentia found 8% of users tied to 80% of incidents.
- Compliance is a floor. Regulators and insurers ask for training delivery, not behavior change.
- Price the work, not the courseware. Simulation, triage, and reporting are the billable parts.
Why the Two Big Training Datasets Disagree
Start with the number every vendor deck leads with. Before training, roughly one employee in three clicks a simulated phish. After a year of ongoing training and simulations, it's closer to one in twenty-five. That's the 86% reduction from KnowBe4's 2025 benchmark report, drawn from 14.5 million users, and the story it tells is simple: keep training, and the clicking stops.
The second number comes from a hospital system in San Diego. Grant Ho, Ariana Mirian, and eight colleagues talked UC San Diego Health into something no vendor benchmark can offer: a randomized controlled trial across its more than 19,500 employees. For eight months they sent ten simulated campaigns designed in-house and measured what the training changed. Recent annual training made no statistically detectable difference, and embedded training after a failed simulation moved click rates by about two percentage points. The engagement logs explain why. Three-quarters of employees spent under a minute on the training page. A third closed it without reading a word.
The researchers gave the full talk at Black Hat 2025:
So the industry's biggest dataset says training cuts clicks by 86%, and the most rigorous trial yet published says it barely moves them. Both are honest. They measure different things.
The 86% is improvement against the platform's own simulations. Send times, sender domains, and template families repeat, and a workforce that has seen a few dozen of them can learn the tell. The two points came from a controlled experiment with nothing to sell, measuring whether the training itself changed what people click. Read together, they say the standard program teaches people to pass the test more than it changes the behavior.
That's a design problem, not a reason to drop training. It also gives you the most useful sentence you can say to a client: the goal isn't a lower click rate on your simulations. It's a faster, louder report when something real lands.
When a client waves the vendor number at you, don't argue with it. An 86% drop in simulated clicks is real, and it's the number a carrier will accept on a renewal form. It just doesn't tell them whether their staff will catch the next invoice fraud attempt. Two numbers, two jobs, and you should be reporting both.
Practitioners run this exact argument in public:
What Compliance Requires, and What It Leaves Open
One of those two jobs is paperwork, and paperwork has rules. Every framework a client operates under asks for training. None of them ask whether it changed anyone's behavior.
| Requirement | What it asks for | What it leaves to you |
|---|---|---|
| FTC Safeguards Rule (16 CFR 314.4) | Security awareness training for personnel, updated as risks change | Frequency, content, and any measurement at all |
| HIPAA Security Rule (164.308(a)(5)) | A security awareness and training program, with periodic reminders | What counts as periodic, and what a reminder contains |
| PCI DSS 4.0 (Req. 12.6) | Training on hire and at least annually, covering phishing and social engineering | Simulation cadence, targeting, and remediation |
| CMMC Level 2 / NIST SP 800-171 (3.2.1 to 3.2.3) | Role-based training and insider threat awareness | How role-based is defined and evidenced |
| Cyber insurance applications | Whether training and simulated phishing run, and how often | Everything about quality |
Two things follow. First, produce the compliance artifact flawlessly; it's cheap, and a missing training record is what turns a claim into an argument. Our breakdown of cyber insurance requirements for MSP clients covers how carriers word these questions and what evidence they accept.
Second, the artifact is where competitors stop. A program that reports behavior alongside completion is a different conversation at renewal when a client asks why they're paying you instead of buying seats direct.
Security Awareness Training Metrics That Predict Risk, and the Ones That Don't
Completion rate and quiz score measure whether a module was opened. Neither predicts what happens when a real lure arrives on a Tuesday afternoon. Swap them for signals tied to behavior.
| Vanity metric | Behavioral replacement | What it tells you |
|---|---|---|
| Completion rate | Report rate on simulations | Whether staff know the reporting path and use it |
| Quiz score | Median time to report | How long an attacker has before your SOC hears about it |
| Click rate on templates | Repeat-clicker concentration | Which users need a different intervention |
| Modules assigned | Report rate on real, unsimulated phish | Whether the behavior transferred outside the test |
| Annual training date | Time from failure to remediation | Whether the program closes its own loop |
Report rate is the anchor, the only metric on that list that produces something operationally useful. Click rate tells you a user failed after the fact. A report tells your team a campaign is live right now, which is why the reporting path deserves more design attention than the courseware.
Median time to report is the second number to track. If a client's median is four hours, an attacker gets a comfortable head start. Under fifteen minutes and your triage queue becomes a live detection feed. That's a number a carrier, an auditor, and a CFO all understand.
Verizon's 2025 Data Breach Investigations Report put the human element in roughly 60% of breaches, and its simulation data argues for the metric swap better than any vendor deck could: recent training barely moved click rates but roughly quadrupled report rates, from about 5% to 21%. Add the concentration finding from Elevate Security and the Cyentia Institute, 8% of employees behind 80% of security incidents. Risk isn't spread evenly across a workforce. Almost nothing in a standard platform rollout uses that.
Human Risk Management Means Aiming at the Users Who Get Targeted
A uniform annual module spends the same budget on the warehouse supervisor who never touches a wire transfer as it does on the controller who approves them. Targeting fixes that without adding cost. The annual module even has its own fatigue threads:
Build the risk tier from data the client already has. Finance and payroll staff, anyone with delegated mailbox access, domain admins, executive assistants, and anyone named on the public website get a different program: shorter, more frequent, and pointed at the lures they receive. Business email compromise attempts on a controller look nothing like the generic credential-harvest page in a standard template library.
Then treat repeat clickers as an operations problem, not a discipline problem. A user who fails three simulations in a row is making a judgment call twenty times a day at speed, and losing a few. Reduce the number of calls they have to make. Tighten their mail filtering rules, add external-sender banners with real detail, remove standing permissions they don't need daily, and route their unknown-sender traffic through an extra check. Technical controls scale better than willpower, and email security solutions that carry their weight do more for a high-risk user than another module ever will.
Targeting matters more now than it did three years ago because the lures got cheaper to personalize. Verizon's 2025 report found the share of malicious emails carrying AI-generated text doubled in two years, and the lures that budget buys read like internal mail, with the sender's writing style, a real project name, and correct timing. If a client wants an analyst name on where this goes, Gartner called it in 2024: enterprises pairing generative AI with a platform-based security behavior and culture program would see 40% fewer employee-driven security incidents by 2026. Treat the precision with the suspicion any vendor benchmark deserves; the mechanism is what holds. Personalization is another way of saying the uniform annual module is the part being replaced.
The spelling-error tells that older training modules teach are gone from that class of attack. A controller trained to look for bad grammar and a suspicious domain is holding a checklist that no longer matches their inbox. One more reason to train the reporting reflex instead of the detection heuristic.
Make the Report Button the Product
The part of the program that pays best is what happens in the ninety seconds after someone reports something. Get that right and users keep reporting. Get it wrong and the button goes quiet within a quarter.
Wrong looks like silence, or a canned auto-reply, or worse, a reply three days later telling a user that a legitimate invoice was fine and they should be more careful. Every one of those trains people not to bother. Right looks like an acknowledgment inside a few minutes, a real answer within the business hour, and public credit when someone catches something genuine.
That means the reporting workflow has to land in your PSA as a real ticket with an owner and an SLA, not in a shared mailbox nobody watches. Auto-triage the obvious noise, group duplicate reports of the same campaign into one incident, and push a short all-staff note when a campaign is confirmed live. Three people reporting the same lure inside ten minutes is a detection event, and treating it that way is what turns awareness training from a cost line into part of the client's security operations.
This is where the tooling question gets practical. Report triage is high-volume, low-complexity work with a clear pattern, which makes it a natural fit for AI-native handling inside the ticket flow rather than a technician reading each one. That's the shape of what we build at Flamingo, an AI-native all-in-one MSP and IT platform where agents work the ticket queue instead of sitting beside it. But the requirement holds whatever you run it on: reports become tickets, tickets get answered fast, and campaigns get correlated.
A 90-Day Phishing Simulation Rollout You Can Sell
Ninety days is long enough to produce a defensible baseline and short enough to fit inside a quarterly business review. Run it identically for every client.
- Days 1 to 14. Deploy the report button across all mailboxes, wire it into the PSA with an owner and a one-hour response target during business hours, and run a single baseline simulation, unannounced to staff but approved in writing by the client owner, before any training goes out. Record baseline click rate, report rate, and median time to report.
- Days 15 to 45. Assign short role-based modules, ten minutes rather than forty, weighted toward the high-risk tier. Run a second simulation using a lure family that differs from the baseline, so you're testing transfer instead of familiarity.
- Days 46 to 90. Move to monthly simulations with rotating lure families, add just-in-time coaching at the moment of a click, apply technical controls to repeat clickers, and publish the first client-facing report against the day-one baseline.
The unannounced baseline in week one is the part people skip, and it's the part that makes everything after it credible. Without it there's no before, and the entire program becomes an assertion.
What the Client Report Should Say
A three-page report beats a platform dashboard nobody logs into. Keep it to what a non-technical owner can read in five minutes and hand to their carrier.
Lead with report rate and median time to report, both trended against the day-one baseline. Follow with the real phishing campaigns your team caught because someone reported them, named by date and outcome. That's the number that makes the program feel worth paying for. Then the compliance block: who completed what, when, and against which requirement. Close with the high-risk tier, what changed for those users this quarter, and what you're recommending next.
Leave the raw click rate in an appendix. It's the number every competitor leads with, and the one that tempts a client to judge the program on how easy the simulations were.
Pricing It So the Program Survives Renewal
Courseware and simulation licensing is the small part of the cost. The work is deployment, triage, incident correlation, remediation, and the quarterly report.
Burying training inside a flat per-seat security line hides the labor, and hidden labor is the first thing that quietly stops getting delivered when margins tighten. A separate line, priced per user per month with the triage SLA written into it, gives you an ROI case for security awareness training that you can defend with numbers: campaigns caught, hours of exposure removed, an insurance question answered cleanly. Price it to cover the triage and reporting hours (figure two to four tech hours per client per month), not the license, and template the simulations and the report across your whole book, because a 25-seat client can't fund custom work. When you rebuild the stack pricing around it, the same overlap analysis applies here as anywhere else in the MSP security stack, where duplicate spend across bundled products is common enough to fund the layers a client is short on.
Plenty of clients already own a platform they bought direct and never ran properly. That's a better starting position than a greenfield sale. Take over administration of what they have, with admin access and co-management responsibility written into the agreement, run the baseline against it, and bill for the operating layer: simulation design, triage, correlation, remediation, reporting. They're not writing off a license they just renewed, and the first quarterly report against a real baseline usually settles who should be running it.
One caution on vendor claims. A platform reporting an 86% improvement is reporting improvement on its own simulations. Ask any vendor what their customers' report rates look like on unsimulated phish, and how they measure transfer to lure families outside the template library. The answers separate the products built for behavior change from the ones built for a completion certificate. MSPs compare platform notes in the open:
What Effective Security Awareness Training Comes Down To
Forget the module library, and forget the click rate on a test the vendor writes. The program you can defend at renewal is a workforce that reports fast, a triage queue that answers faster, and a report that proves both. Build that and the compliance artifact comes free, and the completion report that says 100% finally sits next to a number that means something. Build only the artifact and you're one claim denial away from finding out what it was worth.
Content Marketing Lead
Ohayo! I'm Kristina, and I'm doing good things with content, SEO, social, and community at Flamingo. Before IT, I worked as a correspondent for Ukraine's Public Broadcasting Company and have a Master's in journalism.
