simulated phishing training risk measurement and security controls

Using Simulated Phishing to Train Security Teams

by

What You Are Actually Buying With Simulated Phishing Training

Buying simulated phishing training does not buy proof that employees are secure, that the security function is fully trained, or that real phishing risk has fallen by a measurable amount. It buys a benign behavioral exercise: a controlled message, a set of observable actions, an opportunity to deliver relevant learning, and a test of whether people and systems respond as intended.

The phrase “train security teams” needs a clearer division. The security team usually operates the program: it selects audiences, designs scenarios, monitors delivery, receives reports, and changes controls or procedures. The wider workforce is normally the simulation recipient. Microsoft describes Attack Simulation Training in those terms: simulations are delivered to organizational users and their actions are reported back to the organization.

That distinction changes the buying decision. A platform may be useful even if its simulated click rate is not a credible proxy for enterprise-wide phishing resilience. Conversely, an attractive training library has limited operational value if employees cannot easily report a suspicious message or if those reports never reach the right queue.

Before deploying, renewing, redesigning, or pausing a program, define the outcome you actually want. There are four separate outcomes, and they require different campaign designs, metrics, owners, and safeguards.

  • Test employee behavior. A harmless invoice-change lure for accounts payable can reveal whether recipients inspect a sender, follow a verification process, click a link, or enter data on a mock sign-in page. Treat click and credential-entry rates as observations tied to a particular audience and scenario, not as a score for personal vigilance or a direct estimate of compromise risk.
  • Improve employee phishing reporting. The desired behavior may be a fast report through the approved button, not simply deletion of the message. Measure reporting rate, time to report, and whether reports arrive with enough context to be useful. A reporting-heavy program should also account for legitimate messages reported in error and the workload they create.
  • Validate the security operations and mail-flow workflow. A pilot simulation can test delivery, filtering, link rewriting, browser or proxy behavior, mailbox routing, alert creation, analyst acknowledgement, and escalation. This is often the most concrete value of an exercise: it can expose a broken reporting pipeline or an untested handoff before a genuine lure does.
  • Deliver targeted learning. A simulation can trigger short, role-specific guidance after a risky action, or a program can assign training directly without first making someone fail a test. Completion, comprehension, and later behavior should be measured separately; opening a landing page is not evidence that learning occurred.

For example, a finance-focused exercise may test both an employee’s response to a payment-change request and the organization’s vendor-verification workflow. A QR-code scenario may reveal gaps in mobile reporting that desktop email exercises never expose. Neither result, by itself, establishes that the organization has reduced real-world phishing incidents.

That restraint matters when interpreting dashboards. NIST’s Phish Scale User Guide cautions that click and reporting rates provide only a partial view of phishing risk. A lower rate may reflect an easier lure, a familiar template, changed filtering, or a different recipient group. Use difficulty, audience, delivery conditions, and the reporting path to interpret results rather than celebrating a single “failure rate.”

Viewed this way, simulated phishing training belongs inside a broader security awareness training program with mail protections, identity controls, clear reporting, and incident-response ownership. Purchase it as an operational capability with governance, data-access rules, tested delivery dependencies, and defined decisions that its results will inform. Promise reduced phishing risk only when a credible local evaluation can support that claim.

What Simulated Phishing Can and Cannot Prove

Simulated phishing training produces useful evidence about behavior inside a defined exercise, not a verdict on whether the organization has reduced real-world phishing compromise. The distinction matters because a dashboard can accurately show who reported, clicked, or entered data while still leaving open whether the lure was comparable to an attacker’s message, whether technical controls changed exposure, and whether employees would act the same way in a live incident.

NIST’s 2023 Phish Scale User Guide defines click and reporting rates clearly, but cautions that both are partial indicators of phishing risk. A 5% click rate on an obvious generic prize message does not mean the same thing as 5% on a credible invoice-approval lure sent to accounts payable. Difficulty, audience relevance, delivery conditions, and the reporting channel affect the result.

Claim or use caseWhat a simulation directly observesWhat requires local evaluationWhat remains unverified without a credible causal design
Behavior testingFor the recipients and lure delivered, the platform can record delivery, clicks, attachment or QR-code interaction, mock credential entry, deletion, and use of the approved reporting route.Whether the campaign’s NIST Phish Scale difficulty, business premise, target group, and mail delivery conditions resemble meaningful threats to the organization.That the observed credential-entry rate or click rate predicts behavior against real attacker messages, or that a lower rate represents greater phishing resilience.
Embedded trainingWhether training was assigned, opened, started, and—where telemetry supports it—completed after a simulated action.Whether the material is role-relevant, accessible, understood, retained, and delivered without creating shame or alert fatigue.That completing a module caused safer later behavior or reduced real-world compromise. Completion is not evidence of learning or retention.
Reporting improvementEmployee phishing reporting volume, reporting rate, median time to report, use of the correct button, and the messages reaching the intended mailbox or queue.Whether analysts can triage reports promptly, false-positive reports create manageable workload, and users receive clear acknowledgment or feedback.That improved reporting metrics alone caused earlier containment of genuine phishing, unless the full reporting pipeline is measured against comparable operational outcomes.
Technical-control validationWhether a test message is filtered, rewritten, quarantined, delivered, reported, routed to the SOC, and visible in relevant tools.Whether mail-flow rules, proxy behavior, browser reputation services, SIEM ingestion, ticketing, and escalation procedures work as designed.That controls will stop materially different live attacks, including attacks using other senders, payloads, identities, or channels.
Reduction in real-world compromiseA simulation can show change in exercise outcomes over repeated campaigns with documented conditions.Comparable populations, lure difficulty, exposure, filtering, reporting workflow, incident definitions, and a pre-specified evaluation method.That simulated phishing training reduced actual account takeover, fraud, malware infection, or phishing incidents. Such a claim needs a credible control group or another causal design that addresses competing explanations.

The 2025 UC San Diego Health randomized operational study is useful historical evidence against simplistic claims. Across ten campaigns involving more than 19,500 employees over eight months, it found no significant relationship between recent annual-training completion and simulation failure. It also reported a statistically significant but small average reduction in failure associated with embedded training, reported as 2%.

Those results should not be generalized beyond the healthcare organization and the annual and embedded interventions evaluated there. The study also found that more than half of training sessions ended within ten seconds, while fewer than 24% of users formally completed embedded material. For people who both received and completed interactive training, later clicking was lower, but the researchers noted possible selection effects. Read the findings as a reason to test local program design, not as proof that all training fails or succeeds.

The practical question is whether the exercise improves the security awareness training program and its security operations workflow. Track report-to-triage time, routing failures, analyst workload, repeatable lure difficulty, and changes made after each campaign. A finance-themed simulation, for example, can test whether a suspicious payment-change request reaches both the SOC and the vendor-verification process; it cannot establish that payment fraud will decline.

The FTC frames phishing preparedness as layered: employee simulations and a clear reporting path sit alongside email authentication and other technical safeguards. That is the right procurement standard. Keep filtering, MFA, identity protections, and incident response in scope; use simulations to test and improve the human and operational parts of that system rather than treating favorable phishing reporting metrics as proof that those controls are no longer needed.

simulated phishing training: What Simulated Phishing Can and Cannot Prove

Why Click Rate Alone Is a Weak Buying Metric

A simulated phishing training platform can produce a clean click-rate chart, but the chart alone says little about whether people will recognize, report, and help contain a credible real attack. NIST defines click rate as the number of people who clicked a simulated malicious link or attachment divided by the number sent the message. That is a useful campaign measure, not a complete measure of organizational phishing risk.

Consider two campaigns with the same 5% click rate. One uses an obvious generic “you won a prize” lure sent broadly across the company. The other imitates a document-sharing workflow that a particular department uses every week. The percentages match, yet the second message may be far more difficult for its intended audience to identify as suspicious. Treating those outcomes as equivalent rewards easy tests and can create false confidence.

Lure difficulty is only one source of distortion. A message may be irrelevant to recipients, while a finance-themed invoice request may be highly relevant to accounts-payable staff. Repeated templates can become recognizable as the organization’s own exercise rather than as phishing. A falling click rate may therefore reflect familiarity with the simulation library, not stronger judgment when attackers change their pretexts.

Technical conditions also affect the denominator and the observed behavior. Mail filters, link rewriting, browser reputation services, proxies, and mail-flow rules can block, alter, or interfere with delivery and telemetry; Microsoft documents these deployment dependencies for Attack Simulation Training. If only some intended recipients receive a usable message, or if tracking fails after a click, the reported rate needs qualification rather than a celebratory headline.

Reporting-process changes can make before-and-after comparisons equally misleading. Adding a prominent report button, changing mailbox routing, or teaching a new escalation path may increase employee phishing reporting even if message-recognition behavior has not changed. That can still be a valuable outcome: faster reports can improve the security operations workflow. But it measures a changed reporting pipeline, not necessarily a change in susceptibility.

NIST published the Phish Scale research in 2020 to rate human phishing-detection difficulty according to the message and its target audience. Its practitioner guide, published November 15, 2023, explains click and reporting measures while cautioning that they provide only a partial view of phishing risk. This is evidence-backed guidance for adding context to campaign results.

The editorial recommendation is more specific: use a documented difficulty model before comparing campaigns or claiming improvement. Record the lure’s premise, audience relevance, recognizable cues, delivery conditions, and the denominator used for every metric. A campaign sent to 2,000 people is not directly comparable with one measured only among the 1,600 recipients whose messages were delivered and tracked, unless that difference is disclosed.

For a phishing awareness program, review a balanced set of measures rather than a single “failure rate.” The most useful set links employee actions to operational response:

  • reporting rate and median reporting speed;
  • correct escalation through the approved reporting channel;
  • click and credential-entry behavior, with clearly stated denominators;
  • lure difficulty and relevance to the target group; and
  • control-group-adjusted change where the organization intends to make a causal claim.

That approach makes simulated phishing training more useful as an operational exercise: it can test whether people report suspicious mail, whether the report reaches the right queue, and whether the team responds effectively. It should sit alongside the organization’s phishing awareness program, mail protections, identity controls, and incident-response process—not serve as proof that real-world phishing risk has fallen.

The Measurement Framework to Put in the Procurement Brief

A procurement brief for simulated phishing training should define the measurement model before comparing template libraries or dashboard features. Require the vendor and internal program owner to label every percentage with its denominator: intended recipients, delivered messages, clickers, training assignees, training starters, or completers. A percentage without that denominator cannot be compared reliably across campaigns, departments, or platforms.

Use the NIST Phish Scale User Guide to contextualize behavior by lure difficulty and audience relevance. A low click rate on an obvious generic message is not equivalent to the same rate on a plausible invoice-approval request sent to accounts payable. Record the lure type, target population, delivery conditions, and difficulty rating alongside every outcome.

MetricPrecise definition and required denominatorWhat it revealsCommon interpretation error
Delivery rateDelivered simulation messages ÷ intended recipients. State whether mail filtering, quarantine, proxy rewriting, or exclusions altered delivery.Whether the exercise actually reached its planned audience and whether mail-flow controls interfered.Treating non-delivery as user resilience, or comparing click rates when one campaign was heavily filtered.
Click rateUnique people who clicked the simulated link or opened the payload ÷ delivered messages. If a vendor uses recipients sent instead, label that explicitly.Susceptibility to that specific lure under those delivery conditions.Calling it a real-world compromise rate or comparing lures with different difficulty, relevance, or repeated templates.
Credential-entry rateUnique people who submitted data on a harmless landing page ÷ delivered messages; also report credential entries ÷ clickers.How often a click progressed to the more consequential simulated action.Reporting only the clicker-based figure, which can hide population-level exposure. Never collect or retain genuine passwords.
Reporting rateUnique people whose report reached an approved mailbox, ticketing system, or SOC workflow ÷ delivered messages.Whether employees use the approved reporting path.Counting a forwarded email, deleted message, or report sent to an unmonitored address as an operational report.
Report-to-click ratioUnique approved reports ÷ unique clickers. Publish both underlying counts and state whether the numerator is limited to delivered messages.The balance between reporting behavior and risky interaction within one campaign.Using the ratio alone: a high value can result from very few clicks, very few reports, or both.
Median time-to-reportMedian elapsed time from message delivery to the first approved report, calculated only among users with an approved report.How quickly the workforce can surface suspicious mail for triage.Using an average distorted by late reports, or confusing employee report time with analyst response time.
Report-routing accuracyApproved reports that create the correct mailbox item, ticket, or SOC case ÷ all attempted employee reports.Whether the reporting button and phishing awareness program work as an operational pipeline.Assuming button use proves the SOC received usable telemetry or that the case contained enough information to investigate.
Training assignment, start, and completionReport separately: assigned ÷ eligible users; started ÷ assigned users; completed ÷ assigned users; and completed ÷ starters.Where learners disengage after a simulation or direct campaign.Calling assignment “training delivered,” or calling a start event evidence of completion or learning.
Retention and later behavior changeFor a defined follow-up interval, compare knowledge checks or comparable later behavior among completers and an appropriate comparison group. State the denominator for each outcome.Whether learning persists and is associated with later behavior.Assuming completion proves retention, or assuming lower later clicks were caused by training rather than selection, campaign difficulty, or technical controls.
Help-desk or SOC workloadIncremental tickets, reports, analyst minutes, and escalations attributable to the campaign ÷ delivered messages and ÷ approved reports.Whether reporting volume is manageable and whether the exercise exposes triage bottlenecks.Viewing all extra reports as failure. Some additional volume may be useful if it identifies routing or staffing gaps.
Control-group-adjusted change(Post-intervention outcome − pre-intervention outcome) for the intervention group minus the same change for a comparable control group. Use the same denominator, comparable lure difficulty, and stated observation period.Stronger evidence about whether a specific intervention is associated with change beyond broad trends.Making causal claims from an unadjusted before-and-after decline in clicks.

Separate employee behavior from system performance. A report counts operationally only when it reaches the approved mailbox, ticketing system, or SOC workflow and can be triaged. Track the next steps as separate measures: time from report to analyst acknowledgment, routing into the correct queue, and completion of any required containment or user follow-up. Microsoft documents that filters, proxies, browser reputation services, and mail-flow rules can affect delivery and telemetry, so pilot results should be checked before broad rollout.

Training requires the same discipline. Assignment means a person was designated to receive material; start means they opened it; completion means they met the platform’s completion rule; retention requires later evidence; and behavior change requires a comparable later measure. The 2025 operational study by Liu et al. found that embedded training effects were limited on average and cautioned that differences among people who completed interactive training may reflect selection effects. That is a reason to measure locally, not to promise a fixed reduction in phishing risk.

Finally, require exportable, anonymized campaign-level data rather than dashboard screenshots alone. The export should include audience and exclusion counts, delivery status, timestamps, lure type and difficulty, clicks, credential-entry events, approved reports, routing outcomes, training stages, and workload data. If the platform supports API reporting, such as Microsoft’s simulation report overview endpoint, specify the fields and retention period needed for independent analysis. That makes simulated phishing training auditable as a reporting-pipeline and learning exercise, rather than a vendor-defined scorecard.

The Measurement Framework to Put in the Procurement Brief

How to Design a Safe, Relevant Campaign

A simulated phishing training campaign should begin with a verified organizational threat model, not a vendor template library. Review reported messages, incident records, fraud attempts, help-desk patterns, email-security telemetry, and the business processes attackers would plausibly exploit. This keeps the exercise relevant while avoiding generic lures that measure familiarity with training rather than behavior under realistic work conditions.

Choose a scenario because it maps to an observed workflow. Finance teams may need an invoice-approval or vendor-payment-change exercise; collaboration-heavy groups may need a document-sharing prompt; identity teams may test a mock authentication notice. Payroll, package-delivery, and QR-code lures can be appropriate only where internal evidence supports that risk. Claims about the prevalence of any lure in a particular sector require verification from organization-specific or documented sector evidence.

Rate each message for difficulty before comparing outcomes. The NIST Phish Scale User Guide, published November 15, 2023, is useful for considering the cues in the message and its relevance to the intended audience. A 5% click rate on an implausible prize notice does not mean the same thing as 5% on a credible document-sharing request used in everyday work.

  • Use fictional senders, domains, invoice numbers, and attachments that cannot be mistaken for live business records.
  • Do not imitate a real supplier, customer, executive, or colleague so closely that recipients contact them, delay payments, alter records, or create a genuine dispute.
  • Exclude lures involving traumatic events, layoffs, illness, compensation uncertainty, discrimination, emergencies, or personal hardship.
  • Set written approval rules with security, HR, privacy, legal, employee relations, and business-process owners before launch.

Delivery conditions matter as much as copywriting. Schedule messages within legitimate local work hours and account for distributed teams, shift workers, public holidays, and time zones. Test how the campaign renders in managed mobile email clients as well as desktop browsers; a QR-code exercise, for example, may behave very differently on a phone. Use accessible HTML, meaningful link text, readable contrast, and landing pages that work with keyboard navigation and assistive technology.

Every payload must be harmless. A mock sign-in page may record that a participant attempted credential entry, but it must never collect, transmit, or retain a genuine password, MFA code, recovery code, or other sensitive input. Test mail delivery, link rewriting, browser reputation controls, proxies, and report telemetry in a pilot first. Microsoft notes that these controls can affect simulation delivery and tracking in its deployment guidance.

The post-click page should be constructive, immediate, and specific: identify the cues the recipient could have checked, show the approved reporting route, and explain what happens after a report. Avoid scoreboards, public naming, or language that frames a click as personal failure. The more useful design question is whether the campaign improved the phishing reporting workflow: did the report reach the right queue, with usable context, quickly enough for triage?

Finally, compare two learning models rather than assuming failure-triggered content is best. A direct training campaign gives relevant groups instruction before any simulated lure; embedded training appears after a click or other action. Microsoft documents both simulation-led and direct training campaigns in Attack Simulation Training. Test locally whether either approach changes reporting quality, reporting speed, and later behavior, while recognizing that simulated phishing training alone does not prove reduced real-world phishing risk.

How to Design a Safe, Relevant Campaign

Governance, Privacy, and Employee Trust Before Launch

Before sending a simulated phishing training campaign, treat it as an operational exercise involving employee data—not as a routine email test. Security should own the threat model, campaign design, technical safeguards, reporting-path validation, and interpretation of results. That ownership does not eliminate the need for formal review by HR, legal, privacy, employee relations, and, where applicable, works councils or unions.

The approval should produce written rules rather than informal assurances. Define who may view individual-level results; which roles receive aggregated reports; how long event data, training records, and exported files are retained; and whether exports are restricted, encrypted, or prohibited. Apply data minimization: collect the actions needed to assess the exercise, but do not collect real passwords on mock sign-in pages or retain unnecessary personal data.

  • State which populations are excluded or require adjusted treatment, such as employees on leave, contractors, accessibility accommodation recipients, or teams handling a sensitive business event.
  • Set boundaries for scenarios: no lures built around layoffs, illness, emergencies, compensation uncertainty, or other personal hardship unless the approval group has a compelling, documented reason and accepts the potential harm.
  • Specify whether individual results can influence performance management. If the answer is yes, define the review process, appeal route, evidence threshold, and role of HR and legal before launch.
  • Confirm that the “Report Phishing” route reaches the intended help desk or security operations workflow, rather than generating telemetry that nobody can act on.

The default use of results should be coaching and workflow improvement, not public rankings or automatic punishment. NIST’s human-centered cybersecurity guidance encourages organizations to focus less on punishment and more on helping people understand and report phishing. That approach is also more likely to reveal practical causes of poor outcomes: an unclear reporting button, an inaccessible message format, a mobile workflow that obscures warning signs, or a finance process that makes a dubious invoice request seem normal.

Build an employee-feedback loop into every campaign. Offer a short, non-retaliatory channel where recipients can explain that a scenario was confusing, unrealistic, inaccessible, poorly timed, or harmful. Security and employee-relations reviewers should examine that feedback alongside delivery failures, reporting delays, and help-desk workload. A campaign that produces more reports but also creates widespread confusion may need redesign, even if its phishing reporting metrics look favorable.

Keep sentiment evidence separate from behavioral evidence. A lower click or credential-entry rate does not demonstrate employee acceptance, trust, or psychological safety; it may reflect an easier lure, familiarity with a template, or changes in filtering. Use the NIST Phish Scale User Guide to contextualize campaign difficulty, and gather survey or interview feedback if the organization wants to make claims about workforce experience.

Finally, obtain legal review for the jurisdictions in which recipients work. Privacy, labor, employee-monitoring, works-council, union, records-retention, and disciplinary requirements vary by location and employment arrangement. A governed simulated phishing training program should document those decisions before launch, then connect findings to the organization’s phishing reporting process and incident-response improvements rather than treating a dashboard score as a verdict on individual employees.

Buy a Platform or Use Existing Microsoft 365 Capability?

For buyers evaluating simulated phishing training, the first question is not which product has the most persuasive dashboard. It is whether an existing Microsoft deployment can meet the organization’s operational requirements, or whether a specialist awareness-training platform fills material gaps in campaign design, learning delivery, reporting integration, administration, and support.

Microsoft documents that Attack Simulation Training requires Microsoft 365 E5 or Microsoft Defender for Office 365 Plan 2. Its Attack Simulation Training documentation, updated July 3, 2026, describes simulations, training campaigns, user actions, QR-code payloads, roles, and customer-data handling. Verify licensing terms, prices, bundle availability, assigned-user requirements, and live documentation immediately before publication or purchase; these details can change.

Decision areaUsing Microsoft capabilityAssessing a specialist platformEvidence to request
Existing licensing and population costEstablish whether E5 or Defender for Office 365 Plan 2 is already assigned to the intended recipients and administrators.Price the full intended population, including contractors, frontline staff, seasonal workers, and administrators.Current licensing terms, quote assumptions, minimum commitments, renewal terms, and costs for add-ons.
Campaign flexibilityTest whether available payloads, landing pages, scheduling, targeting, and training campaigns fit the threat model.Test the same requirements rather than assuming broader template libraries produce better outcomes.A documented pilot using comparable lures, including difficulty, delivery, reporting, and learning workflows.
Language and mobile coverageConfirm supported languages, accessibility, mobile-email behavior, and QR-code use for the actual workforce.Confirm equivalent coverage for every required language, device type, and regional delivery context.Named language list, mobile test results, accessibility information, and local support arrangements.
Reporting and automationAssess report-button routing and whether Microsoft Graph reporting meets central reporting needs.Assess reporting-button compatibility, ticketing or SIEM connections, APIs, exports, and webhook options.End-to-end test: employee report, mailbox routing, analyst triage, acknowledgement, export, and dashboard reconciliation.
Administration and dataUse assigned roles and evaluate least-privilege access to simulations and individual results.Review role granularity, tenant separation, administrator logging, customer-data location, and subcontractors.Data-processing terms, retention controls, deletion process, audit logs, export format, and support-access policy.
Delivery dependenciesValidate interactions with Microsoft mail controls and endpoint protections in the production environment.Validate the supplier’s sending domains, tracking methods, allowlisting needs, and security exceptions.Pilot evidence showing delivery, link behavior, and telemetry without weakening ordinary protections.

The table is a procurement framework, not a vendor ranking. A specialist product may be justified where the documented requirements include content governance, languages, mobile delivery, integrations, delegated administration, or service support that the Microsoft option does not meet in a pilot. Conversely, an integrated option may reduce another supplier relationship when it meets the intended use case without creating workarounds.

Do not compare template counts or headline analytics in isolation. Ask each option to run the same small, governed test: a work-relevant but harmless lure, a defined recipient group, the approved reporting route, and a short targeted learning path. Record delivery rate, report rate, reporting time, credential-entry rate where applicable, training assignment and completion, analyst workload, and any telemetry gaps. Apply the NIST Phish Scale User Guide when interpreting results; identical click rates from lures of different difficulty do not mean the same thing.

Reporting-button integration deserves separate scrutiny. The useful outcome is not simply that a user clicked “Report,” but that the message reaches the correct mailbox, queue, or analyst with enough context for triage. Microsoft documents a Graph endpoint for retrieving an attack-simulation report overview; buyers who centralize metrics should confirm its permissions, fields, export limits, retention implications, and fit with their reporting model. The same test should connect to the organization’s phishing reporting process, not a parallel workflow that staff will never use during a real event.

Make least privilege a contractual and configuration requirement. Identify who can create campaigns, upload recipient lists, view individual outcomes, change landing pages, export results, and access support tools. Also ask how long employee-level data is retained, whether results can be pseudonymized or aggregated, how data is deleted at contract end, what exports remain available, and whether customer data is used for product development or model training. Microsoft identifies simulation-related information as Microsoft 365 customer data; an external provider should provide equally specific answers for its service.

Delivery and measurement can fail for reasons unrelated to employee behavior. Microsoft’s deployment guidance warns that browser reputation services, proxies, mail-flow rules, filtering, and related controls can interfere with message delivery, payload tracking, or reporting telemetry. Treat these as implementation risks to test, not problems assumed away through broad allowlisting. A security exception that makes simulations work but weakens normal email or web protections can invalidate the exercise and create avoidable exposure.

Finally, compare the total operating cost of simulated phishing training, not only subscription price. Include identity and licensing prerequisites, implementation time, campaign approval, localization, help-desk preparation, SOC triage, HR and privacy review, administrator training, API work, support escalation, and the effort required to turn findings into better controls or targeted learning. The defensible choice is the option that supports a governed reporting-pipeline exercise for the intended population, with data and permissions the organization can responsibly manage.

Buy a Platform or Use Existing Microsoft 365 Capability?

Pilot the Reporting Pipeline and Decide Whether to Deploy

Treat a simulated phishing training purchase as a pilot of a socio-technical process, not a test of whether employees can be caught. Start with a written threat model: identify the business processes attackers would plausibly target, such as invoice approval, document sharing, payroll changes, or account sign-in. Include the intended audience, the harmless payload, the approved reporting route, the controls being exercised, and scenarios that are out of bounds.

Before sending anything, define success and its denominators. NIST defines click rate and reporting rate using people sent as the denominator, but cautions that neither metric fully represents organizational phishing risk. Record delivery rate separately, and specify whether credential-entry rate is calculated against all recipients or only people who clicked. A result cannot be interpreted if filtering, link rewriting, or mail-flow rules prevented part of the audience from receiving the same exercise.

Obtain HR, privacy, legal, employee-relations, and labor review where applicable before the pilot. The written rules should limit access to identifiable results, set retention periods, prohibit public ranking, and establish how coaching differs from performance management. Exclude harmful themes: personal illness, layoffs, compensation uncertainty, emergencies, and other lures likely to create distress or disrupt legitimate work.

Then validate the mechanics with a very small technical test. Microsoft documents that proxies, browser reputation services, filters, and mail-flow rules can interfere with simulation delivery, tracking, and report telemetry. That makes preflight testing necessary whether the organization uses Microsoft Attack Simulation Training or another platform. Test on managed desktop and mobile paths when both are in scope.

  • Deliver the benign message and confirm who actually received it, who had it quarantined, and whether links or attachments were altered.
  • Use the approved “Report Phishing” channel rather than an informal forwarding address, then verify mailbox, ticket, or SIEM routing.
  • Measure analyst acknowledgment, triage time, escalation to the accountable team, and closure or post-exercise remediation.
  • Confirm that the simulated landing page cannot collect or retain a real password, payment detail, or other sensitive employee input.

Run the first campaign with a limited, representative group and a safe, work-relevant lure. For an accounts-payable pilot, a mock invoice-review request can test both reporting and vendor-payment escalation without impersonating an actual supplier or interrupting a payment. Ask participants and SOC analysts for structured feedback: Was the message plausible? Was reporting easy? Did the explanation help? Did report volume create avoidable workload?

Collect anonymized campaign-level artifacts rather than relying on a dashboard screenshot: audience and exclusions, delivery rate, lure type, difficulty rating, clicks, credential entries, reports, median time-to-report, training assignment and completion, tickets created, and control changes prompted by the exercise. The NIST Phish Scale User Guide provides definitions and explains why click and reporting measures are only partial indicators. Rate lure difficulty and relevance before comparing results.

If leaders want to claim that training caused a change, pre-register a baseline or use a carefully governed control group. Keep the comparison fair: similar populations, comparable delivery conditions, and similarly difficult lures. The NIST Phish Scale was developed to contextualize differences in human phishing-detection difficulty; a lower failure rate on an easier or more familiar message is not evidence of greater resilience. The 2025 operational study by Liu and colleagues also found only a small average reduction in failure associated with embedded training and warned against broad causal conclusions from completion behavior.

A useful pilot can be successful even when few people click. It may reveal that the report button routes nowhere useful, a filter blocks the exercise, analysts lack ownership for acknowledgments, or a modest number of reports overwhelms the queue. Those findings are operational value: they show where the phishing awareness program and security operations workflow need repair before a real phishing message arrives.

Deploy simulated phishing training only when reporting ownership, safe scenario design, approved governance, reliable telemetry, and a measurement plan are in place. Redesign or pause when failure rate is the sole KPI, individual results are broadly exposed, security exceptions have not been tested, lures are harmful, or leadership promises unverified reductions in real compromises.

Finally, keep the decision in proportion. Simulations can test reporting pipeline behavior and support targeted learning, but they do not replace MFA, phishing-resistant MFA where available, email authentication, email filtering, or practiced incident response. NIST’s phishing guidance recommends training and reporting alongside MFA, while the FTC also frames phishing preparedness with reporting and email-authentication measures. Review the NIST guidance and the FTC guidance as layered-control references, not as proof that any one simulation program prevents real compromise.

Beyond Compliance: Real-World Impact of Nis2 on Small Enterprises

Beyond Compliance: Real-World Impact of Nis2 on Small Enterprises

NIS2 for small businesses: what changes in practice?Is the company legally in scope for NIS2 for small businesses?What is the difference between direct duties and supplier pressure?Which NIS2 for small businesses controls create useful evidence rather than compliance...

The Phishing Red Flags Checklist Every Employee Needs

The Phishing Red Flags Checklist Every Employee Needs

Phishing remains one of the most common and dangerous cyber threats facing organizations today. According to industry reports, over 80% of security breaches involve phishing in some form. The good news? Employees who know what to look for can stop these attacks before...

Step-by-Step Guide to Securing Shared Office Printers

Step-by-Step Guide to Securing Shared Office Printers

A Common Office Scene: How Printers Leak Sensitive Data — securing shared office printers You’re rushing between meetings in a busy shared office when you notice a stack of invoices and HR forms sitting unattended in the printer tray. Anyone walking by can pick them...

Can You Outsmart AI? A Cybersecurity Quiz for Managers

Can You Outsmart AI? A Cybersecurity Quiz for Managers

When an Email Looks Real: Start the AI cybersecurity quiz You open your inbox first thing and see a message from your IT director asking you to approve an urgent access request. The sender's signature, tone, and even the avatar look familiar—but the message was...

There’s no reason to postpone training your employees

Get a quote based on your organization’s needs and start building a strong cyber security infrastructure today.