ISO27001

Root Cause Analysis for ISO 27001 Corrective Action

Root Cause Analysis for ISO 27001 Corrective Action
Loading advertisement...
13

Priya Nkemelu had closed the same nonconformity twice before she realized she'd never actually closed it at all.

She was the information security manager at Halvorsen Freight Analytics, a 260-person logistics-data company in Rotterdam that had been certified to ISO 27001 for just under two years. The first time a certification auditor flagged it, the finding was almost trivial: a departed contractor's VPN account was still active three weeks after his last day. Priya's team disabled the account, wrote up a short correction, attached a screenshot of the deactivated user in the identity provider, and the auditor closed the nonconformity in the surveillance audit report. Done, she thought.

Eleven months later, at the next surveillance visit, a different auditor pulled a sample of ten departed employees and found two more accounts still live — one for nineteen days, one for six weeks. Same root cause, different names. This time the auditor didn't just write it up as a repeat — she escalated it from minor to major, because the evidence now showed the control had failed systemically, not once. A major nonconformity meant Halvorsen had 90 days to close it with evidence of effectiveness or risk suspension of certification, right in the middle of a due-diligence data room review by a potential acquirer who had made "clean ISO 27001 status" a term in the letter of intent. Priya's CEO put a number on it in the hallway outside the audit closing meeting: the deal was worth an incremental $4.2 million in valuation premium tied to that unbroken certification. It was now genuinely at risk over an offboarding checkbox.

What went wrong wasn't the deactivation — that correction was fine, both times. What went wrong is that nobody ever asked why the account stayed active. Was it a missed step in a manual checklist? A broken integration between HR's termination workflow and the identity provider? A ticketing SLA nobody was tracking? Priya had fixed the symptom twice without ever finding the cause once. That's the gap this article closes.

Who this is for

You're reading this because you (or someone on your team) has a nonconformity sitting open — flagged in an internal audit, a Stage 2 certification audit, or a surveillance visit — and you need to write a root cause analysis and corrective action that will survive scrutiny at the next audit. This article assumes you already understand the basic mechanics of Clause 10 nonconformity and corrective action requirements and have probably already read about the common nonconformities auditors raise; this piece goes one layer deeper into the how of root cause analysis itself. By the end, you'll be able to run a proper 5 Whys or fishbone session, tell the difference between a correction and a corrective action in your own documentation, and write a CAPA record that an auditor reads and closes on the first pass instead of bouncing back to you with "insufficient root cause identified."

Correction vs. corrective action vs. root cause analysis

These three terms get flattened into one another constantly, and that flattening is precisely what produces nonconformities like Priya's. ISO 27001 Clause 10.2 (Nonconformity and corrective action) is explicit that an organization shall react to a nonconformity — take action to control and correct it, and deal with the consequences — and separately evaluate the need for action to eliminate the cause(s) of the nonconformity so that it does not recur or occur elsewhere. Those are two different obligations, and the standard names them differently on purpose.

Term

What it means

Halvorsen example

Does it satisfy Clause 10.2 alone?

Correction

Immediate fix to the specific instance of the problem

Disabling the one departed contractor's VPN account

No — this is the "react" half only

Root cause analysis

Structured investigation into why the nonconformity occurred

Discovering the HR-to-IdP termination handoff had no automated trigger and depended on a manual email

Not itself an action, but the required input to one

Corrective action

Action taken to eliminate the identified root cause, preventing recurrence

Automating account deprovisioning as part of the HR offboarding workflow, with a control to catch failures

Yes, when informed by genuine root cause analysis

A correction with no root cause analysis behind it is exactly what got Halvorsen a repeat finding. Correction addresses "this specific thing that happened." Corrective action addresses "the mechanism that let this kind of thing happen." Root cause analysis is the bridge between the two — it's the investigative step that tells you what the corrective action actually needs to change. Skip it, and your corrective action is really just a second correction wearing a corrective-action label. If any of this vocabulary is still fuzzy across your own team, it's worth getting everyone aligned using our ISO 27001 terminology and glossary primer before you're mid-audit trying to explain the difference on the fly.

It's worth noting explicitly here: ISO 27001:2022 does not carry a standalone "preventive action" clause the way the 2005 version of ISO 27001 and many older management-system standards did. In the current structure, that forward-looking, before-something-breaks thinking is folded into the risk-based approach that runs through Clauses 4 through 6 — your risk assessment and risk treatment plan are effectively doing the preventive-action job continuously, rather than as a reactive clause triggered by a nonconformity. Corrective action, by contrast, is inherently reactive: something already happened (or was found not to conform), and Clause 10.2 governs what you do about it.

Why auditors reject shallow root cause analysis

Certification body auditors see thousands of CAPA records a year, and they develop a very fast nose for the ones where "root cause" was filled in as an afterthought — usually in the last ten minutes before a document was due to the audit team. The tells are remarkably consistent across auditors and certification bodies, which is useful, because it means you can pattern-match against what gets rejected before you submit anything.

Shallow RCA pattern

What the auditor sees

Why it gets rejected

Root cause = "human error"

No investigation into what made the error possible or likely

Doesn't explain why the same error hasn't happened to every other employee doing the same task

Root cause = the nonconformity restated

"Root cause: the account wasn't deactivated"

That's the finding, not the cause — it answers "what," not "why"

Root cause = a training gap, with no analysis of why training failed

"Employee needed more training"

Training gaps are themselves symptoms of a system issue (no onboarding checklist, no competency verification) unless proven otherwise

One "why" and stop

5 Whys chain with a single iteration

Auditors are trained to ask "and why did that happen?" back at you in the closing meeting

No evidence the RCA process was actually followed

A conclusion with no supporting notes, interview record, or data

Looks fabricated after the fact to satisfy the paperwork requirement

Corrective action doesn't map back to the stated root cause

Action plan addresses a different problem than the one identified

Suggests the RCA and the CAPA were written independently, not as cause-and-effect

The underlying principle auditors are applying is simple: Clause 9.2 internal audit and Clause 9 performance evaluation exist to generate real signal about whether the ISMS is working, and Clause 10.2 exists to make sure that signal actually changes something. If your RCA doesn't survive a skeptical second read, the auditor has no basis to believe your corrective action will prevent recurrence — and preventing recurrence is the entire point of the clause. An auditor who accepts weak RCA on your file today is the same auditor who has to explain a repeat major nonconformity to their own accreditation body eighteen months from now. They have every incentive to push back hard, and they will.

"I can tell within thirty seconds of reading a CAPA record whether root cause analysis actually happened or whether someone reverse-engineered a plausible-sounding cause to justify the fix they'd already decided on. The giveaway is always the same: the 'why' and the 'what we're doing about it' don't quite line up." — Marcus Oduya, Lead Auditor, Kestrel Certification Partners

What Clause 10.2 actually requires — and what it leaves to you

It helps to separate, line by line, what the standard mandates from what is left to practitioner judgment, because a lot of RCA anxiety comes from organizations assuming ISO 27001 prescribes a specific methodology. It doesn't.

Clause 10.2 element

What the standard requires

What it does NOT specify

React to the nonconformity

Take action to control and correct it; deal with the consequences

Which correction tool or ticketing system to use

Evaluate the need for action on the cause

Review the nonconformity; determine its causes; determine if similar nonconformities exist or could potentially occur

Which RCA technique (5 Whys, fishbone, etc.) to use — this is entirely your choice

Implement any action needed

Carry out the corrective action determined necessary

A specific implementation timeline (though certification bodies typically expect 90 days for major NCs)

Review the effectiveness of any corrective action taken

Confirm the action actually worked

A specific effectiveness-review method — you define what "worked" means and how you'll check

Make changes to the ISMS if necessary

Update the management system if the RCA reveals a broader issue

Which documents must change — that's case by case

Retain documented information

Keep records as evidence of the nature of the nonconformities, actions taken, and results

A specific record template — you build your own

That last column matters enormously. Auditors are not checking your work against a hidden ISO-mandated checklist of RCA steps; they're checking whether your process was rigorous enough that a reasonable person would believe the cause you identified is the real cause. That's a judgment call, which is exactly why the practitioner techniques in the next section exist — they're not ISO requirements, they're the tools experienced practitioners use to make that judgment call defensible.

A quick word on where these techniques come from

None of the following methods — 5 Whys, Ishikawa/fishbone diagrams, fault tree analysis, Pareto analysis, or barrier analysis — are named or required anywhere in ISO/IEC 27001 or ISO/IEC 27002. They are quality-management and reliability-engineering techniques with decades of use across manufacturing, aviation safety, healthcare, and IT service management, long predating and existing entirely independent of ISO 27001. I'm including them here because in 15+ years of running ISMS implementations and audits across 200-plus organizations, these are the tools that actually produce defensible root causes — not because a clause anywhere tells you to use them. Pick the one that fits the nonconformity in front of you; auditors don't grade you on which acronym you used, only on whether the cause you landed on holds up.

Technique

Best suited for

Typical time investment

Output format

5 Whys

Single, relatively linear nonconformities with one clear causal chain

20–45 minutes

A short chain of cause-and-effect statements

Ishikawa / fishbone diagram

Nonconformities that could stem from multiple contributing categories (people, process, technology, etc.)

45–90 minutes, often in a group workshop

A visual diagram with branching causes

Fault tree analysis

Complex or safety-critical failures with multiple possible failure paths and combinations

Several hours to a full day

A logic diagram with AND/OR gates showing failure combinations

Pareto analysis

Recurring or high-volume nonconformities where you need to prioritize which causes to tackle first

1–2 hours, requires existing data

A ranked bar chart of causes by frequency or impact

Barrier analysis

Incidents where a control existed but failed to stop the event

30–60 minutes

A table of intended barriers vs. what actually happened at each one

The 5 Whys: fast, cheap, and the right first move for most single-instance nonconformities

The 5 Whys technique is exactly what it sounds like: you state the problem, ask "why did that happen," write down the answer, and then ask "why" again about that answer — repeating until you hit a cause that, if fixed, would plausibly have prevented the original nonconformity. Five is a convention, not a rule; some chains resolve in three iterations, others need seven. The discipline is not stopping at the first answer that sounds sufficient, which is the single most common way 5 Whys sessions fail.

Here's how it should have gone at Halvorsen the first time, instead of stopping at "we disabled the account":

Step

Question

Answer

Problem statement

A departed contractor's VPN account was still active 3 weeks after his last day

Why 1

Why was the account still active?

Because IT was never notified that the contractor's engagement had ended

Why 2

Why wasn't IT notified?

Because offboarding notification depends on the hiring manager submitting a manual offboarding form, which the manager forgot to submit

Why 3

Why does notification depend on a manual form with no backstop?

Because the HR system and the identity provider are not integrated — there's no automated trigger tied to contract end dates

Why 4

Why isn't there an automated trigger, given HR already stores contract end dates?

Because the ISMS implementation in 2023 treated offboarding as a policy/procedure control (Control 6.5) rather than a system integration project, and no one revisited it after go-live

Why 5 (root cause)

Why did no one revisit it?

Because there is no periodic review step that checks whether manual controls tied to HR events are still adequate as headcount grows — the control was never re-risk-assessed after the contractor population tripled

That fifth answer is a genuinely actionable root cause: the organization scaled its contractor headcount without re-evaluating whether a manual, memory-dependent control could keep up. Contrast that with stopping at Why 2 ("the manager forgot") — a corrective action built on that shallow answer would probably just be "retrain the hiring manager," which does nothing for the next manager who forgets, and nothing structural changes. This is precisely the failure pattern that produced Halvorsen's repeat finding.

5 Whys works best on nonconformities with a single, fairly linear causal thread. When you find the chain branching — "well, it could be A or it could be B" — that's your signal to switch to a fishbone diagram instead, which is built to hold multiple branches at once.

"5 Whys is deceptively simple, which is exactly why teams botch it. The failure mode I see constantly is treating the fifth 'why' as sacred — stopping because you hit the number five, not because you hit an actual system-level cause. Sometimes the real root cause is at why three. Sometimes it's at why eight. Count the whys, don't worship them." — Elena Farkas, ISMS Program Lead, Brindlewood Health Systems

The Ishikawa (fishbone) diagram: when the cause could be coming from several directions at once

Named for Kaoru Ishikawa, the quality-management pioneer who popularized it, the fishbone diagram gets its name from its shape: a horizontal spine pointing at the problem, with diagonal "bones" branching off representing categories of possible cause. The classic manufacturing categories are the "6 Ms" — Machine, Method, Material, Manpower, Measurement, Environment — but for information-security nonconformities, most practitioners adapt the categories to something like People, Process, Technology, Policy, Third Parties, and Physical Environment.

The value of a fishbone over 5 Whys is that it doesn't force a single linear chain. It's a brainstorming structure that lets a cross-functional team throw multiple candidate causes onto the board simultaneously, across categories, before converging on which branch actually explains the evidence. This matters for nonconformities where the cause genuinely could be more than one thing — say, a finding that vulnerability scan remediation SLAs were routinely missed (a common flag under technical vulnerability management, Control 8.8).

Category

Candidate causes raised in the workshop

Retained after evidence review?

People

Security team understaffed relative to ticket volume

Yes — headcount data confirmed a 40% increase in scan findings with no corresponding staffing increase

Process

No formal SLA escalation process when a ticket ages past threshold

Yes — ticketing system showed zero automated escalations configured

Technology

Scanning tool generating excessive false positives, causing fatigue and deprioritization

Partially — contributed, but not the primary driver

Policy

Vulnerability management policy states 30-day SLA for critical findings but doesn't define enforcement mechanism

Yes — policy silent on what happens when SLA is breached

Third parties

Patch dependencies on vendor release cycles outside the org's control

No — sampled tickets showed in-house-patchable systems missing SLA too

Physical environment

Not applicable to this finding

N/A

Notice the workshop generated six candidate branches but the evidence review only retained three as genuine contributing causes, discarded one as an outsider factor, and ruled one irrelevant entirely. That discipline — going back to evidence after the brainstorm, not just writing down whatever the loudest person in the room proposed — is what separates a fishbone diagram that survives audit scrutiny from one that's really just a meeting artifact. A good fishbone session usually feeds directly into a Pareto analysis (below) to rank which of the retained branches to fix first.

Fault tree analysis: for nonconformities where failure requires a combination of things going wrong

Fault tree analysis (FTA) comes out of aerospace and nuclear reliability engineering, and it earns its keep on ISO 27001 nonconformities that clearly involved more than one independent failure lining up — the kind of finding where, if any single layer had held, the nonconformity wouldn't have occurred. It's a top-down logic diagram: you start with the undesired top event, then work downward through intermediate events connected by AND gates (all conditions must be true) and OR gates (any one condition is sufficient) until you reach basic events — the lowest-level, independently verifiable causes.

FTA is more effort than 5 Whys or fishbone, and it's overkill for a single-cause finding like a missed patch. It earns its place when the nonconformity looks like a near-miss that could easily have become an incident — for example, a finding that a terminated employee retained badge access to a data center for eleven days, which only didn't result in unauthorized physical access because the employee happened not to attempt entry.

Level

Event

Gate type

Basic causes underneath

Top event

Terminated employee retained physical data center access

Intermediate event A

Badge deactivation did not occur at termination

OR

Facilities team not notified (see People branch) OR badge system sync job failed silently

Intermediate event B

No compensating control caught the gap before day 11

AND

Access review report exists AND was not reviewed within SLA

Basic event 1

Facilities notification depends on a Slack message from HR, not a system trigger

Confirmed via process walkthrough

Basic event 2

Badge system's nightly sync with HR system had been failing silently for six weeks

Confirmed via system logs; no alerting configured on sync failures

Basic event 3

Weekly access review report was generated but sat unread in a shared mailbox

Confirmed via mailbox timestamps

What FTA surfaces that 5 Whys typically wouldn't is the AND-gate structure: the badge stayed active and the compensating control (the weekly access review) also failed to catch it. A corrective action that only fixes the badge sync job leaves the second failure — an access review nobody reads — completely unaddressed, and the next similarly independent failure sails right through. FTA forces you to write a corrective action for every branch, not just the most visible one.

Pareto analysis: prioritizing causes when you have volume, not just a single incident

Not every nonconformity is a single event like a stuck VPN account. Some are volume findings — an auditor samples 25 access requests and finds 9 without documented manager approval, or an internal audit reviews 40 security awareness training completions and finds 12 overdue. When you have enough data points to categorize, Pareto analysis (built on the 80/20 principle that a small number of causes typically account for most occurrences) tells you where a corrective action will have the most leverage.

Cause category (from 9 unapproved access requests sampled)

Count

Cumulative %

Requester used a legacy paper form no longer routed to the approval workflow

5

55.6%

Approval granted verbally, never logged in the ticketing system

2

77.8%

Approver was on leave; request auto-escalated to someone without approval authority

1

88.9%

Genuine one-off process bypass, no systemic pattern

1

100%

Five of nine cases trace to a single cause: a legacy paper form still circulating somewhere in the organization despite the workflow having moved to a ticketing system two years earlier. That's a Pareto-classic result — fixing the top cause (retiring the legacy form and blocking its use) resolves 56% of the sampled failures in one corrective action, versus writing four separate fixes for four separate low-frequency causes. Pareto analysis is what tells a resource-constrained security team where to spend its 90-day corrective action window first, and it's also excellent audit evidence — it shows the auditor you didn't just react to the specific 9 samples they found, you quantified the underlying pattern across your own broader population.

"The question I ask every client staring down a volume nonconformity is: did you fix the nine cases the auditor found, or did you fix the thing that's going to produce case number ten? Pareto data is how you prove it's the second one." — Tomasz Wieczorek, Principal Consultant, Ferro Compliance Advisory

Barrier analysis: for incidents where a control existed but didn't work

Barrier analysis assumes you already had a defense in place — a "barrier" meant to prevent or detect the failure — and asks specifically why that barrier didn't do its job. It's a natural fit for nonconformities discovered because an incident happened despite existing controls, since it maps cleanly onto incident learning (relevant to incident management under Controls 5.24–5.28) as well as pure audit findings.

Intended barrier

Was it present?

Did it function as designed?

What actually happened

DLP rule blocking bulk export of customer records to personal email

Yes, configured

No

Rule only covered attachments over 10MB; export was split into three smaller emails, each under threshold

Manager approval required for bulk data exports

Yes, documented in policy

No

Policy existed but no technical control enforced it — reliant on employee self-reporting

Quarterly access review to catch unnecessary export permissions

Yes

Partially

Review occurred but export permission wasn't in scope of the review template

Security awareness training on data handling

Yes

Yes, in principle

Employee had completed training 14 months earlier; content didn't cover email-based exfiltration specifically

Barrier analysis is valuable precisely because it resists the temptation to conclude "the employee was careless." Every barrier in this table either had a design gap (the DLP threshold), an enforcement gap (policy with no technical backstop), a scope gap (the access review template), or a currency gap (training that didn't cover the actual technique used). None of those are "blame the person" conclusions, and all four point to specific, fixable corrective actions. This is the technique I reach for most often when a nonconformity originates from an actual security incident rather than a routine audit sample, because it naturally produces a barrier-by-barrier action list instead of one vague fix.

Choosing the right technique for the nonconformity in front of you

You don't need to master all five techniques before you're allowed to close a nonconformity — most ISMS managers lean on 5 Whys for 70% or more of findings and reach for the others situationally. Use this table as a quick decision aid the next time a nonconformity lands on your desk.

If the nonconformity looks like...

Reach for...

Why

A single, isolated instance with an apparent linear cause

5 Whys

Fast, requires no special facilitation, usually sufficient

Something that could plausibly stem from several unrelated directions (people, tech, process all plausible)

Fishbone / Ishikawa

Structures a cross-functional brainstorm before you commit to one theory

A near-miss or incident where multiple independent things had to go wrong together

Fault tree analysis

Forces you to address every failure path, not just the most visible one

A pattern across many samples, not a single event

Pareto analysis

Tells you which cause to fix first for maximum coverage

An incident where an existing control failed to do its job

Barrier analysis

Evaluates each defense layer individually instead of assuming one generic failure

A recurring or escalated nonconformity (like Halvorsen's)

5 Whys or fishbone, but go one layer deeper than last time

The first-round RCA almost certainly stopped too early — recurrence is itself evidence of that

It's also entirely legitimate — and often the right move — to combine techniques: run a fishbone session to identify candidate categories, then Pareto the evidence within each category, then 5 Whys the top-ranked branch down to a specific fixable cause. Auditors don't penalize you for layering methods; they penalize you for a conclusion that isn't supported by the evidence you show them.

Scaling RCA rigor to the severity of the nonconformity

Not every finding deserves a two-hour fishbone workshop, and treating every nonconformity with maximum rigor is its own kind of failure — it burns goodwill with the business, trains people to see RCA as bureaucratic theater, and slows down the findings that genuinely need deep investigation. Certification bodies classify nonconformities as minor or major, and that classification is a reasonable proxy for how much RCA effort the finding warrants, though recurrence and potential impact should also weigh in even when the formal classification is minor.

Nonconformity characteristic

Typical classification

Recommended RCA depth

Isolated, first-time, low-impact (e.g., one missing signature on a document)

Minor

Quick 5 Whys, single owner, documented in a few sentences

Isolated but touches a security-relevant control (e.g., one overdue access review)

Minor

5 Whys with at least one other stakeholder consulted; check for similar instances

Pattern across a sample (e.g., 20% of a sampled population fails)

Minor or major, depending on sample size and control criticality

Pareto analysis to quantify the pattern, then 5 Whys or fishbone on the top cause

Repeat finding of any kind

Often escalated to major on the second occurrence

Fishbone or fault tree — the first RCA is now known to have been insufficient

Tied to an actual security incident or near-miss

Frequently major, especially if data or availability was affected

Fault tree or barrier analysis, cross-functional, formally documented

Systemic (affects multiple business units, sites, or control domains)

Major

Fishbone plus Pareto to prioritize, likely requiring ISMS document updates and possibly a scope reassessment

A practical rule I give clients: if you're not sure whether a finding warrants a deeper technique, ask what the cost of being wrong is. A minor nonconformity around a single overdue document review costs you very little if your RCA turns out to be superficial — worst case, it recurs and you do a better RCA next time. A nonconformity that touches access control, incident response, or anything with a plausible path to data exposure costs you a great deal if the RCA is superficial, because the "next time" might be an actual breach rather than another audit finding. Calibrate effort to consequence, not to how the finding happened to be worded in the audit report.

From nonconformity to closed corrective action: the flow

The diagram below lays out the sequence auditors expect to see reflected in your documented information, from the moment a nonconformity is identified through to closure. Each stage maps to a specific Clause 10.2 obligation.

The loop-back arrows are the important part of this diagram, and the part most organizations' internal processes skip. If effectiveness verification fails, the correct response is not to write a new correction and call it done — it's to go back into root cause analysis, because an ineffective corrective action is strong evidence that the cause you identified wasn't the real one, or wasn't the only one.

Turning a root cause into a CAPA record that passes

Once you have a defensible root cause, the corrective action and preventive action (CAPA) record is where you convert analysis into a documented commitment. Auditors are reading this record specifically to check that the stated action addresses the stated cause — not a related cause, not a symptom, the actual identified cause. Use the template structure below as a baseline for your own CAPA record; the column headers are the fields an auditor will look for.

Field

Purpose

Halvorsen example (post-escalation)

Nonconformity reference

Ties the CAPA to a specific audit finding number

Surveillance Audit NC-2026-03

Nonconformity description

The specific finding, in the auditor's or internal audit's own words

2 of 10 sampled departed employees retained active VPN accounts beyond termination date

Correction taken

Immediate fix, dated

Both accounts disabled same day; confirmed via IdP audit log, dated

Root cause (from RCA)

The specific, evidenced cause identified through 5 Whys/fishbone/etc.

Manual, memory-dependent offboarding notification with no automated trigger; control never re-risk-assessed as contractor headcount tripled

Similar nonconformities check

Whether the same cause could affect other areas

Reviewed all systems relying on manual HR-triggered deprovisioning (badge access, SaaS app licenses); found 2 additional systems with the same gap

Corrective action

The specific action(s) that eliminate the root cause

Implement automated HR-to-IdP deprovisioning trigger on contract end date; extend same trigger to badge system and SaaS app provisioning tool

Action owner

Named individual accountable

IT Operations Manager

Target completion date

Realistic date, tied to the nature of the fix

60 days from NC issuance (within the 90-day major NC window)

ISMS document updates

Any policy/procedure changes required

Updated offboarding procedure to reflect automated trigger; updated risk register entry for termination-related access risk

Effectiveness verification method and date

How and when you'll confirm it worked

Sample of all terminations in the 90 days following implementation, reviewed at day 120

Status

Open / in progress / effectiveness pending / closed

Tracked through to closure with dated evidence at each stage

Two rows deserve emphasis because they're the ones organizations skip most often even when the root cause itself is solid. The "similar nonconformities check" row is a direct requirement of Clause 10.2 — the standard explicitly calls for reviewing whether similar nonconformities exist or could occur elsewhere, not just fixing the specific instance found. Halvorsen's second pass caught two additional systems with the identical manual-trigger gap; had they done that check the first time, badge access and SaaS licensing wouldn't have been separate future findings. The "effectiveness verification method and date" row is the one that turns a corrective action from a promise into evidence — and it's the subject of the next section.

"A corrective action without a stated effectiveness verification date is just a to-do item with extra paperwork. I want to see, in the same record, exactly how and when you're going to prove this worked — not a vague 'we'll monitor it.'" — Renata Sowinski, ISO 27001 Lead Auditor, Aldergate Assurance Group

Who should be in the room: RCA roles and responsibilities

A recurring reason RCA sessions produce weak conclusions is that the wrong mix of people is in the room — either too narrow (one analyst guessing at causes outside their visibility) or too broad and unstructured (a dozen people with no clear facilitator, drifting into blame or speculation). The table below reflects the role structure that tends to produce defensible RCA outcomes, adaptable to your organization's size; in a smaller company, one person may reasonably wear two or three of these hats.

Role

Responsibility in the RCA process

Who typically fills it

Facilitator

Runs the session, keeps the group from stopping at the first plausible answer, ensures evidence backs each conclusion

ISMS manager, internal auditor, or an outside consultant for major/sensitive findings

Process owner

Provides ground-truth detail on how the affected process actually works day to day (not how it's documented to work)

The manager or team lead who owns the process where the nonconformity occurred

Technical subject matter expert

Confirms or refutes technical hypotheses with system data (logs, configurations, tickets)

IT operations, engineering, or the relevant control owner

Evidence gatherer / scribe

Pulls the actual data referenced during the session (timestamps, ticket counts, configuration exports) and documents the session

Internal audit support, compliance analyst

Corrective action owner

Accountable for implementing and tracking the resulting action once the root cause is agreed

Named individual, ideally someone with authority to actually make the change, not just report on it

Sponsor / approver

Reviews and signs off on the final RCA and CAPA record, especially for major nonconformities

ISMS manager, CISO, or equivalent — the person accountable to leadership and the certification body

Two staffing mistakes show up constantly. The first is having the same person serve as facilitator and process owner on a finding that touches their own area — it's very hard to rigorously interrogate a process you designed and are also defending, which is exactly why an outside facilitator (internal audit, or an external consultant for major findings) adds real value even in organizations that otherwise run lean. The second is skipping the sponsor/approver step for anything beyond a trivial minor finding — a CAPA record that never had a second set of eyes above the person who wrote it is far more likely to contain the reverse-engineered-conclusion problem auditors are trained to spot.

Effectiveness verification: proving the corrective action actually worked

Clause 10.2 requires the organization to review the effectiveness of any corrective action taken — this is not optional, and it's not satisfied by simply implementing the fix and moving on. Effectiveness verification is a distinct, later step that asks a specific question: after enough time has passed for the new control to be exercised under real conditions, did the nonconformity (or the pattern behind it) actually stop occurring?

The two mistakes I see most often here are checking too early and checking the wrong thing. Checking too early means verifying the day after implementation, which only proves the fix was deployed, not that it works under real operating conditions — a new automated deprovisioning trigger might work perfectly for the one test case you ran and then fail silently on the next edge case (a contractor extension, a rehire, a name change) three weeks later. Checking the wrong thing means confirming the specific instance that triggered the nonconformity is resolved, without checking whether the pattern is resolved — which is really just correction dressed up as effectiveness verification.

Verification approach

What it checks

When to use it

Re-sample the same population

Pull a fresh sample from the same control area the original nonconformity came from (e.g., another batch of terminated employees)

Best default for most access- and process-related corrective actions

Monitor a defined metric over a set window

Track a quantitative indicator (SLA breach rate, overdue training count, failed sync alerts) over 30–90 days

Best for volume-pattern nonconformities identified via Pareto analysis

Targeted re-test of the specific control

Re-run the exact scenario that failed (e.g., simulate a termination end-to-end)

Best for technical/system controls where you can safely simulate the failure condition

Independent internal audit follow-up

A separate internal audit cycle specifically re-checks the control area

Best for major nonconformities or ones tied to a systemic ISMS gap

Management review confirmation

Effectiveness data presented and confirmed at a management review meeting

Appropriate as a closing formality after one of the above methods has produced evidence

A reasonable rule of thumb: give the corrective action at least one full operating cycle before verifying — for something that happens continuously (like terminations), 60–90 days is typical; for something that happens on a fixed schedule (like quarterly access reviews), wait for at least one full quarter to pass. If the corrective action is time-critical because it closes a major nonconformity against a 90-day certification-body deadline, you can present interim evidence (the fix is implemented, initial spot-checks are clean) while committing to a follow-up effectiveness check at the next scheduled internal audit or surveillance visit — auditors generally accept this staged approach as long as it's explicit in the record rather than implied.

Documenting root cause analysis so an auditor can follow your reasoning

Clause 10.2 requires you to retain documented information as evidence of the nature of the nonconformities, the actions taken, and the results. In practice, that means your RCA needs to leave a paper trail an auditor who wasn't in the room can reconstruct and evaluate on their own — not just a one-line conclusion.

Documentation element

Why an auditor looks for it

Common gap

Which technique was used (5 Whys, fishbone, etc.)

Shows a structured method was applied, not a guess

Often omitted entirely — record jumps straight to "root cause: X"

Who participated in the analysis

Cross-functional input increases credibility, especially for fishbone/FTA

Records show one person's name, suggesting no challenge or peer review occurred

Evidence reviewed during the analysis (logs, interviews, tickets, data)

Distinguishes evidence-based conclusions from speculation

Conclusions stated with no supporting artifact referenced

The full chain or diagram, not just the final answer

Lets the auditor see whether the analysis actually reached a system-level cause or stopped early

Only the final "root cause" line is kept; working notes discarded

Date the RCA was performed relative to the nonconformity being raised

Confirms RCA happened as part of the corrective action process, not retrofitted before an audit

RCA dated suspiciously close to the audit date it's meant to satisfy

Link to the resulting corrective action record

Shows traceability from cause to action

RCA and CAPA stored in different systems with no cross-reference

None of this needs to be elaborate. A one-page 5 Whys table with participant names, a date, and a note on what evidence was checked at each "why" is entirely sufficient for the majority of minor nonconformities. Save the heavier documentation (full fishbone diagrams, fault trees) for major nonconformities or recurring findings where the auditor's scrutiny — reasonably — will be higher. The goal is always the same: someone who reads the file six months from now, with no memory of the meeting, should be able to see exactly how you got from the finding to the fix.

Common RCA mistakes that keep nonconformities open (or bring them back)

Across 200-plus ISMS engagements, the same handful of RCA failure patterns account for the overwhelming majority of rejected CAPAs and repeat findings. Worth checking your own draft against this list before you submit anything to an auditor.

Mistake

What it looks like

Why it fails

What to do instead

Stopping at the symptom

"Root cause: the patch wasn't applied"

Restates the finding instead of explaining it

Ask why the patch process allowed this to happen, not just that it did

Blaming the individual

"Root cause: employee error / employee didn't follow procedure"

Rarely survives a second "why" — ask why the system let one person's error cause a nonconformity with no catch

Ask what would need to be true for this error to be structurally impossible or automatically caught

Treating training as a root cause rather than a symptom

"Root cause: insufficient training"

Training gaps are usually themselves caused by something (no onboarding trigger, no competency check)

Ask why the training gap existed and whether a competency verification step, not just more content, is the real fix

Confusing correlation with causation

Assuming the most recent change caused the failure because it's the most visible variable

Can send the corrective action in a completely wrong direction

Verify the causal link with actual evidence (logs, timelines, reproduction) before committing

Single-person RCA on a cross-functional problem

One security analyst writes the RCA alone for a finding that spans HR, IT, and a business unit

Misses causes outside that person's visibility

Pull in the process owners actually involved, even briefly

RCA performed after the corrective action was already decided

The "root cause" conveniently matches whatever fix was already budgeted or planned

Auditors can usually tell when the causal chain was reverse-engineered

Do the analysis first, let the evidence lead to the cause, then design the action

Ignoring the "similar nonconformities" check

Fixing only the exact instance found, not checking adjacent systems/processes for the same gap

Leaves Clause 10.2's explicit requirement unmet and invites a related repeat finding

Always ask "where else could this same cause be operating?" before closing

The individual-blame pattern deserves its own callout because it's both the most common and the most consequential. It's not that individual error never happens — people do make mistakes — but a well-designed ISMS is supposed to be resilient to individual error, and a root cause analysis that ends at "the person made a mistake" has, by definition, failed to examine why the system had no way to catch that mistake. If your corrective action for a training-related finding is "retrain the employee" and nothing else, ask yourself honestly whether the next employee in that role, on their busiest day, would make the same mistake. If the honest answer is yes, you haven't found the root cause yet.

"The moment someone in the room says 'the employee should have known better,' I know we're not done. That sentence is a signpost pointing straight at the actual root cause, which is almost always a process or system gap wearing a person's name." — David Achterberg, VP of Information Security, Loomcraft Insurance Group

Case study one: Halvorsen Freight Analytics — from repeat major NC to closed within the deadline

Back to Priya's situation. With the acquisition-linked deadline pressure, Halvorsen brought in an outside facilitator to run a proper fishbone session within 48 hours of the major nonconformity being raised, pulling in IT operations, HR, and the security team together rather than leaving Priya to write the RCA alone. The session confirmed the manual, memory-dependent notification chain as the dominant branch, and a Pareto pass across two years of offboarding tickets showed 71% of any offboarding-related delay traced back to that same manual handoff, regardless of which system was affected (VPN, badge, or SaaS licensing).

The corrective action built an automated trigger from the HR system's contract-end-date field directly to the identity provider's deprovisioning API, with badge access and the SaaS license management tool wired to the same trigger — addressing the "similar nonconformities" requirement in one build rather than three separate projects. Effectiveness verification was scheduled for day 75 of the 90-day window: a re-sample of every termination processed since go-live, cross-checked against system logs for deprovisioning timestamps. Zero exceptions across 34 terminations in that window. The certification body closed the major nonconformity on schedule, the acquisition's due-diligence data room reflected a clean status, and the $4.2 million valuation premium Priya's CEO had flagged stayed intact. The total cost of the fix — engineering time plus a modest identity-governance module — was $38,000, a number Priya now cites internally every time someone questions whether RCA rigor is "worth the time."

Case study two: Verdant Analytics — the fishbone session that found three causes, not one

Verdant Analytics, a 90-person SaaS company preparing for its first Stage 2 audit, hit a nonconformity during a mock audit dry run: security awareness training completion sat at 68% against a policy commitment of 95% within 30 days of hire. The instinctive first draft of the RCA, written by the compliance lead alone, concluded "employees are too busy and deprioritize non-urgent training" — a classic individual-blame conclusion that the outside auditor prepping them for Stage 2 flatly rejected as insufficient.

A fishbone workshop with People Ops, engineering management, and the compliance lead surfaced three genuine contributing causes instead of one: the training platform's completion reminders were going to a company email alias that new hires didn't have access to until their second week; managers had no visibility into their own reports' completion status, so there was no natural escalation path; and the training assignment itself was triggered manually by HR rather than automatically at account provisioning, causing a lag of up to 12 days for some hires. Verdant's corrective action addressed all three: reminders redirected to personal onboarding email until company access was active, a manager dashboard added to the training platform, and assignment automated at the identity-provisioning step tied to their endpoint and authentication controls. Completion rates hit 97% within the following quarter, verified at the 90-day mark, and the auditor closed the finding during Stage 2 itself rather than requiring a follow-up visit — saving Verdant an estimated $6,500 in additional audit fees and roughly six weeks of schedule risk on their certification target date.

Case study three: Northfield Public Utilities — when 5 Whys wasn't enough

Northfield Public Utilities, a regional infrastructure operator, had a near-miss: a decommissioned test server, believed to be fully offline, was found still reachable on the internal network eight months after its supposed retirement, with an outdated, unpatched OS and no monitoring coverage. A quick 5 Whys by the network team concluded "the decommissioning checklist item for network isolation was skipped" and proposed adding a checklist verification step — a correction masquerading as a corrective action, and one that ignored the fact this was the second time a "decommissioned" asset had turned up live in eighteen months.

Given the recurrence and the potential severity (an unpatched, unmonitored server on an internal utility network), the ISMS manager escalated to fault tree analysis. The FTA revealed an AND-gate structure that 5 Whys had missed entirely: the server stayed reachable AND was invisible to monitoring AND was absent from the asset inventory used for vulnerability management scanning — three independent gaps, each of which alone might have been caught by one of the others, none of which were. Basic events underneath included a decommissioning process that removed assets from the ticketing system before confirming network isolation, an asset inventory that was updated manually and had drifted from reality, and a monitoring tool onboarding process that depended on the same asset inventory being accurate. The corrective action rebuilt asset decommissioning as a single automated workflow — network isolation confirmed via active scan before ticket closure, inventory updated automatically from the scan result, and monitoring de-enrollment triggered from the same event — closing all three gates at once instead of patching one checklist item. Effectiveness verification six months later, using a full network discovery scan reconciled against the asset inventory, found zero orphaned assets, and the internal audit team flagged the case as a model example in Northfield's next management review.

Case

RCA technique used

Root cause found

Outcome

Halvorsen Freight Analytics

Fishbone + Pareto

Manual, memory-dependent HR-to-IT offboarding handoff, never re-assessed as headcount scaled

Major NC closed within 90-day deadline; $4.2M valuation premium preserved

Verdant Analytics

Fishbone

Three independent causes: email routing, manager visibility gap, manual assignment lag

Finding closed during Stage 2 itself; ~$6,500 and 6 weeks saved

Northfield Public Utilities

Fault tree analysis

Three independent gaps (isolation, inventory, monitoring) that each masked the others

Zero orphaned assets on 6-month re-verification; used as internal training example

Where RCA fits into the bigger picture of ongoing certification

Root cause analysis doesn't happen in a vacuum — it's one node in a continuous-improvement loop that runs for as long as your certification does. Nonconformities surface through several different channels, and each one carries slightly different expectations for how fast and how deep your RCA needs to go.

Where the nonconformity comes from

Typical RCA expectation

Notes

Internal audit finding

Full RCA before the next management review cycle

Internal audits are your own early-warning system; use them to practice RCA rigor before an external auditor ever sees the finding

Stage 2 certification audit

RCA and CAPA usually required before the certificate is issued (for major findings) or within an agreed follow-up window (for minor)

The certification body will not issue or maintain certification against unresolved major nonconformities

Surveillance audit

Same 90-day major / next-visit minor convention as Stage 2, but with added scrutiny for anything resembling a repeat finding

This is exactly the channel that escalated Halvorsen's finding from minor to major

Security incident, independent of any audit

RCA typically expected as part of incident closure, feeding into the same CAPA process

Ties directly to the "learning from incidents" intent behind information security incident management

Internal complaint or self-identified gap

No external deadline, but Clause 10.2 still applies once you've classified it as a nonconformity

Don't let the absence of an audit deadline become an excuse to skip the RCA step

Whichever channel surfaces the nonconformity, the resulting root cause and corrective action should also inform whether your Statement of Applicability or risk register needs updating — a root cause that traces back to a control being applied inconsistently, or a risk that was under-assessed, is exactly the kind of signal that should flow back into your next risk review rather than staying siloed in a closed audit finding. If any of the terminology in this article — nonconformity, correction, corrective action, effectiveness — still feels slippery in your own documentation, PentesterWorld's ISO 27001 Glossary of Terms is a fast way to get your team using consistent language, which matters more than it sounds like it should when three different departments are all writing sections of the same CAPA record.

It's also worth knowing that this discipline isn't unique to ISO 27001. If your organization holds or is pursuing SOC 2 alongside ISO 27001, or maps its controls to the NIST Cybersecurity Framework, you'll find the underlying logic of "fix the symptom, then find and fix the cause" shows up under different labels — SOC 2 exception remediation and NIST's Respond/Recover functions both expect essentially the same rigor. A mature RCA process built for ISO 27001 corrective action transfers almost directly to those frameworks, which is a meaningful efficiency if you're managing more than one certification or attestation at once.

Making RCA repeatable instead of reinventing it every time

The organizations that struggle most with RCA aren't the ones lacking technical knowledge of 5 Whys or fishbone diagrams — it's usually a five-minute explanation away. They struggle because RCA isn't a habit yet; every nonconformity triggers a scramble to remember what template to use, who should be in the room, and where the last one was even filed. Building light infrastructure around the process pays for itself the second or third time you use it.

Infrastructure element

What it solves

Minimum viable version

A standing RCA template

Prevents every analyst from inventing their own format, which makes records inconsistent and harder for auditors to follow

A single shared document or spreadsheet tab with the CAPA fields from earlier in this article

A technique quick-reference

Removes the "which method do I even use" hesitation that delays getting started

The choosing-technique table from this article, pinned somewhere your team actually looks

A defined severity-to-rigor mapping

Stops both over-investing in trivial findings and under-investing in serious ones

The severity table from this article, adapted to your own minor/major definitions

A tracking log of open and closed CAPAs

Makes the "similar nonconformities" check possible — you can't check for patterns you haven't logged

Even a basic spreadsheet with nonconformity ID, root cause, status, and effectiveness verification date is sufficient at small scale

A cadence for reviewing open CAPAs

Prevents corrective actions from quietly stalling past their target date

A standing agenda item in your regular security or ISMS meeting, escalating to management review for anything overdue

None of this requires purpose-built GRC software, though a GRC platform does make the tracking and cross-referencing meaningfully easier once you're running more than a handful of nonconformities a year across multiple frameworks. What matters more than the tooling is that the habit exists at all — that the next nonconformity doesn't start from a blank page. PentesterWorld's Internal Audit Report Template is built to capture findings in a format that flows directly into this kind of RCA tracking, and pairing it with the ISO 27001 Mandatory Documents Checklist helps confirm your CAPA records are sitting alongside the other documented information a certification body will expect to see retained.

Root cause analysis as a business advantage, not just an audit requirement

It's tempting to treat RCA purely as an audit-survival exercise — the thing you do because a nonconformity report demands it. That framing undersells what's actually happening when you do it well. Every rigorous root cause analysis is, in effect, free organizational diagnostics: it tells you where a process depends on someone remembering something, where a control's design doesn't match how the business actually operates today, and where two teams have quietly stopped talking to each other about a handoff that used to work. Halvorsen didn't just close a nonconformity — they found and fixed a scaling problem in their offboarding process that would have kept generating security exposure (and eventually a genuine incident, not just an audit finding) regardless of whether an auditor ever caught it. That's the case worth making to a CFO who sees RCA as compliance overhead: it's risk reduction with an audit trail attached, and the common risk assessment mistakes that create nonconformities in the first place are frequently the same gaps RCA surfaces.

The organizations that get the most value from ISO 27001 — the ones for whom certification becomes a genuine competitive differentiator rather than a checkbox — are consistently the ones that treat every nonconformity as a free audit of their own operational assumptions. If your RCA process is mature enough that a customer's security questionnaire, a SOC 2 auditor's request for evidence of continuous improvement, or your own board asking "how do you know this won't happen again" all get answered by the same disciplined CAPA record, you've turned a Clause 10.2 obligation into genuine organizational muscle.

If you're building or refining your own root cause analysis and corrective action process, PentesterWorld's Internal Audit Report Template gives you a structured place to capture findings that feed directly into RCA, and the Internal Audit Checklist helps make sure your internal audits are surfacing the real issues before a certification body finds them for you. If you're heading toward your first certification audit and want to know where nonconformities are most likely to originate, the Certification Readiness Checklist and The Complete ISO 27001 Implementation Guide eBook are both built for exactly that gap-closing work, and the ISO 27001 Glossary of Terms is a quick reference if any of the Clause 10.2 terminology in this article needs a second look.

Whichever technique you reach for next time a nonconformity lands on your desk, the test is always the same one Priya learned the hard way: if the corrective action you write down could plausibly fail to prevent the exact same finding twelve months from now, you haven't found the root cause yet — you've found another correction.


Frequently asked questions

Does ISO 27001 require a specific root cause analysis method like 5 Whys?

No. Clause 10.2 requires that you determine the causes of a nonconformity, but it does not mandate any particular technique. 5 Whys, fishbone diagrams, fault tree analysis, Pareto analysis, and barrier analysis are all practitioner tools borrowed from broader quality and reliability engineering practice, not ISO requirements. Auditors evaluate whether your conclusion is well-supported, not which named method produced it.

What's the actual difference between a correction and a corrective action?

A correction fixes the specific instance of the problem you found — disabling one account, patching one server. A corrective action eliminates the underlying cause so that the same type of problem doesn't recur or occur elsewhere. Clause 10.2 requires both: react and correct, then separately evaluate and act on the cause. A CAPA record that only documents the correction hasn't satisfied the clause.

How long should we spend on root cause analysis before an audit deadline forces our hand?

It scales with severity and recurrence, not with how much time is left on the clock. A straightforward, first-time minor nonconformity might reasonably get a 30-minute 5 Whys session. A major nonconformity, a repeat finding, or anything tied to an actual security incident deserves a facilitated, cross-functional session — often fishbone or fault tree — even if that means requesting a short extension from the certification body, which most bodies will grant if you can show the extra time is going toward a more rigorous analysis rather than delay for its own sake.

Can one person do the root cause analysis alone, or does it need to be a group exercise?

Simple, linear nonconformities are often fine with one knowledgeable person running a 5 Whys, provided they genuinely investigate rather than guess. Anything more complex — where the cause could plausibly span multiple teams or systems — benefits substantially from cross-functional input, because a single person's blind spots become the RCA's blind spots. Fishbone and fault tree analysis are specifically designed as group exercises for this reason.

What happens if our corrective action turns out not to be effective?

You go back into root cause analysis. An ineffective corrective action is itself evidence that either the root cause was misidentified or there was more than one contributing cause. This isn't a failure state to hide from an auditor — documenting an honest "this didn't fully resolve it, here's what we found on the second pass" is generally viewed more favorably than a corrective action that was quietly declared successful without real verification.

How does root cause analysis for ISO 27001 relate to root cause analysis in incident response?

They're the same discipline applied to different triggers. A nonconformity found in an audit and a security incident are both, fundamentally, evidence that something in the system didn't work as intended — the RCA techniques in this article apply equally to both. Many organizations run their incident post-mortems and their nonconformity RCAs through the same process, which also naturally satisfies the "learning from information security incidents" intent behind Control 5.27.

Do we need to document every single "why" or just the final root cause?

Document the chain, not just the endpoint. Auditors are assessing whether you actually investigated or just asserted a conclusion, and a bare final answer with no supporting reasoning looks identical whether it came from a rigorous session or a five-minute guess. Keep the working notes — who was involved, what evidence was checked at each step — as part of your retained documented information.

Is "human error" ever an acceptable root cause?

Rarely as a final answer. Human error is almost always a symptom of a system that lacked a control to prevent, catch, or make the error harmless — insufficient automation, no verification step, unclear procedure, excessive workload. Ask why the system allowed one person's mistake to become a nonconformity with no backstop; that question usually leads to the actual root cause within one or two more iterations.

Will a certification auditor accept a corrective action that's still "in progress" at the time of the audit?

Generally yes, provided the CAPA record shows a clearly defined root cause, a specific action plan with an owner and target date, and — for major nonconformities — a realistic path to closure within the typical 90-day window. What auditors won't accept is a vague commitment with no root cause behind it; the analysis needs to be complete even if the implementation is still underway.

13

About the author

Cybersecurity Expert

Satish Kumar writes about cybersecurity, offensive security, and practical defense strategies on PentesterWorld.

Related Articles

Comments (0)

No comments yet. Be the first to share your thoughts!