The Saturday That Cost Solandra Logistics $4.2 Million
Derek Voss got the call at 2:47 a.m. on a Saturday in March. He was VP of Infrastructure and Security at Solandra Logistics, a regional freight brokerage moving about $310 million a year in contracted trucking capacity across nine states. The on-call network engineer's voice had the particular flatness of someone reading a screen full of bad news out loud: every file server at the Ohio data center was throwing ransom notes. So was the backup repository.
That second sentence was the one that mattered. Solandra's backup target was a set of network-attached storage appliances that lived on the same flat network as production, reachable with the same domain credentials the attackers had already harvested. The ransomware operator had waited eleven days after the initial phishing compromise, quietly mapping the environment, before pulling the trigger — and when it pulled the trigger, it hit the backups first. Derek's team had a disaster recovery runbook. It had never accounted for the backups being gone too.
What happened next is the part that turns an expensive incident into a genuinely bad one. To get freight-matching systems back online before Monday's dispatch cycle, the recovery team rebuilt a domain controller from an eight-month-old golden image, skipped the normal hardening checklist "just for now," and — because the identity provider integration for MFA wasn't cooperating on the rebuilt box — disabled multi-factor authentication for administrative accounts "temporarily." Endpoint detection and response agents on the rebuilt servers were left in a passive install-but-don't-enforce mode because the team was worried about false positives interfering with the recovery. Nobody wrote any of this down as a formal risk acceptance. It was just what got done, fast, by exhausted people, at 5 a.m.
Ninety-six hours later, someone used a dormant helpdesk account — never rotated, never re-verified — to log back into the rebuilt environment and exfiltrate roughly 40,000 customer shipment manifests before anyone noticed the second intrusion. By the time forensics closed the loop, Solandra was staring at a $1.1 million ransom payment its insurer partially covered, $1.6 million in lost freight contracts from carriers who moved volume elsewhere during the outage, $900,000 in legal, forensic, and regulatory notification costs, and roughly $600,000 in incident response and rebuild labor. Call it $4.2 million, and that's before reputational damage that doesn't show up on a single invoice.
Solandra had been mid-certification for ISO 27001 when this happened. Its Statement of Applicability listed both Control 5.29 (information security during disruption) and Control 5.30 (ICT readiness for business continuity) as applicable, with tidy references to a disaster recovery plan sitting in a document management folder. Nobody had ever tested it. Nobody had defined, in writing, which security controls were allowed to be relaxed during recovery and under what compensating conditions — which is exactly what left Derek's team improvising security decisions under duress, with predictable results.
I got called in two weeks after the dust settled, initially to help with the insurer's post-incident report, and ended up rebuilding Solandra's entire business continuity and ICT readiness program from the ground up. What follows in this article is that program — how 5.29 and 5.30 actually work, how they differ from (and connect to) a full disaster recovery capability, and how to produce evidence an ISO 27001 auditor — and your own leadership team, the next time something goes wrong at 2:47 a.m. — will actually accept.
Who This Is For
This article is written for CISOs, IT directors, business continuity managers, and ISO 27001 project leads who are responsible for Controls 5.29 and 5.30 in their Statement of Applicability and need to move from a document nobody has read to a program that survives contact with a real incident. You should walk away knowing exactly what evidence auditors expect for each control, how to run a defensible Business Impact Analysis that produces real RTOs and RPOs instead of guessed numbers, how to keep confidentiality, integrity, and availability controls intact while your organization is in crisis mode rather than quietly switching them off, and how 5.29/5.30 relate to — and fall meaningfully short of — a full ISO 22301 business continuity management system. If you've been treating "we have a DR plan" as equivalent to satisfying these two controls, this is the article that closes that gap before an auditor — or an attacker — finds it for you.
"The mistake I see constantly is treating 5.30 as an IT problem to solve with a runbook. It's a business problem. If the business hasn't told you which processes matter and by when they need to be back, IT is just guessing at recovery targets — and guesses don't hold up under an actual outage, or under audit." — Marcus Ilić, Head of IT Resilience, NordBridge Financial
Why ISO 27001 Has Two Separate Continuity Controls
It's worth pausing on why the 2022 revision of ISO/IEC 27001 treats continuity as two distinct controls rather than one. Control 5.29, information security during disruption, existed conceptually in the 2013 version of the standard (as part of the old A.17 business continuity clause) and asks a narrower question: when something goes wrong — a cyberattack, a flood, a pandemic, a prolonged power outage — does information security keep functioning at a level appropriate to the situation, or does it collapse the moment the crisis team takes over? Control 5.30, ICT readiness for business continuity, is new in the 2022 revision and asks a different, more specific question: is your ICT infrastructure actually capable of supporting the business's recovery time objectives, and have you proven that capability through testing rather than assumed it?
That distinction matters because organizations routinely nail one and fail the other. Solandra had an ICT recovery capability of sorts — servers could, eventually, be rebuilt — but it had no plan at all for maintaining information security discipline during the process, which is precisely how a ransomware incident turned into a data breach. Other organizations run the opposite failure mode: rigid security policies stay bolted down during a crisis, access reviews and change control gates remain fully enforced, and the business misses its recovery window because nobody was authorized to make a fast, documented, risk-accepted exception. Both controls are needed, and they pull in complementary directions: 5.30 is about capability and speed; 5.29 is about discipline and control even when everyone wants to move fast and skip steps.
For readers who want the fuller map of how these sit alongside the rest of the organizational control set, the ISO 27001 Annex A Organizational Controls overview covers all 37 controls in the 5.x family, and if you haven't yet mapped how 2022 changed continuity requirements specifically, that's covered in our piece on what changed between ISO 27001:2013 and ISO 27001:2022.
Control 5.29: Information Security During Disruption — What the Auditor Wants to See
Control 5.29 requires the organization to plan how it will maintain information security at an appropriate level during disruption. The operative phrase is "appropriate level" — ISO 27002's guidance doesn't demand that every control remain fully enforced come what may (that's often not realistic during a genuine crisis), but it does demand that any relaxation of controls be deliberate, documented, risk-assessed, and time-boxed, rather than accidental and permanent, which is exactly what went wrong at Solandra.
Requirement Element | What "Good" Looks Like | Common Evidence Auditors Accept |
|---|---|---|
Continuity plans reference information security | Business continuity / disaster recovery plans explicitly name which security controls (access control, logging, encryption, authentication) must be maintained during disruption | Plan documents with a dedicated "security during disruption" section, not a silent gap |
Defined fallback/degraded operating modes | Written criteria for which controls may be temporarily relaxed, by whom, and for how long, with compensating controls specified | Approved degraded-mode procedures, delegation-of-authority matrix, time-boxed exception log |
Roles assigned for crisis-mode security decisions | A named person (not "IT," a person) with authority to approve security exceptions during an incident | RACI chart, crisis management team charter naming a security decision-maker |
Integration with incident management | 5.29 procedures link explicitly to the incident response process rather than existing as a separate, disconnected document | Cross-references between BC/DR plans and the incident management plan under Controls 5.24–5.28 |
Post-disruption review | Security posture is formally reassessed and restored to normal operating levels after the disruption ends, with any exceptions closed out | Post-incident review report, exception closure sign-off, lessons-learned log |
Awareness and training | Staff involved in crisis response know the degraded-mode security rules before an incident, not during one | Training attendance records, tabletop exercise notes referencing 5.29 procedures |
The single biggest audit finding I see against 5.29 is a complete absence of documentation addressing what happens to security controls specifically during a disruption — the plan talks about restoring servers and notifying customers, but says nothing about authentication, logging continuity, or who's allowed to grant a temporary access exception. That silence is what an auditor will flag as a gap, and it's exactly the silence that let Solandra's recovery team make ad hoc security decisions with no accountability trail.
Control 5.30: ICT Readiness for Business Continuity — What the Auditor Wants to See
Control 5.30 is the newer, more technically demanding of the pair. It requires that ICT continuity be planned, implemented, maintained, and tested based on business continuity objectives and ICT continuity requirements — meaning the IT recovery capability has to be derived from what the business actually needs (recovery time objectives and recovery point objectives coming out of a Business Impact Analysis), not from whatever the infrastructure happens to be capable of.
Requirement Element | What "Good" Looks Like | Common Evidence Auditors Accept |
|---|---|---|
ICT continuity requirements derived from BIA | RTOs/RPOs for each critical system are documented and traceable to a Business Impact Analysis, not invented by IT in isolation | BIA report, RTO/RPO register mapped to systems and business processes |
ICT continuity strategy documented | A defined approach — backup, redundancy, failover, alternate site, cloud failover — matched to each system's criticality tier | ICT continuity/DR plan naming specific technical strategies per tier |
Plans are implemented, not just written | Technical capability (replicated infrastructure, tested backup restores, standby capacity) actually exists and matches the documented plan | Architecture diagrams, infrastructure inventories, backup configuration exports |
Regular testing against objectives | Recovery capability is tested at a defined cadence and measured against the stated RTO/RPO, with results recorded | Test/exercise reports showing actual recovery time vs. target, pass/fail outcomes |
Maintenance and change alignment | ICT continuity plans are updated when systems, architecture, or business criticality change | Change management records showing continuity plan updates tied to Control 8.32 changes |
Management review | Senior management reviews test results and residual risk, and approves any gaps between actual and target recovery capability | Management review minutes, risk acceptance sign-off for RTO/RPO shortfalls |
Notice how much of 5.30's evidence chain depends on capabilities documented elsewhere in Annex A. Recovery point objectives are meaningless without a working information backup regime (Control 8.13), and failover strategies are meaningless without actual redundant processing capacity (Control 8.14) — both of which we'll come back to. This is one of the clearest examples in the whole standard of a control that cannot be satisfied by policy alone; an auditor who's done this before will ask to see a real test report with a real number in it, not a narrative paragraph asserting that recovery "would work."
"5.30 is the control that separates organizations that have thought about disaster recovery from organizations that have practiced it. I've sat across from clients with a beautifully worded DR plan and a four-hour RTO on paper, and when I asked for the last test report, there wasn't one. That's an automatic major nonconformity in my experience — you can't claim readiness you've never demonstrated." — Lena Popova, ISO 27001 Internal Auditor
BC, DR, ICT Readiness, and ISO 22301: Getting the Vocabulary Straight
Half the confusion I clean up in client engagements isn't technical — it's terminology. "Business continuity," "disaster recovery," and "ICT readiness" get used interchangeably in meetings, and then everyone is surprised when the ISO 27001 auditor asks a precise question about a precise term. It's worth being pedantic here because the precision is what makes evidence defensible.
Business continuity (BC) is the organization-wide discipline of keeping critical business processes running — or restoring them within acceptable timeframes — during and after disruption, covering people, facilities, suppliers, and technology together. Disaster recovery (DR) is the narrower technical discipline of restoring IT systems and data after an incident; it's a subset of BC focused specifically on infrastructure. ICT readiness for business continuity, as ISO 27001 Control 5.30 frames it, sits between the two: it's the assurance that the ICT/DR capability is actually sized, planned, and tested to meet the recovery objectives the business continuity process has defined — the connective tissue between "the business needs this back in four hours" and "the infrastructure can actually deliver that."
ISO 22301 is a different animal altogether: it's the international standard for a full Business Continuity Management System (BCMS), with its own certification, its own clause structure, and requirements covering facilities, staffing, supply chain continuity, crisis communications, and organizational resilience far beyond information security. ISO 27001's 5.29 and 5.30 are not a substitute for ISO 22301 — they are the information-security-relevant slice of continuity that an ISMS needs to cover. An organization can be fully ISO 27001 certified with excellent 5.29/5.30 evidence and still have no formal BCMS at all; conversely, an organization with full ISO 22301 certification will usually satisfy most of 5.29/5.30's intent as a byproduct, though an auditor will still expect to see the information-security-specific detail (which controls stay active, how ICT recovery objectives were derived) called out explicitly.
Dimension | Business Continuity (BC) | Disaster Recovery (DR) | ICT Readiness (ISO 27001 5.30) | ISO 22301 BCMS |
|---|---|---|---|---|
Scope | Whole organization: people, process, facilities, suppliers, tech | IT systems and data specifically | ICT capability aligned to BC objectives | Whole organization, formally certifiable management system |
Primary question | "Can the business keep operating?" | "Can we restore the servers and data?" | "Is ICT actually capable of meeting recovery objectives?" | "Is there a governed, auditable system managing all of this?" |
Owner | Business continuity manager / COO | IT / infrastructure team | IT + information security jointly | Executive management, dedicated BCM function |
ISO 27001 relevance | Referenced context, not itself certified | Referenced context, not itself certified | Directly certifiable under Annex A 5.30 | Separate standard, separate certification |
Typical artifact | Business Continuity Plan (BCP) | Disaster Recovery Plan (DRP) | ICT Continuity Plan tied to BIA/RTO/RPO | Full BCMS documentation set, risk assessment, exercise program |
Does ISO 27001 require it? | Referenced via 5.29/5.30 but not mandated in full | Referenced via 5.30 but not mandated in full | Yes — 5.30 mandates ICT continuity be planned and tested | No — ISO 22301 is entirely optional/separate |
If your organization is regulated under frameworks that expect enterprise-wide resilience — financial services under DORA is the clearest current example — 5.29/5.30 alone won't be enough, and it's worth reading our companion comparison, ISO 27001 vs ISO 22301: Business Continuity, before deciding whether a full BCMS is warranted. That title isn't published yet in this series, so for now, treat this article's comparison table above as the working reference.
The Business Impact Analysis: Where RTOs and RPOs Actually Come From
Everything in Control 5.30 depends on one upstream artifact: the Business Impact Analysis (BIA). A BIA is the structured exercise of identifying critical business processes, determining the impact of their disruption over time, and deriving the recovery targets that ICT continuity must be built to meet. Skip the BIA, and every RTO/RPO number in your DR plan is a guess dressed up as a requirement — which is exactly the kind of guess that collapses under a real incident, or under an auditor's first follow-up question ("how was this four-hour figure determined?").
Three terms do the heavy lifting here, and precise definitions matter because auditors — and your own incident responders — will use them literally:
Recovery Time Objective (RTO) is the maximum acceptable length of time a system or process can be unavailable before the impact becomes unacceptable to the business. If order processing has an RTO of four hours, the business has decided that being down for four hours and one minute crosses from "manageable" into "unacceptable."
Recovery Point Objective (RPO) is the maximum acceptable amount of data loss, measured in time, between the last good backup or replication point and the moment of failure. An RPO of fifteen minutes means the business accepts losing at most fifteen minutes of transactions — anything more is unacceptable.
Maximum Tolerable Period of Disruption (MTPD), sometimes called Maximum Acceptable Outage (MAO), is the absolute outer limit beyond which the disruption threatens organizational survival — regulatory license, contractual default, or existential financial damage. RTO should always be set comfortably inside MTPD, never equal to it; if your RTO and MTPD are the same number, you have zero margin for a recovery that runs even slightly long.
Business Process | Criticality Tier | MTPD | RTO (Target) | RPO (Target) | Basis for Figure |
|---|---|---|---|---|---|
Customer order processing / payment capture | Tier 1 — Critical | 6 hours | 2 hours | 15 minutes | Contractual SLA penalties begin at 3 hours; revenue loss ~$40K/hour |
Freight/shipment tracking platform | Tier 1 — Critical | 8 hours | 4 hours | 30 minutes | Carrier contract terms allow rebooking after 6 hours of unavailability |
Customer relationship management (CRM) | Tier 2 — High | 24 hours | 8 hours | 4 hours | Sales operations degrade but do not halt; workaround via spreadsheets exists |
HR / payroll system | Tier 2 — High | 48 hours | 24 hours | 24 hours | Payroll cycle is biweekly; short outages absorbed without missed pay |
Internal document management / intranet | Tier 3 — Standard | 5 days | 72 hours | 24 hours | No direct revenue or contractual impact; productivity loss only |
Marketing website (informational only) | Tier 3 — Standard | 10 days | 5 days | 7 days | Reputational, not operational, impact |
Building this table is the actual work of a BIA — interviewing process owners, quantifying financial and contractual impact at escalating time intervals, and forcing a genuine prioritization conversation rather than declaring everything "critical" (which is the single most common BIA failure I see; when everything is Tier 1, nothing is, and your recovery budget gets spread so thin that the truly critical systems don't get the redundancy they need). The BIA is also where 5.30 connects back into the ISMS risk process generally — the same likelihood-and-impact thinking that drives your risk register under Clause 6 risk assessment and treatment should be reused here rather than reinvented, and BIA findings should feed directly into that risk register as identified continuity risks requiring treatment.
"Boards love a headline like 'we can recover in four hours.' What they need to understand is that the four hours is a business decision with a price tag attached — tighter RTOs cost more in redundant infrastructure and standby capacity. The BIA is where you have that conversation honestly, before an outage forces it." — Aisha Kone, Business Continuity Manager, Solent Marine Insurance
How 5.29, 5.30, Backup, and Redundancy Interlock
The BIA produces targets. The next question is which technical strategy actually meets them, and this is where Control 5.30 leans directly on two other Annex A technological controls: information backup and redundancy of information processing facilities. It's important to log these by their correct numbers rather than treat "backup" and "DR" as generic ideas — an auditor mapping your Statement of Applicability will expect to see the cross-references.
Information backup (Control 8.13) governs whether copies of data actually exist, are tested for restorability, and are protected from the same failure that destroyed the original — Solandra's core failure was a backup architecture that violated this last point by sitting reachable on the same compromised network. Redundancy of information processing facilities (Control 8.14) governs whether there is sufficient duplicate processing capacity — a secondary data center, a failover cloud region, standby hardware — to keep services running or to resume them quickly, independent of backups. Control 5.30 is the layer that ties both of these technical capabilities to the business-derived RTO/RPO from the BIA and mandates that the whole chain be tested, not just built.
flowchart TD
A[Business Impact Analysis - BIA] --> B[Set RTO / RPO / MTPD per critical process]
B --> C[ICT Continuity Requirements - Control 5.30]
C --> D1[Information Backup - Control 8.13]
C --> D2[Redundancy of Processing Facilities - Control 8.14]
D1 --> E[Documented ICT Continuity Plan]
D2 --> E
E --> F[Test: Tabletop / Simulation / Full Failover]
F --> G{RTO / RPO Met?}
G -- No --> H[Remediate Strategy or Escalate Risk Acceptance]
H --> C
G -- Yes --> I[Maintain Information Security During Disruption - Control 5.29]
I --> J[Incident Management 5.24-5.28 Captures Lessons Learned]
J --> AThe diagram's loop back to the BIA is deliberate: continuity is not a project you finish, it's a cycle that re-runs every time business priorities, systems, or threat landscape shift materially — typically reviewed at least annually and after any significant change.
Strategy | What It Provides | Typical RTO Achievable | Typical RPO Achievable | Relative Cost | Best Suited For |
|---|---|---|---|---|---|
Cold backup, offsite/offline (immutable or air-gapped) | Data recoverability after total environment loss, including ransomware | Days | 24 hours or more | Low | Tier 3 systems, ransomware-resilient baseline for all tiers |
Warm backup with periodic replication | Faster restore than cold backup; some infrastructure pre-staged | Hours | 1–4 hours | Medium | Tier 2 systems |
Hot standby / active-passive redundancy (Control 8.14) | Duplicate infrastructure ready to take over with manual or semi-automated failover | Under 1 hour | Minutes | High | Tier 1 systems with strict RTO |
Active-active multi-site / cloud region failover | Continuous availability, near-zero downtime, automated failover | Seconds to minutes | Near-zero | Very high | Tier 1 systems where downtime is contractually catastrophic |
Manual workaround / degraded operating procedure | No technology recovery — business continues via paper/manual process | Immediate (but limited capacity) | N/A | Low | Bridging gap during any tier's recovery window, not a replacement for it |
A pattern I push clients toward: don't buy the same continuity strategy for every system. The instinct after a bad incident (like Solandra's) is often to over-correct and demand active-active redundancy everywhere, which is both unaffordable and unnecessary — a Tier 3 internal wiki does not need the same investment as a Tier 1 payment system. The BIA-driven tiering table from the previous section exists precisely to stop that overcorrection and target spend where the business impact actually justifies it.
Maintaining Information Security During Disruption: Where Control 5.29 Actually Bites
This is the control most organizations get wrong not through ignorance but through good intentions under pressure. When a crisis hits, the instinct of every capable technical team is to do whatever gets the business back online fastest — and security controls, which feel like friction, are often the first thing quietly bypassed. Control 5.29 exists specifically to prevent that instinct from operating unsupervised.
Security Area | Risk If Abandoned During Disruption | How to Maintain It (5.29 Approach) |
|---|---|---|
Authentication and MFA | Rebuilt systems deployed without MFA "temporarily," creating an open door exactly when attackers are watching | Pre-approved emergency-build images with MFA enforced by default; no manual disable without named sign-off and time limit |
Access control / least privilege | Broad "just give everyone admin so we can move fast" access grants during recovery | Pre-defined emergency access roles scoped to the recovery task, auto-expiring, logged |
Logging and monitoring | Rebuilt infrastructure brought online with logging/EDR agents not yet configured or in passive mode | Logging and monitoring agents baked into recovery build templates, verified before systems are reconnected to the network |
Credential hygiene | Reused or unrotated credentials on rebuilt systems, including dormant service/helpdesk accounts | Mandatory credential rotation as a required step in the recovery runbook, not an optional afterthought |
Change management | Emergency changes made with no record, bypassing normal approval entirely | Expedited but still-logged emergency change process, retroactively reviewed within a defined window |
Data classification and handling | Sensitive data copied to unmanaged locations (personal devices, unencrypted shares) to "just get it working" | Emergency procedures that specify approved temporary storage/handling that still meets classification requirements |
Physical and personnel security | Facility access controls relaxed for contractors/vendors brought in urgently during recovery | Visitor/contractor vetting and escort procedures retained even in crisis mode, with expedited but not eliminated checks |
The pattern across every row is the same: 5.29 doesn't ask you to never relax a control during a crisis — sometimes that's genuinely the right call — it asks you to decide that in advance, assign someone the authority to approve it, log it, and reverse it afterward. Solandra's recovery team disabled MFA and let EDR run passively with nobody's name attached to that decision and no plan to reverse it before reconnecting fully to the network. That absence of ownership, not the technical decision itself, is what an ISO 27001 auditor treats as the nonconformity, and it's what let a ransomware incident escalate into a second intrusion.
"The plans that hold up in a real crisis are the boring ones — the ones that already answered 'who is allowed to turn off MFA and for how long' months before anyone needed the answer. The exciting, improvised version of that decision is always the expensive one." — Grace Whitfield, ISO 27001 Lead Implementer, Correlate Consulting
This is also where 5.29 connects directly to your incident response chain. The incident management process under Controls 5.24–5.28 is what activates the crisis, assesses severity, and coordinates response — 5.29's degraded-mode security procedures should be a named, referenced appendix or companion process to that incident plan, not a standalone document nobody connects to the actual moment it's needed.
Testing and Exercises: The Difference Between a Plan and Readiness
A continuity plan that has never been tested is a hypothesis, not a capability. ISO 27002's guidance for Control 5.30 is explicit that ICT continuity has to be tested, and auditors increasingly treat "we test annually" as table stakes rather than a differentiator — the real question is whether the test type matches the maturity of the system it's validating.
Exercise Type | What It Validates | Typical Cadence | Effort/Disruption | Evidence Produced |
|---|---|---|---|---|
Plan review / document walkthrough | Plan is current, contact details correct, roles assigned | Quarterly | Very low | Reviewed/updated plan document, sign-off log |
Tabletop exercise | Team understands roles, decision points, and escalation paths under a simulated scenario | Semi-annually | Low | Facilitation notes, participant list, gaps/actions log |
Simulation / functional exercise | Specific technical procedure (e.g., backup restore, failover script) works as documented, in isolation | Quarterly to semi-annually per critical system | Medium | Test log with start/end times, measured RTO/RPO vs. target, defects found |
Full interruption / failover test | Entire system or site fails over under conditions resembling a real event, with production or production-equivalent load | Annually for Tier 1 systems | High | Full test report, actual recovery time achieved, business sign-off, remediation plan for gaps |
Unannounced/red-team-style continuity test | Response readiness without advance preparation bias | Annually or biennially for highest-criticality systems | High | After-action report, comparison of unannounced vs. announced test performance |
A mature program uses all five, layered by criticality tier — you would not run a full unannounced failover test on your internal wiki, and you should not settle for a tabletop walkthrough as the only evidence for your payment processing platform's four-hour RTO. The evidence auditors want most is the gap between target and actual: a test report showing you aimed for a 4-hour RTO and achieved 3 hours 42 minutes is a stronger artifact than a report that simply says "test successful," because it demonstrates measurement, not assertion.
Test/Exercise Element | Weak Evidence (Avoid) | Strong Evidence (Auditor-Ready) |
|---|---|---|
Scope | "DR test completed" | Named systems tested, scenario description, start/end timestamps |
Result | "Test passed" | Measured recovery time vs. RTO target, measured data loss vs. RPO target |
Participants | Not recorded | Named participants, roles, sign-off |
Findings | Not recorded, or verbal only | Written gaps/defects log with owners and remediation deadlines |
Follow-up | None | Evidence that prior test's findings were remediated before the next test |
Management involvement | IT-only sign-off | Business process owner and management review/approval of results |
Roles and Responsibilities in the Continuity Program
Neither 5.29 nor 5.30 will hold up if ownership is ambiguous, and ambiguous ownership is the single most common root cause I find when a continuity program looks good on paper but fails during a real event — everyone assumed someone else had picked up a task.
Role | Primary Responsibility | Typical Owner |
|---|---|---|
Business Continuity Sponsor | Overall accountability for continuity readiness; approves RTO/RPO trade-offs and residual risk | CEO/COO or designated executive |
Business Impact Analysis Lead | Runs the BIA process, interviews process owners, produces criticality tiers | Business continuity manager or risk manager |
ICT Continuity Owner | Designs and maintains the technical recovery strategy per Control 5.30 | IT/infrastructure director |
Information Security Lead (5.29) | Defines degraded-mode security procedures, approves emergency exceptions during live incidents | CISO or information security manager |
Incident Commander | Coordinates the live response, activates the continuity plan, declares recovery complete | Designated on-call incident commander (rotating) |
Process/System Owners | Validate that recovery meets their process's actual needs; participate in and sign off on tests | Named business unit leads per system |
Internal Audit | Independently verifies test evidence and plan currency against ISO 27001 requirements | Internal audit function |
Assigning these roles formally — with names, not just job titles, in the RACI — is itself evidence an auditor will ask to see, and it directly supports the broader control set on roles and responsibilities in information security that governs how accountability is assigned across the ISMS generally. If your organization already has a mature asset management program under Controls 5.9–5.14, that asset inventory is exactly the input the BIA lead needs to identify which systems exist and who owns them — building the BIA from scratch without that inventory is one of the slower, more error-prone ways to do this work.
Evidence and Mandatory Documentation Auditors Expect
Document/Record | Purpose | Typically Reviewed By Auditor? |
|---|---|---|
Business Impact Analysis report | Establishes criticality tiers and derives RTO/RPO/MTPD | Yes — foundational evidence for 5.30 |
ICT Continuity Plan | Documents technical recovery strategy per system/tier | Yes |
Information Security During Disruption procedure (5.29) | Documents degraded-mode security rules, approval authority, exception handling | Yes |
Test/exercise reports (all types) | Demonstrates recovery capability was actually validated | Yes — often the most scrutinized artifact |
RTO/RPO register | Traceable mapping of targets to systems and their BIA justification | Yes |
Risk register entries for continuity risk | Shows continuity gaps are tracked as managed risk, not ignored | Yes, cross-checked against Clause 6 risk process |
Emergency access/exception logs | Shows 5.29 exceptions were authorized, time-boxed, and closed out | Yes, especially post-incident |
Management review minutes covering continuity | Shows senior oversight of test results and residual gaps | Yes |
Statement of Applicability entries for 5.29/5.30 | Confirms applicability rationale and control implementation status | Yes — first document requested |
The Statement of Applicability is where this whole program gets formally declared, and it's usually the first document an auditor pulls before asking to see the underlying BIA and test reports behind it — so the SoA entry for 5.29/5.30 should never just say "implemented," it should reference the specific plan and test evidence by name and date.
Common Mistakes I See Implementing 5.29 and 5.30
Mistake | Why It Happens | Consequence | Fix |
|---|---|---|---|
Treating "we have a DR plan" as satisfying 5.30 | DR and ICT readiness get conflated as the same thing | Plan exists but was never traced to actual business RTO/RPO or tested against them | Run a real BIA; explicitly document the traceability from business objective to technical strategy |
No BIA, RTOs picked by IT intuition | BIA feels like a slow, bureaucratic exercise; IT just wants a number to build against | Recovery targets don't match what the business actually needs, discovered only during a real outage | Run the BIA properly, even a lightweight version, before setting any RTO/RPO |
Backups stored reachable from production | Convenience, cost, legacy architecture | Ransomware destroys backups along with production, as at Solandra | Immutable/offline/air-gapped copies per Control 8.13, isolated credentials |
Security controls silently disabled during recovery with no record | Pressure to move fast; nobody assigned to think about security during the crisis | Undocumented exposure persists after recovery, sometimes indefinitely | Pre-approved degraded-mode procedures under 5.29 with named approval authority and auto-expiry |
Testing only via tabletop, never technically | Full failover tests are disruptive and expensive to run | Untested technical assumptions fail exactly when needed | Layer test types by tier; require at least one technical test annually for Tier 1 systems |
Plans never updated after architecture changes | Continuity plan maintenance isn't tied to change management | Plan describes infrastructure that no longer exists | Link continuity plan review to the change management process so infrastructure changes trigger a plan review |
Everything declared "Tier 1 critical" | Political reluctance to deprioritize any business unit's system | Redundancy budget spread too thin to protect what's truly critical | Force a genuine ranking exercise; tie tiers to quantified financial/contractual impact |
No link between 5.29/5.30 and supplier-hosted systems | Continuity planning assumes only in-house infrastructure | Outsourced/SaaS systems have no verified recovery capability at all | Extend BIA and continuity requirements to supplier-hosted services per supplier relationship security controls 5.19–5.23 |
Continuity program owned entirely by IT | Perceived as a technology problem | Business impact and priorities never properly represented; IT optimizes for the wrong things | Assign a named business continuity sponsor at executive level, not just an IT lead |
That second-to-last row deserves emphasis: a growing share of the systems in most BIAs today are not run on infrastructure the organization controls at all — they're SaaS platforms, cloud-hosted services, or outsourced processing. Control 5.30 still applies to those dependencies, but your leverage is different: you're validating a supplier's continuity commitments contractually and through due diligence rather than building the redundancy yourself. That's exactly the terrain covered by supplier security controls 5.19–5.23, and any organization with significant SaaS dependency should treat that control family as a direct extension of its continuity program, not a separate exercise.
How 5.29/5.30 Map to Other Frameworks
Organizations rarely implement continuity controls for ISO 27001 alone — most are also answering to SOC 2, NIST CSF, or sector-specific resilience regulation, and it's worth understanding where the requirements overlap and where they diverge so you're not duplicating evidence-gathering across frameworks.
Framework | Relevant Requirement Area | Relationship to ISO 27001 5.29/5.30 |
|---|---|---|
SOC 2 | Availability criteria (Trust Services Criteria) | Requires evidence of system availability commitments and incident recovery; overlaps closely with 5.30's testing evidence, though SOC 2's Availability criteria are narrower on the "maintain security during disruption" dimension that 5.29 covers |
NIST Cybersecurity Framework | Recover function (RC) | Conceptually parallel to 5.30 — recovery planning, improvements, and communications — though NIST CSF's Recover function doesn't separately mandate the 5.29-style "keep security controls intact during disruption" requirement as explicitly |
DORA (Digital Operational Resilience Act) | ICT operational resilience testing, business continuity policies | Considerably more prescriptive than ISO 27001 for in-scope financial entities — mandatory testing frequency, threat-led penetration testing for some entities, and regulatory reporting that 5.29/5.30 alone will not satisfy |
ISO 22301 | Full BCMS | Broader scope than 5.29/5.30 across the whole organization; satisfying ISO 22301 well typically covers 5.29/5.30's intent, but not vice versa |
For organizations selling into regulated financial services or comparing framework overlap generally, our comparison of ISO 27001 against NIST, SOC 2, and PCI DSS is the right companion reference for scoping which evidence can be reused across audits versus which needs framework-specific work. The practical takeaway for continuity specifically: build your BIA, RTO/RPO register, and test reports once, in enough technical depth that they satisfy ISO 27001 5.30, and they'll cover roughly 70–80% of what SOC 2 Availability and NIST CSF Recover ask for too — the remaining gap is almost always DORA-style prescriptive testing cadence and regulatory reporting, which needs to be layered on top rather than assumed.
Case Study: The Test That Worked — Bramwell Trust Bank
Bramwell Trust Bank, a community bank with about $2.1 billion in assets under management, had spent eighteen months building its ISO 27001 continuity program after a near-miss the year before — a fiber cut that took its primary data center offline for ninety minutes and nearly breached its four-hour RTO for core banking services. In response, Bramwell's technology leadership built a proper BIA, formally tiered its systems, invested in an active-passive redundant site (Control 8.14) for its Tier 1 core banking platform, and — critically — committed to a genuine quarterly failover test schedule rather than an annual paper walkthrough.
That investment paid off eleven months later when a regional flood knocked out utility power to Bramwell's primary data center for a projected 30-plus hours. The crisis team activated the documented ICT continuity plan, failed core banking services over to the secondary site, and had customer-facing services restored in 3 hours and 42 minutes against a stated RTO of 4 hours — inside target, with eighteen minutes to spare. Because the failover procedure had been rehearsed four times in the preceding year, the team executing it during the real event had already made and corrected their mistakes in a test environment rather than live. Bramwell's own estimate, based on contractual SLA penalty clauses with corporate clients, put the avoided cost at roughly $890,000 in penalties and remediation credits that a longer outage would have triggered. Their ISO 27001 surveillance audit that year cited the test reports and the real-event recovery time as strong, specific evidence for Control 5.30 — exactly the kind of measured, dated, comparable evidence auditors are looking for rather than a general assertion of readiness.
Case Study: The Tabletop That Found the Single Point of Failure — Verdant Cloud
Verdant Cloud, a mid-sized SaaS analytics provider, believed its disaster recovery posture was solid: a fully built secondary environment in a different cloud region, replicated data, and a documented failover runbook. During a routine semi-annual simulation exercise — not even a full failover test, just a functional walkthrough of the failover procedure — the exercise facilitator asked a simple question: "Where does the failover script authenticate to reach the secondary environment?" Nobody in the room could answer with confidence, so they checked.
It turned out the failover automation authenticated using an administrative credential vaulted in the same primary identity provider tenant that the production environment depended on. If the incident that triggered failover were an identity-provider-level compromise or outage rather than an infrastructure failure, the "secondary" environment would have been just as unreachable as the primary — a single point of failure hiding inside a plan that looked, on paper, like proper redundancy. Verdant Cloud's disaster recovery architect, Sam Okafor, flagged it as a critical finding, and the team spent six weeks building an independent, out-of-band authentication path for failover operations before the next test cycle. No customer was ever affected — the entire value of the finding was that it surfaced during a low-stakes tabletop exercise instead of during an actual crisis, which is precisely the argument for running these exercises even when the team is confident nothing will be found.
"The best outcome of a DR test isn't 'everything passed.' It's finding the thing that would have quietly ruined your day during a real event, while the cost of finding it is just an afternoon and some awkward questions." — Sam Okafor, Disaster Recovery Architect, Verdant Cloud
Case Study: Security That Held Through a Real Disruption — Halcyon Regional Medical Center
Halcyon Regional Medical Center, a 340-bed hospital system, faced a genuine test of Control 5.29 rather than 5.30 when a transformer failure knocked out grid power to its main campus for over 30 hours during a heat wave, well beyond what its generator fuel contracts had originally assumed. Clinical operations continued on backup power, but IT leadership faced real pressure to simplify authentication and loosen access logging on the emergency systems being stood up to keep electronic health records available to clinicians working extended, exhausted shifts.
Because Halcyon had a documented 5.29 procedure — including a pre-approved, time-boxed emergency access role for clinical staff that preserved full audit logging and required re-authentication every four hours rather than disabling authentication outright — the hospital maintained continuous access logging and role-based restrictions on patient records throughout the outage. When the event was over, the security team's post-disruption review confirmed no unauthorized access had occurred and closed out all emergency access grants on schedule, with the full audit trail intact. Halcyon's compliance officer later noted this was the first time the hospital could show a regulator concrete evidence that emergency operating procedures did not create a gap in PHI access controls — turning what could have been a HIPAA exposure narrative into a genuine compliance success story, and directly satisfying the kind of evidence an ISO 27001 auditor looks for under Control 5.29.
The 90-Day Roadmap: Standing Up a 5.29/5.30 Program from Scratch
Phase | Weeks | Key Activities | Output |
|---|---|---|---|
1. Scoping and sponsorship | 1–2 | Secure executive sponsor; confirm which systems/processes are in ISMS scope | Signed continuity program charter |
2. Business Impact Analysis | 3–6 | Interview process owners; quantify impact by time interval; assign criticality tiers | BIA report, draft RTO/RPO register |
3. Gap assessment | 6–8 | Compare current backup (8.13) and redundancy (8.14) capability against BIA targets | Gap register with remediation priorities |
4. Plan and procedure drafting | 8–12 | Draft ICT Continuity Plan (5.30) and Information Security During Disruption procedure (5.29) | Reviewed, approved draft plans |
5. Initial testing | 12–14 | Run first tabletop exercise per Tier 1/2 system; document gaps | First exercise report, action log |
6. Remediation | 14–16 | Close identified gaps: technical, procedural, or documentation | Updated plans, closed action items |
7. Formal test cycle established | 16–18 | Schedule recurring test cadence per the testing table above; assign ownership | Standing test calendar, ownership RACI |
8. Management review and SoA update | 18–20 | Present results to management; update Statement of Applicability with evidence references | Signed management review minutes, updated SoA |
Twenty weeks is an aggressive but realistic timeline for a mid-market organization starting from nothing more than an informal DR runbook — larger, more complex environments with extensive supplier dependencies will run longer, particularly the gap assessment and remediation phases if redundant infrastructure needs to be procured and built rather than merely configured.
Budgeting for Continuity: What This Actually Costs
One reason 5.29/5.30 programs stall is that nobody puts a realistic number in front of leadership until the redundancy bill arrives unexpectedly. It's worth budgeting continuity work in phases so the investment tracks the criticality tiering the BIA produces, rather than either starving Tier 1 systems of protection or gold-plating systems that don't need it.
Cost Category | Tier 1 (Critical) | Tier 2 (High) | Tier 3 (Standard) |
|---|---|---|---|
Backup infrastructure (immutable/offsite, Control 8.13) | Continuous or near-continuous replication | Daily incremental with periodic full backups | Weekly full backup, longer retention |
Redundant processing capacity (Control 8.14) | Hot/active-active standby site or region | Warm standby, manual or semi-automated failover | None required beyond backup restore capability |
Testing program cost | Full annual technical failover test plus quarterly simulations | Semi-annual simulation exercises | Annual tabletop/document review only |
Staffing/coordination overhead | Dedicated ICT continuity owner, on-call rotation | Shared responsibility within infrastructure team | Assigned as part of broader IT operations duties |
Typical relative annual spend | Highest — often 3–5x the Tier 2 figure per system | Moderate | Low |
The financial argument I make to skeptical finance leaders is straightforward: price the redundancy investment against the BIA's own quantified impact figures. If a Tier 1 system's MTPD analysis shows $40,000 an hour in lost revenue and contractual penalty exposure, a six-figure annual investment in hot standby capacity for that one system is a straightforward return calculation, not a discretionary IT expense — and having that BIA-derived dollar figure in hand is exactly what turns a budget request into an approved line item instead of a deferred one.
"Every continuity budget conversation gets easier once you stop asking for money for 'disaster recovery' in the abstract and start asking for money to protect a specific, quantified number the BIA already produced. Finance understands return on a defined risk. They don't approve vague resilience spending." — Tom Reyes, CISO, Solandra Logistics (successor role established post-incident)
Metrics That Matter: Measuring Continuity Program Maturity
Beyond pass/fail test results, a handful of ongoing metrics tell you whether a 5.29/5.30 program is maturing or quietly decaying between audits. I ask every client to track these on a standing dashboard reviewed at least quarterly by the continuity sponsor.
Metric | What It Tells You | Healthy Target |
|---|---|---|
Percentage of Tier 1 systems with a BIA-derived, documented RTO/RPO | Coverage of the foundational requirement | 100% |
Percentage of Tier 1 systems tested in the last 12 months | Whether testing is actually happening, not just scheduled | 100% for Tier 1, 80%+ for Tier 2 |
Average variance between target RTO and actual tested recovery time | Whether recovery capability is keeping pace with stated objectives | Actual within target, or documented risk acceptance if not |
Number of open findings from the last continuity test, by age | Whether remediation is closing gaps before the next test cycle | Zero findings older than one test cycle |
Number of 5.29 emergency exceptions granted vs. closed out on schedule | Whether degraded-mode procedures are actually being followed and reversed | 100% closure within defined time-box |
Continuity plan currency (days since last review vs. last significant architecture change) | Whether plans are being kept aligned with a changing environment | Reviewed within 90 days of any material change |
A dashboard like this converts continuity from an annual audit scramble into an ongoing management discipline, and it gives the continuity sponsor something concrete to bring into management review meetings rather than a qualitative "things seem fine" update — which is exactly the kind of evidence trail that supports the Clause 9 performance evaluation and internal audit process auditors expect to see feeding back into continual improvement.
The Strategic Case: Continuity as Competitive Advantage, Not Just Compliance
It's tempting to treat 5.29 and 5.30 as defensive box-checking — the controls you implement so an auditor doesn't flag a gap. That framing undersells what a genuinely tested continuity capability is worth commercially. Enterprise procurement teams increasingly ask for specific evidence of recovery testing, not just a policy statement, before signing contracts with meaningful data or availability dependencies; being able to hand over a dated test report showing you met your stated RTO is a sales asset, not just an audit artifact. Cyber insurers are asking the same questions during underwriting, and organizations that can demonstrate tested, documented continuity capability are increasingly seeing that reflected in premium and coverage terms. And internally, a leadership team that has actually watched a tabletop exercise surface a gap — the way Verdant Cloud did — walks away with a materially different, more realistic sense of organizational risk than one that has only ever read a plan document.
The broader business case for ISO 27001 generally — and continuity controls specifically as one of its more commercially visible components — is covered in more depth in our piece on ISO 27001 certification benefits and business case, and if you're still building the foundational case for why an ISMS matters at all, the core concepts behind an Information Security Management System is the right starting point before diving into individual controls like these.
Solandra Logistics, eighteen months after its $4.2 million incident, now runs quarterly failover tests, has immutable offline backups isolated from production credentials entirely, and has a documented, rehearsed 5.29 procedure that names exactly who can authorize a security control exception during a live incident and for how long. Derek Voss still keeps the original forensic timeline from that March incident pinned above his desk — not as a trophy, but as the answer to the question every new hire on his team eventually asks: why does this recovery runbook have so many steps that seem like they'd slow us down during an emergency? Every one of those steps has a name and a date attached to the incident that taught it.
If your organization is heading into a certification audit, or has 5.29/5.30 marked as "implemented" in a Statement of Applicability that nobody has stress-tested, don't wait for your own version of that 2:47 a.m. phone call to find out whether the plan actually works. Start with a real BIA, build the RTO/RPO register it produces, and put a test date on the calendar before an auditor — or an incident — puts one there for you. PentesterWorld's ISO 27001 Gap Analysis Tool is a practical starting point for benchmarking exactly where your current continuity documentation and testing evidence stand against what Controls 5.29 and 5.30 actually require, and our ISO 27001 Risk Register Template gives you a structured place to log the continuity gaps a BIA surfaces so they get tracked and treated rather than forgotten in a slide deck.
For teams building the underlying documentation set from scratch, our ISO 27001 Mandatory Documents Checklist will show you exactly where a BIA report, ICT continuity plan, and 5.29 disruption procedure fit alongside the rest of your required ISMS documentation, and the Annex A — All 93 Controls at a Glance cheat sheet is a useful quick reference for keeping 5.29 and 5.30 correctly positioned against the backup, redundancy, and incident management controls they depend on. If you're building an entire ISO 27001 program rather than just the continuity slice of it, The Complete ISO 27001 Implementation Guide walks through the full certification path end to end.
Continuity, done properly, is one of the few compliance investments that pays for itself the very first time you don't have to make Derek Voss's 2:47 a.m. phone call — or worse, the second call, ninety-six hours later, telling him someone got back in.
