Business Continuity for IT
Business continuity for IT is the planning and preparation that keeps critical technology services running, or recovers them quickly, during disruptive events like outages, disasters, or cyberattacks. It connects recovery objectives to organizational priorities and tests that the plans actually work.
itPlatform engineering and SRE | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic — Business Continuity for IT
Business continuity is the discipline of keeping essential work going when normal conditions have taken the afternoon off. It is not a promise that nothing fails. It is the plan for deciding what matters first, what can run in a reduced form, and what proof is needed before declaring the situation less alarming.
The trap is to begin with a server list. Servers are very keen to look important, particularly when arranged in a spreadsheet. But payroll depends on approvals, employee data, banking connectivity, and identity. An order service depends on payment, inventory, fulfillment, and support. The useful starting point is the business impact analysis, or BIA: identify the outcome, the interruption it can tolerate, and every person, service, data source, and supplier needed to keep it moving.
Three clocks keep the argument honest. Maximum tolerable downtime says when a business interruption becomes unacceptable. Recovery time objective, or RTO, says how long restoration may take. Recovery point objective, or RPO, says how much recent data may be lost. They are not decorative numbers. If identity, data, application recovery, and validation take longer than the stated RTO, the RTO has made a brave suggestion rather than described a capability.
Recovery also comes in an awkwardly sensible order. First comes activation and notification: decide what happened and who may act. Then recover prerequisites before dependent services. Finally comes reconstitution, the controlled return to normal work, including reconciliation of anything done through a manual workaround. A recovered application with no usable identity service is a beautifully restored door with no handle.
The other surprise is that a backup job is not a recovery test. Training prepares people for roles. An exercise lets them rehearse decisions. A technical test shows whether systems and components work. The evidence that matters includes timestamps, validation results, decisions, gaps, owners, and retests. A document with a recent date is pleasant. A measured restore and a current dependency map are more convincing.
Read the Introduction when the distinction between continuity, disaster recovery, contingency planning, and incident response is still doing a small dance. Use the Slides for the dependency path and time-objective checks. Keep the Cheatsheet nearby when building procedures or exercises. The Practice Reference turns the idea into a BIA and tabletop method; the Exercise asks you to make the chain fit. After that, the plan has become less of a document and more of a capability, which is what it was trying to be all along.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://nvlpubs.nist.gov/nistpubs/legacy/sp/nistspecialpublication800-34r1.pdf
Supports
- Relationship among business continuity, disaster recovery, incident response, and information system contingency plans
- Seven-step contingency planning process
- Business impact analysis, maximum tolerable downtime, recovery time objective, and recovery point objective
- Dependency-aware recovery priorities and recovery strategy tradeoffs
- Preventive controls, backups, alternate processing, alternate sites, and recovery capabilities
- Activation and notification, recovery, and reconstitution phases
- Plan roles, contacts, procedures, storage, testing, exercises, and maintenance
- Quiz answers concerning impact analysis, objectives, dependencies, restore evidence, and reconstitution
- https://csrc.nist.gov/topics/security-and-privacy/security-programs-and-operations/contingency-planning
Supports
- Contingency planning as coordinated procedures and technical measures for recovering systems, operations, and data
- Alternate equipment, alternate processing, and alternate locations as recovery approaches
- Link rationale for the NIST contingency planning topic
- https://www.ready.gov/business/emergency-plans
Supports
- Business continuity, crisis communications, emergency response, and IT disaster recovery as related preparedness concerns
- IT disaster recovery planning developed with the business continuity plan
- Technology recovery timed to business recovery needs
- Quiz answer distinguishing IT disaster recovery from broader business continuity
- Link rationale for the Ready.gov overview
- https://www.ready.gov/sites/default/files/2020-03/business-continuity-plan.pdf
Supports
- Business continuity plan sections for BIA results, recovery objectives, continuity strategies, and IT restoration
- Link rationale for applying course concepts to an official template
- https://csrc.nist.gov/pubs/sp/800/84/final
Supports
- Training personnel, exercising IT plans, and testing IT systems as distinct readiness activities
- Design, conduct, and evaluation of tabletop and functional exercises
- Using events to identify deficiencies and improve readiness
- Quiz answers concerning tabletop scope, restore evidence, and corrective action
- Link rationale for advancing from planning to exercises and tests
- 2006 publication milestone for IT plan testing, training, and exercising
- https://www.cisa.gov/topics/critical-infrastructure-security-and-resilience
Supports
- Availability of CISA tabletop packages for cyber, physical, and cyber-physical scenarios
- Package documentation for exercise planner, facilitator, evaluator, and participant roles
- Link rationale for extended exercise practice
- https://www.nist.gov/publications/nist-cybersecurity-framework-csf-20
Supports
- CSF 2.0 as organization-wide guidance for managing cybersecurity risk
- High-level outcomes for governance, communication, response, and recovery
- Link rationale connecting continuity work to cybersecurity risk management
- https://csrc.nist.gov/nist-cyber-history/risk-management/chapter
Supports
- 1981 FIPS 87, 1982 SP 500-85, and 1985 SP 500-134 contingency-planning milestones
- NIST's early automatic-data-processing contingency-planning history
- https://www.nist.gov/publications/contingency-planning-guide-information-technology-systems-recommendations-national
Supports
- 2002 publication of NIST SP 800-34
- Modern contingency-planning guidance for information technology systems
- https://www.iso.org/standard/75106.html
Supports
- ISO 22301:2012 as the prior published edition
- October 2019 publication of ISO 22301:2019 edition 2
- ISO 22301:2019 Amendment 1:2024
- Business continuity management-system requirements
- https://docs.aws.amazon.com/drs/latest/userguide/what-is-drs.html
Supports
- AWS Elastic Disaster Recovery replication, recovery drills, recovery, and failback capabilities
- https://azure.microsoft.com/en-us/products/site-recovery
Supports
- Azure Site Recovery as a cloud recovery orchestration product
- https://www.zerto.com/
Supports
- Zerto continuous data protection, testing, and recovery orchestration capabilities
- https://www.veeam.com/products/veeam-data-platform.html
Supports
- Veeam Data Platform backup, recovery, and recovery-orchestration capabilities
- https://www.commvault.com/security-cyber-readiness-and-recovery
Supports
- Commvault Cloud Cleanroom recovery testing and isolated recovery environment
- https://aws.amazon.com/blogs/mt/how-to-test-your-aws-elastic-disaster-recovery-implementation/
Supports
- Difference between isolated tooling validation and application-level recovery testing
- https://aws.amazon.com/blogs/architecture/minimizing-dependencies-in-a-disaster-recovery-plan/
Supports
- Testing disaster recovery plans with unavailable Route 53 and IAM control-plane access
