Disaster Recovery
Disaster recovery plans and implements the procedures for restoring IT systems after a catastrophic failure: natural disaster, cyberattack, or major outage. It defines recovery objectives, replication strategies, failover sites, and the testing that proves recovery actually works under pressure.
itPlatform engineering and SRE | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don’t Panic — Disaster Recovery
Disaster recovery is the arranged return of a technology service after a severe disruption. It exists because a running server, a successful backup job, and a working business service are three different creatures that only occasionally attend the same meeting. The aim is not to rescue every machine first. It is to restore the business outcome that depends on them.
The first useful pair of terms is RTO and RPO. The recovery time objective is the longest acceptable interruption. The recovery point objective is how far back recovered data may be. Together they stop recovery planning from becoming an expensive collection of reassuring nouns. Four hours and fifteen minutes, for example, are requirements for the path, not compliments for the plan.
That path begins with business impact. Name the process that must return, then trace its application, data, identity, network, facilities, suppliers, and the people authorized to decide. A database restored in splendid isolation does not help an authenticated customer complete a transaction. The surprise is that backup and replication are useful, but neither is a complete recovery capability. Replication can copy corruption with great enthusiasm.
Next comes the recovery posture. Backup and restore has low standby cost but takes longer. Pilot light keeps the core ready. Warm standby runs a smaller copy. Active-active serves from multiple locations and adds coordination and consistency demands. The label is not the decision. The complete path must meet the RTO, RPO, capacity, risk, and cost constraints, which is less glamorous but considerably more helpful when the primary environment is unavailable.
A recovery plan is the part that makes those choices executable. Activation criteria say when an authorized team begins. Recovery restores dependencies in order. Reconstitution validates a known state and returns to stable operation. The plan must remain reachable when the primary environment is not, because a recovery document trapped inside the disaster has made a brief but decisive career change.
Testing supplies the evidence. A document review finds stale instructions. A tabletop exposes decisions and coordination. A simulation exercises procedures. A full recovery test checks the complete path and its timing. Record the actual recovery time, recovered data point, failed steps, manual work, and hidden dependencies. Then correct the design and test it again.
Read the Course introduction for the full map and glossary. Use Slides for the dependency flow and recovery-posture comparisons. Keep Cheatsheet open while designing a runbook or reviewing a test. The Quiz is where a component check has to defend itself against a business-function check.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://csrc.nist.gov/pubs/sp/800/34/r1/upd1/final
Supports
- Official status, scope, history, and purpose of NIST SP 800-34 Revision 1
- Availability of business impact analysis and contingency-plan templates
- Relationship of contingency planning to resilience, risk management, and incident response
- https://nvlpubs.nist.gov/nistpubs/legacy/sp/nistspecialpublication800-34r1.pdf
Supports
- Contingency planning as coordinated recovery of systems, operations, and data after disruption
- Seven-step planning process covering policy, business impact analysis, controls, strategy, plans, testing, and maintenance
- Business impact analysis, maximum tolerable downtime, recovery time objective, and recovery point objective
- Recovery priorities, dependencies, alternate processing, backup, and recovery strategies
- Activation and notification, recovery, and reconstitution phases
- Roles, activation criteria, outage assessment, ordered procedures, validation, and return to a known state
- Checklist, tabletop, simulation, parallel, full-interruption, and full-recovery test concepts
- https://csrc.nist.gov/CSRC/media/Projects/risk-management/800-53%20Downloads/800-53r5/SP_800-53_v5_1-derived-OSCAL.pdf
Supports
- Contingency plan ownership, coordination, distribution, review, protection, and maintenance
- Testing plan effectiveness and readiness, reviewing results, and initiating corrective action
- Alternate processing and storage sites, capacity planning, system backup, and recovery
- Test methods from walk-through and tabletop exercises through simulations and full recovery
- https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/disaster-recovery-workloads-on-aws.html
Supports
- Disaster recovery as preparation for and recovery from a business-impacting workload disruption
- Relationship among business continuity, high availability, RTO, RPO, detection, and testing
- Cloud-provider features as inputs to workload recovery design rather than a complete plan
- https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/disaster-recovery-options-in-the-cloud.html
Supports
- Backup and restore, pilot light, warm standby, and multi-site active-active recovery patterns
- Cost, complexity, standby state, and recovery-speed tradeoffs among patterns
- Need to recover data, infrastructure, configuration, application code, and capacity
- Regular assessment and testing against RTO and RPO
- https://learn.microsoft.com/en-us/azure/reliability/concept-business-continuity-high-availability-disaster-recovery
Supports
- Distinction among business continuity, high availability, and disaster recovery
- RTO as acceptable downtime and RPO as acceptable data loss, set for workload flows
- Cost and difficulty of near-zero recovery objectives
- Need for documented and testable recovery plans, communications, escalation, and service guidance
- Separate backup locations, restore testing, integrity checks, and alignment of backup intervals with RPO
- Disaster scenarios including infrastructure loss, human error, data corruption, and security incidents
- https://csrc.nist.gov/nist-cyber-history/risk-management/chapter
Supports
- NIST contingency-planning publications from 1981 through SP 800-34 in 2002 and SP 800-53 in 2005.
- https://csrc.nist.gov/csrc/media/publications/shared/documents/itl-bulletin/itlbul2002-06.pdf
Supports
- Publication and scope of the 2002 NIST SP 800-34 Contingency Planning Guide for Information Technology Systems.
- https://aws.amazon.com/blogs/aws/amazon-web-services-for-backup-and-disaster-recovery/
Supports
- AWS discussion of cloud backup and disaster recovery published in 2010.
- https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/document-revisions.html
Supports
- Initial 2021 AWS disaster-recovery whitepaper and its 2022 active-passive failover and Elastic Disaster Recovery update.
- https://github.blog/news-insights/company-news/oct21-post-incident-analysis/
Supports
- GitHub 2018 database recovery, restore throughput constraints, replication catch-up, and the distinction between tested procedures and whole-cluster restoration.
- https://cloud.google.com/blog/products/management-tools/sre-principles-in-practice-for-business-continuity
Supports
- Google SRE disaster recovery testing exercises covering technical systems, processes, and people.
- https://aws.amazon.com/disaster-recovery/
Supports
- AWS Elastic Disaster Recovery recovery orchestration.
- https://learn.microsoft.com/en-us/azure/site-recovery/site-recovery-overview
Supports
- Azure Site Recovery replication, failover, and failback.
- https://cloud.google.com/backup-disaster-recovery
Supports
- Google Cloud Backup and DR service.
- https://www.veeam.com/products/data-platform.html
Supports
- Veeam Data Platform backup, replication, and recovery orchestration.
- https://www.rubrik.com/products/cyber-recovery
Supports
- Rubrik recovery plans, isolated recovery simulation, and clean-state restoration.
- https://www.zerto.com/
Supports
- Zerto continuous data protection and disaster-recovery orchestration.
- https://www.druva.com/
Supports
- Druva managed data protection and recovery services.
