Disaster recovery, layer by layer
A DR plan usually starts with servers and storage, because that is where the money goes. But a site failover only works if every layer has its own answer: how users find the DR site, how the network reaches it, how each database is promoted, and how anyone runs the recovery when the primary site is gone. This guide walks through each layer, what has to be true at failover, and the common way to do it on premises and in AWS, Azure and Oracle Cloud.
The layers at a glance
What each layer needs, and how it is usually done
Names change between platforms; the job of each layer does not. The right-hand columns are the usual building blocks, not the only options.
| Layer | What must be true at failover | On premises or private cloud | AWS | Azure | Oracle Cloud |
|---|---|---|---|---|---|
| Traffic steering | Send users to the healthy site automatically, or with one decision | GSLB on F5 BIG-IP DNS or Citrix ADC; DNS failover with short TTLs | Route 53 health-check failover; Global Accelerator | Traffic Manager; Front Door | DNS traffic management steering policies |
| Network | DR has the same firewall rules, routes and partner connectivity, already in place | Two diverse data-centre interconnects; MPLS or leased line with IPsec VPN backup; BGP | Direct Connect with Site-to-Site VPN backup; Transit Gateway | ExpressRoute with VPN Gateway backup; Virtual WAN | FastConnect with IPsec VPN backup; Dynamic Routing Gateway |
| Compute and applications | Servers exist at DR, or can be built from tested images within the RTO | VM replication (vSphere Replication, Zerto, Veeam); OpenStack images with Heat or Terraform | Elastic Disaster Recovery; AMIs with Auto Scaling groups | Azure Site Recovery; VM images with scale sets | Full Stack Disaster Recovery; custom images with instance pools |
| Databases | A consistent copy within the RPO, and a tested way to promote it | Oracle Data Guard; SQL Server Always On availability groups; PostgreSQL and MySQL replication | RDS cross-region read replicas; Aurora Global Database; DynamoDB global tables | SQL Database failover groups; SQL Managed Instance failover groups; Cosmos DB multi-region | Autonomous Data Guard across regions; Base Database Data Guard |
| Storage, files and backup | Data replicated to DR, plus a backup copy that ransomware cannot delete | Array replication; Ceph RBD mirroring; backup copies to a second site and to immutable storage | S3 Cross-Region Replication; AWS Backup cross-region copy with Vault Lock | Geo-redundant storage; Azure Backup cross-region restore with immutable vaults | Object Storage replication; block volume cross-region replication |
| Cache and messaging | Queues and topics exist at DR; decide which in-flight messages may be lost | Redis replicas; Kafka MirrorMaker 2; IBM MQ or RabbitMQ federation | ElastiCache Global Datastore; MSK Replicator; SQS and SNS set up per region | Azure Cache for Redis geo-replication; Service Bus geo-disaster recovery; Event Hubs geo-replication | OCI Cache rebuilt at DR; Streaming handled at application level |
| Operations | Monitoring, runbooks and access still work when the primary site is gone | Monitoring and identity hosted outside the primary site; runbooks in an orchestration tool | CloudWatch cross-account, cross-region dashboards; AWS Config; Systems Manager | Azure Monitor; Azure Policy; Automation runbooks | Monitoring; Cloud Guard; Resource Manager stacks |
The database decision: synchronous or asynchronous
Synchronous replication confirms each write at both sites before telling the application it is done. Nothing committed is lost, so the RPO is zero. The cost is latency on every write, which is why it is only practical when the sites are close: in practice within a metro, with a round trip of a few milliseconds.
Asynchronous replication confirms the write locally and ships it to DR a moment later. Writes stay fast and the sites can be hundreds of kilometres apart, but anything still in flight when the primary fails is lost. The RPO is the replication lag, usually seconds to a few minutes, so that lag needs an alert.
Many regulated environments use both: a synchronous copy in a near site for zero data loss, and an asynchronous copy in a far site for regional events. Our core banking DR scenario shows this three-site design in detail.
What breaks failover that is not in any layer
These are the items that most often stop a drill, because nobody owns them.
Identity
Active Directory, SSO and privileged access must run at DR independently. If login depends on the primary site, nobody can start the recovery.
DNS TTLs and hard-coded addresses
A 24-hour DNS TTL or an IP address written into a config file can add hours to an RTO measured in minutes.
Certificates and licences
Certificates issued only for primary hostnames, and licence servers that exist only in production, fail quietly on the day.
Partner allow-lists
Banks, payment gateways and SaaS providers often allow traffic only from known IP addresses. The DR addresses need to be on those lists before a disaster.
Start-up order
Applications started before their database or message broker fail and need manual restarts. The runbook should start layers bottom up.
Failback
Coming back to the primary site is a second migration. Plan how data written at DR returns, or the team stays at DR far longer than intended.
Keeping DR the same as production
A DR site that matched production on the day it was built drifts within months: new firewall rules, a new application server, a patched database version, a changed certificate. Each change made only in production is a failover problem waiting to be found.
- Build both sites from the same code. Infrastructure as code (Terraform, Ansible, Heat or cloud templates) means a change is applied to both sites or to neither.
- Compare configurations on a schedule. A weekly automated comparison of firewall rules, server versions and database parameters between the sites catches drift early.
- Add “and at DR?” to change approval. One question on the change form prevents most drift.
- Prove it with drills. See testing DR for drill types and a checklist.
Related guides
Choosing a DR pattern
Backup and restore, pilot light, warm standby, active-passive and active-active.
Read the guide →What DR really costs
Where the money goes, with planning ranges and a worked example.
Read the guide →DR whitepaper
Tiers, patterns, sizing, ransomware, testing and a 40-point checklist.
Read or download →Not sure where your DR stands?
Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.