Home / Disaster Recovery / DR layer by layer
DR guide

Disaster recovery, layer by layer

A DR plan usually starts with servers and storage, because that is where the money goes. But a site failover only works if every layer has its own answer: how users find the DR site, how the network reaches it, how each database is promoted, and how anyone runs the recovery when the primary site is gone. This guide walks through each layer, what has to be true at failover, and the common way to do it on premises and in AWS, Azure and Oracle Cloud.

The layers at a glance

Scroll sideways to see the whole diagram →
USERS AND TRAFFIC STEERINGPRIMARY SITEDR SITE OR CLOUD REGIONOPERATIONS, OUTSIDE BOTH SITESUsers and partnersbranches, web, mobile, APIsGlobal traffic managerDNS or GSLB with health checksNetwork edgefirewalls, load balancers, WAN linksApplication serversVMs or containers, activeDatabaseprimary, takes all writesStorage and backupsblock, file, object, backup copiesCache and messagingRedis, Kafka, MQNetwork edgesame rules, pre-built and testedApplication serverswarm, or built from images on demandDatabase standbyreplica, promoted at failoverStorage replicaplus an immutable backup copyCache and messagingmirrored, or rebuilt at failoverMonitoringreplication lag, health, alertsDR orchestrationrunbooks, ordered start-upConfiguration drift checkis DR still the same as prod?2activeon failoverDCI + backup VPNimages, IaC3log shipping4replicationmirroring5start-up order6lag, health1User or API trafficControl / API callScheduled copyData / replicationLogging / management
1. Users reach the service through a global traffic manager that health-checks both sites. 2. In normal running it sends everyone to the primary site; on failover it switches to the DR site. 3. Database changes are shipped continuously to a standby that is promoted at failover. 4. Storage and backups are replicated, with at least one copy that cannot be altered. 5. An orchestration runbook starts the DR layers in the right order, bottom up. 6. Monitoring and drift checks run outside both sites, so they still work when one site is lost.

What each layer needs, and how it is usually done

Names change between platforms; the job of each layer does not. The right-hand columns are the usual building blocks, not the only options.

LayerWhat must be true at failoverOn premises or private cloudAWSAzureOracle Cloud
Traffic steeringSend users to the healthy site automatically, or with one decisionGSLB on F5 BIG-IP DNS or Citrix ADC; DNS failover with short TTLsRoute 53 health-check failover; Global AcceleratorTraffic Manager; Front DoorDNS traffic management steering policies
NetworkDR has the same firewall rules, routes and partner connectivity, already in placeTwo diverse data-centre interconnects; MPLS or leased line with IPsec VPN backup; BGPDirect Connect with Site-to-Site VPN backup; Transit GatewayExpressRoute with VPN Gateway backup; Virtual WANFastConnect with IPsec VPN backup; Dynamic Routing Gateway
Compute and applicationsServers exist at DR, or can be built from tested images within the RTOVM replication (vSphere Replication, Zerto, Veeam); OpenStack images with Heat or TerraformElastic Disaster Recovery; AMIs with Auto Scaling groupsAzure Site Recovery; VM images with scale setsFull Stack Disaster Recovery; custom images with instance pools
DatabasesA consistent copy within the RPO, and a tested way to promote itOracle Data Guard; SQL Server Always On availability groups; PostgreSQL and MySQL replicationRDS cross-region read replicas; Aurora Global Database; DynamoDB global tablesSQL Database failover groups; SQL Managed Instance failover groups; Cosmos DB multi-regionAutonomous Data Guard across regions; Base Database Data Guard
Storage, files and backupData replicated to DR, plus a backup copy that ransomware cannot deleteArray replication; Ceph RBD mirroring; backup copies to a second site and to immutable storageS3 Cross-Region Replication; AWS Backup cross-region copy with Vault LockGeo-redundant storage; Azure Backup cross-region restore with immutable vaultsObject Storage replication; block volume cross-region replication
Cache and messagingQueues and topics exist at DR; decide which in-flight messages may be lostRedis replicas; Kafka MirrorMaker 2; IBM MQ or RabbitMQ federationElastiCache Global Datastore; MSK Replicator; SQS and SNS set up per regionAzure Cache for Redis geo-replication; Service Bus geo-disaster recovery; Event Hubs geo-replicationOCI Cache rebuilt at DR; Streaming handled at application level
OperationsMonitoring, runbooks and access still work when the primary site is goneMonitoring and identity hosted outside the primary site; runbooks in an orchestration toolCloudWatch cross-account, cross-region dashboards; AWS Config; Systems ManagerAzure Monitor; Azure Policy; Automation runbooksMonitoring; Cloud Guard; Resource Manager stacks

The database decision: synchronous or asynchronous

Synchronous replication confirms each write at both sites before telling the application it is done. Nothing committed is lost, so the RPO is zero. The cost is latency on every write, which is why it is only practical when the sites are close: in practice within a metro, with a round trip of a few milliseconds.

Asynchronous replication confirms the write locally and ships it to DR a moment later. Writes stay fast and the sites can be hundreds of kilometres apart, but anything still in flight when the primary fails is lost. The RPO is the replication lag, usually seconds to a few minutes, so that lag needs an alert.

Many regulated environments use both: a synchronous copy in a near site for zero data loss, and an asynchronous copy in a far site for regional events. Our core banking DR scenario shows this three-site design in detail.

What breaks failover that is not in any layer

These are the items that most often stop a drill, because nobody owns them.

Identity

Active Directory, SSO and privileged access must run at DR independently. If login depends on the primary site, nobody can start the recovery.

DNS TTLs and hard-coded addresses

A 24-hour DNS TTL or an IP address written into a config file can add hours to an RTO measured in minutes.

Certificates and licences

Certificates issued only for primary hostnames, and licence servers that exist only in production, fail quietly on the day.

Partner allow-lists

Banks, payment gateways and SaaS providers often allow traffic only from known IP addresses. The DR addresses need to be on those lists before a disaster.

Start-up order

Applications started before their database or message broker fail and need manual restarts. The runbook should start layers bottom up.

Failback

Coming back to the primary site is a second migration. Plan how data written at DR returns, or the team stays at DR far longer than intended.

Keeping DR the same as production

A DR site that matched production on the day it was built drifts within months: new firewall rules, a new application server, a patched database version, a changed certificate. Each change made only in production is a failover problem waiting to be found.

  • Build both sites from the same code. Infrastructure as code (Terraform, Ansible, Heat or cloud templates) means a change is applied to both sites or to neither.
  • Compare configurations on a schedule. A weekly automated comparison of firewall rules, server versions and database parameters between the sites catches drift early.
  • Add “and at DR?” to change approval. One question on the change form prevents most drift.
  • Prove it with drills. See testing DR for drill types and a checklist.

Not sure where your DR stands?

Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.