Home / Disaster Recovery / Keeping DR the same as production
DR guide

Keeping DR the same as production

A DR site that matched production on the day it was built starts to drift within weeks. A firewall rule is added in production only. A server is patched in one place. A new application goes live without a DR copy. None of these shows up on a replication dashboard, and each one becomes a failure on the day you need DR. This guide covers where drift comes from and how to catch it before a drill or a disaster does.

How it fits together

Scroll sideways to see the whole diagram →
SINGLE SOURCE OF TRUTHPRODUCTIONDR SITECHECKS, OUTSIDE BOTH SITESCode repositoryIaC, images, firewall rulesDeployment pipelineapplies each change to both sitesChange approvalasks: and at DR?Productionservers, network, databasesDR siteshould match productionDrift scannercompares both sites weeklyDrift reporteach gap with an owner and dateDR readiness figureslag, open drift, last drill2apply1345Control / API callLogging / managementException / alert
1. Infrastructure, images and firewall rules are defined as code in one repository. 2. A pipeline applies every approved change to production and DR together, so the two cannot diverge through normal work. 3. A scanner outside both sites reads the actual configuration of each site on a schedule. 4. Any difference becomes a report item with an owner and a date. 5. Open drift, replication lag and the date of the last successful drill are reported as DR readiness figures.

What drifts, and how to catch it

What driftsHow it usually happensHow to catch it
Firewall and network rulesAn urgent rule is added in production during an incident and never copiedCompare rule sets between sites weekly; keep rules in code
Server and OS versionsPatching runs in production first and DR is forgottenPatch both sites in the same change; report version differences
Database parameters and versionsTuning or an upgrade is done on the primary onlyCompare parameters and versions between primary and standby
New applicationsA new system goes live with no DR tier assignedMake a DR tier a go-live requirement; list systems with no DR copy
Certificates and secretsA certificate is renewed for production hostnames onlyTrack expiry dates for both sites in one place
CapacityProduction grows; DR is still sized for last yearCompare CPU, memory and storage between sites each quarter
Access and identitiesNew administrators or service accounts exist only in productionCompare privileged accounts and groups between sites

Four habits that prevent most drift

  1. Build both sites from the same code. Terraform, Ansible, Heat or cloud templates mean a change is applied to both sites or to neither. Golden images do the same for servers.
  2. Add one question to change approval. “And at DR?” on every change form catches most drift at the source, at almost no cost.
  3. Compare automatically, on a schedule. A weekly comparison of rules, versions, parameters and accounts, run from outside both sites, finds what slipped through.
  4. Report DR readiness like any other service measure. Three numbers are enough: replication lag against the RPO, open drift items, and days since the last successful drill.

A monthly DR readiness check

Replication

  • Lag within RPO for every tier
  • No paused or failed replication jobs
  • Backup copies completed and locked

Configuration

  • No open drift items older than 30 days
  • Every production system has a DR tier
  • Certificates valid at both sites

People and process

  • Runbooks updated after the last change
  • Contact list checked
  • Next drill booked

Not sure where your DR stands?

Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.