Keeping DR the same as production
A DR site that matched production on the day it was built starts to drift within weeks. A firewall rule is added in production only. A server is patched in one place. A new application goes live without a DR copy. None of these shows up on a replication dashboard, and each one becomes a failure on the day you need DR. This guide covers where drift comes from and how to catch it before a drill or a disaster does.
How it fits together
What drifts, and how to catch it
| What drifts | How it usually happens | How to catch it |
|---|---|---|
| Firewall and network rules | An urgent rule is added in production during an incident and never copied | Compare rule sets between sites weekly; keep rules in code |
| Server and OS versions | Patching runs in production first and DR is forgotten | Patch both sites in the same change; report version differences |
| Database parameters and versions | Tuning or an upgrade is done on the primary only | Compare parameters and versions between primary and standby |
| New applications | A new system goes live with no DR tier assigned | Make a DR tier a go-live requirement; list systems with no DR copy |
| Certificates and secrets | A certificate is renewed for production hostnames only | Track expiry dates for both sites in one place |
| Capacity | Production grows; DR is still sized for last year | Compare CPU, memory and storage between sites each quarter |
| Access and identities | New administrators or service accounts exist only in production | Compare privileged accounts and groups between sites |
Four habits that prevent most drift
- Build both sites from the same code. Terraform, Ansible, Heat or cloud templates mean a change is applied to both sites or to neither. Golden images do the same for servers.
- Add one question to change approval. “And at DR?” on every change form catches most drift at the source, at almost no cost.
- Compare automatically, on a schedule. A weekly comparison of rules, versions, parameters and accounts, run from outside both sites, finds what slipped through.
- Report DR readiness like any other service measure. Three numbers are enough: replication lag against the RPO, open drift items, and days since the last successful drill.
A monthly DR readiness check
Replication
- Lag within RPO for every tier
- No paused or failed replication jobs
- Backup copies completed and locked
Configuration
- No open drift items older than 30 days
- Every production system has a DR tier
- Certificates valid at both sites
People and process
- Runbooks updated after the last change
- Contact list checked
- Next drill booked
Related guides
DR layer by layer
Traffic, network, databases, storage and messaging, with on-premises and cloud equivalents.
Read the guide →Testing DR
Drill types, how often to run them, and a drill-day checklist.
Read the guide →DR in the public cloud
Using a cloud region as your DR site: patterns, cost and the catches.
Read the guide →Not sure where your DR stands?
Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.