Executive summary
Most organisations have a disaster recovery plan. Far fewer have disaster recovery that works when it is needed. The gap is rarely the technology. It is that recovery targets were never agreed with the business, every application was protected the same way, backups shared the same failure as production, and the plan was never rehearsed end to end.
This paper sets out a practical way to close that gap. It is written for the people who have to make DR decisions and defend them: CIOs and IT heads, infrastructure architects, and the risk and compliance teams who sign off.
Five things to take away
- Set targets per tier, not per company. Agree recovery time (RTO) and acceptable data loss (RPO) for each group of applications, starting from what an hour of downtime actually costs.
- Use two or three DR patterns, not one. A handful of critical systems need a warm or hot copy. Most applications do not. Tiering typically cuts DR running cost by 40 to 60 percent against protecting everything the same way.
- Replication is not backup. Replication copies ransomware and mistakes as fast as good data. You need immutable, versioned backups that attackers cannot reach.
- The runbook is part of the architecture. Most of the time lost in a real disaster is spent deciding, finding people and fixing what was never written down.
- Only a tested recovery counts. Measure RPO and RTO in drills, report them against targets, and fix what breaks. Untested DR is a guess.
Why DR fails in practice
Power failures, storage faults, network outages, fire, flooding, ransomware and plain human error all take data centres down. None of these are rare across a five-year horizon. The cost shows up as lost revenue, penalties, customer dissatisfaction, compliance findings and damage to reputation, and it grows with every hour of downtime.
When we review DR set-ups, the same failure modes appear again and again:
| What we find | What happens on the day |
|---|---|
| Targets nobody agreed | IT built for 4 hours; the business assumed 30 minutes. The argument happens during the outage. |
| Everything protected the same way | Either the budget runs out before critical systems are covered, or money is wasted keeping test servers hot. |
| Backups in the same failure domain | Backup server in the same room, on the same power feed, or reachable with the same admin password as production. |
| Replication treated as backup | Ransomware-encrypted files replicate to DR within minutes. Both sites are now encrypted. |
| Only the data is copied | Volumes are at DR, but firewall rules, certificates, DNS, licence servers and identity are not. |
| No one decides | There is no named person or agreed criteria for declaring a disaster. Hours pass before failover starts. |
| Never tested end to end | The first full failover happens during a real disaster, and it is also the first time failback is attempted. |
RPO and RTO, set properly
Recovery Point Objective (RPO) is the maximum amount of data, measured in time, the business can afford to lose. If data is replicated every five minutes, up to five minutes of work can be lost. Recovery Time Objective (RTO) is how long a service can be unavailable before the impact becomes unacceptable. An RPO of 15 minutes and an RTO of 2 hours means: lose at most the last 15 minutes of work, and be running again within 2 hours.
Two rules of thumb follow. A lower RPO generally requires more replication and more bandwidth. A lower RTO generally requires more infrastructure running at the DR site. Both cost money, so the targets must come from business impact, not from preference.
A short business impact analysis
For each application, ask its business owner four questions and write the answers down:
- What does one hour of downtime cost? Revenue, penalties, patient or citizen impact, regulatory exposure.
- How much recent data can be re-entered by hand from paper, email or another system?
- Does a regulator, contract or customer SLA set a number?
- What else depends on this system, and what does it depend on?
Then group applications into tiers. Three tiers are enough for most organisations:
| Tier | Typical systems | RPO target | RTO target |
|---|---|---|---|
| Tier 1: critical | Core banking, payments, ERP transactions, hospital information system, citizen-facing portals | Near zero to 15 minutes | 15 minutes to 2 hours |
| Tier 2: important | Customer portals, MIS, email, document management, HR and payroll | 15 minutes to 4 hours | 4 to 12 hours |
| Tier 3: deferrable | Development and test, archives, analytics sandboxes, internal tools | 24 hours | 1 to 3 days |
The five DR patterns
There are five common ways to set up disaster recovery. They differ in what is already running at the DR site before a disaster, and therefore in recovery time, data loss and cost.
| Pattern | What is ready at DR | Typical RPO | Typical RTO | DR running cost vs production |
|---|---|---|---|---|
| Backup and restore | Backup copies only. Servers are rebuilt after the disaster. | Hours (time since last backup) | Hours to days | About 10 to 20% |
| Pilot light | Data replicated. Core services such as database and directory run small. Application servers are off. | Minutes | A few hours | About 20 to 35% |
| Warm standby | The full stack runs at reduced size and is scaled up on failover. | Seconds to minutes | Under an hour to a few hours | About 40 to 60% |
| Active-passive | A full-size copy is ready to take over. | Near zero (synchronous) to minutes | Minutes | About 80 to 100% |
| Active-active | Both sites serve users at the same time. | Near zero | Near zero for users | 100% or more, plus application changes |
Backup and restore
Scheduled backups or volume snapshots are copied to storage at another site, ideally object storage with immutability. After a disaster, infrastructure is rebuilt and data restored. Cheap to run, slow to recover, and the RPO is the time since the last good backup. Right for tier 3 and for long-term retention behind every other pattern.
Pilot light
The data layer is kept live at the DR site: database replication, for example asynchronous log shipping or a Galera or Always On replica, and storage mirroring on a schedule. Application and web servers exist only as templates and are started on failover. Recovery takes hours because servers must boot, attach data and be checked, but running cost stays low.
Warm standby
A scaled-down copy of the whole stack runs continuously at DR, with continuous block or database replication. On failover it is scaled up and traffic is redirected. Because everything is already running, it can be tested without disruption, which is the pattern’s biggest practical advantage.
Active-passive
A full-size copy stands ready. With synchronous replication, every write is confirmed at both sites, which gives near-zero data loss but limits the distance between sites to roughly 100 km and adds write latency. Network gateways and load balancers are pre-configured so failover is a decision, not a build.
Active-active
Both sites serve users at the same time behind a global load balancer, and the loss of one is close to invisible. It is the most expensive pattern and usually needs application changes, because two sites writing to the same data must avoid conflicts and split-brain. Worth it only for services designed for it from day one.
Choosing a strategy: a four-question framework
The choice of pattern follows from four questions, asked per tier:
- What does an hour of downtime cost? If nobody can put a number or a consequence on it, the system probably does not need the expensive patterns.
- How much recent data can be re-entered? If staff can re-key the last 30 minutes, asynchronous replication is usually enough and synchronous replication is not needed.
- Does a regulator or contract set the number? Banks, insurers, market intermediaries and many government systems have written continuity requirements. Start from those.
- Can the application be changed? If not, active-active is off the table, and active-passive or warm standby is the realistic ceiling.
Most organisations land on the same shape: active-passive or warm standby for a handful of critical systems, pilot light for important business applications, and backup and restore for everything else, with immutable backups behind all three.
| Sector | Typical tier 1 pattern | What drives it |
|---|---|---|
| Banking and financial services | Active-passive, sometimes three sites | Round-the-clock transactions, regulator DR drills, near-zero data loss |
| Government services | Warm standby or active-passive | Citizen services during emergencies, data residency, audit trail |
| Healthcare | Warm standby | Clinical systems need fast recovery; records must stay on-premises |
| Manufacturing | Pilot light, with warm standby for MES | Plant continuity at balanced cost; plants often keep running locally |
| AI and GPU platforms | Backup of data and checkpoints; warm standby for inference | GPUs are too costly to duplicate idle; training restarts from checkpoints |
Reference architecture
The diagram shows the tiered design most organisations end up with. Users reach applications through a global DNS or load balancer with short time-to-live values, so traffic can be moved in minutes. Tier 1 runs warm at the DR site, tier 2 is a pilot light, and tier 3 relies on backups.
Numbered flows: (1) database changes replicate asynchronously to the DR replica, (2) storage is mirrored on a schedule, with a shorter interval for tier 1 volumes, (3) immutable backups are written nightly to a vault that production credentials cannot reach. Restores from the vault are tested regularly in an isolated area.
What must exist at the DR site besides data
- Identity: a directory controller and the same roles and groups, so people can log in.
- Network: firewall rules, VPN and partner links, IP plan or DNS changes, all kept as code.
- Certificates and secrets: renewed at both sites on the same schedule.
- Licences: licence servers and any standby licence entitlements, checked against contracts.
- Monitoring and logging: so the team can see what is happening after failover.
- Runbooks and contact lists: stored somewhere reachable when the primary site is not.
Sizing and cost
Replication bandwidth
Bandwidth is driven by how much data changes, not by how much data exists. A useful first estimate:
For example, 400 GB of changes a day, most of it in a 10-hour business window, needs about 400 × 8,000 ÷ 36,000 × 1.3, or roughly 115 Mbps of sustained replication bandwidth. Measure change rates from storage or database logs over at least two weeks, including month end, before buying links. Provision two links from different providers.
Storage
- DR replica: at least one full copy of protected data, plus room for snapshots at the DR site.
- Backups: retention multiplied by daily change, after deduplication. A 30-day retention commonly needs 1.5 to 3 times the protected data.
- Seeding: tens of terabytes can take weeks over a WAN link. Shipping encrypted disks for the first copy is often faster.
A worked example
Production has 40 virtual machines and 30 TB of data. After business impact analysis, 8 VMs are tier 1, 20 are tier 2 and 12 are tier 3.
| Option A: everything active-passive | Option B: protected by tier | |
|---|---|---|
| What runs at DR every day | All 40 VMs at full size | About 7 VM-equivalents: tier 1 at half size, tier 2 databases small |
| Tier 1 recovery | Minutes | Same: minutes |
| Tier 2 recovery | Minutes | A few hours |
| Tier 3 recovery | Minutes | One to two days from backup |
| DR running cost vs production | About 80 to 100% | About 35 to 50% |
Where the DR site is on-premises or in colocation, hardware for tiers 1 and 2 must still be ready, because servers cannot be bought on the day; the saving is in power, licences and tier 3 capacity. In cloud DR, full capacity is created only on drill or disaster day, which is where tiering saves the most.
Costs that are often missed
- Failback: moving data back after a cloud failover can carry large egress charges.
- Licences: some vendors charge in full for a standby database or hypervisor. Read the contract.
- Drills: a full failover test takes several people most of a day, every year.
- Keeping DR current: patching templates, renewing certificates and updating firewall rules at DR is ongoing work.
Ransomware changes the design
Classic DR assumes the primary site is lost but the data is good. Ransomware breaks that assumption: the data itself is the problem, and attackers deliberately go after backups and DR systems first. Replication makes it worse, because encrypted files replicate to the DR site within minutes.
A design that survives ransomware adds four things:
| Control | Why it matters |
|---|---|
| Immutable backups | Object lock or write-once storage, so backups cannot be changed or deleted for the retention period, even by an administrator. |
| Separate credentials and management | Backup and DR systems use their own accounts and multi-factor authentication, not the production directory. |
| An isolated copy | At least one copy in a vault that opens only for scheduled replication windows, or offline media. |
| Clean restore points and a clean room | Backups are scanned, and restores happen first into an isolated environment where identity and critical apps are rebuilt and checked before reconnecting. |
Testing: turning a plan into evidence
DR that has never been tested is a guess. Most problems appear the first time someone actually fails over: a missing firewall rule, an expired certificate, a licence server that only exists in production. Testing finds these while there is no real disaster.
| Test | What happens | Suggested frequency |
|---|---|---|
| Tabletop walk-through | The team reads the runbook step by step and checks every step, name and phone number. | Every quarter |
| Restore test | Restore one server or database from backup and confirm it works. | Every month, rotating systems |
| Isolated failover | Bring a tier up at DR on an isolated network, without touching production users. | Twice a year for each tier |
| Full failover and failback | Move real users to DR for an agreed window, then move back. | Once a year for tier 1, or as a regulator requires |
On drill day
- Before: freeze changes, confirm replication lag is within RPO, tell users and the helpdesk, and write down what success means.
- During: time every step, record every manual fix however small, have the application owner test, and agree a hard stop for rolling back.
- After: report measured RPO and RTO against targets, turn every issue into a fix with an owner and a date, update the runbook that week, and book the next drill.
Keeping DR working between drills
- Alert on replication lag. Lag that quietly grows past the RPO is the most common silent failure.
- Add “and at DR?” to every change: new applications, firewall changes, certificate renewals.
- Patch DR on the same schedule as production, and upgrade both sites together.
- Review tiers once a year. Last year’s tier 3 system may now be critical.
A practical roadmap
For a mid-sized environment, a DR programme typically runs in four stages. Procurement of links, hardware and DR site space usually sets the pace.
| Stage | Typical duration | Outcome |
|---|---|---|
| 1. Assess | 3 to 6 weeks | Application inventory, dependencies, business impact analysis, agreed tiers and targets, gap report. |
| 2. Design | 3 to 4 weeks | Pattern per tier, reference architecture, sizing, bandwidth, cost model, runbook outline. |
| 3. Build | 6 to 16 weeks | DR site, links, replication, backup vault, automation, monitoring of replication lag. |
| 4. Prove | 2 to 4 weeks, then ongoing | First isolated failover per tier, first full drill for tier 1, evidence pack, test calendar. |
Ten common mistakes
- One RPO and RTO for the whole company. Either too expensive or not protective enough.
- Choosing a product before agreeing targets. The product then defines the targets.
- Treating replication as backup. It copies ransomware and mistakes just as fast.
- DR site sharing risks with production. Same building, power feed, flood zone or admin accounts.
- Copying data but not configuration. Networks, certificates and licences rebuilt by hand under pressure.
- No named decision-maker. Hours lost before anyone declares the disaster.
- Long DNS time-to-live values. Failover is done but users keep going to the dead site for hours.
- Never testing failback. Getting back is often the harder half.
- Not monitoring replication lag. The RPO is silently missed for weeks.
- Letting the two sites drift. Different software versions cause surprises on failover day.
DR readiness checklist (40 points)
Use this list to score your current position. Anything you cannot tick with evidence is a gap worth closing.
Governance
- A named owner for DR and a named person who can declare a disaster
- Agreed criteria for declaring a disaster
- RPO and RTO agreed and signed off per tier by business owners
- Regulatory and contractual DR requirements listed
- DR included in change management (“and at DR?”)
The remaining 35 points are in the PDF, laid out as a printable checklist.
Get the full checklistGlossary
| Term | Meaning |
|---|---|
| RPO | Recovery Point Objective: maximum acceptable data loss, measured in time. |
| RTO | Recovery Time Objective: maximum acceptable downtime for a service. |
| Synchronous replication | Each write is confirmed at both sites before it completes. Near-zero data loss, distance-limited. |
| Asynchronous replication | Writes are copied to the second site shortly after. Small possible data loss, any distance. |
| Failover / failback | Moving service to the DR site, and later moving it back. |
| Immutable backup | A backup that cannot be changed or deleted for its retention period. |
| Split-brain | Both sites believe they are primary and accept writes, causing conflicting data. |
| Tabletop exercise | A walk-through of the runbook without touching systems. |
About Vakratron Systems
Vakratron Systems is a vendor-neutral infrastructure design firm. We design data centre, disaster recovery, cloud, GPU and AI platforms for enterprises and government buyers, write our assumptions down, and stay with a design until it is running and tested.